Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method

The method addresses the challenges of high throughput and latency in mesh data transmission by decoding and converting displacement and attribute data, enhancing 3D service quality and enabling applications like autonomous driving.

WO2026084533A1PCT designated stage Publication Date: 2026-04-23LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

The challenge lies in efficiently transmitting and receiving large amounts of mesh data due to the high throughput and latency requirements, as well as the complexity of encoding and decoding processes associated with 3D data such as point cloud or mesh data, which is crucial for applications like VR, AR, and autonomous driving.

Method used

A method and apparatus for decoding mesh data by separating and converting displacement and attribute data, utilizing a processor to decode base mesh, displacement, and packed data, and performing nominal format conversions to restore the mesh efficiently.

Benefits of technology

This approach enables high-quality 3D services and general-purpose applications such as autonomous driving by optimizing mesh data transmission and decoding processes, reducing latency and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016523_23042026_PF_FP_ABST
    Figure KR2025016523_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a decoding method and a decoding device. A mesh data decoding method according to the embodiments comprises the steps of: receiving a bitstream including mesh data; and decoding the mesh data, wherein the step of decoding the mesh data includes a step of decoding basemesh from a basemesh bitstream included in the mesh data, and the step of decoding the mesh data may further comprise the steps of: decoding displacement data from a displacement bitstream if same is included in the mesh data; and decoding packed data from a packed bitstream if same is included in the mesh data.
Need to check novelty before this filing date? Find Prior Art

Description

Mesh data transmission device, mesh data transmission method, mesh data reception device and mesh data reception method

[0001] The embodiments provide a method for providing 3D content to provide various services to users, such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.

[0002] Among 3D content, point cloud data or mesh data is a set of points in 3D space. However, there is a problem in that it is difficult to generate point cloud data or mesh data because there is a large amount of points in 3D space.

[0003] In other words, there is a problem in that a large amount of throughput is required to transmit and receive 3D data with a large amount of points, such as point cloud data or mesh data.

[0004] The technical problem according to the embodiments is to provide an apparatus and method for efficiently transmitting and receiving mesh data in order to solve the aforementioned problems, etc.

[0005] The technical problem according to the embodiments is to provide an apparatus and method for solving the latency and encoding / decoding complexity of mesh data.

[0006] The technical problem according to the embodiments is to provide an apparatus and method for improving the restoration performance of mesh data.

[0007] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that can be inferred by a person skilled in the art based on the entire content of this document.

[0008] To achieve the above-mentioned purpose and other advantages, the decoding method according to the embodiments may include the step of receiving a bitstream containing mesh data and the step of decoding said mesh data.

[0009] According to embodiments, the step of decoding the mesh data may include the step of decoding a base mesh from a base mesh bitstream included in the mesh data.

[0010] According to embodiments, the step of decoding the mesh data may further include, if the mesh data includes a displacement data bitstream, a step of decoding the displacement data from the displacement bitstream, and if the mesh data includes a packed data bitstream, a step of decoding the packed data from the packed bitstream.

[0011] According to embodiments, the decoding method may further include the step of separating displacement data and attribute data from the decoded packed data.

[0012] According to embodiments, the method may further include the step of performing a nominal format conversion on one or more of the data included in the decoded base mesh, the decoded displacement data, or the decoded packed data, and the step of restoring the mesh based on the data converted to the nominal format.

[0013] According to the embodiments, the nominal format conversion step can derive the number of components of the displacement data to 3 if the chroma format of the displacement data is 4:4:4, and derive the number of components of the displacement data to 1 if the chroma format of the displacement data is not 4:4:4.

[0014] According to embodiments, the nominal format conversion step may perform at least one of map extraction, bit depth conversion, resolution conversion, and chroma format conversion of displacement data packed into one or more video planes based on the number of components of the derived displacement data.

[0015] According to embodiments, the bitstream may further include first flag information capable of identifying whether the mesh data includes a displacement data bitstream and second flag information capable of identifying whether the mesh data includes a packed data bitstream.

[0016] According to the embodiments, the restoration step determines whether displacement data exists in the input data based on the first flag information and the second flag information, and if the displacement data exists, the restoration of the displacement data can be performed.

[0017] According to embodiments, the decoding device includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to receive a bitstream containing mesh data and to decode the mesh data.

[0018] According to embodiments, the at least one processor may include a base mesh decoder that decodes a base mesh from a base mesh bitstream included in the mesh data.

[0019] According to embodiments, the at least one processor may further include a displacement decoder that decodes displacement data from a displacement bitstream if the mesh data includes a displacement data bitstream, and a packed decoder that decodes packed data from a packed bitstream if the mesh data includes a packed data bitstream.

[0020] According to embodiments, the at least one processor can separate displacement data and attribute data from the decoded packed data.

[0021] According to embodiments, the at least one processor may further include a nominal format conversion unit that performs nominal format conversion on one or more of the data among the decoded base mesh, the decoded displacement data, or the displacement data included in the decoded packed data, and a restoration unit that restores the mesh based on the data converted to the nominal format.

[0022] According to the embodiments, the nominal format conversion unit can derive the number of components of the displacement data to 3 if the chroma format of the displacement data is 4:4:4, and derive the number of components of the displacement data to 1 if the chroma format of the displacement data is not 4:4:4.

[0023] According to embodiments, the nominal format conversion unit can perform at least one of map extraction, bit depth conversion, resolution conversion, and chroma format conversion of displacement data packed into one or more video planes based on the number of components of the induced displacement data.

[0024] According to embodiments, the bitstream may further include first flag information capable of identifying whether the mesh data includes a displacement data bitstream and second flag information capable of identifying whether the mesh data includes a packed data bitstream.

[0025] According to embodiments, the restoration unit determines whether displacement data exists in the input data based on the first flag information and the second flag information, and if the displacement data exists, it can perform restoration of the displacement data.

[0026] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide a high-quality 3D service.

[0027] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can achieve various video codec methods.

[0028] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide general-purpose 3D content such as autonomous driving services.

[0029] The mesh data receiving method and mesh data receiving device according to the embodiments can perform restoration by considering a packing method according to the chroma format of the displacement video during the nominal format conversion process of the geometry video.

[0030] The mesh data receiving method and mesh data receiving device according to the embodiments perform a reconstruction process by considering PVD, so that the reconstruction process can be performed even when the V3C unit type is PVD.

[0031] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments. For a better understanding of the various embodiments described below, one must refer to the description of the embodiments below in relation to the following drawings, which include parts corresponding to similar reference numerals throughout the drawings.

[0032] FIG. 1(a) and FIG. 1(b) are drawings showing examples of encoders and decoders according to embodiments.

[0033] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.

[0034] FIG. 3 shows a V-MESH compression method according to embodiments.

[0035] FIG. 4 shows the pre-processing of V-MESH compression according to the embodiments.

[0036] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.

[0037] Figure 6 shows a displacement generation process according to embodiments.

[0038] FIG. 7 illustrates the encoding process of mesh data according to embodiments.

[0039] FIG. 8 illustrates the lifting conversion process for displacement according to the embodiments.

[0040] FIG. 9 illustrates the process of packing conversion coefficients according to embodiments into a 2D image.

[0041] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.

[0042] FIG. 11 illustrates the decoding process of mesh data according to embodiments.

[0043] FIG. 12 is a drawing showing an example of a transmitting device according to embodiments.

[0044] FIG. 13 is a drawing showing an example of a receiving device according to embodiments.

[0045] FIG. 14 is a diagram showing an example of a dynamic mesh bitstream structure transmitted by a transmitting device of the present disclosure.

[0046] FIG. 15 is a drawing showing an example of the syntax structure of a V3C unit payload (V3C_unit_payload) according to embodiments.

[0047] FIG. 16 is a drawing showing another example of a transmitting device according to embodiments.

[0048] FIG. 17 is a drawing showing another example of a receiving device according to embodiments.

[0049] FIG. 18 is a drawing showing an example of the syntax structure of a VPS according to embodiments.

[0050] FIG. 19 is a diagram showing examples of nominal chroma formats of geometry video signaled to a VPS according to embodiments.

[0051] FIGS. 20a and FIGS. 20b are drawings showing an example of the syntax structure of an atlas sequence parameter set according to embodiments.

[0052] FIG. 21 is a diagram showing examples of nominal chroma formats of geometry video signaled to ASPS according to embodiments.

[0053] FIG. 22 is a flowchart showing an example of an encoding method according to embodiments.

[0054] FIG. 23 is a flowchart showing an example of a decoding method according to embodiments.

[0055] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments may be practiced without these details.

[0056] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.

[0057] With the recent advancement of 3D data modeling and rendering technologies, research on generating and processing 3D data is being conducted in various fields such as Virtual Reality (VR), Augmented Reality (AR), autonomous driving, CAD (Computer-Aided Design) / CAM (Computer-Aided Manufacturing), and GIS (Geographic Information System). Depending on the representation format, 3D data can be represented as a point cloud, a mesh, etc. Among these, a mesh consists of geometry information representing coordinate values ​​for each vertex (or point), connectivity information representing the relationship between vertices, a texture map representing color information of the mesh surface as 2D image data, and texture coordinates representing mapping information between the mesh surface and the texture map. In this disclosure, a case where one or more of the elements constituting the mesh change over time is defined as a dynamic mesh, and a case where they do not change is defined as a static mesh. That is, dynamic mesh data may refer to mesh data that has an object or movement.

[0058] Because dynamic mesh data has a large amount of data for the elements that make up the mesh compared to 2D image data, technologies to efficiently compress it have been developed to store and transmit massive amounts of mesh data.

[0059] FIGS. 1(a) and FIGS. 1(b) show a V-DMC-based encoder and decoder according to embodiments. In particular, FIGS. 1(a) shows an encoder, and FIGS. 1(b) shows a decoder.

[0060] The basic structure of the currently ongoing V-DMC (v-mesh) is as shown in FIG. 1 (a) and FIG. 1 (b). The encoder according to FIG. 1 (a) and the decoder according to FIG. 1 (b) perform the encoding and decoding processes of media representing the dynamic mesh using Visual Volumetric Video-based Coding (V3C) technology. The pre-processor converts the input dynamic mesh representation into several V3C components (base mesh, displacement set, 2D representation of attributes, and atlas). The original mesh is simplified into a base mesh. The base mesh can be encoded using any mesh codec. The displacement vector can be encoded into a V3C geometry video component using any video codec based on a profile or SEI (supplemental enhancement information) message. For example, depending on the profile, displacement vectors (or displacement data) may be encoded through a video codec-based encoder, a zero-run length encoder, an arithmetic encoder, etc. Attribute data may include additional attributes. For example, texture or material information may be included as additional attributes and may be encoded based on any video codec. Atlas data contains information on how to perform inverse reconstruction and is provided to the V3C (or v-mesh) decoding and / or rendering system of the receiving device. For example, atlas data may include methods for subdividing the base mesh, methods for applying displacement vectors to the vertices of the subdivided mesh, methods for applying attributes to the reconstructed mesh, etc.

[0061] An encoder according to the embodiments may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a pre-processor, an atlas encoding unit, a basemesh encoding unit, a displacement vector encoding unit, a video encoding unit, and a multiplexer.

[0062] The atlas encoding unit generates an atlas bitstream by encoding the atlas of the mesh data. The basemesh encoding unit generates a basemesh bitstream by encoding the basemesh of the mesh data. The displacement vector encoding unit generates a displacement vector bitstream by encoding the displacement vector of the mesh data. The video encoding unit generates an attribute bitstream by encoding the attributes of the mesh data. The encoder according to the embodiments may generate parameter information related to each encoding (which may be referred to as signaling information, metadata, etc.). The encoder according to the embodiments may generate a compressed bitstream including parameter information, an atlas, a basemesh, a displacement vector, and / or attributes, etc.

[0063] The decoder according to the embodiments may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a demultiplexer, an atlas decoding unit, a basemesh decoding unit, a displacement vector decoding unit, and a video decoding unit.

[0064] The atlas decoding unit decodes the atlas within the bitstream. The basemesh decoding unit decodes the basemesh within the bitstream. The displacement vector decoding unit decodes the displacement vector within the bitstream. The video decoding unit decodes the attributes within the bitstream. The decoder according to the embodiments may perform each decoding operation based on parameter information within the bitstream. In the decoder according to the embodiments, the basemesh processing unit restores the current basemesh from the decoded basemesh based on the atlas and / or parameter information. In the decoder according to the embodiments, the displacement processing unit restores the displacement vector by performing coordinate system transformation, etc., of the decoded displacement vector based on the atlas and / or parameter information. In the decoder according to the embodiments, the mesh restoration unit restores the final mesh by combining the restored basemesh and the restored displacement vector based on the atlas and / or parameter information. The restored mesh processing unit of the decoder according to the embodiments can generate and render a reconstructed dynamic mesh image by combining the decoded attribute (or texture map) with the restored final mesh. That is, the reconstructed dynamic mesh image can be displayed to a user.

[0065] Below, the operation of the V-DMC encoder and decoder of FIG. 1 is explained in more detail.

[0066] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.

[0067] The system of FIG. 2 includes a transmitting device (100) and a receiving device (110) according to embodiments. The transmitting device (100) may include a mesh video acquisition unit (101), a mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The receiving device (110) may include a receiving unit (111), a file / segment decapsulator (112), a mesh video decoder (113), and a renderer (114). Each component of FIG. 2 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the mesh data transmitting device according to embodiments may be interpreted as a term referring to a 3D data transmitting device or a transmitting device (100), or a mesh video encoder (hereinafter, encoder) (102). The mesh data receiving device according to the embodiments may be interpreted as a term referring to a 3D data receiving device or a receiving device (110), or a mesh video decoder (hereinafter, decoder) (113).

[0068] The system of Fig. 2 can perform video-based dynamic mesh compression and decompression.

[0069] With advancements in 3D capture, modeling, and rendering, users can access various forms of 3D content, such as AR, XR, the metaverse, and holograms, across multiple platforms and devices. 3D content represents objects more sophisticatedly and realistically to enable users to enjoy immersive experiences, and for this purpose, the creation and use of 3D models require a large amount of data. Among the various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. The embodiments include a series of processing steps in a system that uses such mesh content.

[0070] First, the method for compressing dynamic mesh data originates from the V-PCC (Video-based point cloud compression) standard technology for point cloud data. Point cloud data consists of data containing color information in addition to the coordinates (X, Y, Z) of vertices (or points). In this disclosure, the coordinates of a vertex (i.e., location information) are referred to as geometry information, and the color information of a vertex is referred to as attribute information; the geometry information and attribute information combined are referred to as vertex information or point cloud data. Mesh data refers to vertex information to which connectivity information between vertices has been added. Content can be created in the form of mesh data from the beginning when generating content. Alternatively, it can be converted into mesh data by adding connectivity information to point cloud data for use.

[0071] Currently, the MPEG standards organization defines the data types of dynamic mesh data as the following two types.

[0072] Category 1: Mesh data containing a texture map with color information.

[0073] Category 2: Mesh data with vertex colors as color information.

[0074] Currently, mesh coding standards for Category 1 data are in progress, and standardization work for Category 2 data is also planned for the future. The entire process for providing mesh content services may include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process, as shown in Fig. 2.

[0075] To provide mesh content services, 3D data acquired through multiple cameras or special cameras can be processed into a mesh data type through a series of processes and then generated as a video. The generated mesh video is transmitted after undergoing a series of processes, and at the receiving end, the received data can be processed back into a mesh video and rendered. Through this, the mesh video is provided to the user, and the user can use the mesh content according to their intention through interaction.

[0076] A mesh compression system may include a transmitting device (100) and a receiving device (110) as shown in FIG. 2. The transmitting device (100) may encode a mesh video and output a bitstream, which may be transmitted to the receiving device (110) via a digital storage medium or network in the form of a file or streaming (streaming segment). The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0077] In the transmitting device (100), the encoder may be referred to as a mesh video / image / picture / frame encoding device, and in the receiving device (110), the decoder may be referred to as a mesh video / image / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer (114) may include a display unit, and the renderer and / or the display unit may be composed of separate devices or external components. The transmitting device (100) and the receiving device (110) may further include separate internal or external modules / units / components for the feedback process.

[0078] Mesh data represents the surface of an object using multiple polygons. Each polygon is defined by a vertex in 3D space and connectivity information indicating how those vertices are connected. It may also include vertex attributes such as vertex color and normals. Mapping information, which enables the surface of the mesh to be mapped to a 2D planar area, may also be included in the mesh attributes. Mapping can generally be described by a set of parametric coordinates, referred to as UV coordinates or texture coordinates, associated with the mesh vertices. The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as textures, normals, and displacements. Here, "displacement" may be used interchangeably with "displacement information" or "displacement vector."

[0079] The mesh video acquisition unit (101) may include processing three-dimensional object data acquired through a camera, etc., into a mesh data type having the attributes described above through a series of processes, and generating a video composed of such mesh data. In the mesh video, the attributes of the mesh, namely vertices, polygons, connectivity information between vertices, color, normals, etc., may change over time. A mesh video having attributes and connectivity information that change over time in this way can be described as a dynamic mesh video.

[0080] A mesh video encoder (102) can encode an input mesh video into one or more video streams. A single video may contain multiple frames, and a single frame may correspond to a still image / picture. In this document, the term mesh video may include mesh images / frames / pictures, and mesh video may be used interchangeably with mesh images / frames / pictures. The mesh video encoder (102) can perform a video-based dynamic mesh (V-Mesh) compression procedure. The mesh video encoder (102) can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0081] The file / segment encapsulation module (103) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file or the like. Here, the mesh video-related metadata may be received from a metadata processing unit or the like. The metadata processing unit may be included in the mesh video encoder (102) or may be configured as a separate component / module. The file / segment encapsulation module (103) can encapsulate the data into a file format such as ISOBMFF or process it into other forms such as DASH segments. Depending on the embodiment, the file / segment encapsulation module (103) may include mesh video-related metadata in the file format. The mesh video metadata may be included, for example, in various levels of boxes in the ISOBMFF file format or as data within a separate track in the file. According to an embodiment, the file / segment encapsulator (103) can encapsulate the mesh video-related metadata itself into a file.

[0082] The transmission processing unit may apply processing for transmission to the mesh video data encapsulated according to the file format. The transmission processing unit may be included in the transmission unit (104) or may be configured as a separate component / module. The transmission processing unit may process the mesh video data according to any transmission protocol. Processing for transmission may include processing for transmission via a broadcast network and processing for transmission via broadband. According to an embodiment, the transmission processing unit may receive not only the mesh video data but also mesh video-related metadata from the metadata processing unit and apply processing for transmission to it.

[0083] The transmission unit (104) can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit (111) of the receiving device (110) via a digital storage medium or network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (104) may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit (111) can extract the bitstream and transmit it to a decoding device.

[0084] The receiver (111) can receive mesh video data transmitted by the mesh data transmission device. Depending on the transmission channel, the receiver (111) may receive mesh video data through a broadcast network or through broadband. Alternatively, it may receive mesh video data through a digital storage medium.

[0085] The receiving processing unit can perform processing according to the transmission protocol on the received mesh video data. The receiving processing unit may be included in the receiving unit (111) or may be configured as a separate component / module. In order to correspond to the processing for transmission performed at the transmission side, the receiving processing unit may perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit may transmit the acquired mesh video data to a file / segment decapsulator (112) and transmit the acquired mesh video related metadata to a metadata parser. The mesh video related metadata acquired by the receiving processing unit may be in the form of a signaling table.

[0086] The file / segment decapsulator (112) can decapsulate mesh video data in the form of a file received from the receiving processing unit. The file / segment decapsulator (112) can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder (113), and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder (113) or may be configured as a separate component / module. The mesh video-related metadata obtained by the file / segment decapsulator (112) may be in the form of boxes or tracks within the file format. The file / segment decapsulator (112) may receive metadata required for decapsulation from the metadata processing unit if necessary. The mesh video-related metadata may be passed to the mesh video decoder (113) to be used in the mesh video decoding process, or passed to the renderer (114) to be used in the mesh video rendering process.

[0087] The mesh video decoder (113) can receive a bitstream and perform an inverse operation corresponding to the operation of the mesh video encoder (102) to decode the video / image. The decoded mesh video / image can be displayed through the display unit of the renderer (114). The user can view all or part of the rendered result through a VR / AR display or a general display.

[0088] The feedback process may include the process of transmitting various feedback information, which can be obtained during the rendering / display process, to the transmitting side or to the decoder of the receiving side. Interactivity in mesh video consumption may be provided through the feedback process. According to an embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted during the feedback process. According to an embodiment, the user may interact with elements implemented in a VR / AR / MR / autonomous driving environment, and in this case, information related to such interaction may be transmitted to the transmitting side or the service provider side during the feedback process. According to an embodiment, the feedback process may not be performed.

[0089] Head orientation information can refer to information regarding the user's head position, angle, movement, etc. Based on this information, information about the area the user is currently viewing within the mesh video—that is, viewport information—can be calculated.

[0090] Viewport information may be information about the area currently being viewed by the user in the mesh video. Through this, gaze analysis can be performed to determine how the user consumes the mesh video and which areas of the mesh video they gaze at for how long. Gaze analysis may be performed at the receiving end and transmitted to the transmitting end via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0091] According to the embodiment, the aforementioned feedback information may not only be transmitted to the transmitting side but may also be consumed at the receiving side. That is, decoding and rendering processes at the receiving side may be performed using the aforementioned feedback information. For example, using head orientation information and / or viewport information, only the mesh video for the area currently viewed by the user may be preferentially decoded and rendered.

[0092] This document relates to embodiments of dynamic mesh video compression as described above. The methods / embodiments disclosed in this document may be applied to the MPEG (Moving Picture Experts Group) Video-based Dynamic Mesh Compression Method (V-Mesh) standard or next-generation video / image coding standards. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.

[0093] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.

[0094] In this document, "picture" or "frame" generally refers to a unit representing a single image of a specific time period.

[0095] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.

[0096] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0097] As mentioned above, the encoding process of Fig. 2 is as follows.

[0098] In other words, the Video-based Dynamic Mesh Compression (V-Mesh) compression method can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding). In the V-Mesh compression process, compression is performed by receiving the following data as input.

[0099] Input mesh: Contains the 3D coordinates of the vertices constituting the mesh, normal information for each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between the vertices constituting the surface. The surface of the mesh can be represented by triangles or polygons of greater size, and connection information between the vertices constituting each surface is stored according to a defined shape. The input mesh can be saved in the OBJ file format.

[0100] Attribute map: (Hereafter, Texture map is used with the same meaning): It contains information on the attributes (color, normals, displacement, etc.) of a mesh and stores data in the form of mapping the mesh surface onto a 2D image. Mapping which part of the mesh (surface or vertex) corresponds to each data point in this attribute map is based on the mapping information contained in the input mesh. Since the attribute map holds data for each frame of the mesh video, it can also be referred to as an attribute map video. In the V-Mesh compression method, the attribute map primarily contains the mesh's color information and is stored in image file formats (PNG, BMP, etc.).

[0101] Material Library File: Contains material attribute information used in the mesh, specifically information that links the input mesh with its corresponding attribute map. This is saved in the Wavefront Material Template Library (MTL) file format.

[0102] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0103] Base Mesh: By simplifying (decimating) the input mesh through a pre-processing process, objects in the input mesh are represented using the minimum number of vertices determined according to user criteria.

[0104] Displacement: Displacement information used to represent the input mesh as similarly as possible using the base mesh, and is expressed in the form of 3D coordinates.

[0105] Atlas information: This is metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be generated and utilized as sub-units (sub-mesh, patch, etc.) that constitute the mesh.

[0106] Referring to FIGS. 3 to 7, a method for encoding mesh position information (or vertex position information) is described, and referring to FIGS. 7 to 10, a method for restoring mesh position information and encoding attribute information (attribute map) is described.

[0107] FIG. 3 shows a V-MESH compression method according to embodiments.

[0108] FIG. 3 illustrates the encoding process of FIG. 2, and the encoding process may include a pre-processing process and an encoding process. The mesh video encoder (102) of FIG. 2 may include a pre-processor (200) and an encoder (201) as in FIG. 3. Additionally, the transmitting device of FIG. 2 may be broadly referred to as an encoder, and the mesh video encoder (102) of FIG. 2 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing process (Pre-processing, 200) and an encoding process (Encoding, 201) as in FIG. 3. The pre-processor (200) of FIG. 3 may be located in front of the encoder (201) of FIG. 3. The pre-processor (200) and the encoder (201) of FIG. 3 may be referred to as a single encoder.

[0109] The pre-processor (200) can receive a static of dynamic mesh (M(i)) and / or an attribute map (A(i)). The pre-processor (200) can generate a base mesh (m(i)) and / or a displacement (d(i)) through pre-processing. The pre-processor (200) can receive feedback information from the encoder (201) and generate the base mesh and / or the displacement based on the feedback information.

[0110] The encoder (201) may receive a base mesh (m(i)), a displacement (d(i)), a static (M(i)) of a dynamic mesh, and / or an attribute map (A(i)). In the present disclosure, at least one of the base mesh (m(i)), the displacement (d(i)), the static (M(i)) of a dynamic mesh, and / or an attribute map (A(i)) may be referred to as mesh-related data. The encoder (201) may encode the mesh-related data to generate a compressed bitstream.

[0111] Figure 4 shows the pre-processing process of V-MESH compression according to the embodiments.

[0112] FIG. 4 illustrates the configuration and operation of the pre-processor of FIG. 3. In FIG. 4, the input mesh may include a static of dynamic mesh (M(i)) and / or an attribute map (A(i)). Additionally, the input mesh may include 3D coordinates of vertices constituting the mesh, normal information for each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between vertices constituting the surface.

[0113] FIG. 4 illustrates a process of performing pre-processing on an input mesh. The pre-processing process (200) may include four main steps: 1) GoF (Group of Frame) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). According to the embodiments, GoF generation may be referred to as the GoF generation process or GoF generation section, Mesh Decimation as the Mesh Decimation process or Mesh Decimation section, UV parameterization as the UV parameterization process or UV parameterization section, and Fitting subdivision surface as the Fitting subdivision surface process or Fitting subdivision surface section. The pre-processor (200) can generate displacement and / or base mesh from the received input mesh and transmit it to the encoder (201). The pre-processor (200) can transmit GoF information associated with GoF generation to the encoder (201).

[0114] Below, each step of Fig. 4 is explained.

[0115] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, the number of texture coordinates, vertex connection information, and texture coordinate connection information of the mesh of the previous frame and the current mesh are all identical, the previous frame can be set as the reference frame. That is, if only the vertex coordinate values ​​differ between the current input mesh and the reference input mesh, the encoder (201) can perform inter-frame encoding. Otherwise, intra-frame encoding is performed for the corresponding frame.

[0116] Mesh Decimation: This is the process of simplifying the input mesh to generate a simplified mesh, or base mesh. After selecting vertices to remove from the original mesh based on user-defined criteria, the selected vertices and the triangles connected to them can be removed.

[0117] In the process of performing mesh decimation, information regarding the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) is passed as input, and a decimated mesh can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.

[0118] UV Parameterization: This is the process of mapping 3D surfaces to a texture domain for a decimated mesh. Parameterization can be performed using a UV Atlas tool. Through this process, mapping information is generated regarding where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process.

[0119] OrthoAtlas technology is a technique that generates texture coordinates using orthographic projection. In orthoAtlas technology, the processes of patch generation and patch packing are performed sequentially. First, Connected Components (CCs) are generated by splitting adjacent triangles, and then the optimal CCs are merged using a cost function to generate the patch. The cost function can measure the cost based on the degree of distortion that occurs when orthographically projecting the patch in each direction. Finally, texture coordinates can be calculated by packing the patch that minimizes the cost function into the texture domain. With orthoAtlas technology, texture coordinates can be derived in the base mesh decoder without compressing texture coordinate and texture connection information during the base mesh encoding process.

[0120] Fitting subdivision surface (300): This is a process of performing subdivision on a decimated mesh (i.e., a decimated mesh having texture coordinates). The displacement generated through this process and the base mesh are output to the encoder (201). A user-defined method, such as the mid-edge method, may be applied as the subdivision method. A fitting process is performed so that the input mesh and the mesh that has undergone subdivision become similar to each other. In this disclosure, the mesh on which the fitting process has been performed is referred to as the fitted subdivided mesh (or fitted subdivided mesh). This process is a process of performing fitting so that the mesh that has undergone subdivision on the base mesh becomes similar to the surface of the input mesh. As for the subdivision method, a user-defined method such as the Mid-edge method (see Fig. 5), Loop method, and LS3 method may be applied.

[0121] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.

[0122] Figure 5 illustrates the mid-edge method of the fitting subdivision surface described in Figure 4. Referring to Figure 5, an original mesh containing four vertices is subdivided to create a sub-mesh. A sub-mesh can be created by generating a new vertex at the midpoint of the edges between vertices. Then, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, thereby creating a fitted subdivided mesh.

[0123] When a fitted subdivided mesh (hereinafter referred to as the fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The positional difference between this result and the fitted subdivided mesh for each vertex becomes the displacement for each vertex. Since displacement represents a positional difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of the Cartesian coordinate system. Depending on user input parameters, (x, y, z) coordinate values ​​can be converted into (normal, tangential, bi-tangential) coordinate values ​​of the local coordinate system.

[0124] FIG. 6 illustrates a displacement generation process according to embodiments. The displacement generation process of FIG. 6 may be performed in a pre-processor (200) or in an encoder (201).

[0125] FIG. 6 illustrates in detail the method of calculating the displacement of the fitting subdivision surface (300) as described in FIG. 5.

[0126] The encoder and / or pre-processor according to the embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may perform subdivision on the restored base mesh to generate a subdivided restored base mesh. Here, the restoration of the base mesh may be performed in the pre-processor (200) or in the encoder (201). The local coordinate system calculation unit receives the fitted subdivided mesh and the subdivided restored base mesh, and may convert the coordinate system regarding the mesh to a local coordinate system based on them. The local coordinate system calculation operation may be optional. The displacement calculation unit calculates the position difference between the fitted subdivision mesh and the subdivided restored base mesh. For example, it may generate a position difference value between the vertices of the two input meshes. The vertex position difference value becomes the displacement.

[0127] The mesh data transmission method and apparatus according to the embodiments can encode mesh data as follows. Mesh data is a term that includes point cloud data. Point cloud data according to the embodiments (which may be referred to as point cloud for short) may refer to data including vertex coordinates (or referred to as geometry information) and color information (or referred to as attribute information). Additionally, geometry images, attribute images, accusation maps, and additional information (or referred to as patch information) generated through patch generation and packing based on vertex coordinates and color information are also referred to as point cloud data. Therefore, point cloud data including connection information may be referred to as mesh data. In this document, point cloud and mesh data may be used interchangeably.

[0128] The V-Mesh compression (restoration) method according to the embodiments may include intra-frame encoding and inter-frame encoding.

[0129] Intra-frame encoding or inter-frame encoding is performed based on the results of the aforementioned GoF generation. In the case of intra-frame encoding, the data to be compressed may include the base mesh, displacement, and attribute map. In the case of inter-frame encoding, the data to be compressed may include displacement, attribute map, and the motion field between the reference base mesh and the current base mesh.

[0130] Figure 7 illustrates a V-DMC encoding process according to embodiments.

[0131] The elements of the transmitting device illustrated in FIG. 7 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the transmitting device of FIG. 7 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the transmitting device of FIG. 7 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the transmitting device of FIG. 7. The execution order of each element in FIG. 7 may be changed, some elements may be omitted, and some elements may be newly added.

[0132] In the present disclosure, the operation process of a transmitting end for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in FIG. 7. The transmitting device of FIG. 7 may support both an intra-frame encoding (or intra-encoding or intra-frame encoding) process and / or an inter-frame encoding (or inter-encoding or inter-frame encoding) process.

[0133] The encoding process of FIG. 7 illustrates in detail the encoding of the mesh video encoder (102) of FIG. 2. The encoder of FIG. 7 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of FIG. 7 may correspond to the pre-processor (200) and encoder (201) of FIG. 4.

[0134] The pre-processor (200) receives an input mesh and can perform the aforementioned pre-processing. Through pre-processing, a base mesh and / or a fitted subdivided mesh can be generated.

[0135] The quantizer (411) of the encoder (201) can quantize the base mesh and / or the fitted subdivided mesh.

[0136] According to embodiments, the base mesh quantized in the mesh quantization unit (411) may be output to a static mesh encoder (413) or a motion vector encoder (414) through a switching unit (412). According to embodiments, the base mesh is output to a motion vector encoder (414) through the switching unit (412) when inter-encoding is performed on the corresponding mesh frame, and is output to a static mesh encoder (413) through the switching unit (412) when intra-encoding is performed on the corresponding mesh frame. The motion vector encoder (414) may be referred to as a motion encoder.

[0137] For example, when intra-frame encoding is performed on the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (413). In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. That is, the vertex coordinates, vertex connection information, texture coordinates, texture connection information, etc. of the mesh can be encoded in the static mesh encoder (413). The base mesh bitstream generated through encoding is transmitted to a multiplexer (not shown).

[0138] As another example, when performing inter-frame encoding for the corresponding mesh frame, the motion vector encoder (414) may take the current base mesh and the reference reconstructed base mesh (or reconstructed quantized reference base mesh) as inputs, calculate the motion vector between the two meshes, and encode the value. Additionally, the motion vector encoder (414) may perform a prediction based on connectivity information using a previously encoded / decoded motion vector as a predictor, and entropy-encode the difference motion vector (or residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector. Depending on the embodiments, motion vector encoding may be performed at the vertex level or at the subgroup level. The motion vector bitstream generated through motion vector encoding is transmitted to a multiplexer (not shown) as a base mesh bitstream. That is, in the case of in-frame encoding, the static mesh bitstream is input to the multiplexer as a base mesh bitstream, and in the case of inter-frame encoding, the motion vector bitstream is input to the multiplexer as a base mesh bitstream.

[0139] In FIG. 7, the base mesh restoration unit (415) can generate a reconstructed base mesh by receiving a base mesh encoded in the static mesh encoder (413) or a motion vector encoded in the motion vector encoder (414). The base mesh restoration unit (415) performs the restoration of the base mesh according to the encoding type of the current mesh (inter-frame encoding or intra-frame encoding). For example, the base mesh restoration unit (415) can restore the base mesh by performing static mesh decoding on the base mesh encoded in the static mesh encoder (413). At this time, quantization can be applied before static mesh decoding, and inverse quantization can be applied in the inverse quantization unit (416) after static mesh decoding. That is, when intra-frame encoding is performed, the current base mesh can be restored by performing inverse quantization in the inverse quantization unit (416) on the base mesh quantized through the mesh quantization unit (411). As another example, the base mesh restoration unit (415) can restore the base mesh based on the restored quantized reference base mesh and the motion vector encoded by the motion vector encoder (414). That is, when cross-frame encoding is performed, the motion vector can be decoded using a motion vector decoding method, and then the decoded motion vector can be applied (i.e., added) to the reference restored base mesh to generate the current base mesh. In this case, if the motion vector is not quantized, the motion vector restoration process is omitted, and the current base mesh can be restored using the motion vector calculated by the motion vector encoder (414). The restored base mesh is output to the displacement vector calculation unit (417) and the mesh restoration unit (425).

[0140] According to the embodiments, the displacement vector calculation unit (417) can perform mesh subdivision on the restored base mesh. Additionally, the displacement vector calculation unit (417) can calculate a displacement vector, which is the difference in vertex positions between the subdivided restored base mesh and the fitted subdivided (or subdivided) mesh generated by the pre-processor (200). That is, the displacement vector is the difference in positions between the vertices of the two meshes so that the fitted subdivided (or subdivided) mesh becomes similar to the original mesh. At this time, the displacement vector can be calculated for the number of vertices of the subdivided mesh. That is, the displacement vector for the number of vertices of the subdivided (subdivided) mesh can be calculated through the displacement vector calculation unit (417).

[0141] The lifting transformation unit (418) can perform a lifting transformation on the input displacement vector to generate a lifting coefficient (or displacement vector transformation coefficient). The quantizer (419) can quantize the lifting coefficient, i.e., the displacement vector transformation coefficient.

[0142] In the present disclosure, the displacement vector or quantized displacement vector transformation coefficient can be encoded through a 2D video codec-based encoding method, and / or a zero-run length encoding method and / or an arithmetic encoding method, etc.

[0143] If an arithmetic encoding method is used, the displacement vector or quantized displacement vector transformation coefficient is encoded based on an arithmetic codec in the arithmetic encoding unit (421) after inter prediction in the inter prediction unit (420), and if a 2D video codec-based encoding method is used, the displacement vector or quantized displacement vector transformation coefficient is encoded based on a 2D video codec in the video encoding unit (423) after image packing in the image packing unit (422) and can be output as a displacement bitstream (i.e., compressed displacement bitstream). For example, the image packing unit (422) can pack an image based on quantized lifting coefficients (i.e., displacement vector transformation coefficients). The video encoding unit (423) can encode the packed image. That is, the quantized lifting coefficients are packed into a frame as a 2D image by the image packing unit (422), compressed through the video encoding unit (423), and output as a displacement bitstream (i.e., compressed displacement bitstream).

[0144] The displacement vector restoration unit (424) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (424) performs decoding on the encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (425). The mesh restoration unit (425) restores the deformed mesh based on the base mesh restored in the base mesh restoration unit (415) and the displacement vector restored in the displacement vector restoration unit (424). That is, the mesh restoration unit (425) restores the reconstructed and deformed mesh through the restored displacement output from the displacement vector restoration unit (424) and the restored base mesh (or subdivided restored base mesh) output from the inverse quantization unit (416). The present disclosure refers to the reconstructed and deformed mesh as the restored deformed mesh. The restored mesh (or restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.

[0145] The attribute transfer (426) receives an input mesh and / or an input attribute map and regenerates an attribute map based on the restored deformed mesh. An attribute map refers to a texture map corresponding to attribute information among the mesh data components, and in this disclosure, attribute map and texture map may be used interchangeably. Push-pull padding (427) can pad data into the attribute map based on a push-pull method. A color space conversion unit (428) can convert the space of the color components of the attribute map. For example, the attribute map can be converted from the RGB color space to the YUV color space. A video encoding unit (429, or referred to as a video encoder) can encode the attribute map and output it as a compressed attribute bitstream.

[0146] According to the embodiments, the atlas encoder (430) can generate a compressed atlas bitstream by encoding atlas information (or atlas data). The atlas bitstream generated through atlas information encoding is transmitted to a multiplexer (431). In the present disclosure, the atlas may be information required in a mesh reconstruction process and may refer to information such as tiles and patches. Additionally, the atlas information may refer to data required in processes such as 2D mapping of 3D objects, texture mapping information, mesh decoding, and mesh reconstruction, and may include additional information such as subdivision methods, transformation methods, quantization methods, and the location and size of patches within the atlas frame. In the present disclosure, the atlas information may be encoded through Exp-Golomb coding, etc., of the atlas encoder (430).

[0147] According to embodiments, a multiplexer (430) can multiplex an input compressed base mesh bitstream, a compressed displacement (or displacement vector) bitstream, a compressed attribute (or texture map) bitstream, and a compressed atlas bitstream to produce a single compressed bitstream. The multiplexed bitstream can be encapsulated into one or more tracks of a file.

[0148] According to the embodiments, the multiplexed bitstream or file from the multiplexer (431) may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, etc., and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0149] In summary, for base meshes, encoding can be performed in different ways depending on the base mesh type (INTRA type, INTER type, SKIP type). If the base mesh is of the INTRA type, it can be encoded using the static mesh encoding method. If the base mesh is of the INTER type, the motion field between the reference base mesh and the current base mesh can be encoded. If the current base mesh is of the SKIP type, the reference base mesh can be guided to the current base mesh.

[0150] After being encoded in an encoder, the decoded base mesh can be processed to obtain a subdivided mesh. Subdivision algorithms such as mid-point subdivision and loop subdivision can be used.

[0151] Static Base Mesh Encoding (Intra Base Mesh Encoding): When performing intra encoding on the current basemesh, the base mesh generated during the pre-processing stage can be encoded using static mesh compression technology after undergoing a quantization process. Static mesh compression utilizes MPEG EdgeBreaker (MEB) technology, and the base mesh's vertex position information, mapping information (texture coordinates), vertex connectivity information, and normals are subject to compression.

[0152] Connection information can be encoded and compressed based on the edgebreaker algorithm. The edgebreaker algorithm is a technique that sequentially traverses triangles according to rules, maps symbols based on the characteristics of each triangle, and then encodes those symbols.

[0153] Techniques for compressing vertex location information can calculate predicted values ​​based on prediction techniques such as multiple parallelogram prediction, and then encode the residual value, which is the difference between the current vertex and the predicted value.

[0154] A technique for compressing mapping information (texture coordinates) can calculate a predicted value based on a prediction technique such as stretching, and then encode the residual value, which is the difference between the current mapping information (texture coordinates) and the predicted value.

[0155] Normal compression techniques can obtain predicted values ​​based on prediction techniques such as delta coding, multiple parallelogram prediction, and cross product-based prediction, and then encode the residual value, which is the difference between the current normal and the predicted value.

[0156] Motion Field Encoding (Inter Basemesh Encoding): Inter basemesh encoding can be performed when a one-to-one correspondence exists between the reference mesh and the current input mesh, differing only in their vertex position information. When performing inter encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh—that is, the motion field—is calculated and this information is encoded. The reference base mesh is the result of quantizing already decoded base mesh data and is determined by the reference frame index.

[0157] The motion field can be encoded as is, or the predicted motion field can be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between this predicted motion field value and the current vertex's motion field value, can be encoded. This value can be encoded using entropy coding.

[0158] Displacement Encoding: After Base mesh encoding, reconstruction and inverse quantization are performed to Recon. A base mesh is generated, and the displacement between the result of performing subdivision on it and the Fitted subdivided mesh can be calculated. For effective encoding, a data transform process such as a wavelet transform can be applied to the displacement information, and Figure 8 shows the process of transforming displacement information using a lifting transform in V-Mesh. The displacement vector transformation coefficients generated through the transformation process are quantized, and the quantized transformation coefficients may be compressed through a video codec or through arithmetic encoding depending on the compression method.

[0159] When compressed through a video codec, the 2D image is packed as shown in Fig. 9. Transform coefficients are organized into one block for every N^2 (N*N) units, and each block can be packed in z-scan order. The number of horizontal blocks is fixed at N, while the number of vertical blocks can be determined by the number of vertices of the subdivided base mesh. Within a single block, the transform coefficients can be packed by aligning them using Morton code. The packed images generate a displacement video for every GoF unit, and this displacement video can be encoded using an existing video compression codec.

[0160] When compressed via arithmetic encoding, cross-frame prediction can be performed on the quantized displacement vector transformation coefficients. When cross-frame prediction is performed on the current quantized displacement vector transformation coefficients, the residual value, which is the difference between the current displacement vector transformation coefficient and the reference displacement vector transformation coefficient, can be encoded, and information about the reference target can be encoded. Depending on the displacement vector type, the quantized displacement vector transformation coefficients can be arithmetic encoded if it is of the INTRA type, and the residual value if it is of the INTER type. Arithmetic encoding can be performed based on Context Adaptive Binary Arithmetic Coding (CABAC). The CABAC process can first binarize the displacement vector data and map it to a bin string. The bin string can be an output binarized into 0s and 1s, where each 0 or 1 can be a bin. Each bin can be arithmetic encoded using context information selected from the context model, and a process of updating probabilities can be performed.

[0161] FIG. 8 illustrates the lifting conversion process for displacement according to the embodiments.

[0162] FIG. 9 illustrates the process of packing a conversion factor (or lifting factor) according to embodiments into a 2D image.

[0163] Figures 8 and 9 respectively show the process of converting the displacement of the encoding process of Figure 7 and the process of packing the conversion coefficients.

[0164] The encoding method according to the embodiments includes displacement encoding.

[0165] After base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through reconstruction and inverse quantization, and the displacement between the result of performing subdivision on this reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated (417 in FIG. 7). For effective encoding, a data transform process such as a wavelet transform can be applied to the displacement information (418 in FIG. 7).

[0166] FIG. 8 shows the process of transforming displacement information using a lifting transform in the lifting transform unit (418) of FIG. 7. For example, a linear wavelet-based lifting transform may be performed. The transformation coefficients generated through the transformation process are quantized in a quantizer (419) and then packed into a 2D image as shown in FIG. 9 through an image packing unit (422). The transformation coefficients are organized into one block for every 256 (=16×16) units, and each block can be packed in z-scan order. The number of horizontal blocks is fixed at 16, while the number of vertical blocks can be determined according to the number of vertices of the subdivided base mesh. The transformation coefficients can be packed by aligning them with a Morton code within a single block. Packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec in a video encoding unit (423, or referred to as a video encoder).

[0167] Referring to FIG. 8, the base mesh (original) may include vertices and edges for Level of Detail (LoD) 0. A first subdivision mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices of the base mesh (LoD0). A second subdivision mesh may be generated by dividing (or subdividing) the first subdivision mesh again. The second subdivision mesh includes LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 which includes vertices further divided (or subdivided) from LoD0, and vertices further divided (or subdivided) from LoD1. LoD is a Level of Detail that indicates the degree of detail in the mesh data content. As the level index increases, the distance between vertices decreases, and the level of detail increases. In other words, a smaller LoD value indicates lower detail in the mesh data content, while a larger LoD value indicates higher detail. LoD N includes the vertices contained in the previous LoD N-1. When a mesh (or vertex) is further subdivided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1 and v2 and the subdivided vertex v. Instead of encoding the information for the current LoD N as is, residuals between the previous LoD N-1 can be generated and encoded using these residuals to reduce the size of the bitstream. The prediction process refers to the operation of predicting the current vertex v based on the previous vertices v1 and v2. Since adjacent subdivided meshes contain similar data, efficient encoding can be achieved by utilizing this property.The current vertex position information is predicted as a residual for the previous vertex position information, and the previous vertex position information is updated through the residual. In this disclosure, vertex, vertex, and point may be used interchangeably. Also, LoDs may be defined during the refinement process of the base mesh. According to embodiments, the refinement process of the base mesh may be performed in a pre-processor (200) or in a separate component / module.

[0168] Referring to FIG. 9, the vertex has a transformation coefficient (or lifting coefficient) generated through a lifting transformation. The transformation coefficient of the vertex related to the lifting transformation can be packed into an image by the image packing unit (422) and then encoded by the video encoding unit (423).

[0169] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.

[0170] According to the embodiments, FIG. 10 shows the detailed operation of the attribute transfer (426) of FIG. 7.

[0171] The encoding according to the embodiments includes attribute map encoding. According to the embodiments, attribute map encoding can be performed in the video encoding unit (429) of FIG. 7.

[0172] According to embodiments, the encoder in the present disclosure compresses information about an input mesh through base mesh encoding (i.e., intra encoding), motion field encoding (i.e., inter encoding), and displacement encoding. The input mesh compressed during the encoding process is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding processes, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress an input attribute map as shown in FIG. 7. The reconstructed deformed mesh has vertex position information, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Accordingly, as shown in FIG. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the restored deformed mesh is regenerated through the attribute transfer process of the attribute transfer (426).

[0173] According to embodiments, attribute transfer (426) first checks for all points P(u, v) in a 2D texture domain whether the corresponding vertex belongs to a texture triangle of the reconstructed deformed mesh, and if it exists within a texture triangle T, the barycentric coordinate of P(u, v) according to that triangle T ( , , Calculate ). And the 3D vertex positions of triangle T and ( , , Calculate the 3D coordinates M(x, y, z) of P(u, v) using ). Find the vertex coordinates M'(x', y', z') corresponding to the location most similar to the calculated M(x, y, z) in the input mesh domain, and triangle T' containing this vertex. Then, find the coordinates of the centroid of M'(x', y', z') in this triangle T' ( ', ', Calculate '). Texture coordinates corresponding to the three vertices of Triangle T' and ( ', ', Texture coordinates (u', v') are calculated using '), and color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is then assigned to the (u, v) pixel location in the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm, such as the push-pull algorithm of push-pull padding (427).

[0174] The new attribute map generated through attribute transfer (426) is grouped into GoF units to form an attribute map video, which is then compressed using the video codec of the video encoding unit (429).

[0175] Referring to Fig. 10, the reference relationships between the input mesh, the input attribute map, the reconstructed deformed mesh, and the regenerated attribute map can be seen.

[0176] The decoding process of Fig. 2 can perform the reverse process of the corresponding process of the encoding process of Fig. 2. The specific decoding process is as follows.

[0177] FIG. 11 illustrates the decoding process of V-Mesh technology according to embodiments.

[0178] FIG. 11 illustrates the configuration and operation of a mesh video decoder (113) of the receiving device of FIG. 2. Additionally, FIG. 11 can restore mesh data by performing the reverse process of the encoding process of FIG. 7. In the present disclosure, the receiving device of FIG. 11 may be referred to as a mesh data receiving device, a decoder, a decoder of a receiving device, a V-Mesh decoder, or a dynamic mesh decoder.

[0179] The elements of the receiving device illustrated in FIG. 11 may be implemented in hardware, software, processors connected to memory, and / or combinations thereof. That is, the elements of the receiving device of FIG. 11 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the receiving device of FIG. 11 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the receiving device of FIG. 11. The execution order of each element in FIG. 11 may be changed, some elements may be omitted, and some elements may be newly added.

[0180] FIG. 11 may largely include a demultiplexer (611), an atlas decoder (612), and a decoding unit (620).

[0181] According to embodiments, a bitstream of mesh data (i.e., a compressed bitstream) received by a receiver (not shown) may be demultiplexed in a demultiplexer (611) into a base mesh bitstream (or base mesh sub-stream), a displacement vector bitstream (or displacement sub-stream), a texture map bitstream (or attribute map sub-stream or texture map sub-stream), and / or an atlas bitstream after file / segment decapsulation. If the bitstream of mesh data is not encapsulated in the form of a file at the transmitting device, the decapsulation process at the receiving device is omitted. If the current mesh is inter-encoded, the base mesh bitstream may be a motion vector bitstream.

[0182] According to embodiments, an atlas bitstream is provided to an atlas decoder (612). The atlas decoder (612) can decode the atlas bitstream to restore atlas information. The restored atlas information can be used in a mesh decoding process. According to embodiments, the process in which the atlas information is used may be a subdivision process, a displacement vector restoration process, etc., and may include information such as tiles and patches.

[0183] According to the embodiments, the atlas bitstream can be decoded through the Exp-Golomb coding of the atlas decoder (612). At this time, the atlas may be information required in the mesh reconstruction process and may refer to information such as tiles and patches. Also, the atlas data may refer to data required in the mesh decoding, mesh reconstruction, etc. process and may include a subdivision method, a transformation method, a quantization method, the location and size of patches within the atlas frame, etc.

[0184] According to the embodiments, the base mesh bitstream is provided to the motion vector decoder (623) or to the static mesh decoder (622) through the switching unit (621).

[0185] For example, if the current mesh is inter-encoded, the base mesh bitstream, i.e., the motion vector bitstream, is received, demultiplexed, and then output to the motion vector decoder (623) through the switching unit (621). As another example, if the current mesh is intra-encoded, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (622) through the switching unit (621). Here, the motion vector decoder (623) may be referred to as a motion decoder.

[0186] According to the embodiments, the motion vector decoder (623) can perform decoding on the motion vector bitstream on a vertex-by-vertex or subgroup-by-subgroup basis.

[0187] According to embodiments, the motion vector decoder (623) can restore the final motion vector by adding the difference motion vector (i.e., residual motion vector) decoded from the bitstream using the previously decoded motion vector as a predictor. That is, the motion vector decoder (623) can decode the difference motion vector (or residual motion vector) in vertex or subgroup (or subblock) units through the motion vector bitstream, and decode the motion vector by adding the residual motion vector by performing a connection information-based prediction using the previously decoded motion vector as a predictor.

[0188] According to the embodiments, the static mesh decoder (622) can decode the base mesh bitstream to restore the connection information, vertex geometry information, texture coordinates (i.e., attribute geometry information), normal information, etc. of the base mesh. That is, the static mesh decoder (622) can restore the connection information, vertex geometry information, vertex texture coordinates, etc. of the restored quantized base mesh, for example, the base mesh.

[0189] According to embodiments, the base mesh restoration unit (631) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh is subjected to inter-frame encoding, the base mesh restoration unit (631) can generate the restored base mesh (i.e., the current base mesh) by adding the decoded (or restored) motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh is subjected to intra-frame encoding, the base mesh restoration unit (631) can generate the restored base mesh (i.e., the current base mesh) by performing inverse quantization on the base mesh decoded (or restored) through the static mesh decoder (622). According to embodiments, inverse quantization may be omitted.

[0190] According to the embodiments, the displacement sub-bitstream is provided to the arithmetic decoding unit (625) or to the video decoding unit (627) through the switching unit (624) depending on the decoding method.

[0191] For example, if the decoding method is an arithmetic codec method, the arithmetic decoding unit (625) decodes the displacement substream based on the arithmetic codec, and the inverse prediction unit (626) performs the inverse prediction process on the decoded displacement information and outputs it to the inverse quantization unit (629). As another example, if the decoding method is a 2D video codec method, the video decoding unit (627) decodes the displacement substream based on the 2D video codec, and the image unpacking unit (628) unpacks the image of the decoded displacement video and outputs it to the inverse quantization unit (629).

[0192] The displacement information provided by the above-mentioned inverse prediction unit (626) or image unpacking unit (628) is inversely quantized in the inverse quantization unit (629) and inversely transformed in the inverse linear lifting unit (630) to be restored as displacement information for each vertex (i.e., Recon. displacements).

[0193] According to the embodiments, the mesh restoration unit (632) restores a reconstructed and deformed mesh through the restored displacement output from the inverse linear lifting unit (630) and the restored base mesh output from the base mesh restoration unit (631) (i.e., decoded mesh). That is, the inversely quantized restored base mesh is combined with the restored displacement information to generate a final decoded mesh. In this disclosure, the final decoded mesh is referred to as a reconstructed deformed mesh.

[0194] According to the embodiments, the attribute map sub-stream is decoded through a video decoding unit (633) corresponding to the video compression codec used in encoding, and then restored to a final attribute map (i.e., decoded attribute map) through processes such as color format conversion and color space conversion in a color conversion unit (634).

[0195] According to the embodiments, the restored decoded mesh and decoded attribute map can be utilized at the receiving end as final mesh data that can be utilized by the user.

[0196] To summarize Fig. 11, the base mesh sub-stream can be decoded through a static mesh decoder (622) based on MEB (MPEG EdgeBreaker) technology, for example, if it is of the INTRA type, depending on the base mesh type, and as a result, connection information, vertex geometry information, vertex mapping information (texture coordinates), etc. of the base mesh can be restored.

[0197] In the case where the texture parameterization method in the encoder according to the embodiments is orthoAtlas, the decoder can derive mapping information (texture coordinates) and attribute information (texture) connection information using vertex coordinates. The process of deriving mapping information (texture coordinates) and connection information can generate mapping information (texture coordinates) and attribute information (texture) connection information by calculating the homography transform of each face and then projecting the vertex based on this.

[0198] According to the embodiments, when the base mesh type is an INTER type, motion information can be decoded through entropy decoding and inverse prediction processes. The decoded motion information is combined with a reference base mesh that has already been restored and stored in a buffer to generate a Reconstructed quantized base mesh for the current frame. An inverse quantization process can be performed on the restored base mesh.

[0199] In other words, when mesh data within the bitstream is encoded based on inter-prediction, the motion vector decoder derives the motion field of the basemesh of the current frame through motion estimation and compensation, based on the basemesh within the reference frame. When mesh data within the bitstream is encoded based on intra-prediction, the static mesh decoder decodes the basemesh.

[0200] Depending on the compression method used by the encoder, the displacement sub-stream is decoded into displacement video through the video compression codec's decoder if compressed, for example, by a video codec, and then an image unpacking process is performed.

[0201] As another example, when compressed through arithmetic coding, the displacement sub-stream can be decoded into binarized syntax elements through arithmetic decoding, and a Contextual Probability Model (CPM) can be adaptively determined according to each bin of the syntax elements, and arithmetic decoding can be performed by predicting the probability of bin occurrence through the CPM. The binarized syntax elements can be decoded through inverse binarization. Quantized displacement vector transformation coefficients can be derived from the decoded syntax elements. As another example, when the displacement information type is INTER (when inter prediction is performed), an inverse inter prediction process is performed using reference information for the quantized displacement vector transformation coefficients.

[0202] The quantized displacement vector displacement coefficients are restored as displacement information for each vertex through inverse quantization, inverse transformation, and coordinate system transformation processes.

[0203] The restored base mesh and restored displacement information are combined to generate the final decoded mesh. The attribute map sub-stream is decoded through the decoder of the video compression codec used in the encoder, and then restored to the final attribute map through processes such as color format conversion.

[0204] The restored decoded mesh and decoded attribute map can be utilized at the receiving end as final mesh data available to the user.

[0205] FIG. 12 shows a mesh data transmission device according to embodiments.

[0206] FIG. 12 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 2, the encoder (preprocessor and encoder) of FIG. 3 or FIG. 7, and / or the corresponding transmitting encoding device. Each component of FIG. 12 corresponds to hardware, software, a processor, and / or a combination thereof.

[0207] The operation process of the transmitting end for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in FIG. 12. The transmitting device of FIG. 12 may perform an intra-frame encoding (or intra-encoding or intra-frame encoding) process and / or an inter-frame encoding (or inter-encoding or inter-frame encoding) process.

[0208] The pre-processor (811) receives the original mesh as input and generates a subdivided mesh that is fitted with the decimated mesh (or base mesh). Decimation can be performed based on the number of target vertices or the number of target polygons constituting the mesh. For the decimated mesh, parameterization (or parameterization) can be performed to generate texture coordinates and texture connection information per vertex. For example, parameterization is the process of mapping a 3D surface to a texture domain for the decimated mesh. If parameterization is performed using a UV Atlas tool, mapping information is generated that can identify where each vertex of the decimated mesh can be mapped to on a 2D image. The mapping information is stored in the form of texture coordinates, and through this process, the final base mesh is generated. Additionally, the task of quantizing floating-point mesh information into fixed-point form can be performed. This result can be output as a base mesh to a motion vector encoder (813) or a static mesh encoder (814) through a switching unit (812). The pre-processor (811) can generate additional vertices by performing mesh subdivision on the base mesh. Depending on the subdivision method, vertex connection information including the added vertices, texture coordinates, and connection information of texture coordinates can be generated. The pre-processor (811) can generate a fitted subdivided mesh by adjusting vertex positions so that the subdivided mesh becomes similar to the original mesh.

[0209] According to the embodiments, when inter-encoding is performed on the corresponding mesh frame, the base mesh is output to the motion vector encoder (813) through the switching unit (812), and when intra-encoding is performed on the corresponding mesh frame, it is output to the static mesh encoder (814) through the switching unit (812). The motion vector encoder (813) may be referred to as a motion encoder.

[0210] For example, when intra-frame encoding is performed on the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (814). In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to a multiplexer (823).

[0211] As another example, when inter-frame encoding is performed on the corresponding mesh frame, the motion vector encoder (813) can take a base mesh and a reference restored base mesh (or restored quantized reference base mesh) as inputs, calculate a motion vector between the two meshes, and encode the value. Additionally, the motion vector encoder (813) can perform a prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to a multiplexer (823).

[0212] The base mesh restoration unit (815) can generate a reconstructed base mesh by receiving a base mesh encoded in the static mesh encoder (814) or a motion vector encoded in the motion vector encoder (813) as input. For example, the base mesh restoration unit (815) can restore the base mesh by performing static mesh decoding on the base mesh encoded in the static mesh encoder (814). At this time, quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding. As another example, the base mesh restoration unit (815) can restore the base mesh based on the restored quantized reference base mesh and the motion vector encoded in the motion vector encoder (813). The restored base mesh is output to the displacement calculation unit (816) and the mesh restoration unit (820).

[0213] The displacement calculation unit (816) can perform mesh subdivision on the restored base mesh. The displacement calculation unit (816) can calculate a displacement vector, which is the difference in vertex positions between the subdivided restored base mesh and the fitted subdivision (or subdivided) mesh generated by the pre-processor (811). At this time, the displacement vector can be calculated for as many vertices as the number of vertices of the subdivided mesh. The displacement calculation unit (816) can convert the displacement vector calculated in a 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.

[0214] The displacement vector video generation unit (817) may include a linear lifting unit, a quantizer, and an image packing unit. That is, in the displacement vector video generation unit (817), the linear lifting unit can transform the displacement vector for effective encoding. Depending on the embodiments, the transformation may be performed as a lifting transformation, a wavelet transformation, etc. Additionally, quantization can be performed in the quantizer on the transformed displacement vector value, i.e., the transformation coefficient. At this time, different quantization parameters can be applied to each axis of the transformation coefficient, and the quantization parameters can be derived by the agreement of the encoder / decoder. The displacement vector information that has undergone transformation and quantization can be packed into a 2D image in the image packing unit. The displacement vector video generation unit (817) can generate a displacement vector video by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh.

[0215] The displacement vector video encoder (818) can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer (823).

[0216] The displacement vector restoration unit (819) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (819) performs decoding on the encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (820). The mesh restoration unit (820) restores a deformed mesh based on the base mesh restored in the base mesh restoration unit (815) and the displacement vector restored in the displacement vector restoration unit (819). The restored mesh (or referred to as the restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.

[0217] The texture map video generation unit (821) can regenerate a texture map based on the texture map (or attribute map) of the original mesh and the restored deformed mesh output from the mesh restoration unit (820). According to embodiments, the texture map video generation unit (821) can assign vertex-specific color information of the original mesh's texture map to the texture coordinates of the restored deformed mesh. According to embodiments, the texture map video generation unit (821) can generate a texture map video by grouping the texture maps regenerated for each frame into GoF units.

[0218] The generated texture map video can be encoded using the video compression codec of the texture map video encoder (822). The texture map video bitstream generated through encoding is transmitted to the multiplexer (823).

[0219] The multiplexer (823) multiplexes the motion vector bitstream (e.g., for inter-encoding), base mesh bitstream (e.g., for intra-encoding), displacement vector bitstream, and texture map bitstream into a single bitstream. The single bitstream can be transmitted to the receiver via the transmitter (824). Alternatively, the motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream can be encapsulated into a file with one or more track data or segments and transmitted to the receiver via the transmitter (824).

[0220] Referring to FIG. 12, the transmitting device (encoder) can encode the mesh using an intra-frame or inter-frame method. The transmitting device according to intra-encoding can generate a base mesh, a displacement vector (or displacement), and a texture map (or attribute map). The transmitting device according to inter-encoding can generate a motion vector (or motion), a displacement vector (or displacement), and a texture map (or attribute map). The texture map obtained from the data input unit is generated and encoded based on the restored mesh. The displacement is generated and encoded through the difference in vertex positions between the base mesh and the divided (or subdivided) mesh. More specifically, the displacement is the difference in position between the fitted subdivided mesh and the subdivided restored base mesh, that is, the difference in vertex positions between the two meshes. The base mesh is generated by simplifying and encoding the original mesh through pre-processing. Motion is generated as motion vectors for the mesh of the current frame based on the reference base mesh of the previous frame.

[0221] FIG. 13 shows a mesh data receiving device according to embodiments.

[0222] FIG. 13 corresponds to the receiving device (110) or mesh video decoder (113) of FIG. 2, the decoder of FIG. 11, and / or a corresponding receiving decoding device. Each component of FIG. 13 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of FIG. 13 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of FIG. 12.

[0223] The bitstream of mesh data received by the receiver (910) is demultiplexed in the demultiplexer (911) into a compressed motion vector bitstream (e.g., inter-decoding) or base mesh bitstream (e.g., intra-decoding), displacement vector bitstream, and texture map bitstream after file / segment decapsulation. For example, if the current mesh is inter-frame encoding (i.e., inter-encoding), the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder (913) via the switching unit (912). As another example, if the current mesh is intra-frame encoding (i.e., intra-encoding), the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (914) via the switching unit (912). Here, the motion vector decoder (913) may be referred to as the motion decoder.

[0224] According to embodiments, if the current mesh has inter-frame encoding applied according to the frame header information, the motion vector decoder (913) can perform decoding on the motion vector bitstream. According to embodiments, the motion vector decoder (913) can use the previously decoded motion vector as a predictor and add it to the residual motion vector decoded from the bitstream to restore the final motion vector.

[0225] According to the embodiments, if the current mesh has in-screen encoding applied according to the frame header information, the static mesh decoder (914) can decode the base mesh bitstream to restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.

[0226] According to the embodiments, the base mesh restoration unit (915) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-frame encoding applied, the base mesh restoration unit (915) can generate the restored base mesh by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh has intra-frame encoding applied, the base mesh restoration unit (915) can generate the restored base mesh by performing inverse quantization on the base mesh decoded through the static mesh decoder (914).

[0227] According to the embodiments, the displacement vector video decoder (917) can decode the displacement vector bitstream as a video bitstream using a video codec or decode it using an arithmetic codec. That is, depending on the encoding codec type, if the displacement vector bitstream is encoded through a video codec, for example, after decoding using a video codec, a reverse packing process can be performed. As another example, if it is encoded through arithmetic coding, arithmetic decoding can be performed on the displacement vector bitstream through a displacement vector arithmetic decoding unit, and if inter-frame prediction is performed, the current displacement vector transformation coefficient can be generated by adding the residual value to the reference displacement vector transformation coefficient through inter-frame prediction.

[0228] According to the embodiments, the displacement vector restoration unit (918) extracts displacement vector transformation coefficients from the decoded displacement vector video and restores the displacement vector by applying inverse quantization and inverse transformation processes to the extracted displacement vector transformation coefficients. To this end, the displacement vector restoration unit (918) may include an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. If the restored displacement vector is a value in the local coordinate system, an inverse transformation process to the Cartesian coordinate system may be performed.

[0229] The mesh restoration unit (916) can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and connection information of the texture coordinates can be generated. At this time, the mesh restoration unit (916) can generate a final restored mesh (or restored deformed mesh) by combining the subdivided restored base mesh with the restored displacement vector.

[0230] According to the embodiments, the texture map video decoder (919) can restore the texture map by decoding the texture map bitstream as a video bitstream using a video codec. The restored texture map has color information for each vertex contained in the restored mesh, and the color value of the corresponding vertex can be obtained from the texture map using the texture coordinates of each vertex.

[0231] According to the embodiments, the mesh restored in the mesh restoration unit (916) and the texture map restored in the texture map video decoder (919) are shown to the user through a rendering process in the mesh data renderer (920).

[0232] Referring to FIG. 13, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding receives a base mesh, a displacement vector (or displacement), and a texture map (or attribute map), and can render mesh data based on the restored mesh and the restored texture map. A receiving device according to inter-decoding receives a motion vector (or motion), a displacement vector (or displacement), and a texture map (or attribute map), and can render mesh data based on the restored mesh and the restored texture map.

[0233] A mesh data transmission device and method according to the embodiments may pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream containing the encoded mesh data. A point mesh data reception device and method according to the embodiments may receive a bitstream containing mesh data and decode the mesh data. A mesh data transmission and reception method / device according to the embodiments may be referred to simply as a method / device according to the embodiments. A mesh data transmission and reception method / device according to the embodiments may also be referred to as a 3D data transmission and reception method / device or a point cloud data transmission and reception method / device.

[0234] As described above, each vertex constituting the mesh in this disclosure represents a position in three-dimensional space and is expressed, for example, by x, y, and z coordinates (i.e., a canonical coordinate system). Also, the polygon may be a triangle or a quadrilateral. In this disclosure, vertex, vertex, and point may be used interchangeably. That is, a vertex has coordinates in 3D space, and a polygon of triangle or quadrilateral can be generated through connections between multiple vertices. Furthermore, the V-DMC referred to in this disclosure may also be referred to as V-mesh below, and the two terms are expressions used interchangeably.

[0235] In the present disclosure, displacement information may be obtained based on a subdivided mesh (or referred to as a sub-mesh). That is, the difference in vertex positions between the subdivided reconstructed base mesh, generated by performing a fitting process to make the input mesh and the sub-mesh similar to each other and performing subdivision on the reconstructed base mesh, is calculated. In the present disclosure, this value of the difference in vertex positions is referred to as a displacement vector. In the present disclosure, the term displacement vector may be used interchangeably with the meanings of displacement or displacement information. Furthermore, the term displacement video may be used interchangeably with the meanings of displacement vector video, displacement vector transformation coefficient video, or geometry video, and the term displacement vector may be used interchangeably with the meanings of displacement vector transformation coefficient or displacement vector coefficient.

[0236] And, the V-DMC encoder of the transmitting device (or referred to as the encoder or encoding device) can encode displacement information (i.e., displacement vector or quantized displacement vector transformation coefficient) through a 2D video codec-based encoding method, and / or a zero-run length encoding method, and / or an arithmetic encoding method, etc. And, the V-DMC decoder of the receiving device (or referred to as the decoder or decoding device) can decode the inverse process of the V-DMC encoder of the transmitting device, that is, the encoded displacement information, through a 2D video codec-based decoding method, and / or a zero-run length decoding method, and / or an arithmetic decoding method, etc.

[0237] Taking 2D video codec-based encoding and decoding as an example, the V-DMC encoder calculates a displacement vector, which is the difference between the mesh restored from the base mesh and the mesh fitted during the pre-processing stage. The calculated displacement vector in the canonical coordinate system (i.e., in the form of x, y, z) is then converted into a displacement vector in the local coordinate system (i.e., in the form of normal, tangential, and bi-tangential). Afterward, the displacement vector in the local coordinate system undergoes lifting transformation and quantization to be encoded into a displacement vector bitstream. In this case, the V-DMC decoder restores the displacement vector by performing the inverse process of the V-DMC encoder.

[0238] For example, when V-DMC technology performs compression on displacement vector information using a 2D video codec, it performs image packing on 3D displacement vector information by applying one of the YUV 4:4:4 format, YUV 4:2:0 format, or YUV 4:0:0 format. In particular, the canonical coordinate system (z,y,z) is converted into a local coordinate system (normal, tangential, bi-tangential), and then the normal, tangential, and bi-tangential components are used to perform image packing on the frame according to the image packing format information (YUV 4:4:4, YUV 4:2:0, YUV 4:0:0).

[0239] In the present disclosure, the YUV 4:4:4, YUV 4:2:0, and YUV 4:0:0 formats represent chroma formats when packing the normal component, tangential component, and bi-tangential component of displacement vector coefficients into an image. In other embodiments, the present disclosure may also pack the x component, y component, and z component of displacement vector coefficients into an image. For convenience of explanation, the present disclosure may refer to the normal component or x component as the first component, the tangential component or y component as the second component, and the bi-tangential component or z component as the third component.

[0240] In the present disclosure, the YUV 4:4:4 format means that the sizes of the Y plane, the U plane, and the V plane are equal. The YUV 4:2:0 format means that the sizes of the U plane and the V plane are smaller by a predetermined multiple (e.g., 4 times) compared to the size of the Y plane. Additionally, the YUV 4:0:0 format means that only the size of the Y plane exists; that is, packing is performed only in the Y plane. In other words, among the Y, U, and V planes, only the Y plane exists.

[0241] For convenience of explanation, the present disclosure uses the YUV 4:4:4 format interchangeably with the first format, the YUV 4:2:0 format interchangeably with the second format, and the YUV 4:0:0 format interchangeably with the third format. Additionally, for convenience of explanation, the present disclosure may refer to the Y plane as the first plane, the U plane as the second plane, and the V plane as the third plane. Furthermore, the planes may be referred to as channels.

[0242] Meanwhile, the present disclosure allows for encoding by packing a texture map (or attribute information) and a displacement information frame into a single frame and then using a video codec (i.e., a video compression codec). For example, by packing displacement information into a video frame and packing attribute information into the remaining area, two types of data can be jointly packed into a single frame. This allows the video codec previously used for displacement information compression and the video codec previously used for texture map compression to be combined into one, thereby achieving the effect of increased coding efficiency. Furthermore, the decoder of the receiving device has the effect of decoding displacement information and / or attribute information of related areas at once by unpacking the data packed into a single frame. For example, assuming that a video codec (e.g., HEVC, SHVC, or VVC) is used for displacement information representing geometry information and another video codec (e.g., HEVC, SHVC, or VVC) is used for texture maps, packing the texture maps and displacement information into a single frame allows the use of a single video codec instead of two. Here, HEVC supports encoding / decoding of a single layer, and SHVC is an extension of HEVC that supports encoding / decoding of multiple layers (i.e., one base layer and one or more enhancement layers). Additionally, VVC is a successor standard to SHVC that supports encoding / decoding of multiple layers (i.e., multilayers).

[0243] In this disclosure, data in which a texture map and displacement information are packed into a single frame is referred to as PVD (Packed Video Data), PVD bitstream, PVD sub-bitstream, or packed video sub-bitstream. Additionally, a frame in which a texture map and a displacement information frame are packed together is referred to as a 'packed video frame,' a 'texture map and displacement information frame,' or a 'texture map and displacement vector frame.' Furthermore, in this disclosure, the texture map may be referred to as texture map data, texture data, attribute data, attribute information, or attribute map. Additionally, the displacement information may be referred to as geometry data or geometry information.

[0244] FIG. 14 is a diagram showing an example of a dynamic mesh (V-DMC) bitstream structure encoded and transmitted by a transmitting device of the present disclosure. That is, dynamic mesh content can be encoded into a bitstream structure such as FIG. 14 and transmitted to a receiving device. In particular, the present disclosure may use a sample stream data unit used when encoding V3C content of the V3C codec specification (ISO / IEC 23090-5) as shown in FIG. 14. That is, in the present disclosure, the V3C bitstream can be transmitted / received in either a V3C unit stream format or a V3C sample stream format. The V-DMC bitstream of the present disclosure may follow the V3C bitstream structure defined in the V3C codec specification (ISO / IEC 23090-5) described later. At this time, the existing V3C bitstream structure may be followed, but some V3C units may not be used, and some structures for V-DMC only, such as V-DMC extensions, may be followed.

[0245] That is, the bitstream transmitted from the transmitting device to the receiving device of the present disclosure (referred to as a V-DMC bitstream or dynamic mesh bitstream) may be composed of a sample stream DMC header and a plurality of sample stream DMC units. In the present disclosure, the sample stream DMC header may be referred to as a sample stream header, and the sample stream DMC unit may be referred to as a sample stream data unit.

[0246] In this case, if the sample stream DMC unit complies with the V3C codec specification (ISO / IEC 23090-5), the sample stream DMC unit may be composed of V3C sample stream size information and a V3C unit. The V3C unit is further composed of a V3C unit header (V3C_unit_header) and a V3C unit payload (V3C_unit_payload).

[0247] The above V3C sample stream size information specifies the size of the subsequent V3C unit in bytes.

[0248] The above V3C unit header includes type information (vuh_unit_type) that indicates the type of data carried by the corresponding V3C unit payload. Depending on the type information (vuh_unit_type), the V3C unit payload can carry one of a V3C / V-DMC parameter set (VPS), atlas data (AD), base mesh data (BMD), displacement data / geometry video data (DD / GVD), attribute video data (AVD), or packing video data (PVD).

[0249] Here, VPS may include decoder configuration information related to mesh encoding / decoding and parameter set information such as sequence headers. Atlas data (AD) may include additional information related to 2D mapping or texture mapping for 3D objects. Base mesh data (BMD) is compressed base mesh data for mesh encoding / decoding. Also, DD / GVD represents displacement data (or displacement information), where DD is when the displacement data is arithmetic coding and GVD is when the displacement data is encoded using a video codec. Attribute video data (AVD) is attribute or texture data (or texture map information) compressed using a video codec. Packed video data (PVD) is packed texture map and displacement information compressed using a video codec.

[0250] FIG. 15 is a diagram showing an example of the syntax structure of a V3C unit payload (V3C_unit_payload) according to embodiments. In FIG. 15, numBytesInV3CPayload represents the size of the V3C unit and can be specified by the V3C sample stream size information.

[0251] The V3C unit payload of FIG. 15 may include one of a V3C parameter set (v3c_parameter_set()), an atlas sub-bitstream (atlas_sub_bitstream()), or a video sub-bitstream (video_sub_bitstream()) depending on the value of the vuh_unit_type field of the V3C unit header.

[0252] For example, if the vuh_unit_type field indicates a V3C parameter set (V3C_VPS), the V3C unit payload includes a V3C parameter set (v3c_parameter_set()) containing overall encoding information of the bitstream, and if it indicates atlas data (V3C_AD) or common atlas data (V3C_CAD), it includes an atlas sub-bitstream (atlas_sub_bitstream()) carrying the atlas data or common atlas data. In addition, as an example, if the vuh_unit_type field indicates accusation video data (V3C_OVD), the V3C unit payload includes an accusation video sub-bitstream (video_sub_bitstream()) that carries the accusation video data; if it indicates geometry video data (V3C_GVD), it includes a geometry video sub-bitstream (video_sub_bitstream()) that carries the geometry video data; and if it indicates attribute video data (V3C_AVD), it includes an attribute video sub-bitstream (video_sub_bitstream()) that carries the attribute video data. Furthermore, if the type information (vuh_unit_type) of the V3C unit header indicates PVD, the V3C unit payload includes a packed video sub-bitstream (video_sub_bitstream()) that carries the packed video data. That is, in the present disclosure, packet video data (i.e., texture and displacement data packed into one frame) is transmitted to a receiving device through a V3C unit corresponding to V3C_PVD (vuh_unit_type == V3C_PVD).

[0253] According to the embodiments, the atlas sub-bitstream is referred to as the atlas substream, the accusative video sub-bitstream as the accusative video substream, the geometry video sub-bitstream as the geometry video substream, the attribute video sub-bitstream as the attribute video substream, and the packed video sub-bitstream as the packed video substream. The V3C unit payload according to the embodiments has a NAL unit structure coded in HEVC or VVC (SHVC or multi-layer VVC).

[0254] FIG. 16 is a drawing showing another example of a transmitting device according to embodiments. The transmitting device of FIG. 16 may correspond to the transmitting device of FIG. 1(a), the transmitting device of FIG. 2, the transmitting device of FIG. 7, or the transmitting device of FIG. 12, etc. Therefore, for parts not described in FIG. 16, reference will be made to the description of the transmitting device of FIG. 1(a), the transmitting device of FIG. 2, the transmitting device of FIG. 7, or the transmitting device of FIG. 12. The elements of the transmitting device shown in FIG. 16 may be implemented in hardware, software, a processor connected to memory, and / or a combination thereof. That is, the elements of the transmitting device of FIG. 16 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not shown in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the transmitting device of FIG. 16 described above. In addition, one or more processors may operate or execute a set of software programs and / or instructions for performing operations and / or functions of the elements of the transmitting device of FIG. 16.

[0255] Referring to FIG. 16, dynamic mesh video data acquired through the dynamic mesh video acquisition unit can be encoded in each encoder according to atlas data, base mesh data, displacement data, and attribute data.

[0256] For example, atlas data is encoded into an atlas sub-bitstream in an atlas data encoder, and basemesh data is encoded into a basemesh sub-bitstream in a basemesh encoder before being input into a multiplexer.

[0257] As another example, if packed video is not supported, displacement data is encoded into a geometry video sub-bitstream in a displacement encoder, and attribute data is encoded into an attribute video sub-bitstream in an attribute encoder. The term attribute video may include texture video. In this case, the displacement encoder can perform conversion and quantization on the displacement data, and then perform encoding through video coding based on a video codec, arithmetic coding based on an arithmetic codec, etc. The attribute encoder can regenerate a texture map by performing a texture transfer process and a padding process on the input texture map, and then encode through video coding after performing color space conversion, etc. on the regenerated texture map. That is, the texture transfer process is a process of regenerating the texture map of the current mesh based on the input mesh (i.e., the texture map of the original mesh (or attribute map)) and the restored mesh (i.e., full-resolution geometry).

[0258] As another example, if packed video is supported, the packed encoder can perform encoding by packing displacement frames and attribute frames into a single frame. In this case, the packed encoder can perform transformation and quantization on the displacement data and then pack the quantized displacement vector transformation coefficients into the frame. Additionally, the packed encoder can perform video coding on the attribute frames and then pack the displacement vector transformation coefficients into the packed frame. Here, the attribute frames may be texture maps generated through texture transfer and padding processes, color space conversion, etc. That is, packed video data (i.e., video data in which displacement data and texture map data are packed into a single frame) is encoded into a packed video sub-bitstream by the packed encoder. According to the embodiments, the packed encoder can perform video codec-based encoding on packed video data in which displacement data and texture map data are packed into a single frame. According to embodiments, the displacement frame and the attribute frame may be packed into a single frame, or packed vertically and horizontally, and information regarding the packed area may be signaled through packing information. In one embodiment, the packing information is included in the VPS.

[0259] According to embodiments, when encoding displacement information based on a video codec in a displacement encoder or a packed encoder, for 3D displacement vector information, one of the YUV 4:4:4 format, YUV 4:2:0 format, or YUV 4:0:0 format is applied to perform image packing on one or more of the first to third planes. For example, if the components of the displacement information are normal, tangential, and bitancial, and the chroma format is 4:2:0 or 4:2:2, the normal component, tangential component, and bitancial component of the displacement information can all be packed into the first plane of the frame (e.g., the Y plane). As another example, if the displacement information components include normal, tangential, and bitancial data, and the chroma format is 4:4:4, the normal, tangential, and bitancial components of the displacement information can be packed into the first plane (e.g., Y plane), second plane (e.g., U plane), and third plane (e.g., V plane), respectively, of the frame. In this way, the packing structure (or plane placement structure or packing method) varies depending on the number of displacement information components and the chroma format. For example, if the number of components is 3 and the chroma format is 4:4:4, packing is performed in the first to third planes, respectively, while if the chroma format is 4:2:0, packing is performed only in the first plane.

[0260] According to the embodiments, if packed video is not supported, geometry video sub-bitstreams and attribute video sub-bitstreams are input to the multiplexer, and if packed video is supported, packed video sub-bitstreams can be input to the multiplexer.

[0261] According to embodiments, if the multiplexer does not support packed video, it receives an atlas sub-bitstream, a basemesh sub-bitstream, a geometry video sub-bitstream, and an attribute video sub-bitstream as shown in FIG. 14 and multiplexes them into a single V3C bitstream (or V-DMC bitstream or bitstream of mesh data).

[0262] According to embodiments, when the multiplexer supports packed video, it receives an atlas sub-bitstream, a basemesh sub-bitstream, and a packed video sub-bitstream as shown in FIG. 14 and multiplexes them into a single V3C bitstream (or V-DMC bitstream or bitstream of mesh data).

[0263] According to embodiments, geometry video sub-bitstreams, attribute video sub-bitstreams, or packed video sub-bitstreams may have encoding and transmission omitted.

[0264] The V3C bitstream output from the above multiplexer may further include a VPS.

[0265] According to embodiments, the V3C bitstream output from the multiplexer may be transmitted as is to a receiving device through a file / segment encapsulation and transmission unit, or it may be encapsulated into a file format such as ISOBMFF through the file / segment encapsulation and transmission unit, or processed into other forms such as DASH segments and then transmitted. According to embodiments, the file / segment encapsulation and transmission unit may include mesh video-related metadata in the file format. Mesh video-related metadata may be included, for example, in boxes at various levels in the ISOBMFF file format or as data within a separate track in the file. In this disclosure, mesh video-related metadata may include packing information. According to embodiments, the file / segment encapsulation and transmission unit may encapsulate the mesh video-related metadata itself into a file. According to embodiments, the bitstream of mesh data multiplexed by the multiplexer may be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcasting networks and / or communication networks, etc., and the digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0266] The receiving device of FIG. 17 divides and decodes each sub-bitstream from the received V3C bitstream (or V-DMC bitstream or mesh data bitstream) and then performs mesh reconstruction. If the V3C bitstream (or V-DMC bitstream or mesh data bitstream) is received encapsulated as a file, it is decapsulated in a file / segment decapsulation module (not shown) and then provided for V3C unit extraction. That is, when the stored and / or received file passes through the file / segment decapsulation module, it can become a form similar to the V3C bitstream structure of FIG. 14. If the V3C bitstream is not encapsulated as a file at the transmitting device, the decapsulation process at the receiving device is omitted.

[0267] FIG. 17 illustrates another example of a receiving device according to embodiments. In the present disclosure, the receiving device of FIG. 17 may be referred to as a mesh data receiving device or a decoder or a decoder of a receiving device or a V-Mesh decoder or a dynamic mesh decoder.

[0268] FIG. 17 corresponds to the receiving device of FIG. 1(b), the receiving device of FIG. 2 or a mesh video decoder, the receiving device of FIG. 11 or the receiving device of FIG. 13 and / or a corresponding receiving device. The receiving (decoding) operation of FIG. 17 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of FIG. 16. The elements of FIG. 17 may be implemented in hardware, software, processors, and / or combinations thereof. That is, the elements of the receiving device of FIG. 17 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not shown in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the receiving device of FIG. 17 described above. In addition, one or more processors may operate or execute a set of software programs and / or instructions for performing operations and / or functions of the elements of the receiving device of FIG. 17. The order of execution of each element in FIG. 17 may be changed, some elements may be omitted, and some elements may be newly added.

[0269] FIG. 17 may be an example of a process for decoding and restoring a V3C bitstream, the V3C bitstream may be the input of a dynamic mesh decoder, and the dynamic mesh may be the output.

[0270] That is, in FIG. 17, the V3C unit extraction unit extracts an atlas subbitstream, a base mesh subbitstream, a geometry video (or displacement) subbitstream, an attribute video (or texture map) subbitstream, or a packed video subbitstream from the input V3C bitstream.

[0271] According to embodiments, the V3C unit extractor parses the V3C unit header and V3C unit payload within the V3C bitstream and extracts an atlas subbitstream, a base mesh subbitstream, a geometry video (or displacement) subbitstream, an attribute video (or texture map) subbitstream, or a packed video subbitstream based on type information (i.e., vuh_unit_type) included in each V3C unit header. Here, the geometry video subbitstream, the attribute video subbitstream, or the packed video subbitstream may be transmitted by the transmitting device or may not be transmitted. That is, they may be received by being included in the V3C bitstream or may not be received. In addition, the V3C unit extractor extracts from the V3C bitstream data corresponding to V3C_VPS where the type information (i.e., vuh_unit_type) included in the V3C unit header is V3C_VPS, i.e., VPS data.

[0272] The atlas subbitstream extracted from the above V3C unit extraction unit is decoded by an atlas data decoder and output as atlas data, and the base mesh subbitstream is decoded by a base mesh decoder and output as base mesh data. In addition, the geometry video (or displacement) subbitstream is decoded by a geometry decoder and output as geometry data, and the attribute video (or texture map) subbitstream is decoded by an attribute decoder and output as atlas data. Furthermore, the packed video subbitstream is decoded by a packed decoder and output as packed data.

[0273] That is, the V3C unit extraction process in the V3C unit extraction unit is a process of extracting subbitstreams according to the V3C unit type. According to the embodiments, depending on the vuh_unit_type of the V3C unit header, if vuh_unit_type is V3C_VPS, a VPS subbitstream can be extracted; if it is V3C_AD, an atlas subbitstream is extracted; if it is V3C_GVD, a geometry video subbitstream is extracted; if it is V3C_AVD, an attribute video subbitstream is extracted; and if it is V3C_PVD, a packed video subbitstream is extracted. Each V3C subbitstream can output atlas data decoded through an atlas data decoder if the input is an atlas subbitstream, base mesh data decoded through a base mesh decoder if the input is a base mesh subbitstream, geometry video decoded through a geometry decoder if the input is a geometry video subbitstream if the input is a geometry video subbitstream, attribute video decoded through an attribute decoder if the input is an attribute video subbitstream if the input is an attribute video subbitstream, and packed video decoded through a packed decoder if the input is a packed video subbitstream.

[0274] According to embodiments, the nominal format converter performs nominal format conversion on at least one of base mesh data, geometry data, attribute data, or packed data output from each decoder based on information and / or atlas data included in the VPS data.

[0275] According to embodiments, the reconstruction unit reconstructs the mesh based on atlas data, base mesh data, geometry data, attribute data, etc., for which nominal format conversion has been performed or has not been performed in the nominal format conversion unit.

[0276] That is, each data output through each decoder may undergo a process of conversion to a nominal format in a nominal format conversion unit before the reconstruction process is performed in the reconstruction unit. Subsequently, the data converted to a nominal format may be used as the input to the reconstruction process of the reconstruction unit, and a mesh frame may be output as the output of the reconstruction process of the reconstruction unit. According to embodiments, a post-reconstruction unit may perform a post-reconstruction process on the mesh frame output from the reconstruction unit. According to embodiments, the post-reconstruction process may be a process such as zippering that adjusts the difference between submesh boundaries.

[0277] A detailed explanation of the above-mentioned nominal format conversion unit and recovery unit will be provided later.

[0278] The present disclosure proposes a method for solving problems that occur during the nominal format conversion and reconstruction process of geometry data (or referred to as geometry video or displacement data). In particular, the present disclosure relates to V-DMC, a method for compressing 3D dynamic mesh data using a 2D video codec. It describes a method for converting geometry data to a nominal format by considering a packing method during nominal format conversion, a method for performing a reconstruction process when the V3C unit type is PVD, and the syntax and semantics related thereto.

[0279] In other words, when video codec-based decoding is performed during the process of decoding displacement data in current V-DMC technology, a process of converting to a nominal format is performed before performing a reconstruction process on the decoded geometry video. At this time, the decoded geometry video may be in various chroma formats, and although the packing method varies depending on each chroma format, these factors are not taken into account during the nominal format conversion process. That is to say, while the displacement encoding / decoding process in existing V-DMC allows for various packing methods by considering the chroma format of the geometry video, errors may occur during the decoding (or reconstruction) process because the geometry video is converted to a nominal format without considering its currently packed state. The present disclosure aims to resolve this by performing nominal format conversion while considering the packing method of the geometry video. Through this, errors that may occur during the nominal format conversion of the geometry video can be resolved.

[0280] In addition, although V-DMC currently supports the PVD type among V3C unit types, a problem arises in that the reconstruction process of displacement data (or geometry video or geometry data) is not performed when the V3C unit type is the PVD type. In other words, although V-DMC currently supports the PVD type among V3C unit types, this is not specified in the reconstruction process. The present disclosure is intended to solve this problem by performing the reconstruction process by considering the PVD type during the reconstruction process related to displacement data. Through this, the reconstruction process can be performed even when the V3C unit type is PVD.

[0281] As such, the present disclosure allows for conversion into a suitable video by considering a packing method during the nominal format conversion process of a geometry video and performing the conversion into a nominal format, and in the case of a PVD type, a reconstruction process related to displacement data can be performed.

[0282] The following is an explanation of the nominal format conversion (or post-decoding conversion to nominal video formats) in the nominal format conversion section.

[0283] That is, after decoding video (i.e., geometry information and / or attribute information) in one or more decoders, the decoded video can be converted into a nominal format in a nominal format conversion unit before the mesh is restored in the restoration unit and used as input for the mesh restoration process. In the present disclosure, the nominal format may be resolution, bit depth, chroma format, etc. Furthermore, the target for nominal format conversion of geometry information may be geometry video, attribute video, packed video, base mesh, arithmetic decoded displacement vector data, etc.

[0284] According to embodiments, the nominal format conversion unit may perform processes such as unpacking, map extraction, bit depth conversion, resolution conversion, output order conversion, atlas composition alignment, attribute dimension packing, and chroma up-sampling. According to embodiments, the target of unpacking may be packed data, the target of map extraction may be geometry data and / or attribute data, and the target of bit depth conversion may be at least one of base mesh data, geometry data, and attribute data. The target of resolution conversion may be geometry data and / or attribute data, the target of output order conversion may be base mesh data, and the target of atlas composition alignment may be at least one of base mesh data, geometry data, and attribute data. The target of attribute dimension packing and chroma up-sampling may be attribute data.

[0285] That is, the process of converting geometry data to a nominal format may be a process of converting the decoded geometry video (i.e., geometry data) into a nominal format. That is, the nominal format converter converts the decoded geometry video into a geometry video in a nominal format. According to embodiments, processes such as map extraction, bit depth conversion, frame resolution conversion, and composition time alignment may be performed on the decoded geometry video (i.e., geometry data). Here, the geometry video may refer to a video in which displacement data is packed, and may be provided from a displacement decoder or provided after unpacking from packed data.

[0286] The above unpacking process may be a process of separating geometry video (or referred to as geometry data, displacement data, or displacement information) and attribute video (or referred to as attribute data, attribute information, or texture map data) from packed data. Then, nominal format conversion may be performed on each of the separated geometry video and attribute video.

[0287] The above bit depth conversion process may be a process of converting the bit depth of the decoded video into a nominal bit depth, and information about the nominal bit depth may be signaled to the geometry information() of the VPS by the transmitting device and transmitted to the receiving device.

[0288] The above resolution conversion process may be a process of converting the resolution of a decoded video to a nominal resolution, and information regarding the nominal resolution may be signaled to the ASPS (atlas sequence parameter set) of the atlas data by a transmitting device and transmitted to a receiving device.

[0289] The above output composition time conversion (or atlas composition arrangement) process may be a process of adjusting the output order index of input frames to match the output composition time.

[0290] In the present disclosure, the packing method may vary depending on the chroma format and the number of displacement components of the displacement video (or geometry video). For example, the displacement components may be normal, tangential, and bitangent, and depending on the selection of the encoder of the transmitting device, only the normal may be transmitted, the normal, tangential, and bitangent may all be transmitted, or only some may be transmitted.

[0291] According to an embodiment, in the encoder of the transmitting device, displacement vector data can be packed in the first plane of the video (or frame) when there are three displacement components (e.g., normal, tangential, bi-tangential, x, y, z, etc.) and the chroma format is 4:2:0 or 4:2:2, or when there is one displacement component (e.g., normal, etc.) and the chroma format is 4:0:0. And, when there are three displacement components and the chroma format is 4:4:4, displacement vector data can be packed by component in each of the three planes of the video (or frame), namely the first to third planes.

[0292] According to embodiments, the nominal format conversion unit may perform a nominal format conversion process for the chroma format of the geometry video. According to embodiments, the process of converting to a nominal chroma format may be performed through a chroma upsampling process. Accordingly, the chroma format of the decoded geometry video can be converted to a nominal chroma format through a chroma upsampling process.

[0293] According to the embodiments, the nominal chroma format may be fixed to a specific chroma format such as 4:4:4, or information about the nominal chroma format may be signaled by a transmitting device using syntax such as VPS, ASPS, etc., and transmitted to a receiving device.

[0294] According to embodiments, the number of components of the geometry video may be fixed to 1, fixed to 3, fixed to a specific value, or derived.

[0295] As described above, the nominal format converter converts the decoded geometry video into a geometry video in nominal format.

[0296] At this time, the geometry video in the nominal format can be 5-dimensional and can correspond to a map index, geometry video frame index, components, column index, and row index, respectively.

[0297] Also, the number of components of the geometry video (geoNumComp variable) can be fixed to a specific value or determined according to the chroma format.

[0298] According to the embodiments, when the number of components of the geometry video is fixed to a specific value, it may be fixed to 3 or 1.

[0299] At this time, if the number of components of the geometry video is fixed to 1, only the area corresponding to the first plane of the video can be used, and the use of the 4:4:4 chroma format among the chroma formats of the displacement video can be restricted.

[0300] In the present disclosure, when the number of components of a geometry video is fixed to 3, the area corresponding to the first to third planes of the video can be used, and in such a case, if the chroma format is not 4:4:4, a chroma upsampling process can be performed to a 4:4:4 chroma format.

[0301] According to embodiments, in the present disclosure, the number of components of a geometry video is determined according to the chroma format, and if the decoded geometry chroma format is 4:4:4, the number of components of the geometry video can be derived to 3, and otherwise, the number of components of the geometry video can be derived to 1. In this case, the map index and the frame index may be considered together. That is, the number of components of the geometry video can be derived to 3 or 1 depending on the map index and the frame index. For example, if the decoded geometry chroma format (DecGeoChromaFormat[mapIdx][frameIdx]) corresponding to the map index (mapIdx) and frame index (frameIdx) is 4:4:4, the number of geometry video components can be set to 3, and if it is not 4:4:4 (e.g., 4:2:2, 4:2:0, etc.), the number of geometry video components can be set to 1.

[0302] The following is an explanation of the geometry nominal format conversion process. In other words, the geometry nominal format conversion process converts the decoded geometry frames (DecGeoFrames) into the nominal format.

[0303] 이때, 다음 변수들인 geoNumComp, geoMultipleMapsPresentFlag, geoMapCount, 그리고 geoMSBAlignFlag은 아래와 같이 설정한다.

[0304] if(DecGeoChromaFormat == 4:4:4)

[0305] geoNumComp = 3

[0306] else

[0307] geoNumComp = 1

[0308] geoMultipleMapsPresentFlag = vps_multiple_map_streams_present_flag[ ConvAtlasID ]

[0309] geoMapCount = vps_map_count_minus1[ ConvAtlasID ] + 1

[0310] if(pin_geometry_present_flag[ ConvAtlasID ])

[0311] geoMSBAlignFlag = pin_geometry_msb_align_flag[ ConvAtlasID ]

[0312] else

[0313] geoMSBAlignFlag = gi_geometry_msb_align_flag[ ConvAtlasID ]

[0314] Let the 1D array geoMapAbsCodingFlag be set as follows:

[0315] for( c = 0; c < geoMapCount; c++ ) {

[0316] geoMapAbsCodingFlag[ c ] =

[0317] vps_map_absolute_coding_enabled_flag[ ConvAtlasID ][ c ]

[0318] }

[0319] In the above, DecGeoChromaFormat is the chroma format of the decoded geometry video (e.g., 4:4:4, 4:2:0), geoNumComp is the number of components of the geometry video (or displacement video), and vps_multiple_map_streams_present_flag[ConvAtlasID] is a flag indicating whether multiple map streams are signaled to the VPS. vps_map_count_minus1[ConvAtlasID] is the number of maps signaled to the VPS minus 1, and pin_geometry_present_flag[ConvAtlasID] is a flag indicating whether geometry information is signaled to the packing information (packing_information()). pin_geometry_msb_align_flag and gi_geometry_msb_align_flag are flags indicating the most significant bit (MSB) alignment policy of the geometry video signaled in the packing information (packing_information()) or geometry information (geometry_information()) within the VPS, vps_map_absolute_coding_enabled_flag[ConvAtlasID][c] indicates whether absolute coding is enabled for map c signaled in the VPS, and ConvAtlasID is the identifier of the atlas to be converted.

[0320] That is, the code above represents the process of deriving control variables required for nominal format conversion, such as the number of components of the geometry video (geoNumComp), whether to use multiple map streams (geoMultipleMapsPresentFlag), the number of maps (geoMapCount), whether to align the most significant bit (geoMSBAlignFlag), and / or the map absolute coding flag (geoMapAbsCodingFlag), based on the chroma format of the decoded geometry video and the syntax elements of the VPS.

[0321] According to embodiments, the nominal format converter checks the chroma format of the decoded geometry video and determines the number of geometry components as follows. In one embodiment, if DecGeoChromaFormat is 4:4:4, the number of geometry video components (geoNumComp) is set to 3. In another embodiment, for other formats (e.g., 4:2:0, 4:2:2, etc.), the number of geometry video components (geoNumComp) is set to 1.

[0322] The output (GeoFramesNF) of the above nominal format converter is as follows:

[0323] That is, a 5-dimensional array GeoFramesNF, which is decoded geometry frames in nominal format. The dimensions of this array correspond to the map index, geometry video frame index, components, column index, and row index of the frames.

[0324] As mentioned above, the chroma format of the decoded geometry video can vary, and the packing method differs depending on each chroma format; however, conventionally, nominal format conversion was performed without considering these factors during the nominal format conversion process. In other words, the conventional nominal format conversion process did not consider the situation of packing data into three planes (i.e., the first plane to the third plane) by fixing the number of video planes to always 1 (i.e., fixing the number of components of the geometry video to 1). In this case, even though the chroma format is 4:4:4 and the number of displacement data components is 3, and each component of the displacement data is packed in the first to third planes at the transmitting device, the nominal format conversion unit performed nominal format conversion only on the components packed in the first plane, and did not perform nominal format conversion on the second and third planes, assuming that there was no packed data therein. And, only the components of the first plane converted to the nominal format were output to the restoration unit. As a result, an error occurs during the decoding (or restoration) process.

[0325] In the present disclosure, when the chroma format is 4:4:4 and the number of components of displacement data is 3, the number of components of the geometry video is set to 3, so that the nominal format conversion unit performs nominal format conversion by considering all of the first to third planes. That is, for each component of the geometry video packed in each plane (e.g., the first to third planes), processes such as map extraction, bit depth conversion, frame resolution conversion, and composition time alignment are performed.

[0326] Meanwhile, the restoration unit restores the mesh by processing BasemeshFramesNF, which is a basemesh component of the nominal format, GeoFramesNF, which is a video component of the nominal format, and AttrFramesNF if applicable, and decoded atlas data. At this time, in the nominal format conversion unit, when the chroma format is 4:4:4 and the number of components of the displacement data is 3, the number of components of the geometry video is set to 3, and since the nominal format conversion is performed by considering all of the first to third planes, the restoration unit can also perform mesh restoration by considering all of the first to third planes.

[0327] According to the embodiments, the mesh restoration process in the restoration unit may be performed on data output after being converted to a nominal format in the nominal format conversion unit. In this case, the inputs of the restoration unit are VPS (V3C parameter set), atlas data, and base mesh data, which are mandatory, and geometry data (or displacement data), attribute data, and packed data are optional. That is, when geometry data (or displacement data) is input, this geometry data (or displacement data) may be data decoded by a displacement decoder, or geometry data (or displacement data) separated after being decoded by a packed decoder.

[0328] In the above-mentioned restoration unit, the mesh restoration process may determine whether to perform some processes depending on the presence or absence of displacement data. During the mesh restoration process, restoration processes related to displacement data may include, according to the embodiments, a process of backpacking the displacement vector transformation coefficient image, a process of inversely quantizing the quantized displacement vector transformation coefficients, a process of inversely transforming the displacement vector transformation coefficients, and a process of generating normal, tangent, and bitangent vectors of the subdivided mesh. Depending on the coding method of the displacement data, for example, if it is coded through a video codec in an encoder, an image backpacking process may be performed, and if it is coded through an arithmetic codec, the image backpacking process may be omitted.

[0329] According to embodiments, whether to perform a displacement data restoration process may be determined based on the existence of displacement data. In the present disclosure, the existence of displacement data may be determined through the existence of a geometry video and the existence of a packed video. That is, the restoration unit determines that displacement data exists if a geometry video exists or a packed video exists.

[0330] In the present disclosure, there may be a constraint that geometry video and packed video cannot exist simultaneously. According to embodiments, the presence of geometry video may be determined by vps_geometry_video_present_flag, a syntax signaled at the VPS level, and the presence of packed video may be determined by vps_packed_video_present_flag, a syntax signaled at the VPS level.

[0331] According to embodiments, the restoration unit determines that displacement data exists in the input of the restoration unit if vps_geometry_video_present_flag is 1 or vps_packed_video_present_flag is 1. Here, the displacement data may be data decoded by a displacement decoder or data separated after decoding by a packed decoder. That is, if packed data exists as an input to the mesh restoration process, the packed data can be used as the input to the restoration unit. And, the existence of packed video can be checked through the packed_video_present_flag signaled at the VPS level.

[0332] In another embodiment, the presence of packed video may be determined using pin_geometry_present_flag signaled in the packing information. For example, if the value of vps_geometry_video_present_flag[RecAtlasID] or pin_geometry_present_flag[RecAtlasID] is equal to 1 and the profile or geometry codec ID indicates the use of a 2D video codec, a 5-dimensional array geoFramesNF[0][compTimeIdx][0][y][x] is used to specify geometry frames decoded in the nominal format. In this case, y is in the range from 0 to asps_frame_height - 1, and x is in the range from 0 to asps_frame_width - 1. In other cases, i.e., when the profile or geometry codec represents an arithmetic codec, a 3D array dispFramesNF[compTimeIdx][dispIdx][dispDimIdx] is used to specify the displacement decoded from the nominal format. In this case, dispIdx is in the range of 0 to DispCountPerFrame[compTimeIdx] - 1, and dispDimIdx is in the range of 0 to DispDimension - 1.

[0333] If the value of vps_geometry_video_present_flag[j] is 0, it indicates that geometry video data is not associated with the atlas with atlas ID j. If the value of vps_geometry_video_present_flag[j] is 1, it indicates that geometry video data must be associated with the atlas with atlas ID j. If vps_geometry_video_present_flag[j] does not exist, its value is inferred to be equal to 1. As a requirement for bitstream conformance, if the value of vps_geometry_video_present_flag[j] is 1 for the atlas with atlas ID j, the value of pin_geometry_present_flag[j] in the packing information for the same atlas ID j must be 0.

[0334] If the value of vps_packed_video_present_flag[j] is 0, it indicates that packed video data is not associated with the atlas with atlas ID j. If the value of vps_packed_video_present_flag[j] is 1, it indicates that packed video data should be associated with the atlas with atlas ID j. If vps_packed_video_present_flag[j] does not exist, its value is inferred to be equal to 0.

[0335] If the value of pin_geometry_present_flag[j] is 0, it indicates that the packed video frames of the atlas with atlas ID j do not have regions containing geometry data. If the value of pin_geometry_present_flag[j] is 1, it indicates that the packed video frames of the atlas with atlas ID j have regions containing geometry data. If pin_geometry_present_flag[j] does not exist, its value is inferred to be equal to 0. As a requirement for bitstream conformance, if the value of pin_geometry_present_flag[j] is 1 for an atlas with atlas ID j, the value of vps_geometry_video_present_flag[j] for the same atlas ID j must be 0.

[0336] In other words, while existing V-DMCs support the PVD type among V3C unit types, a problem arises in that the restoration process is not performed when the V3C unit type is the PVD type during the displacement data restoration process. The present disclosure determines whether a packed video exists based on vps_packed_video_present_flag and / or pin_geometry_present_flag. If it is determined that a packed video exists, the restoration unit performs the following steps for the displacement data partitioned from the packed data: inverse packing of the displacement vector transformation coefficient image, inverse quantization of the quantized displacement vector transformation coefficients, inverse transformation of the displacement vector transformation coefficients, and generation of normal, tangent, and bitangent vectors of the subdivided mesh.

[0337] As such, the present disclosure can perform a restoration process by considering the PVD type during the restoration process related to displacement data. That is, the restoration unit can perform the restoration process of displacement data not only when the V3C unit type is DD / GVD, but also when it is PVD.

[0338] In the present disclosure, the nominal chroma format of the geometry data may be signaled to geometry information (geometry_information()) as in FIG. 18, signaled to extension information (asps_vdmc_extension()) within ASPS as in FIG. 20a and 20b, or derived from a decoder without being signaled. When the nominal chroma format of the geometry data is signaled to VPS or ASPS as in FIG. 18 or FIG. 20a and 20b, the nominal format conversion unit parses the VPS of FIG. 18 or the ASPS of FIG. 20a and 20b to obtain the nominal chroma format and can use it for converting the nominal format of the geometry data.

[0339] FIG. 18 is a diagram showing an example of the syntax structure of geometry_information(atlasID) included in a VPS according to embodiments. That is, FIG. 18 is an example of the syntax structure when signaling the chroma format of a geometry video in a VPS according to embodiments.

[0340] In FIG. 18, gi_geometry_codec_id[atlasID] represents the mapping index of the codec identifier of the video decoder used to decode the geometry video sub-bitstream for the atlas identified by the atlas ID (if present).

[0341] The value obtained by adding 1 to gi_geometry_2d_bit_depth_minus1[atlasID] represents the nominal 2D bit depth at which all geometry videos of the atlas identified by the atlas ID must be converted.

[0342] gi_geometry_msb_align_flag[atlasID] indicates how decoded geometry video samples associated with an atlas identified by the atlas ID are converted into samples of nominal geometry bit depth.

[0343] The value obtained by adding 1 to gi_geometry_3d_coordinates_bit_depth_minus1[atlasID] represents the bit depth of the geometry coordinates of the restored volumetric content (restored 3D content) of the atlas identified by the atlas ID.

[0344] gi_geometry_chroma_format[atlasID] can represent the nominal chroma format of the geometry video of the atlas identified by the atlas ID.

[0345] FIG. 19 is a diagram showing examples of nominal chroma formats of geometry videos according to embodiments. That is, the value of gi_geometry_chroma_format[atlasID] can be derived through FIG. 19. For example, if the value of gi_geometry_chroma_format[atlasID] is 0, the chroma format is 4:2:0; if the value of gi_geometry_chroma_format[atlasID] is 1, the chroma format is 4:4:4; if the value of gi_geometry_chroma_format[atlasID] is 2, the chroma format is 4:0:0; and if the value of gi_geometry_chroma_format[atlasID] is 3, the chroma format is 4:2:2.

[0346] FIGS. 20a and 20b illustrate an example of the syntax structure of asps_vdmc_extension() included in an atlas sequence parameter set according to embodiments. That is, FIGS. 20a and 20b are examples of the syntax structure when signaling the chroma format of a geometry video in ASPS according to embodiments. According to embodiments, when signaling the chroma format of a geometry video in ASPS, the nominal chroma format of the geometry video can be signaled in the atlas syntax signaled at the sequence level.

[0347] In FIG. 20a and FIG. 20b, asve_geometry_chroma_format[atlasID] may represent the nominal chroma format of the geometry video of the atlas identified by the atlas ID.

[0348] FIG. 21 is a diagram showing examples of nominal chroma formats of geometry video according to embodiments. That is, the value of asve_geometry_chroma_format[atlasID] can be derived through FIG. 21. For example, if the value of asve_geometry_chroma_format[atlasID] is 0, the chroma format is 4:2:0; if the value of asve_geometry_chroma_format[atlasID] is 1, the chroma format is 4:4:4; if the value of asve_geometry_chroma_format[atlasID] is 2, the chroma format is 4:0:0; and if the value of asve_geometry_chroma_format[atlasID] is 3, the chroma format is 4:2:2.

[0349] FIG. 22 is a flowchart showing an example of an encoding method according to embodiments. The encoding method according to embodiments may include the step of encoding a base mesh of mesh data (S31011), the step of encoding a displacement of mesh data (S31012), and the step of encoding attributes of mesh data (S31013). Additionally, the encoding method according to embodiments may further include the step of encoding atlas data and the step of encoding packed data.

[0350] In the step of encoding the base mesh of the above mesh data (S31011), if intra-frame encoding is performed for the corresponding mesh frame, the base mesh can be encoded using a static mesh encoder. In this case, encoding can be performed on the connectivity information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. In the step of encoding the base mesh of the above mesh data (S31011), if inter-frame encoding is performed for the corresponding mesh frame, a motion vector encoder can calculate the motion vector between the base mesh and the reference restored base mesh (or the restored quantized reference base mesh) and encode the value. Additionally, a prediction based on connectivity information can be performed using a previously encoded / decoded motion vector as a predictor, and the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector can be encoded.

[0351] The step of encoding the displacement of the mesh data (S31012) may perform video codec-based encoding or arithmetic codec-based encoding on the displacement data. The step of encoding the displacement of the mesh data (S31012) may convert the coordinate system of the displacement data from a 3D Cartesian coordinate system to a local coordinate system before encoding the displacement data.

[0352] The step of encoding attributes of the mesh data (S31013) can perform encoding based on a video codec for the attribute data (or texture map).

[0353] The step of encoding the above atlas data can encode the atlas data.

[0354] The step of encoding the packed data above can be performed by packing the displacement frame and the attribute frame into a single frame.

[0355] According to embodiments, when encoding displacement information based on a video codec in the step of encoding the displacement of mesh data (S31012) or the step of encoding packed data (S31012), for 3D displacement vector information, one of the YUV 4:4:4 format, YUV 4:2:0 format, or YUV 4:0:0 format is applied to perform image packing on one or more of the first to third planes. For example, if the components of the displacement information are normal, tangential, and bitancial, and the chroma format is 4:2:0 or 4:2:2, the normal component, tangential component, and bitancial component of the displacement information can all be packed into the first plane of the frame (e.g., the Y plane). As another example, if the displacement information components include normal, tangential, and bitancial data, and the chroma format is 4:4:4, the normal, tangential, and bitancial components of the displacement information can be packed into the first plane (e.g., Y plane), second plane (e.g., U plane), and third plane (e.g., V plane), respectively, of the frame. In this way, the packing structure (or plane placement structure or packing method) varies depending on the number of displacement information components and the chroma format. For example, if the number of components is 3 and the chroma format is 4:4:4, packing is performed in the first to third planes, respectively, while if the chroma format is 4:2:0, packing is performed only in the first plane.

[0356] According to embodiments, if packed video is not supported, the atlas sub-bitstream, basemesh sub-bitstream, geometry video sub-bitstream, and attribute video sub-bitstream are multiplexed into a single V3C bitstream (or referred to as the V-DMC bitstream or mesh data bitstream) and transmitted.

[0357] According to embodiments, when packed video is supported, an atlas sub-bitstream, a basemesh sub-bitstream, and a packed video sub-bitstream can be multiplexed into a single V3C bitstream (or referred to as a V-DMC bitstream or a bitstream of mesh data) and transmitted.

[0358] According to embodiments, encoding and transmission may be omitted for geometry video sub-bitstreams, attribute video sub-bitstreams, or packed video sub-bitstreams. The V3C bitstream may further include a VPS.

[0359] The encoding method of the present disclosure may be performed by an encoding device (encoder). The encoding device includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to encode a base mesh of mesh data, encode displacements of mesh data, encode attributes of mesh data, and encode atlas data and / or packed data.

[0360] The embodiments further include a computer-readable storage medium that stores a bitstream generated by the method according to FIG. 22.

[0361] The embodiments further include a method comprising the steps of acquiring a bitstream for mesh data, said bitstream being generated based on the steps of encoding a base mesh of the mesh data, encoding a displacement of the mesh data, and encoding attributes of the mesh data, and transmitting data including said bitstream.

[0362] FIG. 23 is a flowchart showing an example of a decoding method according to embodiments. The decoding method according to embodiments may include a step of decoding a base mesh within a bitstream (S32011), a step of decoding a displacement within a bitstream (S32012), a step of decoding an attribute within a bitstream (S32013), a nominal format conversion step (S32014), and a step of restoring a mesh (S32015). The decoding step of FIG. 23 may further include a V3C unit extraction step. The V3C unit extraction step may separate a VPS, an atlas sub-bitstream, a base mesh sub-bitstream, a displacement sub-bitstream, an attribute sub-bitstream (or attribute video sub-bitstream), and / or a packed video sub-bitstream from the bitstream according to type information of a V3C unit header within a V-DMC bitstream (or V3C bitstream) such as FIG. 14.

[0363] The step of decoding the base mesh within the bitstream (S32011) can restore the final motion vector by using the previously decoded motion vector as a predictor and adding it to the residual motion vector decoded from the base mesh sub-bitstream if the current mesh is subject to inter-frame encoding. The step of decoding the base mesh within the bitstream (S32011) can restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh by statically decoding the base mesh sub-bitstream if the current mesh is subject to intra-frame encoding.

[0364] The step of decoding displacement within the bitstream (S32012) performs decoding on the displacement sub-bitstream based on a video codec if the displacement data is encoded based on a video codec, and performs decoding based on an arithmetic codec if the displacement data is encoded based on an arithmetic codec. In this disclosure, the displacement sub-bitstream may be referred to as a displacement vector bitstream. The step of decoding displacement within the bitstream (S32012) may perform decoding on the packed video sub-bitstream based on a video codec.

[0365] The step of decoding attributes within the bitstream (S32013) can perform decoding based on a video codec for the attribute sub-bitstream.

[0366] The decoding method according to the embodiments may further include the step of decoding an atlas sub-bitstream.

[0367] The above nominal format conversion step (S32014) performs nominal format conversion on at least one of the base mesh data, geometry data, attribute data, or packed data output from each decoding step based on the information and / or atlas data included in the VPS data.

[0368] According to embodiments, the nominal format conversion step (S32014) may perform processes such as unpacking, map extraction, bit depth conversion, resolution conversion, output order conversion, atlas composition alignment, attribute dimension packing, and chroma up-sampling. According to embodiments, the target of unpacking may be packed data, the target of map extraction may be geometry data and / or attribute data, and the target of bit depth conversion may be at least one of base mesh data, geometry data, and attribute data. The target of resolution conversion may be geometry data and / or attribute data, the target of output order conversion may be base mesh data, and the target of atlas composition alignment may be at least one of base mesh data, geometry data, and attribute data. The target of attribute dimension packing and chroma up-sampling may be attribute data.

[0369] That is, the process of converting geometry data into a nominal format may be a process of converting the decoded geometry video (i.e., geometry data) into a nominal format. That is, the nominal format conversion step (S32014) converts the decoded geometry video into a geometry video in a nominal format. According to embodiments, processes such as map extraction, bit depth conversion, frame resolution conversion, and composition time alignment may be performed on the decoded geometry video (i.e., geometry data).

[0370] In the present disclosure, the packing method may vary depending on the chroma format and the number of displacement components of the displacement video (or geometry video). For example, the displacement components may be normal, tangential, and bitangent, and depending on the selection of the encoder of the transmitting device, only the normal may be transmitted, the normal, tangential, and bitangent may all be transmitted, or only some may be transmitted.

[0371] According to an embodiment, in the encoder of the transmitting device, displacement vector data can be packed in the first plane of the video (or frame) when there are three displacement components (e.g., normal, tangential, bi-tangential, x, y, z, etc.) and the chroma format is 4:2:0 or 4:2:2, or when there is one displacement component (e.g., normal, etc.) and the chroma format is 4:0:0. And, when there are three displacement components and the chroma format is 4:4:4, displacement vector data can be packed by component in each of the three planes of the video (or frame), namely the first to third planes.

[0372] According to embodiments, the nominal format conversion step (S32014) may perform a nominal format conversion process for the chroma format of the geometry video. According to embodiments, the process of converting to a nominal chroma format may be performed through a chroma upsampling process. Thus, the chroma format of the decoded geometry video can be converted to a nominal chroma format through a chroma upsampling process.

[0373] According to the embodiments, the number of components of the geometry video (geoNumComp variable) may be fixed to a specific value or determined according to the chroma format.

[0374] According to the embodiments, when the number of components of the geometry video is fixed to a specific value, it may be fixed to 3 or 1.

[0375] According to embodiments, in the present disclosure, the number of components of a geometry video is determined according to the chroma format, and if the decoded geometry chroma format is 4:4:4, the number of components of the geometry video can be derived to 3, and otherwise, the number of components of the geometry video can be derived to 1.

[0376] In the present disclosure, when the chroma format is 4:4:4 and the number of components of the displacement data is 3, the number of components of the geometry video is set to 3, so that in the nominal format conversion step (S32014), nominal format conversion is performed by considering all of the first to third planes. That is, for each component of the geometry video packed in each plane (e.g., the first to third planes), processes such as map extraction, bit depth conversion, frame resolution conversion, and composition time alignment are performed.

[0377] The step of restoring the mesh (S32015) restores the mesh based on atlas data, base mesh data, geometry data, attribute data, etc., for which nominal format conversion was performed or not performed in the nominal format conversion step (S31024).

[0378] At this time, in the nominal format conversion step (S31024), when the chroma format is 4:4:4 and the number of components of the displacement data is 3, the number of components of the geometry video is set to 3, and since the nominal format conversion is performed by considering all of the first to third planes, the mesh restoration step (S32015) can also be performed by considering all of the first to third planes.

[0379] According to the embodiments, the inputs of the step of restoring the mesh (S32015) are VPS (V3C parameter set), atlas data, and base mesh data, which are mandatory, and geometry data (or displacement data), attribute data, and packed data, which are optional. That is, when geometry data (or displacement data) is input, this geometry data (or displacement data) may be data decoded by a displacement decoder, or geometry data (or displacement data) separated after decoding by a packed decoder.

[0380] In the step of restoring the mesh (S32015) described above, the mesh restoration process may determine whether to perform some processes depending on the existence of displacement data. During the mesh restoration process, restoration processes related to displacement data may include, according to embodiments, a process of backpacking the displacement vector transformation coefficient image, a process of inversely quantizing the quantized displacement vector transformation coefficients, a process of inversely transforming the displacement vector transformation coefficients, and a process of generating normal, tangent, and bitangent vectors of the subdivided mesh. Depending on the coding method of the displacement data, for example, if it is coded through a video codec in an encoder, an image backpacking process may be performed, and if it is coded through an arithmetic codec, the image backpacking process may be omitted.

[0381] According to embodiments, whether to perform the displacement data restoration process may be determined based on the existence of displacement data. In the present disclosure, the existence of displacement data may be determined by the existence of geometry video and packed video. That is, the step of restoring the mesh (S32015) determines that displacement data exists if geometry video exists or packed video exists. According to embodiments, the existence of geometry video may be determined based on vps_geometry_video_present_flag, a syntax signaled at the VPS level, and the existence of packed video may be determined based on vps_packed_video_present_flag, a syntax signaled at the VPS level.

[0382] According to embodiments, the step of restoring the mesh (S32015) determines that displacement data exists in the input of the restoration unit if vps_geometry_video_present_flag is 1 or vps_packed_video_present_flag is 1. Here, the displacement data may be data decoded in the displacement decoding step, or data separated after decoding in the packed decoding step. In addition, the existence of packed video can be checked through packed_video_present_flag signaled at the VPS level. In another embodiment, the existence of packed video may be determined using pin_geometry_present_flag signaled in the packing information within the VPS.

[0383] In other words, while existing V-DMCs support the PVD type among V3C unit types, a problem arises in that the restoration process is not performed when the V3C unit type is of the PVD type during the displacement data restoration process. The present disclosure determines whether a packed video exists based on vps_packed_video_present_flag and / or pin_geometry_present_flag. If it is determined that a packed video exists, the restoration unit performs the following processes for the displacement data that is partitioned from the packed data: inverse packing of the displacement vector transformation coefficient image, inverse quantization of the quantized displacement vector transformation coefficients, inverse transformation of the displacement vector transformation coefficients, and generation of normal, tangent, and bitangent vectors of the subdivided mesh. In this way, the present disclosure can perform the restoration process by considering the PVD type during the restoration process related to displacement data. That is, the restoration unit can perform the displacement data restoration process not only when the V3C unit type is DD / GVD, but also when it is PVD.

[0384] For details not described in FIG. 23, refer to the descriptions in FIG. 14 to FIG. 21.

[0385] The decoding method of the present disclosure may be performed by a decoding device (decoder). The decoding device includes a memory and at least one processor connected to the memory, and the at least one processor may be configured to decode a base mesh in a bitstream, decode a displacement in a bitstream, and decode an attribute in a bitstream.

[0386] As described above, while the existing displacement encoding / decoding process allows for various chroma formats of geometry video, the nominal format conversion process of geometry video did not take into account various video chroma formats; however, the present disclosure performs nominal format conversion by taking into account the chroma format of geometry video. In addition, when the V3C unit type is PVD, it supports packing displacement data in the geometry area and packing attribute data in the attribute area, but this was not specified in the restoration process of conventional V-DMC; however, the present disclosure performs the restoration process by taking into account PVD.

[0387] Therefore, from the perspective of a decoder, the present disclosure has the following effects. That is, through the present embodiment, restoration can be performed by considering a packing method according to the chroma format of the displacement video during the nominal format conversion process of the geometry video. In addition, through the present embodiment, by performing the restoration process by considering PVD, the restoration process can be performed even when the V3C unit type is PVD.

[0388] Each of the aforementioned parts, modules, or units may be software, processors, or hardware parts that execute successive processes stored in memory (or storage units). Each step described in the aforementioned embodiments may be performed by processors, software, or hardware parts. Each module / block / unit described in the aforementioned embodiments may operate as a processor, software, or hardware. Additionally, the methods presented in the embodiments may be executed as code. This code may be written to a storage medium that is readable by a processor and thus read by a processor provided by the apparatus.

[0389] Furthermore, throughout the specification, when a part is described as “comprising” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Also, terms such as “…part” as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.

[0390] Although the drawings have been described separately for convenience of explanation in this specification, it is also possible to design a new embodiment by combining the embodiments described in each drawing. Furthermore, according to the needs of a person skilled in the art, designing a computer-readable recording medium on which a program for executing the previously described embodiments is recorded falls within the scope of the embodiments.

[0391] The apparatus and method according to the embodiments are not limited to the configurations and methods of the embodiments described above; rather, the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.

[0392] Although preferred embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by those skilled in the art without departing from the essence of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the embodiments.

[0393] Various components of the device according to the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented as a single chip, for example, a single hardware circuit. Components according to the embodiments may each be implemented as separate chips. At least one of the components of the device according to the embodiments may be composed of one or more processors capable of executing one or more programs, and one or more programs may include instructions for performing or executing any one or more of the operations / methods according to the embodiments. Executable instructions for performing the methods / operations of the device according to the embodiments may be stored in non-transient CRMs or other computer program products configured to be executed by one or more processors, or may be stored in transient CRMs or other computer program products configured to be executed by one or more processors. Additionally, memory according to the embodiments may be used as a concept that includes not only volatile memory (e.g., RAM, etc.) but also non-volatile memory, flash memory, PROM, etc. In addition, it may also include implementation in the form of a carrier wave, such as transmission over the Internet. Furthermore, processor-readable recording media may be distributed across networked computer systems, allowing processor-readable code to be stored and executed in a distributed manner.

[0394] In this document, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." Additionally, in this document, "or" is interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." In other words, "or" in this document may mean "additionally or alternatively."

[0395] Various elements of the embodiments may be performed by hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be performed on a single chip, such as a hardware circuit. Depending on the embodiments, the embodiments may optionally be performed on individual chips. Depending on the embodiments, at least one of the elements of the embodiments may be performed within one or more processors that include instructions for performing operations according to the embodiments.

[0396] Additionally, the operation according to the embodiments described herein may be performed by a transceiver device comprising one or more memories and / or one or more processors according to the embodiments. One or more memories may store programs for processing / controlling the operation according to the embodiments, and one or more processors may control the various operations described in this document. One or more processors may be referred to as controllers, etc. The operations of the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in a processor or in memory.

[0397] Terms such as "first," "second," etc., may be used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not imply the same user input signals unless clearly indicated in the context.

[0398] The terms used to describe the embodiments are intended for the purpose of describing specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless explicitly indicated in the context. Expressions of and / or are used to mean including all possible combinations between the terms. Expressions of “include” describe the presence of features, numbers, steps, elements, and / or components and do not imply the exclusion of additional features, numbers, steps, elements, and / or components. Conditional expressions such as “if” or “when” used to describe the embodiments are not limited to being optional. It is intended to be interpreted as when a specific condition is satisfied, to perform a related action in response to a specific condition, or to interpret the related definition.

[0399] As described above, the relevant details have been explained in the best mode for carrying out the embodiments.

[0400] As described above, the embodiments may be applied wholly or partially to mesh data transmission and reception devices and systems. Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments. The embodiments may include changes / modifications, and such changes / modifications do not deviate from the scope of the claims and their equivalents.

Claims

1. A step of receiving a bitstream containing mesh data; and The step of decoding the above mesh data; is included, The step of decoding the above mesh data is, A method comprising the step of decoding a base mesh from a base mesh bitstream included in the above mesh data, Decoding method.

2. In claim 1, the step of decoding the mesh data If the above mesh data includes a displacement data bitstream, the step of decoding displacement data from the displacement bitstream, and A decoding method that further includes the step of decoding packed data from the packed bitstream if the mesh data includes a packed data bitstream.

3. In Paragraph 2, A decoding method further comprising the step of separating displacement data and attribute data from the above-decoded packed data.

4. In Paragraph 3, A step of performing a nominal format conversion on one or more of the data included in the decoded base mesh, the decoded displacement data, or the decoded packed data, and A decoding method further comprising the step of restoring a mesh based on data converted into the above-mentioned nominal format.

5. In claim 4, the nominal format conversion step A decoding method that derives the number of components of the displacement data to 3 if the chroma format of the displacement data is 4:4:4, and derives the number of components of the displacement data to 1 if the chroma format of the displacement data is not 4:4:

4.

6. In claim 5, the nominal format conversion step A decoding method that performs at least one of map extraction, bit depth conversion, resolution conversion, and chroma format conversion of displacement data packed into one or more video planes based on the number of components of the above-described displacement data.

7. In Paragraph 4, A decoding method comprising a first flag information capable of identifying whether the mesh data contains a displacement data bitstream and a second flag information capable of identifying whether the mesh data contains a packed data bitstream.

8. In claim 7, the restoration step A decoding method that determines whether displacement data exists in input data based on the first flag information and the second flag information, and performs restoration of the displacement data if the displacement data exists.

9. Memory; and At least one processor connected to the memory; comprising, The above at least one processor is: Receive a bitstream containing mesh data; and Configured to decode the above mesh data; and The above-mentioned at least one processor is, A base mesh decoder comprising a base mesh decoder that decodes a base mesh from a base mesh bitstream included in the above mesh data, Decoding device.

10. In claim 9, the at least one processor If the above mesh data includes a displacement data bitstream, a displacement decoder that decodes displacement data from the displacement bitstream, and A decoding device further comprising a packed decoder that decodes packed data from the packed bitstream if the mesh data includes a packed data bitstream.

11. In claim 10, the at least one processor A decoding device for separating displacement data and attribute data from the above-decoded packed data.

12. In claim 11, the at least one processor A nominal format conversion unit that performs nominal format conversion on one or more of the data among the decoded base mesh, the decoded displacement data, or the displacement data included in the decoded packed data, and A decoding device further comprising a restoration unit that restores a mesh based on data converted into the above-mentioned nominal format.

13. In claim 12, the nominal format converter A decoding device that derives the number of components of the displacement data to 3 if the chroma format of the displacement data is 4:4:4, and derives the number of components of the displacement data to 1 if the chroma format of the displacement data is not 4:4:

4.

14. In claim 13, the nominal format converter A decoding device that performs at least one of map extraction, bit depth conversion, resolution conversion, and chroma format conversion of displacement data packed into one or more video planes based on the number of components of the above-described displacement data.

15. In Paragraph 12, The bitstream further includes first flag information capable of identifying whether the mesh data includes a displacement data bitstream and second flag information capable of identifying whether the mesh data includes a packed data bitstream. The above restoration unit determines whether displacement data exists in the input data based on the above first flag information and the above second flag information, and if the displacement data exists, the decoding device performs restoration of the displacement data.

Citation Information

Patent Citations

  • Signaling volumetric visual video-based coding content in immersive scene descriptions

    WO2023137229A1

  • Signaling displacement data for video-based mesh coding

    WO2024091593A1

  • Mesh compression texture coordinate signaling and decoding

    WO2024091594A1

  • Visual volumetric video-based coding method, encoder and decoder

    WO2024163690A2

  • Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method

    WO2024186127A1