Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

WO2024205193A3PCT designated stage expired Publication Date: 2025-06-19LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/003771
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-31
Filing Date
2024-03-26
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The challenge lies in efficiently transmitting and receiving point cloud data due to its large size and complexity, which results in high processing requirements and latency issues, particularly in applications like VR, AR, and autonomous driving.

Method used

A point cloud data transmission method and device that employs dynamic mesh compression techniques, such as V-MESH compression, which encodes and decodes mesh data into bitstreams for efficient transmission and reception, utilizing existing 2D video codecs like HEVC and VVC, and image packing formats like YUV 4:2:0 for optimal bit reduction and quality.

Benefits of technology

This approach enables high-quality point cloud services by reducing latency and encoding/decoding complexity, supporting various video codec methods and providing efficient point cloud content for applications like autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024003771_19062025_PF_FP_ABST
    Figure KR2024003771_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A mesh data transmission method according to embodiments may comprise the steps of: encoding mesh data; and transmitting a bitstream including the mesh data. A mesh data reception method according to embodiments may comprise the steps of: receiving a bitstream including mesh data; and decoding the mesh data.
Need to check novelty before this filing date? Find Prior Art

Description

Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method

[0001] The embodiments provide a method for providing Point Cloud content to provide users with various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and autonomous driving services.

[0002] A point cloud is a collection of points in 3D space. The sheer number of points in 3D space makes it difficult to generate point cloud data.

[0003] There is a problem that a lot of processing power is required to transmit and receive point cloud data.

[0004] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device, and reception method for efficiently transmitting and receiving point clouds in order to solve the problems described above.

[0005] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device, and reception method for resolving latency and encoding / decoding complexity.

[0006] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.

[0007] A mesh data transmission method according to embodiments may include a step of encoding mesh data; and a step of transmitting a bitstream including mesh data. A mesh data reception method according to embodiments may include a step of receiving a bitstream including mesh data; and a step of decoding mesh data.

[0008] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide a high-quality point cloud service.

[0009] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can achieve various video codec methods.

[0010] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide general-purpose point cloud content such as autonomous driving services.

[0011] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.

[0012] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.

[0013] Figure 2 illustrates a V-MESH compression method according to embodiments.

[0014] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.

[0015] Figure 4 illustrates a mid-edge subdivision method according to embodiments.

[0016] Figure 5 shows a displacement generation process according to embodiments.

[0017] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.

[0018] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.

[0019] Figure 8 shows a lifting conversion process for displacement according to embodiments.

[0020] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.

[0021] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.

[0022] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.

[0023] Fig. 12 shows a V-MES and Fig. 13 shows a point cloud data transmission device according to embodiments.

[0024] Fig. 13 shows a point cloud data transmission device according to embodiments.

[0025] Fig. 14 shows a point cloud data receiving device according to embodiments.

[0026] Fig. 15 illustrates a dynamic mesh encoder according to embodiments.

[0027] Fig. 16 shows a displacement vector encoding unit according to embodiments.

[0028] Figure 17 shows the structure of displacement vector coefficients according to embodiments.

[0029] Figure 18 shows displacement vector coefficient 2D image packing according to embodiments.

[0030] Figure 19 shows the 2D image packing of displacement vector coefficients by LoD (Level of Detail) according to embodiments.

[0031] Figure 20 illustrates a 2D Moulton code-based packing according to embodiments.

[0032] Figure 21 shows a displacement vector coefficient packing method according to embodiments.

[0033] Figure 22 shows a displacement vector coefficient packing method according to embodiments.

[0034] Figure 23 shows a displacement vector coefficient packing method according to embodiments.

[0035] Figure 24 shows the overall level packing frame configuration according to embodiments.

[0036] Figure 25 shows a level-by-level packing frame configuration according to embodiments.

[0037] Fig. 26 shows a dynamic mesh decoder according to embodiments.

[0038] Figure 27 shows a displacement vector decoding unit according to embodiments.

[0039] Figure 28 shows a displacement vector coefficient inverse packing method according to embodiments.

[0040] Figure 29 shows a displacement vector coefficient inverse packing method according to embodiments.

[0041] Figure 30 shows a displacement vector coefficient inverse packing method according to embodiments.

[0042] Figure 31 shows a displacement vector coefficient inverse packing method according to embodiments.

[0043] Figure 32 shows a displacement vector coefficient inverse packing method according to embodiments.

[0044] Figure 33 shows an atlas sequence parameter set in a bitstream according to embodiments.

[0045] Figure 34 shows an atlas sequence parameter set in a bitstream according to embodiments.

[0046] Figure 35 shows displacement vector coefficient packing and inverse packing according to embodiments.

[0047] Figure 36 shows a sampling method and sampling values ​​of displacement vector coefficient components according to embodiments.

[0048] Figure 37 shows a mesh data transmission method according to embodiments.

[0049] Figure 38 shows a mesh data receiving method according to embodiments.

[0050] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.

[0051] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.

[0052] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.

[0053] The system of FIG. 1 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a reception unit (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or a dynamic mesh video encoder (hereinafter, referred to as an encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to a receiving device (110) or a dynamic mesh video decoder (hereinafter, decoder) (113).

[0054] The system of Fig. 1 can perform video-based dynamic mesh compression and decompression.

[0055] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.

[0056] First, the method of compressing dynamic mesh data starts with the V-PCC (Video-based point cloud compression) standard technology. Point cloud data is data that contains color information at the vertex coordinates (X, Y, Z). Mesh data refers to data in which connectivity information between vertices is added to this vertex information. When creating content, it can be created in the form of mesh data from the beginning. By adding connectivity information to point cloud data, it can be converted into mesh data and used.

[0057] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.

[0058] Mesh coding standards for Category 1 data are currently in development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.

[0059] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.

[0060] A mesh compression system may include a transmitting device and a receiving device. The transmitting device can encode mesh video to output a bitstream, which can be delivered to the receiving device via digital storage media or a network in the form of a file or streaming segment. The digital storage media may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, or SSD.

[0061] The transmitting device may roughly include a mesh video acquisition unit, a mesh video encoder, and a transmitting unit. The receiving device may roughly include a receiving unit, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / video / picture / frame encoding device, and the decoder may be referred to as a mesh video / video / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device and the receiving device may further include separate internal or external modules / units / components for a feedback process.

[0062] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by its vertices in 3D space and connection information that describes how those vertices are connected. It can also contain vertex properties such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included as a mesh property. The mapping is typically described as a set of parametric coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement.

[0063] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc. into a mesh data type with the properties described above through a series of processes and generating a video composed of such mesh data. The mesh video may have properties of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video with properties and connection information that change over time can be expressed as a dynamic mesh video.

[0064] A mesh video encoder can encode an input mesh video into one or more video streams. A single video can include multiple frames, and a single frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and the mesh video can be used interchangeably with the mesh image / frame / picture. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. A mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0065] The encapsulation processing unit (file / segment encapsulation module) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from the metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder, or may be configured as a separate component / module. The encapsulation processing unit may encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The encapsulation processing unit may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. According to an embodiment, the encapsulation processing unit may encapsulate mesh video-related metadata itself in a file.

[0066] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit, or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process the same for transmission.

[0067] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and can include an element for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.

[0068] The receiver can receive mesh video data transmitted by the mesh video transmission device. Depending on the transmission channel, the receiver can receive mesh video data via a broadcast network, via broadband, or via digital storage media.

[0069] The receiving processing unit can perform processing on the received mesh video data according to a transmission protocol. The receiving processing unit can be included in the receiving unit, or can be configured as a separate component / module. In response to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmitting processing unit. The receiving processing unit can transfer the acquired mesh video data to the decapsulation processing unit, and transfer the acquired mesh video-related metadata to a metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.

[0070] A decapsulation processing unit (file / segment decapsulation module) can decapsulate mesh video data in file format received from a receiving processing unit. The decapsulation processing unit can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder, or may be configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation processing unit may be in the form of a box or track within a file format. If necessary, the decapsulation processing unit may receive metadata required for decapsulation from the metadata processing unit. Mesh video related metadata can be passed to a Mesh video decoder for use in the Mesh video decoding process, or passed to a renderer for use in the Mesh video rendering process.

[0071] A mesh video decoder can receive a bitstream and perform operations corresponding to those of a mesh video encoder to decode video / images. The decoded mesh video can be displayed via a display unit. Users can view all or part of the rendered result via a VR / AR display or a general display.

[0072] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.

[0073] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.

[0074] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of ​​the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0075] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.

[0076] This document relates to dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and properties that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-view video, and AR / VR.

[0077] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.

[0078] In this document, picture / frame can generally mean a unit representing one video of a specific time period.

[0079] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.

[0080] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0081] The encoding process of Fig. 1 is as follows.

[0082] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. The V-Mesh compression process receives the following data as input and performs compression.

[0083] Input mesh: Contains the 3D coordinates (geometry) of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.

[0084] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the properties (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video (or attribute for short). The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).

[0085] Material Library File: Contains information about the material properties used in a mesh, particularly information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.

[0086] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0087] Base mesh: The input mesh is simplified (decimated) through a preprocessing process to express the objects of the input mesh using the minimum number of vertices determined by the user's standards.

[0088] Displacement: This is displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.

[0089] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.

[0090] Referring to FIGS. 2 to 7, a method for encoding mesh position information (vertex) is described, and referring to FIGS. 6-10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.

[0091] Figure 2 illustrates a V-MESH compression method according to embodiments.

[0092] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The encoder of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. The transmitting device of Fig. 1 may be broadly referred to as an encoder, and the dynamic mesh video encoder of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing (200) and an encoding (201) process as in Fig. 2. The pre-processor of Fig. 2 may be located in front of the encoder of Fig. 2. The pre-processor and the encoder of Fig. 2 may be referred to as a single encoder.

[0093] The preprocessor can receive a static dynamic mesh and / or an attribute map. The preprocessor can generate a base mesh and / or displacement through preprocessing. The preprocessor can receive feedback information from the encoder and generate the base mesh and / or displacement based on the feedback information.

[0094] The encoder can receive a base mesh, displacement mesh, static dynamic mesh, and / or attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.

[0095] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.

[0096] Figure 3 shows the configuration and operation of the pre-processor of Figure 2.

[0097] Fig. 3 shows a process of performing preprocessing on an input mesh. The preprocessing process (200) can be broadly divided into four steps: 1) Group of Frame (GoF) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). The preprocessor (200) can receive an input mesh, generate a displacement and / or base mesh, and transmit the generated displacement and / or base mesh to the encoder (201). The preprocessor (200) can transmit GoF information related to GoF generation to the encoder (201).

[0098] Below, each step of Figure 3 is described.

[0099] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, number of texture coordinates, vertex connection information, and texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values ​​are different between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame performs intra-frame encoding.

[0100] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.

[0101] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.

[0102] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.

[0103] Fitting subdivision surface: This is the process of performing subdivision on a simplified mesh. The subdivision method can be a user-defined method, such as the mid-edge method. The fitting process ensures that the input mesh and the subdivision mesh are similar to each other.

[0104] Figure 4 illustrates a mid-edge subdivision method according to embodiments.

[0105] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex midway between the edges between the vertices.

[0106] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decrypted base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents a position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values ​​can be converted to (normal, tangential, bi-tangential) coordinate values ​​of the local coordinate system.

[0107] Figure 5 shows a displacement generation process according to embodiments.

[0108] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.

[0109] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a reconstructed base mesh and generate a subdivided reconstructed base mesh. The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and transform a coordinate system of the mesh into a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.

[0110] The method and device for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data (which may be referred to as a point cloud for short) according to the embodiments can refer to data including vertex coordinates and color information. The term "point cloud" includes mesh data, and in this document, point cloud and mesh data can be used interchangeably.

[0111] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).

[0112] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.

[0113] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.

[0114] The encoding process of Fig. 6 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an intra-frame method. The encoder of Fig. 6 may include a preprocessor (200) and / or an encoder (201).

[0115] The preprocessor can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivided mesh. The quantizer can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder can encode the static mesh. The static mesh encoder can generate a bitstream including the encoded base mesh. The static mesh decoder can decode the encoded static mesh. The inverse quantizer can inversely quantize the quantized static mesh. The displacement calculation unit can receive the reconstructed static mesh and generate displacement, which is a position difference, based on the fitted subdivided mesh. The forward linear lifting unit can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. The image packing unit can pack an image based on the quantized lifting coefficients. A video encoder can encode a packed image. A video decoder decodes the encoded video. An image unpacker can unpack a packed image. A dequantizer can inversely quantize an image. An inverse linear lifting unit applies inverse lifting to the image to generate a reconstructed displacement. A mesh restoration unit restores a warped mesh using the reconstructed displacement and the reconstructed base mesh. An attribute transfer unit receives an input mesh and / or an input attribute map, and generates an attribute map based on the reconstructed warped mesh. A push-pull padding unit can pad data in the attribute map based on a push-pull method. A color space transformation unit can transform the space of a color component, which is an attribute. A video encoder can encode an attribute. A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.

[0116] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.

[0117] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a preprocessor (200) and / or an encoder (201).

[0118] Among the encoding operations of Fig. 7, the corresponding configuration of the encoding operation of Fig. 6 refers to the description of Fig. 7. For the inter-frame-based encoding of Fig. 7, the motion encoder can encode motion based on the restored quantized reference base mesh. The base mesh restoration unit can restore the base mesh based on the restored quantized reference base mesh.

[0119] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.

[0120] The encoding method according to the embodiments includes base mesh encoding (intra encoding). When performing intra frame encoding on the current input mesh frame, the base mesh generated in the preprocessing process can be encoded using a static mesh compression technique after going through a quantization process. In the V-Mesh compression method, for example, Draco technology is applied, and the vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh are compressed.

[0121] The encoding method according to the embodiments may include motion field encoding (inter encoding). Inter frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field, may be calculated and this information may be encoded. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may be encoded as a value. Alternatively, the predicted motion field may be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and this predicted motion field The residual motion field, which is the difference between the value and the motion field value of the current vertex, can be encoded. This value can be encoded using entropy coding. The process of encoding the displacement and attribute map, excluding the motion field encoding process of inter-frame encoding, is the same as the structure of the intra-frame encoding method except for the base mesh encoding.

[0122] Figure 8 shows a lifting conversion process for displacement according to embodiments.

[0123] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.

[0124] Figures 8-9 show the process of converting the displacement of the encoding process of Figures 6-7 and the process of packing the conversion coefficients, respectively.

[0125] The encoding method according to the embodiments includes displacement encoding.

[0126] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information.

[0127] Figure 8 shows the process of transforming displacement information using lifting transform in V-Mesh. The transform coefficients generated through the transform process are quantized and then packed into a 2D image as shown in Figure 9. The transform coefficients are organized into one block for each 256 (=16×16) units, and each block can be packed in z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. The transform coefficients can be packed by sorting them with Morton code within a block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec.

[0128] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. The first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices (LoD0) of the base mesh. The first subdivision mesh may be generated by dividing the second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 including the vertices additionally generated from LoD0, and the vertices additionally divided from LoD1. LoD is a level indicating the degree of detail (Level of Detail), and as the level index increases, the distance between vertices becomes closer and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When a vertex is further divided through subdivision, considering the previous vertices v1, v2 and the subdivided vertex v, the mesh can be encoded based on a prediction and / or update method. Instead of still encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated, and the mesh can be encoded using the residual value to reduce the size of the bitstream. The prediction process means the operation of predicting the current vertex v using the previous vertices v1, v2. Since adjacent subdivision meshes have similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted as the residual for the previous vertex position information, and the previous vertex position information is updated using the residual.

[0129] Referring to Figure 9, the vertices have coefficients generated through the lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.

[0130] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.

[0131] Figure 10 shows the detailed operation of attribute transfer of encoding such as Figures 6-7.

[0132] Encoding according to embodiments includes attribute map encoding.

[0133] Information about the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has vertex position information, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is created through the attribute transfer process.

[0134] Attribute transfer first checks whether each point P(u, v) in the 2D texture domain belongs to a texture triangle of the reconstructed deformed mesh, and if it is in the texture triangle T, calculates the barycentric coordinate (α, β γ) of P(u, v) according to the triangle T. Then, using the 3D vertex position and (α, β γ) of triangle T, calculate the 3D coordinate M(x, y, z) of P(u, v). Find the vertex coordinate M'(x', y', z') that corresponds to the position most similar to the calculated M(x, y, z) in the input mesh domain and the triangle T' that contains this point. And in this triangle T', the center of mass coordinates (α', β', γ') of M'(x', y', z') are calculated. Using the texture coordinates corresponding to the three vertices of triangle T' and (α', β', γ'), the texture coordinates (u', v') are calculated, and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.

[0135] The new attribute map generated through attribute transfer is grouped into GoF units to form an attribute map video, which is then compressed using a video codec.

[0136] Referring to Figure 10, the reference relationship between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.

[0137] The decoding process of Fig. 1 can perform the reverse process of the corresponding encoding process of Fig. 1. The specific decoding process is as follows.

[0138] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.

[0139] Fig. 11 shows the configuration and operation of a decoder of a receiving device such as Fig. 1.

[0140] Figure 11 illustrates the intra decoding process of V-Mesh technology. First, the input bitstream can be separated into a mesh sub-stream, a displacement sub-stream, an attribute map sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.

[0141] The mesh sub-stream is decoded by the decoder of the static mesh codec used in encoding, such as Google Draco, and as a result, the connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be restored. The displacement sub-stream is decoded into a displacement video by the decoder of the video compression codec used in encoding, and goes through the image unpacking, inverse quantization, and inverse transform processes to restore displacement information for each vertex. Inverse quantization is applied to the restored base mesh, and this result is combined with the restored displacement information to generate the final decoded mesh.

[0142] The attribute map sub-stream is decoded through the decoder of the video compression codec used in encoding, and then restored to the final attribute map through processes such as color format conversion.

[0143] The restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.

[0144] Referring to FIG. 11, the bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term "substream" is interpreted as referring to a portion of the bitstream included in the bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).

[0145] The decoder performs the following decoding operations within the frame. The static mesh decoder decodes the mesh to generate a reconstructed quantized base mesh, and the inverse quantizer applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder decodes the displacement, the unpacker unpacks the decoded video image, and the inverse quantizer inversely quantizes the quantized image. The linear lifting inverse transform applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit generates a warped mesh based on the base mesh and the displacement. The video decoder decodes the attribute map, and the color transformation unit transforms the color format and / or space to generate the decoded attribute map.

[0146] Figure 12 shows the inter-frame decoding process of the V-MESH compression method.

[0147] Fig. 12 shows the configuration and operation of the decoder of the receiving device of Fig. 1.

[0148] Figure 12 illustrates the inter-decoding process of V-Mesh technology. First, the input bitstream can be separated into a motion sub-stream, a displacement sub-stream, an attribute sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.

[0149] The motion sub-stream is decoded through entropy decoding and inverse prediction processes, and the reconstructed motion information is combined with the already reconstructed reference base mesh to generate a reconstructed quantized base mesh for the current frame. The result of applying inverse quantization to this is combined with the displacement information reconstructed in the same way as the intra decoding described above to generate the final decoded mesh. The attribute map sub-stream is decoded in the same way as the intra decoding. The reconstructed decoded mesh and the decoded attribute map can be utilized by the receiver as the final mesh data that can be utilized by the user.

[0150] Referring to Fig. 12, the bitstream includes motion, displacement, and attribute maps. Since inter-frame decoding is performed, a process of decoding inter-frame motion information is further included. The motion is decoded, and a restored quantized base mesh for the motion is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operations in Fig. 12, which are identical to those in Fig. 11, refer to the description in Fig. 11.

[0151] Fig. 13 shows a point cloud data transmission device according to embodiments.

[0152] Fig. 13 corresponds to the transmitting device (100), the dynamic mesh video encoder (102), the encoder (preprocessor and encoder) of Fig. 2, and / or the transmitting encoding device corresponding thereto of Fig. 13. Each component of Fig. 13 corresponds to hardware, software, a processor, and / or a combination thereof.

[0153] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13.

[0154] The mesh preprocessor receives the original mesh as input and generates a simplified mesh (decimated mesh). Simplification can be performed based on the target number of vertices or the target number of polygons that constitute the mesh. Parameterization can be performed on the simplified mesh to generate texture coordinates and texture connection information per vertex. Additionally, quantization of floating-point mesh information into fixed-point information can be performed. This result can be encoded as a base mesh through a static mesh encoding unit. The mesh preprocessor can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The subdivided mesh can be fitted by adjusting the vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.

[0155] When performing intra-encoding on the corresponding mesh frame, the base mesh generated through the mesh preprocessing unit can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to the multiplexing unit.

[0156] When performing inter encoding for the corresponding mesh frame, a motion vector encoding unit is performed, which can calculate a motion vector between the base mesh and the reference reconstruction base mesh as input and encode the value. The motion vector encoding unit can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexing unit.

[0157] The encoded base mesh and motion vectors can be used to generate a restored base mesh through the base mesh restoration unit.

[0158] The displacement vector calculator can perform mesh refinement on the restored base mesh. The displacement vector can be calculated as the difference in vertex positions between the refined restored base mesh and the fitted subdivision mesh generated in the preprocessing unit. As a result, the number of displacement vectors can be calculated equal to the number of vertices in the refined mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.

[0159] A displacement vector video generator can transform displacement vectors for effective encoding. The transformation can be performed by a lifting transformation, a wavelet transformation, etc., depending on the embodiment. In addition, quantization can be performed on the transformed displacement vector values, i.e., the transform coefficients. Different quantization parameters can be applied to each axis of the transform coefficients, and the quantization parameters can be derived according to the promise of the encoder / decoder. The transformed and quantized displacement vector information can be packed into a 2D image. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh.

[0160] A displacement vector video encoder can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer.

[0161] The displacement vector restored through the displacement vector restorer and the base mesh restored through the base mesh restorer and refined are restored through the mesh restorer, and the restored mesh has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.

[0162] The texture map of the original mesh can be regenerated as a texture map for the restored mesh through the texture map video generation unit. The color information per vertex of the texture map of the original mesh can be assigned to the texture coordinates of the restored mesh. The regenerated texture maps for each frame can be bundled into GoF units to generate a texture map video.

[0163] The generated texture map video can be encoded using a video compression codec through a texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to a multiplexing unit.

[0164] The generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be multiplexed into a single bitstream and transmitted to a receiver via a transmitter. Alternatively, the generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via a transmitter.

[0165] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (displacement), and a texture map (attribute map). A transmitting device according to inter-encoding can generate a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map). A texture map obtained from a data input unit is generated and encoded based on a restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the segmented mesh. The base mesh is generated by preprocessing, simplifying, and encoding the original mesh. Motion is generated as a motion vector for the mesh of the current frame based on the reference base mesh of the previous frame.

[0166] Fig. 14 shows a point cloud data receiving device according to embodiments.

[0167] Fig. 14 corresponds to the receiving device (110), the dynamic mesh video decoder (113), the decoder of Figs. 11-12, and / or the receiving decoding device corresponding thereto of Fig. 1. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.

[0168] The bitstream of the received Mesh is demultiplexed into a compressed motion vector bitstream or base mesh bitstream, displacement vector bitstream, and texture map bitstream after file / segment decapsulation.

[0169] If the current mesh has inter-frame encoding applied according to the frame header information, the motion vector decoding unit can perform decoding on the motion vector bitstream. The final motion vector can be reconstructed by adding the previously decoded motion vector to the residual motion vector decoded from the bitstream using the previously decoded motion vector as a predictor.

[0170] If the current mesh has been encoded within the screen according to the frame header information, the base mesh bitstream can restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh through the static mesh encoding unit.

[0171] In the base mesh restoration unit, if the current mesh has inter-frame encoding applied, the current base mesh can be restored by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. If the current mesh has intra-frame encoding applied, the static mesh decoding unit can perform inverse quantization on the decoded mesh to generate a restored base mesh.

[0172] The displacement vector bitstream can be decoded as a video bitstream using a video codec in a displacement vector video decoding unit.

[0173] In the displacement vector restoration unit, displacement vector transformation coefficients are extracted from the decoded displacement vector video, and the displacement vector is restored through the inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, a process of inverse transformation to the Cartesian coordinate system can be performed.

[0174] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Subdivision can generate vertex connection information, texture coordinates, and texture coordinate connection information, including the added vertices. The subdivided restored base mesh can be combined with the restored displacement vector to generate the final restored mesh.

[0175] The texture map bitstream can be decoded as a video bitstream using a video codec in a texture map video decoding unit. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of the corresponding vertex.

[0176] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.

[0177] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data. A receiving device according to inter-decoding can receive a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data.

[0178] A point cloud data transmission device and method according to embodiments can encode mesh data and transmit a bitstream including the encoded mesh data. A point cloud data reception device and method according to embodiments can receive a bitstream including mesh data and decode the mesh data. The point cloud data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The point cloud data transmission and reception method / device according to embodiments may also be referred to as mesh data transmission and reception method / device according to embodiments.

[0179] The point cloud data transmission method / device according to the embodiments is interpreted as a term including a transmission device (100) of FIG. 1, a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), a transmitter (104), a pre-processor (200) of FIG. 2, an encoder (201), an encoder of FIG. 6-7, a transmission device of FIG. 13, an encoder of FIG. 15-16, a transmission method of FIG. 37, etc.

[0180] The point cloud data receiving method / device according to the embodiments is interpreted as a term including a receiving device (110) of Fig. 1, a receiver (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), a renderer (114), a decoder of Figs. 11-12, a receiving device of Fig. 14, a decoder of Figs. 26-27, a receiving method of Fig. 38, etc.

[0181] The method / device according to the embodiments may include and perform a method for displacement video packing with YUV 420 format.

[0182] The embodiments relate to Video-based Dynamic Mesh Compression (V-DMC), a method for compressing 3D dynamic mesh data using an existing 2D video codec. In a V-DMC decoder, a displacement vector between a reconstructed base mesh and a fitted mesh in a preprocessing step is calculated, transformed, and quantized to be encoded / decoded as a displacement vector bitstream. The embodiments propose a method using the YUV 4:2:0 format in the image packing step for encoding / decoding the transformed displacement vector using a video codec in the above process. In addition, a sampling method and a signaling method are proposed so that points at optimized positions for each component of the transformed displacement vector can be sampled and transmitted. Since the YUV 4:2:0 format is used in most video codec profiles, not only is compatibility good, but also efficient bit savings and better image quality mesh data can be obtained in terms of performance.

[0183] The examples relate to Video-based Dynamic Mesh Compression (V-DMC), a method for compressing 3D dynamic mesh data using an existing 2D video codec. The examples propose a method for encoding / decoding displacement vector units in displacement vector transformation and quantization steps, as well as syntax and semantics information related thereto. In addition, the operations of a transmitter and receiver applying the same are described.

[0184] Recently, V-DMC technology has been actively standardized since the CfP Response in April 2022. In the technologies currently applied to V-DMC, displacement vector information is expressed as 3D or 1D vectors, compressed, and transmitted. When performing displacement vector encoding / decoding using a 2D video codec, images are packed in the YUV 4:4:4 or YUV 4:0:0 formats. Using the YUV 4:4:4 format during the image packing process has the disadvantage of large file sizes due to the lack of subsampling. Conversely, using the YUV 4:0:0 format significantly reduces file size, but degrades image quality due to the discarding of U and V (tangential and bi-tangential) components. Therefore, we propose an image packing method using the YUV 4:2:0 format, which is efficient in terms of both file size and image quality compared to the two formats currently used in V-DMC technology and is the most widely used format in current video codecs. Furthermore, when applying the YUV 4:2:0 format, the sampling method of U and V components (tangential and bi-tangential) may currently be fixed to a single method. This may be inefficient, as better sampling methods may exist depending on the content or sampling area. Therefore, we propose a method to perform sampling by selecting the optimal value during the process of mapping each tangent and bi-tangent component to the U and V channels.

[0185] Accordingly, the embodiments include, as a solution for solving these technical problems, a method of using the YUV 4:2:0 format in the image packing step for video encoding / decoding of a displacement vector, and / or a method of sampling and signaling data for each component of a displacement vector.

[0186] Meanwhile, in this document, the term V-DMC can also be referred to as the term V-Mesh and is an expression used with the same meaning.

[0187] Fig. 15 illustrates a dynamic mesh encoder according to embodiments.

[0188] Fig. 15 shows the configuration of a dynamic mesh encoder corresponding to the Fig. 1 transmitting device (100), the dynamic mesh video acquisition unit (101), the dynamic mesh video encoder (102), the file / segment encapsulator (103), the transmitter (104), the Fig. 2 pre-processor (200), the encoder (201), the Figs. 6-7 encoder, the Fig. 13 transmitting device, the Figs. 15-16 encoder, and the Fig. 37 transmitting method. Each component of Fig. 15 may correspond to hardware, software, a processor, and / or a combination thereof.

[0189] Current V-DMC technology simplifies and parameterizes the original mesh to generate a base mesh. This base mesh is then quantized and transmitted as a base mesh bitstream via a static mesh encoder. Furthermore, the mesh data obtained after refining and fitting the simplified mesh from the original mesh is compared with the mesh data obtained by reconstructing the previously encoded mesh, thereby calculating the displacement vector, which is the difference between each vertex.

[0190] In order to efficiently encode the calculated displacement vector, the displacement vector coordinate system is transformed into a local coordinate system, the transformed vector is transformed and quantized into displacement vector coefficients, and then encoded to transmit the displacement vector bitstream. The embodiments include a method of using the YUV 4:2:0 format in the displacement vector image packing unit as shown in Fig. 15, a method of sampling each component of the displacement vector, and a signaling method. The principles performed at each step are described in detail below.

[0191] The displacement vector calculation unit calculates the vector between the fitted subdivided mesh and the restored base mesh, which is the mesh on which subdivision has been performed. At this time, as many displacement vectors as the number of vertices of the subdivided mesh can be calculated.

[0192] The dynamic mesh video acquisition unit (101) can acquire the original mesh.

[0193] The mesh simplification unit can simplify the vertices and connections between vertices in the original mesh. The simplified original mesh can be called the base mesh.

[0194] The mesh parameterization unit can generate texture coordinates and texture connection information for each vertex within the mesh.

[0195] The mesh quantization unit quantizes the base mesh based on the quantization parameters.

[0196] A dynamic mesh encoder can encode a base mesh using either an intra-frame method or an intra-frame method. The intra-frame method can generate a prediction value for the current mesh by referencing the mesh included in the frame, generate a residual value, and encode only the residual value to generate a base mesh bitstream. At this time, the mesh can be a static mesh that does not change over time. The inter-frame method can generate a prediction value for a similar base mesh within the reference frame by referencing the reference frame for the current frame, generate a residual value, and encode only the residual value to generate a base mesh bitstream. At this time, the mesh is a dynamic mesh that changes over time, and the motion vector can be predicted and encoded. In order to encode detailed vertex-related data like the original mesh separately from the base mesh, the following procedure is performed.

[0197] The mesh refinement unit further refines the simple base mesh into detailed vertices.

[0198] The mesh fitting unit can reconstruct the mesh based on mesh connection information.

[0199] The displacement vector coordinate system transformation part can use the previously calculated vertex displacement vector as is in the canonical coordinate system (x, y, z) or can be transformed into a local coordinate system (normal, tangential, bi-tangential).

[0200] The displacement vector coordinate system transformation unit can determine the transformation coordinate system according to the encoder / decoder agreement and signal the displacement vector coordinate system transformation flag (asps_vmc_ext_displacement_coordinate_system). If the syntax value is 0, the canonical coordinate system is used as is, and if it is 1, the transformation can be performed to the local coordinate system.

[0201] Fig. 16 shows a displacement vector encoding unit according to embodiments.

[0202] Figure 16 illustrates the displacement vector encoding unit in Figure 15 in detail.

[0203] The displacement vector encoding unit can perform displacement vector encoding on 2D video encoders such as H.264, HEVC, and VVC. Displacement vector transformation is performed, and the transformed displacement vector coefficients are quantized and packed into a 2D image, which can then be encoded using a video codec. The detailed process in Fig. 16 is described in detail below.

[0204] The displacement vector transformation unit can transform the calculated vertex displacement vector (x, y, z) or the coordinate system transformed displacement vector (n, t, b) into displacement vector coefficients by performing linear lifting transformation, butterfly lifting transformation, etc.

[0205] The displacement vector coefficient quantization unit can perform quantization on the displacement vector coefficients that have been transformed by performing the displacement vector transformation unit in advance. Quantization can derive a quantized value (quant) for each channel by multiplying the displacement vector coefficient (value) by a scale and adding an offset, as in Equation 1. At this time, each channel can be x, y, z or n, t, b, respectively, depending on the coordinate system of the displacement vector.

[0206] Formula 1.

[0207] Formula 2.

[0208] The offset in Equation 1 is in units of sequence or frame, and can use a fixed value for each channel.

[0209] The scale of Equation 1 can be determined by the quantization parameter (QP) and the level-specific scale (level_scale), as in Equation 2.

[0210] Level_scale in Equation 2 can use a value determined per frame or sequence for each level. In addition, α, β, can be a parameter constant value determined by the encoder.

[0211] Similarly, the scale can have individual values ​​set for each channel of the displacement vector.

[0212] Figure 17 shows the structure of displacement vector coefficients according to embodiments.

[0213] Figure 17 shows the structure of the displacement vector coefficient converted in Figure 16.

[0214] The displacement vector coefficient packing unit packs the quantized displacement vector coefficients into a 2D image of size W x H. A lifting transformation is performed starting from the vertices of the LoD of the high layer, so that the displacement vector coefficients of the low LoD are stored in the front of the packing image, and the displacement vector coefficients of the high LoD are stored in the back of the packing image.

[0215] Figure 18 shows displacement vector coefficient 2D image packing according to embodiments.

[0216] The 1D displacement vector coefficients of Fig. 17 can be configured as in Fig. 18. The size bx×by can be configured as one block, and the number of displacement vector coefficients N can be configured as L×M blocks. At this time, packing can be performed as a 2D image of size W×H according to a 2D Morton code, zig-zag scan order, etc.

[0217] The L×M displacement vector coefficient blocks can be packed into an image of size (bx*L)×(by*M) according to the order defined in the encoder / decoder. At this time, the packing can be performed sequentially according to the scanning order starting from the R_0 (LoD 0) displacement vector coefficient block, and if the total size of the displacement vector coefficients is smaller than L×M, padding can be performed to make the displacement vector size (bx*L)×(by*M) and filled as shown in Figure 4. Alternatively, the packing and padding can be performed in the reverse order of Figure 4 to pack a 2D image.

[0218] L and M are determined based on the number of displacement vector coefficients (N), or L (or M) is defined by the encoder / decoder agreement, and then M (or L) can be derived based on the number of displacement vector coefficients (N) as in Equation 3.

[0219] The following formula derives M given L:

[0220]

[0221]

[0222] Figure 19 shows the 2D image packing of displacement vector coefficients by LoD (Level of Detail) according to embodiments.

[0223] According to an embodiment, packing of displacement vector coefficient blocks (bx×by) may be performed as in 19 to fit the CTU size of a 2D video encoder encoded for each LoD level. At this time, packing may be performed using the median value or the last displacement vector coefficient value of the image to fit the CTU size for each LoD.

[0224] Figure 20 illustrates a 2D Moulton code-based packing according to embodiments.

[0225] A single displacement vector coefficient block can be packed with bx×by displacement vector coefficients within the block through a zig-zag scan sequence or a 2D Morton code sequence such as in Fig. 20.

[0226] The displacement vector coefficients converted to the local coordinate system can be packed into Y, U, and V channels for each of the Normal, Tangential, and Bi-tangential components. In addition, formats such as YUV 4:4:4, YUV 4:2:0, and YUV 4:0:0 can be selected to perform image packing, and displacement vector coefficients can be configured according to each format. In addition, image packing format information (ColourSpace_displacement_video) of the specified displacement vector coefficients can be signaled.

[0227] Displacement vector coefficients per 4 vertex units for each format The sampling process can be performed, and the details can be as follows.

[0228] Figure 21 shows a displacement vector coefficient packing method according to embodiments.

[0229] Figure 21 illustrates a displacement vector coefficient packing method when using the YUV 4:4:4 format.

[0230] The displacement vector coefficient packing section performs packing of the values ​​of the normal, tangential, and bi-tangential components into the Y, U, and V channels, respectively.

[0231] For example, if the N component includes N1 to N4, the N1 to N4 components are packed as is in the Y channel. If the T component is T1 to T4, the T1 to T4 components are packed as is in the U channel. If the B component is B1 to B4, the B1 to B4 components are packed as is in the V channel.

[0232] Figure 22 shows a displacement vector coefficient packing method according to embodiments.

[0233] Figure 22 illustrates a displacement vector coefficient packing method when using the YUV 4:0:0 format.

[0234] All values ​​of the normal component can be packed into the Y channel as is. Tangential and bi-tangential components may not be sampled.

[0235] Figure 23 shows a displacement vector coefficient packing method according to embodiments.

[0236] Figure 23 illustrates a displacement vector coefficient packing method when using the YUV 4:2:0 format.

[0237] All normal components can be packed into the Y channel as is.

[0238] Tangential and bi-tangential components can only sample one component out of four into the U and V channels, respectively.

[0239] At this time, the encoder may select one component based on judgment, or sample the partial sum average of some components or the average value of all components. Alternatively, the values ​​at fixed positions of each component may be sampled.

[0240] The sampling methods for tangential and bi-tangential components into U and V channels can be as follows: 1) selecting one optimal component, 2) selecting by the average of the optimal partial sum of components, 3) selecting by the average of all components.

[0241] 1) How to choose the best ingredient:

[0242] Among the four components of each T, B component, the most optimal component can be selected and sampled as U, V channels. The sampled T, B in Fig. 23 can be T=T_1,T_2,T_3,T_4 and B=B_1,B_2,B_3,B_4, respectively. The optimal vertex can be determined by the encoder or can be a pre-arranged location. The optimal component value and location information can be signaled through the values ​​(value=1,2,3,4) of the t_sample, b_sample syntax for each t, ​​b component selected as the optimal component.

[0243] 2) How to select the optimal component partial sum average:

[0244] Of the four components of each T, B component, two or three components can be selected and the average value can be sampled as U, V channels. The sampled T, B of Fig. 23 can be T=(T1+T2) / 2, (T1+T3) / 2, (T1+T4) / 2, (T2+T3) / 2, (T2+T4) / 2, (T3+T4) / 2, (T1+T2+T3) / 2, (T1+T2+T4) / 2, (T2+T3+T4) / 2 and B=(B1+B2) / 2, (B1+B3) / 2, (B2+B3) / 2, (B2+B4) / 2, (B3+B4) / 2, (B1+B2+B3) / 2, (B1+B2+B4) / 2, (B2+B3+B4) / 2.

[0245] The optimal vertex components can be selected by the encoder or can be pre-arranged locations. By selecting two or three optimal vertices and using the t_sample and b_sample syntax values ​​(values=5,6,7,8,9,10,11,12,13), which represent the average of their values, the optimal component partial sum average and selected location information can be signaled.

[0246] 3) How to select by the average of all components:

[0247] The average of the four vertices of each T, B component can be sampled as U, V channels. The sampled T, B of Fig. 23 are respectively

[0248] T=(T1+T2+T3+T4) / 4, B=(B1+B2+B3+B4) / 4. The syntax values ​​t_sample and b_sample (value=0) that mean the average of four vertices can be signaled.

[0249] Even if the optimal sampling value is calculated by 1) the optimal component selection method and 2) the optimal component partial sum average method among these sampling methods, the encoder can signal the t_sample, b_sample syntax values ​​(value=0) so that all components can be restored to the same value.

[0250] When sampling tangential and bi-tangential components into the U and V channels, the sampling method can be selected independently. Alternatively, once the sampling method for the tangential component is determined, sampling of the bi-tangential component can be performed using the same sampling method.

[0251] Figure 24 shows the overall level packing frame configuration according to embodiments.

[0252] Displacement vector coefficients for each channel can be visualized to form an encoding unit in the form of a packing frame. The entire level within a mesh frame can be packed into one packing frame, or each LoD within the mesh frame can be packed into a separate packing frame. The asps_vmc_ext_displacement_LoD_packing_method syntax, which means the index of the packing method in mesh sequence units, can be signaled, and the packing frame and channel forms derived from one mesh frame according to the packing method can be as follows: entire level packing frame composition (Fig. 24), level-by-level packing frame composition (Fig. 25).

[0253] Full-level packing frame configuration (Fig. 24):

[0254] Each geometric axis (normal, tangential, bi-tangential) of the displacement vector coefficients of the entire level can be packed for each channel (Y Channel, U Channel, V Channel), and the entire packed channel can form a packing frame.

[0255] 1) When using YUV 4:4:4 format: The packing size of each Y, U, and V channel can be the same as hn=ht=hb, wn=wt=wb.

[0256] 2) When using YUV 4:0:0 format: Packing is performed only for the Y channel, and packing data for the U and V channels may not exist.

[0257] 3) When using YUV 4:2:0 format: Packing is performed as is for the Y channel, and for the U and V channels, only one vertex's data is sampled per four vertices, so ht, hb=hn / 2, wt, wb=wn / 2.

[0258] Figure 25 shows a level-by-level packing frame configuration according to embodiments.

[0259] Packing frame configuration by level (Fig. 25):

[0260] Each geometric axis (normal, tangential, bi-tangential) of the displacement vector coefficients of one level can be packed for each channel (Y Channel, U Channel, V Channel), and the channels packed at one level can form a packing frame.

[0261] The size (w, h) packed for each channel of Y, U, and V may differ depending on the format in the same way as the method of configuring the entire frame packing above.

[0262] The displacement vector image / video encoding unit can encode a 2D image packed through a displacement vector coefficient packing unit using a 2D video encoder such as H.264, HEVC, or VVC.

[0263] Fig. 26 shows a dynamic mesh decoder according to embodiments.

[0264] Fig. 26 shows the configuration of a dynamic mesh decoder corresponding to the Fig. 1 receiving device (110), receiver (111), file / segment decapsulator (112), dynamic mesh video decoder (113), renderer (114), decoder of Figs. 11-12, receiving device of Fig. 14, decoder of Figs. 26-27, receiving method of Fig. 38, etc. Each component of Fig. 26 may correspond to hardware, software, processor, and / or a combination thereof.

[0265] When a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream encoded by a dynamic mesh encoder are transmitted, the decoder decodes each bitstream to restore the mesh. First, the base mesh is decoded by a motion vector or static mesh decoding unit depending on whether it is an inter or intra frame, and then the geometric information is restored together with the decoded displacement vector information through subdivision. Embodiments include a method of decoding a 2D image frame transmitted in a YUV 4:2:0 format by parsing the sampling method for each component of the displacement vector in a displacement vector coefficient decoding unit. The principle performed at each step is described in detail below.

[0266] Figure 27 shows a displacement vector decoding unit according to embodiments.

[0267] The displacement vector decoding unit of Fig. 26 receives a displacement vector bitstream in which the displacement vector is encoded from the encoding unit, performs decoding, and can restore the displacement vector through the process of Fig. 27. The process of Fig. 27 is described in detail below.

[0268] The video decoding unit receives a displacement vector bitstream as input and performs decoding on the displacement vector coefficient image / video through a 2D video codec. The restored displacement vector coefficient video restored through the video decoding unit can perform displacement vector coefficient assignment corresponding to each vertex of the restored mesh by performing a displacement vector coefficient depacking unit for each frame.

[0269] The displacement vector coefficient inverse packing unit can perform inverse packing according to the scanning order defined by the sub / decoder agreement or the scanning order parsed in upper-level units (sequences, frames, etc.) from the restored displacement vector coefficient image corresponding to the current mesh frame. As shown in Fig. 19, the displacement vector coefficient block packed according to a specific scanning order in units of bx X by, and the displacement vector coefficient packed according to a specific order (such as 2D Morton Code or Zig-zag scan) within one block can derive the displacement vector transformation coefficient of the k-th vertex according to the sub / decoder agreement or the sizes bx, by and L, M of the parsed block and the scanning order.

[0270] Figure 28 shows a displacement vector coefficient inverse packing method according to embodiments.

[0271] Figure 28 shows an example of a displacement vector coefficient inverse packing method using the YUV 4:4:4 format.

[0272] By parsing the image packing format (ColourSpace_displacement_video) information, if the value is YUV444 (value=3), restoration can be performed for each Y, U, and V channel in YUV 4:4:4 format.

[0273] The displacement vector coefficients of the restored Y, U, and V channels can be restored as Normal, Tangential, and Bi-tangential components, respectively, according to a specific order (displacement_scan_method) within the block.

[0274] Since the YUV 4:4:4 format transmits all values ​​of the N, T, and B components, the data before encoding and the decoded result can be the same.

[0275] Figure 29 shows a displacement vector coefficient inverse packing method according to embodiments.

[0276] Figure 29 shows an example of a displacement vector coefficient inverse packing method using the YUV 4:0:0 format.

[0277] By parsing the image packing format (ColourSpace_displacement_video) information, if the value is YUV400 (value=1), restoration can be performed only for the Y channel in YUV 4:0:0 format.

[0278] The displacement vector coefficients of the restored Y channel can be restored as normal components according to a specific order (displacement_scan_method) within the block.

[0279] In the case of the YUV 4:0:0 format, all values ​​are transmitted for the N component of the Y channel, so the data before encoding the Normal component and the decoded result may be the same.

[0280] Since transmission is not performed on the U and V channels of the T and B components, the decoding results of the tangential and bi-tangential components may not exist.

[0281] Figure 30 shows a displacement vector coefficient inverse packing method according to embodiments.

[0282] Figure 30 illustrates an example of a displacement vector coefficient inverse packing method using the YUV 4:2:0 format.

[0283] By parsing the image packing format (ColourSpace_displacement_video) information, if the value is YUV420 (value=2), restoration can be performed on the N, T, and B components of the Y, U, and V channels in YUV 4:2:0 format.

[0284] The displacement vector coefficients of the restored Y, U, and V channels can be restored as Normal, Tangential, and Bi-tangential components, respectively, according to a specific order (displacement_scan_method) within the block.

[0285] Since all values ​​are transmitted for the N component of the Y channel, the data before encoding and the decoded result of the Normal component can be the same.

[0286] For the U and V channels of the T and B components, the t_smaple and b_sample syntaxes, which indicate the optimal component sampling method selected by the encoder, can be parsed to perform restoration according to the values ​​of the tangential and bi-tangential components and the restoration positions within the T and B channels.

[0287] When selecting one optimal component: For example, if the syntax t_sample value of the U channel is 2 and the syntax b_sample value of the V channel is 3, the displacement vector coefficients can be restored by performing reverse packing for each channel as in Fig. 30. The combination of each component before performing quantization of the displacement vector coordinates (n, t, b) for the four vertices restored in this way can be V1 (N1, 0, 0), V2 (N2, T2, 0), V3 (N3, 0, B3), and V4 (N4, 0, 0).

[0288] Figure 31 shows a displacement vector coefficient inverse packing method according to embodiments.

[0289] Figure 31 illustrates an example of a displacement vector coefficient inverse packing method using the YUV 4:2:0 format.

[0290] For optimal component partial sum average: For example, if the syntax t_sample value of the U channel is 5 and the syntax b_sample value of the V channel is 10, the displacement vector coefficients can be restored by performing depacking for each channel as shown in Fig. 31 below. For optimal component partial sum average, the transmitted average value (T or B) can be positioned at the position of the components used for partial sum by referring to the syntax table. The combination of each component before performing quantization of the displacement vector coordinates (n, t, b) for the four vertices restored in this way can be V1 (N1, T, 0), V2 (N2, T, 0), V3 (N3, 0, B), and V4 (N4, 0, B).

[0291] Figure 32 shows a displacement vector coefficient inverse packing method according to embodiments.

[0292] Figure 32 illustrates an example of a displacement vector coefficient inverse packing method using the YUV 4:2:0 format.

[0293] In the case of the average of all components: For example, if the syntax t_sample, b_sample values ​​of the U and V channels are both 0, the displacement vector coefficients can be restored by performing depacking for each channel as in Fig. 32. In the case of the average of all components, the average value (T, B) can be positioned at the position of all components. The combination of each component before performing quantization of the displacement vector coordinates (n, t, b) for the four vertices restored in this way can be V1 (N1, T, B), V2 (N2, T, B), V3 (N3, T, B), V4 (N4, T, B).

[0294] The displacement vector coefficient inverse quantization unit of Fig. 27 performs inverse quantization of displacement vector coefficients allocated per vertex through the displacement vector coefficient inverse packing unit. Depending on the embodiment, the quantization parameter (QP) for each axis (Normal, Tangential, Bi-tangential) can be transmitted in sequence or frame units, and when only the displacement vector of the Normal component is encoded / decoded (ColourSpace_displacement_video=yuv400), the quantization rate can be determined by parsing only the quantization parameter (QP) for the Normal component.

[0295] Depending on the embodiment, displacement vector coefficients may be quantized through different quantization parameters for each axis, and the quantization rate may be determined for each LoD level by deriving quantization parameters or scaling parameters by a decoder / decoder agreement.

[0296] The displacement vector inverse transform section of Fig. 27 performs an inverse transform on the displacement vector coefficients on which inverse quantization has been performed to calculate a restored displacement vector. Depending on the embodiment, the transform may be a linear lifting transform, a butterfly lifting transform, etc.

[0297] When the lifting inverse transformation is performed, the vertex Rk of the kth subdivision level is used as a predictor when performing prediction, Rt(t <k 또는 t<=k)의 세분화 정점 변위 벡터를 통해 k번째 세분화 레벨의 변위벡터 예측을 수행할 수 있다.

[0298] According to an embodiment, when performing prediction of a displacement vector, an average or distance-based weighted average prediction can be performed on n points near the current vertex based on connection information among vertices with a lower level of detail than the current vertex.

[0299] In some embodiments, prediction can be performed based on the displacement vectors of n vertices used to generate the current vertex in the mesh refinement step.

[0300] When lifting inverse transformation is performed, a process of updating the displacement vector of the vertex used for prediction in the encoder can be performed through the parsed residual signal.

[0301] In the displacement vector coordinate system inverse transformation section of Fig. 26, the coordinate system transformation flag (asps_vmc_ext_displacement_coordinate_system) is parsed by frame sequence or GOF (Group of Frames) or frame or submesh unit, and if the value is 1, the inverse quantized restored displacement vector can be inversely transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z).

[0302] Based on the restored vertex position information of the restored base mesh, a normal vector per vertex is calculated, and for the vertices additionally generated through the subdivision process, the normal value of the newly generated vertex can be assigned by interpolating through the vertex normal vector of the calculated restored base mesh.

[0303] In the case of interpolation, interpolation can be performed by averaging or distance-based weighting the normal information of the base mesh used for subdivision.

[0304] Alternatively, the normal information of the base mesh can be used as is for the subdivided vertices on the same plane.

[0305] Through the calculated normal vector per vertex, the tangential and bi-tangential vectors orthogonal to the normal vector can be calculated, and the displacement vector coordinate system inverse transformation can be performed through Equation 4. dispn[0], dispn[1], and dispn[2] in Equation 4 represent the results of the n, t, and b components that have been inversely quantized and inversely transformed for each Y, U, and V channel.

[0306] Formula 4.

[0307] When packed in YUV 4:0:0 format, the coordinate system inversion equation 4 is It can be expressed as . At this time, the coordinate system inverse transformation can be performed by multiplying the result of performing inverse quantization and inverse transformation of the n components by the calculated normal vector per vertex.

[0308] If there are multiple results of performing inverse transformation for the same vertex coordinates, the sum of all values ​​can become the value of the final displacement vector.

[0309] Depending on the embodiment, it is always possible to perform coordinate system inversion without sending flags.

[0310] The mesh restoration unit of Fig. 26 can calculate and restore vertex geometric information of the restoration mesh by adding a restoration displacement vector to vertices generated through a subdivision process in the mesh subdivision unit.

[0311] The point cloud data transmission method / device according to the embodiments (Fig. 1 Transmission device (100), dynamic mesh video acquisition unit (101), dynamic mesh video encoder (102), file / segment encapsulator (103), transmitter (104), Fig. 2 Pre-processor (200), encoder (201), Figs. 6-7 Encoder, Fig. 13 Transmission device, Figs. 15-16 Encoder, Fig. 37 Transmission method) can encode mesh data and transmit it in the form of a bitstream. In addition, parameter information related to encoding can be generated and transmitted by including it in a bitstream.

[0312] The point cloud data receiving method / device according to the embodiments (receiving device (110) in FIG. 1, receiver (111), file / segment decapsulator (112), dynamic mesh video decoder (113), renderer (114), decoder in FIG. 11-12, receiving device in FIG. 14, decoder in FIG. 26-27, receiving method in FIG. 38) can receive a bitstream and decode mesh data based on parameters in the bitstream.

[0313] Below, FIGS. 33 to 36 describe the syntax and semantics of parameters included in the bitstream.

[0314] Figure 33 shows an atlas sequence parameter set in a bitstream according to embodiments.

[0315] Figure 33 illustrates the syntax of an Atlas Sequence Parameter Set (ASPS) included in a bitstream. ASPS is extended from the Atlas Sequence Parameter Set to include additional parameters related to mesh data encoding.

[0316] Displacement vector coordinate system (asps_vmc_ext_displacement_coordinate_system): Indicates the type of coordinate system for encoding displacement vectors. For example, if this value is 0: Canonical, this value indicates that the coordinate system of the displacement vectors of the mesh data is encoded by transforming it based on the local coordinate system.

[0317] Displacement vector transformation method (asps_vmc_ext_transform_method): Indicates how the displacement vector is transformed. For example, if this value is 0: None, if this value is 1: Linear_Lifting.

[0318] Displacement vector coefficient packing method (ext_packing_method): Indicates the packing method for displacement vector coefficients. For example, if this value is 0, displacement vector coefficients are packed in ascending order, and if it is 1, displacement vector coefficients are packed in descending order.

[0319] Displacement Vector Encoding Method (asps_vmc_ext_displacement_method): Indicates the encoding method of the displacement vector. For example, if this value is 0: None (no displacement vector encoding), 1: Arithmetic Coding, and 2: Video Coding.

[0320] Displacement vector packing method by LoD (asps_vmc_ext_displacement_LoD_packing_method): Indicates the packing method for displacement vectors by LoD. For example, if this value is 0: it indicates that the displacement vectors are packed based on the full-level packing frame configuration, and if it is 1: it indicates that the displacement vectors are packed based on the level-specific packing frame configuration.

[0321] Displacement vector transformation method (asps_vmc_ext_transform_method): Indicates the transformation method for the displacement vector. For example, if this value is 0: None (no transformation for the displacement vector), if it is 1: it indicates that the displacement vector was transformed by linear lifting.

[0322] Figure 34 shows an atlas sequence parameter set in a bitstream according to embodiments.

[0323] Displacement_scan_method: Indicates the scan order method when packing displacement vectors. For example, if it is 0, it indicates that the displacement vectors are scanned using the 2D Morton Code, and if it is 1, it indicates that the displacement vectors are scanned using the Zig-Zag scan order.

[0324] GeometryVideoBlockSize: Indicates the displacement vector video image block size. For example, this value can indicate the number of bx×by blocks, and the default value can be 16.

[0325] geometryVideoBitDepth: Indicates the displacement vector video image bit depth unit. The default value can be 10 bits.

[0326] ColorSpace_displacement_video: Indicates the displacement vector video packing image format. For example, if this value is 0: None (there is no image format for displacement vector video packing), 1: yuv400, 2: yuv420, 3: yuv444, it indicates that the displacement vector is packed based on this.

[0327] Sampling method and sampling value of tangential component (t_sample): Indicates the sampling method of tangential component and the location information of the displacement vector coefficient used for sampling.

[0328] Sampling method and sampling value of bi-tangential component (b_sample): Indicates the sampling method of bi-tangential component and the location information of the displacement vector coefficient used for sampling.

[0329] Figure 35 shows displacement vector coefficient packing and inverse packing according to embodiments.

[0330] Fig. 35 shows the operation of packing and reverse packing displacement vector coefficients according to Fig. 23 and Fig. 32.

[0331] The transmission method / device according to the embodiments, as described in Fig. 23 and the like, packs the components of the NTB channel into a YUV channel with reference to Fig. 35, and the packing method may be a YUV 4:2:0 format. If the packed displacement vector coefficient video image is reverse-packed with Fig. 32 and the like, the components of the NTB channel can be restored from the YUV components again, as shown in Fig. 35. This sampling method and sampling value are signaled by t_sample and b_sample of Fig. 34, and the details of these values ​​are as shown in Figs. 36 and 37.

[0332] Figure 36 shows a sampling method and sampling values ​​of displacement vector coefficient components according to embodiments.

[0333] As described above, the sampling method and sampling value according to the t_sample value are illustrated in Fig. 36.

[0334] If t_sample is 0: When packing T components, sample the average value of each T component and pack it into the U channel component, and when packing back from the U channel to the T channel, restore each T component with the average value.

[0335] If t_sample is 1: When packing T components, sample with T1 (the first T value) and pack it into the components of the U channel, and when reverse packing from the U channel to the T channel, restore the first value of the T channel with T1 (the first T value), and restore the second to fourth values ​​to 0.

[0336] If t_sample is 2: When packing the T component, sample with T1 (the second T value) and pack it into the U channel component, and when reverse packing from the U channel to the T channel, restore the second value of the T channel with T1 (the second T value), and restore the first, third, and fourth values ​​to 0.

[0337] If t_sample is 3: When packing the T component, sample with T1 (the third T value) and pack it into the U channel component, and when reverse packing from the U channel to the T channel, restore the third value of the T channel with T1 (the third T value), and restore the first, second, and fourth values ​​to 0.

[0338] If t_sample is 4: When packing the T component, sample with T1 (the 4th T value) and pack it into the U channel component, and when reverse packing from the U channel to the T channel, restore the 4th value of the T channel with T1 (the 4th T value), and restore the 1st, 2nd, and 3rd values ​​to 0.

[0339] If t_sample is 5: When packing T components, sample with the average value of T1 (first T value) and T2 (second T value) and pack it into the component of the U channel, and when reverse packing from the U channel to the T channel, restore the first and second values ​​with the average value of T1 (first T value) and T2 (second T value), and restore the third and fourth values ​​to 0.

[0340] If t_sample is 6: When packing T components, sample with the average value of T1 (the first T value) and T3 (the third T value) and pack it into the U channel component, and when reverse packing from the U channel to the T channel, restore the first and third values ​​with the average value of T1 (the first T value) and T3 (the third T value), and restore the second and fourth values ​​to 0.

[0341] If t_sample is 7: When packing T components, sample with the average value of T1 (the first T value) and T4 (the fourth T value) and pack it into the U channel components, and when reverse packing from the U channel to the T channel, restore the first and fourth values ​​with the average value of T1 (the first T value) and T4 (the fourth T value), and restore the second and third values ​​to 0.

[0342] If t_sample is 8: When packing the T components, sample with the average value of T2 (the second T value) and T3 (the third T value) and pack it into the U channel components, and when reverse packing from the U channel to the T channel, restore the second and third values ​​with the average value of T2 (the second T value) and T3 (the third T value), and restore the first and fourth values ​​to 0.

[0343] If t_sample is 9: When packing T components, sample with the average value of T2 (the second T value) and T4 (the fourth T value) and pack it into the U channel components, and when reverse packing from the U channel to the T channel, restore the second and fourth values ​​with the average value of T2 (the second T value) and T4 (the fourth T value), and restore the first and third values ​​to 0.

[0344] If t_sample is 10: When packing T components, sample with the average value of T3 (the third T value) and T4 (the fourth T value) and pack it into the U channel components, and when reverse packing from the U channel to the T channel, restore the third and fourth values ​​with the average value of T3 (the third T value) and T4 (the fourth T value), and restore the first and second values ​​to 0.

[0345] If t_sample is 11: When packing T components, sample with the average value of T1 (first T value), T2 (second T value), and T3 (third T value) and pack them into the U channel components, and when reverse packing from the U channel to the T channel, restore the first, second, and third values ​​with the average value of T1 (first T value), T2 (second T value), and T3 (third T value), and restore the fourth value to 0.

[0346] If t_sample is 12: When packing T components, sample with the average value of T1 (the first T value), T2 (the second T value), and T4 (the fourth T value) and pack them into the components of the U channel, and when reverse packing from the U channel to the T channel, restore the first, second, and fourth values ​​with the average value of T1 (the first T value), T2 (the second T value), and T4 (the fourth T value), and restore the third value to 0.

[0347] If t_sample is 13: When packing the T components, sample the average value of T2 (the second T value), T3 (the third T value), and T4 (the fourth T value) and pack it into the U channel components, and when reverse packing from the U channel to the T channel, restore the second, third, and fourth values ​​as the average value of T2 (the second T value), T3 (the third T value), and T4 (the fourth T value), and restore the first value to 0.

[0348] If b_sample is 0: When packing B components, sample the average value of each B component and pack it into the U channel components, and when packing back from the V channel to the B channel, restore each B component with the average value.

[0349] If b_sample is 1: When packing the B component, sample with B1 (the first B value) and pack it into the V channel component, and when reverse packing from the V channel to the B channel, restore the first value of the B channel with B1 (the first B value), and restore the second to fourth values ​​to 0.

[0350] If b_sample is 2: When packing the B component, sample with B1 (the second B value) and pack it into the V channel component, and when reverse packing from the V channel to the B channel, restore the second value of the B channel with B1 (the second B value), and restore the first, third, and fourth values ​​to 0.

[0351] If b_sample is 3: When packing the B component, sample with B1 (the third B value) and pack it into the V channel component, and when reverse packing from the V channel to the B channel, restore the third value of the B channel with B1 (the third B value), and restore the first, second, and fourth values ​​to 0.

[0352] If b_sample is 4: When packing the B component, sample with B1 (the 4th B value) and pack it into the V channel component, and when reverse packing from the V channel to the B channel, restore the 4th value of the B channel with B1 (the 4th B value), and restore the 1st, 2nd, and 3rd values ​​to 0.

[0353] If b_sample is 5: When packing the B component, sample the average value of B1 (first B value) and B2 (second B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the first and second values ​​as the average value of B1 (first B value) and B2 (second B value), and restore the third and fourth values ​​to 0.

[0354] If b_sample is 6: When packing the B component, sample the average value of B1 (the first B value) and B3 (the third B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the first and third values ​​as the average value of B1 (the first B value) and B3 (the third B value), and restore the second and fourth values ​​to 0.

[0355] If b_sample is 7: When packing the B component, sample the average value of B1 (the first B value) and B4 (the fourth B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the first and fourth values ​​as the average value of B1 (the first B value) and B4 (the fourth B value), and restore the second and third values ​​to 0.

[0356] If b_sample is 8: When packing the B component, sample the average value of B2 (the second B value) and B3 (the third B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the second and third values ​​as the average value of BT2 (the second B value) and B3 (the third B value), and restore the first and fourth values ​​to 0.

[0357] If b_sample is 9: When packing the B component, sample the average value of B2 (the second B value) and B4 (the fourth B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the second and fourth values ​​as the average value of B2 (the second B value) and B4 (the fourth B value), and restore the first and third values ​​to 0.

[0358] If b_sample is 10: When packing the B component, sample the average value of B3 (the third B value) and B4 (the fourth B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the third and fourth values ​​as the average value of B3 (the third B value) and B4 (the fourth B value), and restore the first and second values ​​to 0.

[0359] If b_sample is 11: When packing the B component, sample the average value of B1 (the first B value), B2 (the second B value), and B3 (the third B value) and pack it into the V channel component, and when reverse packing from the V channel to the B channel, restore the first, second, and third values ​​with the average value of B1 (the first B value), B2 (the second B value), and B3 (the third B value), and restore the fourth value to 0.

[0360] If b_sample is 12: When packing the B component, sample the average value of B1 (the first B value), B2 (the second B value), and B4 (the fourth B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the first, second, and fourth values ​​as the average value of B1 (the first B value), B2 (the second B value), and B4 (the fourth B value), and restore the third value to 0.

[0361] If b_sample is 13: When packing the B component, sample the average value of B2 (the second B value), B3 (the third B value), and B4 (the fourth B value) and pack it as a component of the V channel, and when reverse packing from the V channel to the B channel, restore the second, third, and fourth values ​​as the average value of B2 (the second B value), B3 (the third B value), and B4 (the fourth B value), and restore the first value to 0.

[0362] Referring to FIG. 15, a transmitting device according to embodiments generates a base mesh by passing an original mesh to be transmitted through a simplification unit and a parameterization unit. The generated base mesh is quantized, and in the case of an inter-frame, a motion vector is calculated from a previously referenced restored base mesh and encoded, and in the case of an intra-frame, a base mesh bitstream is transmitted through a static mesh encoding unit. Then, a displacement vector is calculated between mesh data obtained by segmenting and fitting a mesh that has passed through the mesh simplification unit and mesh data restored from the previously encoded base mesh. In order to efficiently encode the calculated displacement vector, the displacement vector coordinate system can be converted to a local coordinate system, and a displacement vector conversion unit converts and quantizes the displacement vector into a displacement vector coefficient and encodes it into a displacement vector bitstream. A displacement vector coefficient image packing process is performed in a displacement vector encoding unit using a YUV 4:2:0 format according to embodiments.

[0363] The displacement vector encoding unit can be divided into detailed modules as shown in Fig. 15. First, the transformed displacement vectors can be transformed into displacement vector coefficients by performing linear lifting transformation, butterfly lifting transformation, etc. In addition, the displacement vectors of (n, t, b) expressed in the local coordinate system can be transformed for each n, t, b component. Afterwards, the displacement vector transformation unit can be performed to quantize the transformed displacement vector coefficients, and quantize them into individual values ​​for each channel. In the displacement vector coefficient packing unit, the quantized displacement vector coefficients are packed into a 2D image of the size of WXH. As shown in Fig. 17, the displacement vector coefficients composed of the LoD ascending order in a 1D form are packed into an image in a 2D form as shown in Fig. 18. The size of bx X by can be configured as one block, and it can be configured as LXM blocks determined according to the number N of displacement vector coefficients. The basic unit size information of the block (geometryVideoBlockSize) and bit depth information (geometryVideoBitDepth) can be signaled to derive L and M. In addition, displacement vector coefficients can be packed inside the block in a zig-zag scan order or a 2D Morton code order, and displacement vector packing order information (displacement_scan_method) can be signaled. Padding can be performed with the median value or the last displacement vector coefficient value of the image to match the size of the basic block for each LoD or the entire 2D video image. The displacement vector coefficients converted to the local coordinate system can constitute packing frames for the Y, U, and V channels for the Normal, Tangential, and Bi-tangential components, respectively. To perform 2D image packing, the image packing format information (ColourSpace_displacement_video) can be signaled by selecting a format such as YUV 4:4:4, YUV 4:2:0, or YUV 4:0:0.When using the YUV 4:4:4 format, the values ​​of the Normal, Tangential, and Bi-tangential components are packed into the Y, U, and V channels, respectively. When using the YUV 4:0:0 format, the values ​​of the Normal component can all be packed into the Y channel as is, and the Tangential and Bi-tangential components may not be sampled. When using the YUV 4:2:0 format, the values ​​of the Normal component can all be packed into the Y channel as is, and the Tangential and Bi-tangential components may be sampled into the U and V channels, respectively, for some or all of every four vertices, depending on the sampling method of each component. When selecting one optimal component among the T and B components as the sampling method, the position of the corresponding component and the values ​​of the t_sample and b_sample syntax (value=1,2,3,4) meaning sampling one optimal component can be signaled. When calculating the mean of two or three partial sums among the four components of each T, B component, the values ​​of the t_sample, b_sample syntax (value=5,6,7,8,9,10,11,12,13) ​​indicating the optimal partial sum mean and the positions of the selected vertices can be signaled. When sampling the mean values ​​of the four vertices of each T, B component, the values ​​of the t_sample, b_sample syntax (value=0) can be signaled.

[0364] Displacement vector coefficients can be imaged for each channel to form an encoding unit in the form of a packing frame. The entire level within a mesh frame can be packed into one packing frame, or each LoD within a mesh frame can be packed into a separate packing frame, and the asps_vmc_ext_displacement_LoD_packing_method syntax, which indicates the packing method for each mesh sequence, can be signaled. In the displacement vector image / encoding unit, a 2D image packed through the displacement vector coefficient packing unit is encoded through a 2D video encoder such as H.264, HEVC, or VVC to generate a displacement vector bitstream.

[0365] Finally, the texture map generation unit generates a new texture map having color information corresponding to the texture coordinates of the restored mesh, encodes the texture map through a 2D video encoder, and transmits it as a texture bitstream.

[0366] The base mesh bitstream, displacement vector bitstream, and texture bitstream generated through the entire process described above at the transmitter are generated as a single bitstream through a multiplexer and transmitted through the transmitter.

[0367] Referring to FIG. 26, the receiving device receives the bitstream transmitted from the transmitter and performs a process of decoding the base mesh bitstream, displacement vector bitstream, and texture map bitstream, respectively, through a demultiplexing unit.

[0368] First, the base mesh bitstream is decoded through a motion vector decoding unit for inter-frames and a static mesh decoding unit for intra-frames. The decoded base mesh then passes through a restoration unit for mesh refinement.

[0369] The displacement vector bitstream is decoded in the reverse order of encoding to decode the displacement vector coefficients, perform inverse quantization and inverse transformation, and then inversely transform to the coordinate system to restore the mesh geometric information along with the base mesh data. The displacement vector decoding process proposed in the present invention is described in detail for each module as shown in Fig. 27.

[0370] The video decoding unit receives a displacement vector bitstream as input and performs decoding on a displacement vector coefficient image / video through a 2D video codec. Inverse packing can be performed from a restored displacement vector coefficient image corresponding to the current mesh frame. Scan order information (displacement_scan_method), block size information (geometryVideoBlockSize), and bit depth information (geometryVideoBitDepth) during displacement vector packing are parsed to derive block and packing frame sizes (bx, by, and L, M), and inverse packing is performed on the packed displacement vector coefficients. During the inverse packing process, packing image format information (ColourSpace_displacement_video) is parsed, and the method of deriving transformation coefficients may differ depending on the transmitted image packing format. When using the YUV 4:4:4 format (ColourSpace_displacement_video=3), the Normal, Tangential, and Bi-tangential components of the Y, U, and V channels can be restored respectively. Since the YUV 4:4:4 format transmits all values ​​of the N, T, and B components, they can be restored to the same values ​​as the data before encoding. When using the YUV 4:0:0 format (ColourSpace_displacement_video=1), only the Normal component of the Y channel can be restored. When using the YUV 4:2:0 format (ColourSpace_displacement_video=2), the Normal component of the Y channel can be restored in its entirety, but some or all of the Tangent and Bi-tangent components of the U and V channels may be restored depending on the sampling method. The displacement vector coefficients can be restored according to the sampling method of the T and B components by parsing the t_sample and b_sample syntax.When the T_sample, b_sample syntax values ​​are 1 to 4, one optimal component is sampled, and restoration can be performed according to each sampling position. When the T_sample, b_sample syntax values ​​are 5 to 13, two or three optimal components are selected and sampled as an average value, and restoration can be performed according to the selected component positions. When the T_sample, b_sample syntax values ​​are 0, the average value of four vertices is sampled, and restoration can be performed according to the entire position of the component. Or, even if a part of the optimal component is sampled, restoration can be performed according to the entire position if the t_sample, b_sample syntax values ​​are 0. The displacement vector coefficient dequantization unit performs dequantization on the displacement vector coefficients assigned to each vertex through the displacement vector coefficient depacking unit. Displacement vector coefficients can be quantized through different quantization parameters for each axis, and the quantization rate can be determined for each LoD level by deriving the quantization parameter or scaling parameter. The displacement vector inverse transform unit calculates the restored displacement vector by performing an inverse transformation on the displacement vector coefficients on which inverse quantization was performed using the transformation method parsed in asps_vmc_ext_transform_method. When all displacement vectors are inversely transformed, asps_vmc_ext_displacement_coordinate_system is parsed, and if the value is 1, the restored displacement vector can be inversely transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z). The normal vector per vertex is calculated based on the restored vertex position information of the restored base mesh, and the tangential and bi-tangential vectors orthogonal to the normal vector are calculated using the calculated normal vector per vertex, thereby performing an inverse transformation of the displacement vector coordinate system. If the value is 0, the calculated restored displacement vector can become the final restored displacement vector.If there are multiple results of performing inverse transformation on the same vertex coordinates, the sum of all values ​​can become the value of the final displacement vector. In the mesh restoration unit, the vertex geometry information of the restoration mesh can be calculated by adding the restoration displacement vector to the vertices generated through the subdivision process in the mesh refinement unit, thereby restoring the final geometry information. The received texture map bitstream is decoded through the texture map decoding unit, and the final restoration mesh is generated together with the previously restored geometry information.

[0371] Figure 37 shows a mesh data transmission method according to embodiments.

[0372] Fig. 37 shows the mesh data transmission operation of the Fig. 1 transmitting device (100), the dynamic mesh video acquisition unit (101), the dynamic mesh video encoder (102), the file / segment encapsulator (103), the transmitter (104), the Fig. 2 pre-processor (200), the encoder (201), the Fig. 6-7 encoder, the Fig. 13 transmitting device, and the Fig. 15-16 encoder.

[0373] The mesh data transmission method according to the embodiments may include a step (S3700) of encoding mesh data.

[0374] The mesh data transmission method according to the embodiments may further include a step (S3701) of transmitting a bitstream including mesh data.

[0375] The encoding step (S3700), referring to FIG. 15, may include: a step of generating a base mesh by simplifying (decimating) mesh data; a step of encoding the base mesh based on at least one of an inter-frame and an intra-frame method; a step of decompressing the base mesh; a step of subdividing the base mesh; a step of generating a displacement vector based on the restored base mesh and the subdivided base mesh; and a step of encoding the displacement vector.

[0376] The step of encoding a displacement vector may include, with reference to FIGS. 16 to 20, a step of converting a displacement vector into a displacement vector coefficient; and a step of packing the displacement vector coefficient into an image.

[0377] The step of packing displacement vector coefficients into an image may include sampling values ​​included in the first component (Normal), the second component (Tangential), and the third component (Bi-tangential) of the local coordinate system of the displacement vector coefficients and packing them into the first channel (Y), the second channel (U), and the third channel (V), respectively, as shown in FIGS. 21 to 23.

[0378] All values ​​of the first component of the local coordinate system can be packed into the first channel, only one value of the values ​​of the second component of the local coordinate system can be packed into the second channel, and only one value of the values ​​of the third component of the local coordinate system can be packed into the third channel. That is, the optimal one value of the values ​​can be used during packing.

[0379] All values ​​of the first component of the local coordinate system may be packed into a first channel, the average of at least two values ​​of the second component of the local coordinate system may be packed into a second channel, and the average of at least two values ​​of the third component of the local coordinate system may be packed into a third channel. That is, the partial sum average or the total sum average of the values ​​may be used in packing.

[0380] The method of FIG. 37 can be performed by a transmitting device, and the transmitting device includes, with reference to FIG. 1, a memory; and a processor that performs one or more instructions stored in the memory; and the processor can perform: encoding mesh data; and transmitting a bitstream including the mesh data.

[0381] Figure 38 shows a mesh data receiving method according to embodiments.

[0382] Fig. 38 shows the mesh data reception operation of the Fig. 1 receiving device (110), receiver (111), file / segment decapsulator (112), dynamic mesh video decoder (113), renderer (114), Fig. 11-12 decoder, Fig. 14 receiving device, and Fig. 26-27 decoder.

[0383] A method for receiving mesh data according to embodiments may include a step (S3800) of receiving a bitstream including mesh data.

[0384] The mesh data receiving method according to the embodiments may further include a step (S3801) of decoding mesh data.

[0385] In the receiving step (S3800), the bitstream received, with reference to FIGS. 33-34, may include at least one of information indicating an encoding type of a displacement vector (asps_vmc_ext_displacement_method), a method indicating a packing type for each LoD of a displacement vector (asps_vmc_ext_displacement_LoD_packing_method), information indicating a scan order of displacement vector packing (displacement_scan_method), information indicating a displacement vector video packing format (ColourSpace_displacement_video), or information indicating a position of a displacement vector coefficient on a local coordinate system (t_sample, b_sample).

[0386] The decoding step (S3801), referring to FIG. 26, may include a step of decoding a base mesh for mesh data based on at least one of an inter-frame and an intra-frame method; and a step of decoding a displacement vector within a bitstream.

[0387] The step of decoding a displacement vector may include, with reference to FIG. 27, decoding the displacement vector based on a video method to restore the displacement vector coefficients and packing the displacement vector coefficients inversely from the frame.

[0388] The step of de-packing displacement vector coefficients, referring to FIG. 30, may include de-packing values ​​in a first channel with values ​​of a de-packed first channel, de-packing values ​​in a second channel with first values ​​of a de-packed second channel, and the second values ​​of the de-packed second channel may be zero, de-packing values ​​in a third channel with first values ​​of a de-packed third channel, and the second values ​​of the de-packed third channel may be zero. That is, this is a de-packing operation when t_sample: 2, b_sample: 3.

[0389] The step of de-packing displacement vector coefficients, referring to FIG. 31, may include de-packing values ​​in a first channel with values ​​of the de-packed first channel, de-packing values ​​in a second channel with first and second values ​​of the de-packed second channel, and the third values ​​of the de-packed second channel may be zero, de-packing values ​​in a third channel with first and second values ​​of the de-packed third channel, and the third values ​​of the de-packed third channel may be zero. That is, this is a de-packing operation when t_sample: 5, b_sample: 10.

[0390] The step of de-packing displacement vector coefficients may include, with reference to FIG. 32, de-packing values ​​in a first channel with values ​​of the de-packed first channel, de-packing values ​​in a second channel with values ​​of the de-packed second channel, and de-packing values ​​in a third channel with values ​​of the de-packed third channel. That is, this is a de-packing operation when t_sample: 0, b_sample: 0.

[0391] The method of FIG. 38 is performed by a receiving device, the receiving device including a memory; and a processor that performs one or more instructions stored in the memory; wherein the processor is capable of performing: receiving a bitstream including mesh data; and decoding the mesh data.

[0392] The embodiments solve the following technical problems and provide technical effects. Specifically, in the mesh compression technology applied to the existing V-DMC, displacement vector information is expressed in the form of a 3D or 1D vector, compressed, and transmitted. This method requires the use of the YUV 4:4:4 or YUV 4:0:0 format for image packing when encoding / decoding displacement vectors using a 2D video codec. Using the YUV 4:4:4 format can transmit all component data of the displacement vector, but has the disadvantage of large capacity. Using the YUV 4:0:0 format does not transmit the tangential or bi-tangential components, but only the normal component, so the capacity is greatly reduced, but the disadvantage is that the image quality is somewhat degraded. Therefore, the present embodiments solve this problem by packing displacement vector images using the YUV 4:2:0 format. Compared to the two formats used in V-DMC technology, it can be efficient in terms of capacity and image quality, and it is also much more advantageous in terms of compatibility because it uses the YUV 4:2:0 format, which is currently the most widely used 2D video codec. In addition, it has a technical effect in that it provides an efficient and effective sampling method for the tangential and bi-tangential components of the U and V channels when applying the YUV 4:2:0 format to V-DMC technology. By using a method that can perform sampling by selecting the optimal component of each tangential and bi-tangential component and a signaling method that can restore the sampling value at the location of the component involved in the sampling, it is possible to transmit more efficient and high-quality images.

[0393] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.

[0394] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.

[0395] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may be implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium may be distributed across network-connected computer systems, allowing the processor-readable code to be stored and executed in a distributed manner.

[0396] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.

[0397] Terms such as first, second, etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by the above terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a second user input signal. Similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although a first user input signal and a second user input signal are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.

[0398] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.

[0399] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.

[0400] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.

[0401] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.

[0402] As described above, the relevant contents have been described in the best form for carrying out the embodiments.

[0403] As described above, the embodiments may be applied in whole or in part to a point cloud data transmission and reception device and system.

[0404] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.

[0405] Embodiments may include modifications / changes, which do not depart from the scope of the claims and their equivalents.

Claims

1. Step of encoding mesh data; and A step of transmitting a bitstream including the above mesh data; comprising: How to transmit mesh data.

2. In paragraph 1, The steps for encoding the above mesh data are: A step of creating a base mesh by simplifying (decimating) the above mesh data; A step of encoding the base mesh based on at least one of an inter-frame and an intra-frame method; Step of northing the above base mesh; A step of subdividing the above base mesh; A step of generating a displacement vector based on the restored base mesh and the refined base mesh; and A step of encoding the displacement vector; including; How to transmit mesh data.

3. In paragraph 2, The steps of encoding the above displacement vector are: A step of converting the above displacement vector into a displacement vector coefficient; and A step of packing the displacement vector coefficients into an image; including; How to transmit mesh data.

4. In paragraph 3, The steps for packing the above displacement vector coefficients into an image are: It includes sampling the values ​​included in the first component, the second component, and the third component of the local coordinate system of the displacement vector coefficient and packing them into the first channel, the second channel, and the third channel, respectively. How to transmit mesh data.

5. In paragraph 4, All values ​​of the first component of the local coordinate system are packed into the first channel, only one value of the values ​​of the second component of the local coordinate system is packed into the second channel, and only one value of the values ​​of the third component of the local coordinate system is packed into the third channel. How to transmit mesh data.

6. In paragraph 4, All values ​​of the first component of the local coordinate system are packed into the first channel, an average value of at least two values ​​of the second component of the local coordinate system are packed into the second channel, and an average value of at least two values ​​of the third component of the local coordinate system are packed into the third channel. How to transmit mesh data.

7. Memory; and A processor configured to perform one or more instructions stored in the memory, wherein the processor comprises: Encode mesh data; and transmitting a bitstream containing the above mesh data; Mesh data transmitter.

8. A step of receiving a bitstream containing mesh data; and A step of decoding the above mesh data; comprising: How to receive mesh data.

9. In paragraph 8, The steps for decoding the above mesh data are: A step of decoding a base mesh for the above mesh data based on at least one of an inter-frame and an intra-frame method; A step of decoding a displacement vector within the bitstream; comprising: How to receive mesh data.

10. In paragraph 9, The steps for decoding the above displacement vector are: The displacement vector is decoded based on the video method to restore the displacement vector coefficients, Including packing the displacement vector coefficients inversely from the frame, How to receive mesh data.

11. In Article 10 The steps for reverse packing the above displacement vector coefficients are: Repack the values ​​in the first channel with the values ​​in the reverse-packed first channel, The value in the second channel is reverse-packed with the first value of the reverse-packed second channel, and the second values ​​of the reverse-packed second channel become zero. The value in the third channel is reverse-packed with the first value of the reverse-packed third channel, and the second values ​​of the reverse-packed third channel become zero. How to receive mesh data.

12. In Article 10 The steps for reverse packing the above displacement vector coefficients are: Repack the values ​​in the first channel with the values ​​in the reverse-packed first channel, The values ​​in the second channel are reverse-packed with the first and second values ​​of the reverse-packed second channel, and the third values ​​of the reverse-packed second channel become zero. The values ​​in the third channel are reverse-packed with the first and second values ​​of the reverse-packed third channel, and the third values ​​of the reverse-packed third channel become zero. How to receive mesh data.

13. In Article 10 The steps for reverse packing the above displacement vector coefficients are: Repack the values ​​in the first channel with the values ​​in the reverse-packed first channel, Repack the values ​​in the second channel with the values ​​in the reverse-packed second channel, Including reverse packing the values ​​in the third channel with the values ​​in the reverse packed third channel. How to receive mesh data.

14. In paragraph 8, The bitstream includes at least one of information indicating an encoding type of a displacement vector, a method indicating a packing type for each LoD of a displacement vector, information indicating a scanning order of a displacement vector packing, information indicating a displacement vector video packing format, or information indicating a position of a displacement vector coefficient on a local coordinate system. How to receive mesh data.

15. Memory; and A processor configured to perform one or more instructions stored in the memory, wherein the processor comprises: Receive a bitstream containing mesh data; and Decoding the above mesh data; performing; Mesh data receiving device.

Citation Information

Patent Citations

  • APPARATUS AND METHOD FOR coding three dimentional mesh

    KR101669873B1

  • Apparatus and method for encoding 3D mesh, and apparatus and method for decoding 3D mesh

    KR1020120085134A

  • Varifocal lens composition comprising internal plasticized PVC and varifocal lens having the same

    KR102556641B1

  • Methods and apparatuses for dynamic mesh compression

    US20230014820A1

  • 3D data transmission device, 3D data transmission method, 3D data reception device, and 3D data reception method

    WO2023014086A1