Mesh data transmission apparatus, mesh data transmission method, mesh data reception apparatus, and mesh data reception method

By performing mesh data encoding and decoding on point cloud data and utilizing V-Mesh compression technology, the problems of high throughput and encoding complexity in point cloud data transmission were solved, achieving efficient point cloud data transmission and autonomous driving services.

CN121128181APending Publication Date: 2025-12-12LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480028039.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-25
Filing Date
2024-04-22
Publication Date
2025-12-12

Smart Images

  • Figure CN121128181A_ABST
    Figure CN121128181A_ABST
Patent Text Reader

Abstract

A grid data transmission method according to an embodiment may comprise the steps of: encoding grid data; and transmitting the bitstream including the mesh data. A mesh data receiving method according to an embodiment may comprise the steps of: receiving a bitstream including mesh data; and decoding the grid data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments provide a method for providing point cloud content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and self-driving services to a user. BACKGROUND

[0002] A point cloud is a collection of points in a three-dimensional (3D) space. Since the number of points in the 3D space is large, it is difficult to generate point cloud data.

[0003] Data for transmitting and receiving a point cloud requires a large throughput. SUMMARY

[0004] TECHNICAL PROBLEM

[0005] An object of the disclosure conceived to solve the above-described problem is to provide a point cloud data transmitting apparatus, a point cloud data transmitting method, a point cloud data receiving apparatus, and a point cloud data receiving method for efficiently transmitting and receiving point cloud data.

[0006] Another object of the disclosure is to provide a point cloud data transmitting apparatus, a point cloud data transmitting method, a point cloud data receiving apparatus, and a point cloud data receiving method for solving a delay and encoding / decoding complexity.

[0007] Embodiments of the disclosure are not limited to the above-described objects, and the scope of the embodiments can be extended to other objects that can be inferred by those skilled in the art based on the entire contents of the disclosure.

[0008] TECHNICAL SOLUTION

[0009] To achieve these objects and other advantages, a method of transmitting mesh data according to an embodiment can include encoding mesh data, and transmitting a bitstream including the mesh data. A method of receiving mesh data according to an embodiment can include receiving a bitstream including the mesh data, and decoding the mesh data.

[0010] ADVANTAGEOUS EFFECTS

[0011] The point cloud data transmitting method, the point cloud data transmitting apparatus, the point cloud data receiving method, and the point cloud data receiving apparatus according to the embodiments can provide a high-quality point cloud service.

[0012] The point cloud data transmitting method, the point cloud data transmitting apparatus, the point cloud data receiving method, and the point cloud data receiving apparatus according to the embodiments can implement various video codec methods.

[0013] The point cloud data transmitting method, the point cloud data transmitting apparatus, the point cloud data receiving method, and the point cloud data receiving apparatus according to the embodiments can provide general point cloud content such as autonomous driving services. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the present disclosure and serve to explain the principles of the present disclosure. To facilitate a better understanding of the various embodiments described below, reference should be made to the drawings where like reference numerals designate identical items in the figures. In the drawings:

[0015] Figure 1 FIG. illustrates a system for providing dynamic mesh content according to an embodiment;

[0016] Figure 2 FIG. illustrates a V-MESH compression method according to an embodiment;

[0017] Figure 3 FIG. illustrates a pre-processing in V-MESH compression according to an embodiment;

[0018] Figure 4 FIG. illustrates a mid-edge subdivision method according to an embodiment;

[0019] Figure 5 FIG. illustrates a displacement generation process according to an embodiment;

[0020] Figure 6 FIG. illustrates an intra-frame encoding process in V-MESH compression method according to an embodiment;

[0021] Figure 7 FIG. illustrates an inter-frame encoding process for V-MESH data according to an embodiment;

[0022] Figure 8 FIG. illustrates a lifting transform process on displacement according to an embodiment;

[0023] Figure 9 FIG. illustrates a process of packing transform coefficients into 2D image according to an embodiment;

[0024] Figure 10 FIG. illustrates an attribute transfer process in V-MESH compression method according to an embodiment;

[0025] Figure 11 FIG. illustrates an intra-frame decoding process in V-MESH compression method according to an embodiment;

[0026] Figure 12 FIG. illustrates a V-MESH compression inter-frame decoding process according to an embodiment;

[0027] Figure 13 FIG. illustrates a point cloud data transmitting device according to an embodiment;

[0028] Figure 14 FIG. illustrates a point cloud data receiving apparatus according to an embodiment;

[0029] Figure 15 FIG. illustrates displacement information and attribute information for each scalable level of detail according to an embodiment;

[0030] Figure 16 FIG. illustrates a tile according to an embodiment;

[0031] Figure 17 FIG. illustrates tile-based packing of displacement information and attribute information according to an embodiment;

[0032] Figure 18 and Figure 19 FIG. illustrates syntax of tile-related parameter information according to an embodiment;

[0033] Figure 20 FIG. illustrates a mesh data encoding method according to an embodiment; and

[0034] Figure 21 FIG. illustrates a mesh data decoding method according to an embodiment. DETAILED DESCRIPTION

[0035] Reference will now be made in detail embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The detailed description, which will be given below with reference to the accompanying drawings, is intended to explain exemplary embodiments of the present disclosure, rather than to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without such specific details.

[0036] Although most of the terms used in the present disclosure are selected from general terms widely used in the art, some terms are arbitrarily selected by the applicant, and the meanings thereof are described in detail in the following description as needed. Therefore, the present disclosure should be understood based on the intended meanings of the terms rather than their simple names or meanings.

[0037] Figure 1 FIG. illustrates a system for providing dynamic mesh content according to an embodiment.

[0038] Figure 1 The system in FIG. includes a point cloud data transmitting apparatus 100 and a point cloud data receiving apparatus 110. The point cloud data transmitting apparatus can include a dynamic mesh video acquisition unit (or portion) 101, a dynamic mesh video encoder 102, a file / segment encapsulator 103, and a transmitter 104. The point cloud data receiving apparatus 110 can include a receiver 111, a file / segment decapsulator 112, a dynamic mesh video decoder 113, and a renderer 114. Figure 1Each component in the system can correspond to hardware, software, a processor, and / or a combination thereof. In the following description, a point cloud data transmitting device according to an embodiment can be interpreted to refer to the transmitting device 100 or a dynamic mesh video encoder (hereinafter, referred to as an encoder) 102. A point cloud data receiving device according to an embodiment can be interpreted to refer to the receiving device 110 or a dynamic mesh video decoder (hereinafter, referred to as a decoder) 113.

[0039] Figure 1 The system of FIG. 1 can perform video-based dynamic mesh compression and decompression.

[0040] With advancements in 3D capture, modeling, and rendering, users are allowed to access 3D content across multiple platforms and devices in various forms such as AR, XR, metaverse, and hologram, among others, in an enhanced manner. 3D content is increasingly becoming fine and realistic in the representation of objects in order to provide immersive experiences to users. However, this requires a large amount of data for the generation and use of 3D models. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system using mesh content.

[0041] First, a method of compressing dynamic mesh data starts with the Video-based Point Cloud Compression (V-PCC) standard technique. Point cloud data is data having color information in the coordinates (X, Y, Z) of vertices. Mesh data refers to vertex information including inter-vertex connection information. Content can be initially created in the form of mesh data. Connection information can be added to point cloud data, and point cloud data can be transformed into mesh data.

[0042] Currently, the MPEG standards group has defined two data types for dynamic mesh data: mesh data of category 1 having a texture map as color information and mesh data of category 2 having vertex color as color information.

[0043] A mesh coding standard for category 1 data is currently in progress, and standardization for category 2 data is expected to follow. The overall process for providing mesh content services can include a capture, encoding, transmission, decoding, rendering, and / or feedback process, as shown in FIG. 2. Figure 1

[0044] To provide mesh content services, 3D data acquired through a plurality of cameras or a special camera can be processed into a mesh data type through a series of steps to generate a video. The generated mesh video can be transmitted through a series of operations, and the receiving side can process the received data back into a mesh video for rendering. Through this process, a mesh video can be provided to a user, allowing the user to interactively utilize mesh content according to their intent.​

[0045] The mesh compression system can include a transmitting device and a receiving device. The transmitting device can encode a mesh video to output a bitstream, which can be delivered to the receiving device in the form of a file or a stream (stream segment) through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0046] The transmitting device can illustratively include a mesh video acquisition unit, a mesh video encoder, and a transmitter. The receiving device can illustratively include a receiver, a mesh video decoder, and a renderer. The encoder can be referred to as a mesh video / image / picture / frame encoding device. The decoder can be referred to as a mesh video / image / picture / frame decoding device. The transmitter can be included in the mesh video encoder, and the receiver can be included in the mesh video decoder. The renderer can include a display, and the renderer and / or the display can be configured as a separate device or an external component. The transmitting device and the receiving device can further include separate internal or external modules / units / components for a feedback process.

[0047] A mesh data represents a surface of an object using a plurality of polygons. Each polygon is defined by vertices in 3D space and connectivity information indicating how the vertices are connected. In addition, vertex attributes such as color and normal vector can be included in the data. Mapping information that allows the surface of the mesh to be mapped onto a 2D plane can also be included in the attributes of the mesh. The mapping is usually described using a set of parametric coordinates (called UV coordinates or texture coordinates) related to the vertices of the mesh. The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as texture, normal, and displacement.

[0048] The mesh video acquisition unit can include processing 3D object data acquired through a camera or the like into a mesh data type having the above-described attributes through a series of operations and generating a video composed of mesh data. In the mesh video, the attributes (e.g., vertices, polygons, connections between vertices, colors, and normals) of the mesh can vary over time. The mesh video having attributes and connectivity information that vary over time is referred to as a dynamic mesh video.

[0049] A mesh video encoder can encode an input mesh video into one or more video streams. A video can contain multiple frames, each of which can correspond to a still image / picture. In this disclosure, a mesh video can include a mesh image / frame / picture. The term "mesh video" can be used interchangeably with mesh image / frame / picture. The mesh video encoder can perform a video-based dynamic mesh (V-Mesh) compression process. For compression and encoding efficiency, the mesh video encoder can perform a series of processes such as prediction, transform, quantization, and entropy encoding. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0050] A packaging processor (file / segment packaging module) can package the encoded mesh video data and / or mesh video related metadata in the form of a file or the like. The mesh video related metadata can be received from a metadata processor. The metadata processor can be included in the mesh video encoder or can be configured as a separate component / module. The packaging processor can package the data into a file format such as ISOBMFF or process it into a form such as a DASH segment. According to embodiments, the packaging processor can include mesh video related metadata in a file format. For example, the mesh video metadata can be included in a level box in the ISOBMFF file format or as data on a separate track in the file. In some embodiments, the packaging processor can package the mesh video related metadata into a file.

[0051] A transmission processor can apply processing to the packaged mesh video data based on a file format for transmission. The transmission processor can be included in a transmitter or implemented as a separate component / module. The transmission processor can process the mesh video data according to any transmission protocol. The processing for transmission can include transmission processing via a broadcast network and transmission processing via broadband. In some embodiments, the transmission processor can receive mesh video related metadata as well as mesh video data from the metadata processor and process it for transmission.

[0052] A transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter can include elements that generate a media file through a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiver can extract the bitstream and transfer it to a decoding device.

[0053] The receiver can receive the mesh video data transmitted by the mesh video transmitting apparatus. According to a channel for transmission, the receiver can receive the mesh video data via a broadcast network or a broadband network, or can receive the mesh video data via a digital storage medium.

[0054] The reception processor can perform processing on the received mesh video data according to a transmission protocol. The reception processor can be included in the receiver, or can be configured as a separate component / module. To correspond to the processing performed at the transmission side for transmission, the reception processor can perform the inverse of the operations of the above-described transmission processor. The reception processor can transfer the acquired mesh video data to the decapsulation processor and the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the reception processor can be in the form of a signaling table.

[0055] The decapsulation processor (file / segment decapsulation module) can decapsulate the mesh video data in the form of a file received from the reception processor. The decapsulation processor can decapsulate the file according to ISOBMFF or the like to acquire a mesh video bitstream or mesh video-related metadata (metadata bitstream). The acquired mesh video bitstream can be delivered to the mesh video decoder, and the acquired mesh video-related metadata (metadata bitstream) can be delivered to the metadata processor. The mesh video bitstream can include metadata (metadata bitstream). The metadata processor can be included in the mesh video decoder, or can be configured as a separate component / module. The mesh video-related metadata acquired by the decapsulation processor can be in the form of a box or track in a file format. The decapsulation processor can receive metadata required for decapsulation from the metadata processor when needed. The mesh video-related metadata can be delivered to the mesh video decoder for a mesh video decoding process, or to the Tenderer for a mesh video rendering process.

[0056] The mesh video decoder can receive an input bitstream and perform operations corresponding to the operations of the mesh video encoder to decode a video / image. The decoded mesh video / image can be displayed through a display of the Tenderer. A user can view all or a part of the rendering result through a VR / AR display, a general display, or the like.

[0057] The feedback process can include transmitting various types of feedback information obtainable during the rendering / displaying operation to the decoders on the transmitting side or the receiving side. The feedback process can provide interactivity when consuming the mesh video. In some embodiments, the feedback process can include transmitting head orientation information, viewport information indicating the area the user is currently viewing, etc. In some embodiments, the user can interact with objects implemented in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction during the feedback process can be transmitted to the transmitting side or the service provider. In some embodiments, the feedback process can be skipped.

[0058] The head orientation information can refer to information about the user's head position, angle, movement, etc. Based on this information, information about the area within the mesh video that the user is currently viewing (i.e., viewport information) can be calculated.

[0059] The viewport information can be information about the area within the mesh video that the user is currently viewing. Gaze analysis can be performed based on this information to determine how the user consumes the mesh video, how often the user looks at a particular area of the mesh video, etc. The gaze analysis can be performed on the receiving side, and the result can be transmitted to the transmitting side through the feedback channel. A device such as a VR / AR / MR display can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0060] In some embodiments, the above-described feedback information can not only be transmitted to the transmitter but also consumed on the receiving side. In other words, operations such as decoding and rendering can be performed on the receiving side based on the above-described feedback information. For example, based on the head orientation information and / or the viewport information, only the mesh video of the area that the user is currently viewing can be preferentially decoded and rendered.

[0061] The present disclosure relates to dynamic mesh video compression as described above. The methods / embodiments disclosed herein can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or any next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connectivity information and attributes that change over time. It can perform lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, and AR / VR.

[0062] The dynamic mesh video compression method described below is based on the V-Mesh method of the MPEG.

[0063] In the present disclosure, a picture / frame can generally refer to a unit representing one image at a particular time.

[0064] A pixel or a pel can refer to the smallest unit constituting a picture (or a video). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally refer to a pixel or a pixel value. It can refer to only a pixel / pixel value of a luma component, or only a pixel / pixel value of a chroma component, or only a pixel / pixel value of a depth component.

[0065] A unit can denote a basic unit of image processing. A unit can include at least one of a specific region of a picture and / or information related to the region. In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an MxN block can include a set (or an array) of samples (or a sample array) or transform coefficients consisting of M columns and N rows.

[0066] Figure 1 The encoding process of the V-Mesh is performed as follows.

[0067] A compression method based on video-based dynamic mesh compression (V-Mesh) can provide a method of compressing dynamic mesh video data based on a 2D video codec (e.g., High Efficiency Video Surface (HEVC) and Versatile Video Coding (VVC)). In the V-Mesh compression process, the following data is received as input and compressed.

[0068] Input mesh: includes 3D coordinates (geometry) of the vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection between the vertices constituting the surface. The mesh surface can be represented by a triangle or other polygon, and the connection information between the vertices constituting the surface is stored according to a predetermined shape. The input mesh can be stored in an OBJ file format.

[0069] Attribute map (texture map can also be used interchangeably hereinafter): contains information about attributes (color, normal, displacement, etc.) of the mesh, and stores data in the form of mapping of the mesh surface to a 2D image. Mapping indicating which part (surface or vertex) of the mesh corresponds to each piece of data in the attribute map is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be referred to as an attribute map video (or simply an attribute). The attribute map in the V-Mesh compression method mainly contains color information of the mesh, and is stored in an image file format (PNG, BMP, etc.).

[0070] Material library file: contains information on material attributes used in the mesh, particularly information linking the input mesh to the corresponding attribute map. It is stored in a Wavefront Material Template Library (MTL) file format.

[0071] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0072] Base mesh: Decimate the input mesh through a pre-processing procedure, using the minimum number of vertices determined according to user criteria to represent the object in the input mesh.

[0073] Displacement: Displacement information to represent the input mesh as similar as possible using the base mesh, expressed in 3D coordinates.

[0074] Atlas information: Metadata needed to reconstruct the mesh using the base mesh, displacement, and attribute map information. It can be generated and utilized in sub-units (sub-meshes, patches, etc.) that make up the mesh.

[0075] In combination Figures 2 to 7 A method of encoding mesh position information (or vertices) is described, in combination Figures 6 to 10 A method of reconstructing mesh position information to encode attribute information (attribute map) is described.

[0076] Figure 2 A V-MESH compression method according to an embodiment is illustrated.

[0077] Figure 2 An encoding process of Figure 1 is illustrated, where the encoding process can include a pre-processing procedure and an encoding process. As shown in Figure 2 , an encoder of Figure 1 can include a pre-processing period 200 and an encoder 201. Figure 1 A transmitting device of Figure 1 can be broadly referred to as an encoder, and Figure 2 A dynamic mesh video encoder of Figure 2 can be referred to as an encoder. As shown in Figure 2 , a V-Mesh compression method can include a pre-processing and an encoding 201. Figure 2 A pre-processing period of can be located at the front end of an encoder of

[0078] A pre-processing period and an encoder of can be referred to as a single encoder.

[0079] A pre-processing period can receive a static dynamic mesh and / or an attribute map. The pre-processing period can generate a base mesh and / or a displacement through pre-processing. The pre-processing period can receive feedback information from the encoder, and can generate the base mesh and / or the displacement based on the feedback information.

[0080] Figure 3 An encoder can receive a base mesh, a displacement, a static dynamic mesh, and / or an attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.

[0081] Figure 3 An encoding process of Figure 2configuration and operation of a pre-processor of the present disclosure.

[0082] Figure 3 A process of performing pre-processing on an input mesh is illustrated. The pre-processing 200 can include four operations: 1) GoF generation, 2) mesh decimation, 3) UV parameterization, and 4) fitting subdivision surface (300). The pre-processing stage 200 can receive an input mesh, generate a displacement and / or a base mesh, and deliver it to an encoder 201. The pre-processing stage 200 can deliver GoF information related to GoF generation to the encoder 201.

[0083] Each operation in the pre-processing stage 200 is described below. Figure 3

[0084] GoF generation: a process of generating a reference structure of mesh data. When a mesh of a previous frame and a mesh of a current frame have the same number of vertices, the same number of texture coordinates, the same vertex connection information, and the same texture coordinate connection information, the previous frame can be set as a reference frame. In other words, if only vertex coordinate values are different between a current input mesh and a reference input mesh, inter-frame encoding can be performed. Otherwise, intra-frame encoding can be performed on the frame.

[0085] Mesh decimation: a process of simplifying an input mesh to create a simplified mesh (referred to as a base mesh). Selected vertices to be removed can be selected from an original mesh based on a user-defined criterion, and then the selected vertices and triangles connected to the selected vertices can be removed.

[0086] In a process of performing mesh decimation, a voxelized input mesh, a target triangle ratio (TTR), and minimum triangle component (CCCount) information can be delivered as inputs, and a decimated mesh can be obtained as an output. In the process, connected triangle components smaller than a set minimum triangle component (CCCount) can be removed.

[0087] UV parameterization: a process of mapping a 3D surface to a texture domain for a decimated mesh. Parameterization can be performed using a UVAtlas tool. The process generates mapping information indicating where each vertex of a decimated mesh can be mapped on a 2D image. The mapping information is represented as texture coordinates and stored, and a final base mesh is generated through the process.

[0088] Fitting subdivision surface: an operation of performing subdivision on a decimated mesh. A user-defined method such as a mid-edge method can be applied as a subdivision method. The fitting process is performed so that an input mesh and a subdivided mesh are similar to each other.

[0089] Figure 4 A mid-edge subdivision method according to an embodiment is illustrated.

[0090] ​Figure 4 FIG. illustrates a fitting subdivision surface with respect to Figure 3 A fitting subdivision surface is described. FIG. illustrates a fitting subdivision surface with respect to Figure 4 An original mesh, which contains four vertices, is subdivided to create a sub-mesh. The sub-mesh can be created by creating new vertices in the middle of edges between vertices.

[0091] Once the fitting subdivision mesh is generated, displacement is calculated based on the result and the previously compressed and decoded base mesh (hereinafter referred to as reconstructed base mesh). In other words, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The result is the displacement of each vertex between the position difference of each vertex in the fitting subdivision mesh. Since the displacement represents the position difference in 3D space, it is represented as a value in the (x, y, z) space in the Cartesian coordinate system. According to a user input parameter, the (x, y, z) coordinate value can be converted into a (normal, tangent, bitangent) coordinate value in the local coordinate system.

[0092] Figure 5 FIG. illustrates a displacement generation process according to an embodiment.

[0093] Figure 5 FIG. illustrates how to calculate displacement for a fitting subdivision surface 300 as described with respect to Figure 4 FIG. illustrates how to calculate displacement for a fitting subdivision surface 300 as described with respect to

[0094] An encoder and / or pre-processor according to an embodiment can include 1) a subdivider, 2) a local coordinate system calculator, and 3) a displacement vector calculator. The subdivider can receive a reconstructed base mesh and generate a subdivided reconstructed base mesh. The local coordinate system calculator can receive a fitting subdivision mesh and a subdivided reconstructed base mesh, and can transform the coordinate system related to the mesh into a local coordinate system. The local coordinate system calculation can be optional. The displacement calculator calculates the position difference between the fitting subdivision mesh and the subdivided reconstructed base mesh. For example, it can generate the position difference between vertices in the two input meshes. The position difference between vertices is displacement.

[0095] A point cloud data transmission method and apparatus according to an embodiment can encode a point cloud as follows. Point cloud data (which can be simply referred to as a point cloud) according to an embodiment can refer to data including vertex coordinates and color information. A point cloud is a term including mesh data. Here, the terms point cloud and point cloud data can be used interchangeably.

[0096] According to an embodiment, a V-Mesh compression (reconstruction) method can include intra-frame encoding (I) Figure 6 ) and inter-frame encoding (P) Figure 7 ).

[0097] Based on the results generated by the GoF described above, intra-frame encoding or inter-frame encoding is performed. In intra-frame encoding, the data to be compressed can be the base mesh, the displacement, the attribute map, etc. In inter-frame encoding, the data to be compressed can be the displacement, the attribute map, and the motion field between the reference base mesh and the current base mesh.

[0098] Figure 6 An intra-frame encoding process in the V-MESH compression method according to an embodiment is illustrated.

[0099] Figure 6 The encoding process refines the encoding of Figure 1 . That is, it represents the configuration of the encoder when the encoding of Figure 1 is intra-frame encoding. Figure 6 The encoder of

[0100] The pre-processor can receive an input mesh and perform the pre-processing described above. A base mesh and / or a fitting subdivision mesh can be generated by the pre-processing. A quantizer can quantize the base mesh and / or the fitting subdivision mesh. A static mesh encoder can encode a static mesh. The static mesh encoder can generate a bitstream containing the encoded base mesh. A static mesh decoder can decode the encoded static mesh. A dequantizer can dequantize the quantized static mesh. A displacement calculator can receive the reconstructed static mesh and can generate a displacement that is a position difference based on the fitting subdivision mesh. A forward linear lifting unit can receive the displacement and generate lifting coefficients. A quantizer can quantize the lifting coefficients. An image packer can pack images based on the quantized lifting coefficients. A video decoder decodes the encoded video. An image unpacker can unpack the packed images. A dequantizer can dequantize the images. An inverse linear lifting unit applies inverse lifting to the images to generate reconstructed displacements. A mesh reconstructor reconstructs a deformed mesh based on the reconstructed displacements and the reconstructed base mesh. An attribute transfer receives an input mesh and / or an input attribute map and generates an attribute map based on the reconstructed deformed mesh. A push-pull padding unit can pad data to the attribute map based on a push-pull method. A color space converter can convert the space of color components that are attributes. A video encoder can encode the attributes. A multiplexer can multiplex the compressed base mesh, the compressed displacement, and the compressed attributes to generate a bitstream.

[0101] Figure 7 An inter-frame encoding process in the V-MESH compression method according to an embodiment is illustrated.

[0102] Figure 7 The encoding process details the encoding of Figure 1 . That is, it represents the configuration of the encoder when the encoding of Figure 1 is inter-frame encoding.Figure 7 An encoder of the encoding method can include a pre-processor 200 and / or an encoder 201.

[0103] For the encoding operation corresponding to the inter-frame-based encoding in Figure 6 For the components of the encoding operation of the inter-frame-based encoding in Figure 7 See the description of Figure 7 For the inter-frame-based encoding in Figure 7 A motion encoder can encode the motion based on the reconstructed quantized reference base mesh. A base mesh reconstructor can reconstruct the base mesh based on the reconstructed quantized reference base mesh.

[0104] Figure 6 An encoder in compresses the base mesh, displacement, and attributes within a frame to generate a bitstream, while Figure 7 An encoder in compresses the motion, displacement, and attributes between a current frame and a reference frame to generate a bitstream.

[0105] The encoding method according to embodiments includes base mesh encoding (intra-frame encoding). When intra-frame encoding is performed on a current input mesh frame, the base mesh generated during pre-processing can be quantized and then encoded using static mesh compression techniques. For example, in the V-Mesh compression method, Draco techniques are applied, and the targets to be compressed include vertex position information related to the base mesh, mapping information (texture coordinates), vertex connection data, etc.

[0106] The encoding method according to embodiments can include motion field encoding (also referred to as inter-frame encoding). When a reference mesh and a current input mesh have a one-to-one vertex correspondence relationship, and they only differ in terms of vertex position information, inter-frame encoding can be performed. When inter-frame encoding is performed, the base mesh can not be compressed. Instead, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field, can be calculated and encoded. The reference base mesh is the result of quantized decoded base mesh data, and is determined by the reference frame index determined in the GoF generation. The motion field can be encoded as is. Alternatively, a predicted motion field can be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and a residual motion field can be encoded as the difference between the value of the predicted motion field and the value of the motion field of the current vertex. The value of this field can be encoded using entropy encoding. The process of encoding the displacement and attribute maps is the same as the structure of the intra-frame encoding method except for the base mesh encoding, except for the motion field encoding in inter-frame encoding.

[0107] Figure 8 A lifting transform process for displacement according to embodiments is illustrated.

[0108] Figure 9 A process of packing transform coefficients into a 2D image according to embodiments is illustrated.

[0109] Figure 8 and Figure 9 respectively illustrate Figure 6 and Figure 7 the process of transforming the displacement and packing the transform coefficients in the encoding process.

[0110] The encoding method according to embodiments comprises displacement encoding.

[0111] After the base mesh encoding and / or the motion field encoding, a reconstructed base mesh can be generated by reconstruction and inverse quantization, and the displacement can be calculated between the tessellated result of the reconstructed base mesh and the fitted tessellated mesh generated by fitting the tessellated surface. A data transform process (e.g. wavelet transform) can be applied to the displacement information for efficient encoding.

[0112] Figure 8 The process of transforming the displacement information in V-Mesh using lifting transform is illustrated. The transform coefficients generated by the transform process are quantized and then packed into 2D images. The transform coefficients can be organized into blocks, one block for every 256 (= 16 x 16) cells. Each block can be packed in z-scan order. The number of rows in a block is fixed to 16, but the number of columns in a block can be determined by the number of vertices in the tessellated base mesh. Within a block, the transform coefficients can be sorted by Morton code and packed. For the packed images, a displacement video can be generated per GoF. The displacement video can be encoded using a conventional video compression codec.

[0113] Reference is made to Figure 8, the base mesh (original) can include vertices and edges for LoD0. The first subdivision mesh generated by subdividing the base mesh includes vertices generated by further splitting the edges of the base mesh. The first subdivision mesh contains vertices for LoD0 and vertices for LoD1. LoD1 includes subdivided vertices and vertices from the base mesh (LoD0). The first subdivision mesh can be split to generate a second subdivision mesh. The second subdivision mesh contains LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 containing vertices generated additionally from LoD0, and vertices further split from LoD1. LoD is a level of detail, which indicates the degree of detail. As the index of the level increases, the distance between vertices is shortened, and the level of detail is increased. LoD N contains vertices contained in LoD N-1. In the case where vertices are further split by subdivision, the previous vertices v1 and v2 and the subdivided vertex v can be encoded based on a prediction and / or update method. Instead of encoding the information of the current LoD N as it is, a residual with respect to the previous LoD N-1 can be generated. Thus, the mesh can be encoded using the residual to reduce the bitstream size. The prediction process refers to the operation of predicting the current vertex v from the previous vertices v1 and v2. Since adjacent subdivision meshes have similar data, this characteristic can be utilized for efficient encoding. The current vertex position information is predicted from the residual of the previous vertex position information, and the previous vertex position information is updated based on the residual.

[0114] Referring to Figure 9 , the vertices have coefficients generated by lifting transformation. The vertex coefficients related to the lifting transformation can be packed into an image and then encoded.

[0115] Figure 10 An attribute transfer process in a V-MESH compression method according to an embodiment is illustrated.

[0116] Figure 10 Detailed operations of attribute transfer in encoding of Figure 6 , Figure 7 , etc. are illustrated.

[0117] Encoding according to an embodiment includes attribute map encoding.

[0118] Information about the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. The input mesh compressed in the encoding process is reconstructed through base mesh decoding (intra-frame), motion field decoding (inter-frame), and displacement video decoding, and as a result of reconstruction, the reconstructed morphed mesh (hereinafter referred to as reconstructed morphed mesh) is used to compress the input attribute map, as Figure 6 and Figure 7The reconstructed morphing mesh has position information about vertices, texture coordinates, and corresponding connectivity information, but no color information corresponding to the texture coordinates. Thus, as shown in FIG. 2, the new attribute map is generated by an attribute transfer process. Figure 10 In the V-Mesh compression method, as shown in FIG. 2, a new attribute map with color information corresponding to the texture coordinates of the reconstructed morphing mesh is regenerated by an attribute transfer process.

[0119] The attribute transfer first checks for each point P(u, v) in the 2D texture domain whether the corresponding point is inside a texture triangle of the reconstructed morphing mesh. When the corresponding vertex is in a texture triangle T, the attribute transfer computes barycentric coordinates (a, b, g) of P(u, v) from triangle T. Then, it computes 3D coordinates M(x, y, z) of P(u, v) based on the 3D vertex positions of triangle T and (a, b, g). In the input mesh domain, the vertex coordinates M'(x', y', z') and the triangle T' containing this point corresponding to the position closest to the computed M(x, y, z) are searched. Then, barycentric coordinates (a', b', g') of M'(x', y', z') in triangle T' are computed. Based on the texture coordinates corresponding to the three vertices of triangle T' and (a', b', g'), texture coordinates (u', v') are computed and the color information corresponding to the coordinates is searched in the input attribute map. The color information found in this way is then assigned to the (u, v) pixel position in the new input attribute map. If P(u, v) does not belong to any triangle, a fill algorithm (e.g., a flood fill algorithm) is used to fill the pixel at this position in the new input attribute map with a color value.

[0120] The new attribute map generated by the attribute transfer is bundled into the GoF to construct an attribute map video, which is compressed using a video codec.

[0121] From Figure 10 The reference relationships between the input mesh, the input attribute map, the reconstructed mesh, and the generated attribute map can be seen.

[0122] Figure 1 The decoding process of the V-Mesh compression method can perform the inverse process of the encoding process of Figure 1 In particular, the decoding process is performed as disclosed below.

[0123] Figure 11 FIG. 1 illustrates an intra-frame decoding process of the V-MESH compression method according to an embodiment.

[0124] Figure 11 FIG. 1 illustrates a configuration and operation of a decoder of a receiving device of Figure 1

[0125] Figure 11 ​An intra decoding process of the V-Mesh technique according to the embodiment is shown. First, the bitstream can be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information about the mesh such as V3C / V-PCC.

[0126] For example, the mesh substream can be decoded by a decoder of a static mesh codec used in encoding such as Google Draco. As a result, connectivity information, vertex geometry information, vertex texture coordinates, and the like related to the base mesh can be reconstructed. The displacement substream can be decoded into a displacement video by a decoder of a video compression codec used in encoding. Then, image unpacking, inverse quantization, and inverse transform are performed to reconstruct the displacement information of each vertex. Subsequently, inverse quantization is applied to the reconstructed base mesh, and then its result is combined with the reconstructed displacement information to generate the final decoded mesh.

[0127] The attribute map substream is decoded by a decoder of a video compression codec used in encoding, and then a final attribute map is reconstructed through color format conversion and the like.

[0128] The reconstructed decoded mesh and the decoded attribute map can be utilized at the receiving side as final mesh data that can be used by a user.

[0129] Reference Figure 11 The bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term substream is interpreted to refer to a partial bitstream included in the bitstream. The bitstream contains patch information (data), mesh information (data), displacement information (data), and attribute map information (data).

[0130] The decoder performs intra decoding as follows. A static mesh decoder decodes the mesh to generate a reconstructed quantized base mesh, and an inverse quantizer inversely applies a quantization parameter of a quantizer to generate a reconstructed base mesh. A video decoder decodes the displacement, an unpacker unpacks the decoded video image, and an inverse quantizer inversely quantizes the quantized image. An inverse linear lifting unit applies a lifting transform in the inverse process of the encoder to generate a reconstructed displacement. A mesh reconstructor 630 generates a deformed mesh based on the base mesh and the displacement. A video decoder decodes the attribute map, and a color transformer transforms the color format and / or space to generate a decoded attribute map.

[0131] Figure 12 An inter decoding process of the V-MESH technique is illustrated.

[0132] Figure 12 The configuration and operation of a decoder of a receiving device of Figure 1 are illustrated.

[0133] Figure 12The inter-frame decoding process of V-MESH technique is illustrated. First, the input bitstream can be separated into motion sub-stream, displacement sub-stream, attribute sub-stream and sub-stream containing patch information about the mesh (e.g. V3C / V-PCC).

[0134] The motion sub-stream is decoded by entropy decoding and inverse prediction, and the reconstructed motion information is combined with the pre-reconstructed and stored reference base mesh to generate the reconstructed quantized base mesh for the current frame. Then, inverse quantization is applied to the reconstructed quantized base mesh, and its result is combined with the displacement information reconstructed using the same method as the intra-frame decoding described above to generate the final decoded mesh. The reconstructed decoded mesh and the decoded attribute map can be utilized at the receiving side as the final mesh data that can be used by the user.

[0135] Referring to Figure 12 , the bitstream contains motion information, displacement and attribute map. Since inter-frame decoding is performed, the process further includes decoding the inter-frame motion information. The reconstructed base mesh is generated by decoding the motion information and generating the reconstructed quantized base mesh for the motion based on the reference base mesh. For Figure 12 , the same operations as Figure 11 are referred to the description of Figure 11 .

[0136] Figure 13 A point cloud data transmitting device according to an embodiment is illustrated.

[0137] Figure 13 Corresponding to the transmitting device 100 or dynamic mesh video encoder 102 of Figure 1 , the encoder (pre-processor and encoder) of Figure 2 and / or the corresponding transmitting encoding device. Figure 13 Each component of corresponds to hardware, software, a processor and / or a combination thereof.

[0138] The operation process for compressing and transmitting dynamic mesh data using V-Mesh compression technique at the transmitting side can be configured as shown in Figure 13 .

[0139] The mesh pre-processor receives the original mesh and generates a decimated mesh. Decimation can be performed based on a target number of vertices or a target number of polygons constituting the mesh. Parameterization can be performed on the decimated mesh to generate texture coordinates and texture connectivity information for each vertex. The mesh information in floating point form can be quantized to fixed point form. The result is a base mesh, which can be encoded by the static mesh encoder. The mesh pre-processor can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connectivity information, texture coordinates, and connectivity information about the texture coordinates including the additional vertices can be generated. A fit subdivision mesh can be generated by adjusting the vertex positions such that the subdivision mesh is similar to the original mesh.

[0140] When intra coding (intra coding) is performed on the mesh frame, the base mesh can be compressed by the static mesh encoder. In this case, connectivity information related to the base mesh, vertex geometry information, vertex texture information, normal information, etc. can be encoded. The base mesh bitstream generated by encoding is transmitted to the multiplexer.

[0141] When inter coding (inter prediction) is performed on the mesh frame, the motion vector encoder can operate to receive the base mesh and the reference reconstructed base mesh as input, calculate the motion vector between the two meshes, and encode its value. In addition, the motion vector encoder can perform connectivity information-based prediction using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by encoding is transmitted to the multiplexer.

[0142] The reconstructed base mesh can be generated from the encoded base mesh and the motion vector by the base mesh reconstructor.

[0143] The displacement vector calculator can perform mesh subdivision on the reconstructed base mesh. The vectors can be calculated as the values of the vertex position difference between the subdivision reconstructed base mesh generated by the pre-processor and the fit subdivision mesh. As a result, as many displacement vectors as the vertices in the subdivision mesh can be calculated. The displacement vector calculator can transform the calculated displacement vectors in the 3D Cartesian coordinate system (3D Cartesian coordinate system) to the local coordinate system based on the normal vector of each vertex.

[0144] The displacement vector video generator can transform the displacement vectors for efficient encoding. According to embodiments, the transformation can be a lifting transform, a wavelet transform, etc. Furthermore, quantization can be performed on the transformed displacement vector values (i.e., transform coefficients). In this case, different quantization parameters can be applied to the axes of the transform coefficients, respectively. The quantization parameters can be derived by convention between the encoder / decoder. After the transformation and quantization, the displacement vector information can be packed into 2D images. The displacement vector video can be generated by grouping the packed 2D images of each frame. The displacement vector video can be generated for each group of frames (GoF) of the input mesh.

[0145] The displacement vector video encoder can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is sent to the multiplexer.

[0146] The reconstructed displacement vectors by the displacement vector reconstructor are reconstructed with the base mesh reconstructed and subdivided by the base mesh reconstructor through the mesh reconstructor. The reconstructed mesh has reconstructed vertices, inter-vertex connectivity information, texture coordinates, and inter-texture coordinate connectivity information.

[0147] The texture map of the reconstructed mesh can be regenerated from the texture map of the original mesh by the texture map video generator. The vertex-by-vertex color information in the texture map of the original mesh can be assigned to the texture coordinates of the reconstructed mesh. The texture map video can be generated by grouping the frame-level regenerated texture maps into GoF.

[0148] The generated texture map video can be encoded by the texture map video encoder using a video compression codec. The encoded generated texture map video bitstream is sent to the multiplexer.

[0149] The generated motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream can be multiplexed into a single bitstream to be sent to the receiving side through the transmitter. Alternatively, for the generated motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream, a file having one or more track data can be generated, or the bitstreams are encapsulated into segments and sent to the receiving side through the transmitter.

[0150] Reference Figure 13, the sender (encoder) can encode the mesh in intra- or inter-frame manner. According to intra-frame encoding, the sending device can generate a base mesh, displacement vectors and texture maps. According to inter-frame encoding, the sending device can generate motion vectors, a base mesh, displacement vectors and texture maps. The texture maps obtained from the data input unit are generated and encoded based on the reconstructed mesh. The displacements are generated and encoded based on the difference in vertex positions between the base mesh and the tessellated mesh. The base mesh is generated by pre-processing, extracting and encoding the original mesh. For motion, the motion vectors are generated for the mesh in the current frame based on a reference base mesh in a previous frame.

[0151] Figure 14 A point cloud data receiving device according to an embodiment is illustrated.

[0152] Figure 14 The receiving device 110 or mesh video decoder 113, Figure 1 corresponds to the decoder and / or the corresponding receiving decoding device of Figure 11 .Each component of the receiving device 110 or mesh video decoder 113, Figure 14 corresponds to hardware, software, a processor and / or a combination thereof. The receiving (decoding) operation of the receiving device 110 or mesh video decoder 113, Figure 14 may follow the inverse process of the corresponding process of the sending (encoding) operation of the sending device 10 or mesh video encoder 11. Figure 13

[0153] The received mesh bitstream is subjected to file / segment un-encapsulation and then de-multiplexed into a compressed motion vector bitstream or a base mesh bitstream, displacement vector bitstream and texture map bitstream.

[0154] In case of applying inter-frame encoding to the current mesh based on the frame header information, the motion vector decoder can decode the motion vector bitstream. The previously decoded motion vector can be used as a predictor and added to the residual motion vector decoded from the bitstream to reconstruct the final motion vector.

[0155] In case of applying intra-frame encoding to the current mesh based on the frame header information, the static mesh decoder can decode the base mesh bitstream to reconstruct connectivity information, vertex geometry information, texture coordinates, normal information, etc. related to the base mesh.

[0156] In case of applying inter-frame encoding to the current mesh, the base mesh reconstructor can add the decoded motion vector to the reference base mesh and perform inverse quantization to reconstruct the current base mesh. In case of applying intra-frame encoding to the current mesh, inverse quantization can be performed on the base mesh decoded by the static mesh decoder to generate the reconstructed base mesh.

[0157] The displacement vector video decoder can decode the displacement vector bitstream into a video bitstream using a video codec.

[0158] The displacement vector reconstructor extracts the displacement vector transform coefficients from the decoded displacement vector video and reconstructs the displacement vector through inverse quantization and inverse transform. If the reconstructed displacement vector is a value in a local coordinate system, an inverse transform to Cartesian coordinates can be performed.

[0159] The mesh reconstructor can subdivide the base mesh to generate additional vertices. Subdivision generates vertex connectivity information, including additional vertices, texture coordinates, and connectivity information about the texture coordinates. The subdivided base mesh can be combined with the reconstructed translation vectors to generate the final reconstructed mesh.

[0160] A texture map video decoder can use a video codec to decode a texture map bitstream into a video bitstream. The reconstructed texture map has color information about each vertex in the reconstructed mesh, and the texture coordinates of each vertex can be used to obtain the vertex's color value from the texture map.

[0161] The reconstructed mesh and texture map are presented to the user through a rendering process using a mesh data renderer, etc.

[0162] Reference Figure 14 The receiving device (decoder) can decode the mesh in an intra-frame or inter-frame manner. In intra-frame decoding, the receiving device can receive the base mesh, translation vectors, and texture map, and decode the reconstructed mesh and texture map to render the mesh data. In inter-frame decoding, the receiving device can receive motion vectors, the base mesh, translation vectors, and texture map, and decode the reconstructed mesh and texture map to render the mesh data.

[0163] The point cloud data transmitting apparatus and method according to the embodiments can encode grid data and transmit a bit stream containing the encoded grid data. The point cloud data receiving apparatus and method according to the embodiments can receive a bit stream containing grid data and decode the grid data. The point cloud data transmitting / receiving method / apparatus according to the embodiments may be referred to as the method / apparatus according to the embodiments. The point cloud data transmitting / receiving method / apparatus may also be referred to as the grid data transmitting / receiving method / apparatus.

[0164] The mesh encoding method and encoder according to the embodiments include and perform... Figure 1 The transmitter 100, the dynamic mesh video acquisition unit 101, the dynamic mesh video encoder 102, the file / segment encapsulator 103, and the transmitter 104 are all included. Figure 2 Preprocessor 200, encoder 201, Figure 6 and Figure 7 encoder, Figure 13 Transmitting equipment Figures 15 to 17 Patch-based grid data encoding Figure 18 andFigure 19 Patch-based parameter generation and Figure 20 The corresponding operation for the encoding.

[0165] The grid decoding method and decoder according to the embodiments include and perform... Figure 1 The receiver 110, receiver 111, file / fragment decapsulator 112, dynamic mesh video decoder 113, and renderer 114 are included. Figure 11 and Figure 12 decoder Figure 14 The receiving device, as Figures 15 to 17 Decoding based on the inverse process of tile encoding Figure 18 and Figure 19 Patch-based parameter parsing and Figure 21 The corresponding operation for decoding.

[0166] The method / apparatus according to the embodiments may include and perform a method for combining and packaging geometric data and texture data based on tiles in a video codec. Geometric data refers to displacement information associated with mesh data, and texture data refers to attribute information associated with mesh data.

[0167] The embodiments relate to video-based dynamic mesh compression (V-Mesh), which is a method for compressing 3D dynamic mesh data using a 2D video codec. The embodiments include scalable mesh coding and decoding methods that apply scalability based on a mesh coding structure based on standard video.

[0168] The scalable mesh coding method according to embodiments includes methods for reflecting scalability in displacement (also referred to as geometric information or displacement information) video and texture map (also referred to as attribute information or texture information) video, which are compressed using a 2D video codec in video-based mesh coding.

[0169] The embodiments include a method for integrating displacement video and texture map information based on tiles and packaging them into a single frame. Tiles may include, for example, HEVC tiles, but are not limited to this.

[0170] Displacement and texture map information associated with the dynamic mesh can be sent at the level of detail (LoD) using a single frame. Based on the characteristics of the tiles, displacement and texture map information can be selectively decoded at the LoD level.

[0171] According to the embodiments, dynamic grid data can be scalably encoded and decoded, and grid data with optimal resolution and quality can be selectively reconstructed by the user or the system according to the user's receiving environment.

[0172] The embodiments relate to V-DMC, a method for compressing 3D dynamic mesh data using a 2D video codec. To efficiently perform scalable mesh coding, the embodiments include a method for integrating displacement video and texture map information based on tiles and packaging it into a single frame (image).

[0173] The embodiments include methods for effectively encoding and decoding dynamic mesh data (referring to mesh data including moving objects or people). The embodiments include methods for compressing and reconstructing dynamic mesh data based on the V-PCC standard, as a method for compressing and reconstructing point cloud data.

[0174] In the current V-DMC standard, displacement information generated during encoding is converted into video for compression using existing 2D video codecs. Texture map information (meaning attribute information, texture maps, or attribute maps) from the input mesh data is also processed into video to be compressed using a 2D video codec. Displacement video and texture map video are generated and encoded based on a user-defined resolution and reconstructed at the same resolution at the receiving end.

[0175] The current V-DMC standard includes scalable encoding and decoding. The video codec (HEVC) is used to represent geometric information through displacement. The video codec (HEVC) is also used for texture map coding. In this state, performing scalable encoded displacement and texture map coding based on LoD to support scalable structures may be inefficient. Furthermore, it may introduce various problems, such as synchronization-related issues.

[0176] Figure 15 The illustration shows the displacement and attribute information for each scalable level of detail according to an embodiment.

[0177] According to the embodiments, the grid encoding method and encoder (corresponding to) Figure 1 The transmitter 100, the dynamic mesh video acquisition unit 101, the dynamic mesh video encoder 102, the file / segment encapsulator 103, and the transmitter 104 are all included. Figure 2 Preprocessor 200, encoder 201, Figure 6 and Figure 7 encoder, Figure 13 Transmitting equipment Figures 15 to 17 Patch-based grid data encoding Figure 18 and Figure 19 Patch-based parameter generation and Figure 20 The encoding can be configured into multiple layers for the displacement and attribute information generated at each level of detail to support scalability.

[0178] In video-based dynamic mesh compression, displacement information can be packed separately into an image frame and encoded with HEVC, and texture map information can be encoded with HEVC into a texture image. The method / device according to the embodiment includes a method of packing displacement information and texture map information (attribute information) of each LoD into a tile in a single frame for more efficient support of scalable structure. An example of data generated for scalable support of V-DMC is shown in Figure 15 .

[0179] For example, when a scalable layer includes one base layer and two enhancement layers, displacement information corresponding to a level in each layer and a texture map corresponding to the level can be packed together into a tile in a frame. The scalable layer can be referred to as a first layer, a second layer, etc.

[0180] A level can mean a unit of a subdivided region of mesh data, such as a level of detail, depth, level, or layer.

[0181] When there is displacement information for three levels and attribute information for three levels, displacement information for a first level and attribute information for the first level can be packed into a frame for a first layer, displacement information for a second level and attribute information for the second level can be packed into a frame for a second layer, and displacement information for a third level and attribute information for the third level can be packed into a frame for a third layer.

[0182] When displacement information for a first level and displacement information for a second level have a small difference, they can be packed together as displacement information for a single level in a frame.

[0183] The operation of packing displacement information (i.e., displacement vectors) can include converting displacement vectors into displacement vector coefficients, quantizing displacement vector coefficients, packing displacement vector coefficients, and encoding displacement vector images / videos.

[0184] The method / device according to the embodiment can present displacement information of mesh data in a level of detail and encode and decode displacement information of mesh data in a level.

[0185] The method / device according to the embodiment can present attribute information of mesh data in a level of detail and encode and decode attribute information of mesh data in a level.

[0186] For scalable encoding / decoding, mesh data can be represented and encoded / decoded based on one or more layers. The layers can include a base layer and / or two or more enhancement layers. The base layer and the enhancement layers can be referred to as a first layer, a second layer, etc.

[0187] The first layer can include attribute information for level 0. The first layer can include displacement information for level 0. According to an embodiment, the first layer can include displacement information for level 0 and level 1 (as shown in Figure 15

[0188] The second layer can include attribute information for level 1. The second layer can include displacement information for level 2.

[0189] The third layer can include attribute information for level 2. The third layer can include displacement information for level 3.

[0190] The blank area within the image (frame) in which each layer of displacement information is packed can be a padding area.

[0191] Here, according to the current V-DMC CTC, the number of geometry (displacement information) LoD (level) is determined, and according to system or user preference, there can be various LoD combinations, including Figure 15 examples in Figure 15 . That is, the LoD level of displacement and the LoD level of texture map do not need to be exactly matched as in

[0192] For example, when there are a texture map of LoD 0, a texture map of LoD 1, a texture map of LoD 2, a texture map of LoD 3, displacement information of LoD R0, displacement information of LoD R1, displacement information of LoD R2, and displacement information of LoD R3, the displacement information of LoD R0 and LoD R1 can be included in the first layer, and the texture map of LoD 0 can be included in the first layer. The displacement information of LoD R1 can be included in the second layer, and the texture map of LoD 1 can be included in the second layer. The displacement information of LoD R2 can be included in the third layer, and the texture map of LoD 2 can be included in the third layer. The displacement information of LoD R3 can be included in the fourth layer, and the texture map of LoD 3 can be included in the fourth layer.

[0193] Figure 16 A tile according to an embodiment is illustrated.

[0194] Figure 16 A frame structure for packing displacement information and attribute information of each layer and each level (LoD) as described in Figure 15 is illustrated. Specifically, an image (frame) can be composed of at least one tile, as shown in Figure 16

[0195] Figure 16 ​​The tiles of the picture can be, for example, HEVC tiles. Embodiments are not limited to HEVC tiles. An area obtained by dividing a rectangular area by vertical and horizontal lines can be referred to as a tile. Tiles can have different sizes. Figure 16 Tiles 0, 1, 2, and 3 are illustrated. Based on these tiles, data included in the tiles can be compressed and reconstructed.

[0196] Figure 17 Tile-based packing of attribute information and displacement information according to an embodiment is illustrated.

[0197] Figure 17 Tile-based packing of attribute information and displacement information according to an embodiment is illustrated. Figure 15 Figure 16 Figure 6 Figure 7 Figure 11 Figure 12

[0198] For data supporting scalable function of V-DMC, attribute information (texture map) and displacement information can be integrated per LoD, and can be packed into each tile based on a tile structure, as shown in Figure 17

[0199] Tile 0 can include attribute information for level 3 of the third layer and displacement information for level 3, and tile 1 can include attribute information for level 2 of the second layer and displacement information for level 2. Tile 2 can include attribute information for level 1 of the first layer and displacement information for level 1, and tile 3 can include attribute information for level 0 of the first layer and displacement information for level 0.

[0200] A remaining area within each tile other than attribute information and displacement information can be a padding area.

[0201] A mesh data decoding method according to an embodiment can be implemented in the following manner:

[0202] For example, assuming that a 3D mesh image service is provided with four different resolutions, each 3D mesh image corresponding to each resolution (LoD) includes texture (attribute) information and geometry data (displacement).

[0203] Depending on a device specification or a network connection condition, an appropriate resolution or LoD is selected.

[0204] ​​​​​​​In this case, for example, when selection is performed only up to LoD 2, tiles corresponding to LoD 0, LoD 1, and LoD 2 are selected and partially decoded.

[0205] Based on the decoding result, a 3D mesh image corresponding to the image quality of LoD 2 is provided as a service.

[0206] Figure 18 and Figure 19 A syntax of tile-related parameter information according to an embodiment is illustrated.

[0207] A mesh encoding method and an encoder according to an embodiment (corresponding to a transmitting device 100, a dynamic mesh video acquisition unit 101, a dynamic mesh video encoder 102, a file / segment packager 103, a transmitter 104, a preprocessor 200, an encoder 201, a parameter of a mesh, and an encoding of a mesh based on tiles according to an embodiment of the disclosure) Figure 1 of the disclosure) generate a parameter of a mesh Figure 2 and transmit the parameter in a bitstream. Figure 6 and Figure 7 An encoder of the disclosure, Figure 13 a transmitting device of the disclosure, Figures 15 to 17 a mesh data encoding based on tiles according to an embodiment of the disclosure, and Figure 20 an encoding of the disclosure) generate a parameter of a mesh Figure 6 and transmit the parameter in a bitstream. Figure 7 A mesh decoding method and a decoder according to an embodiment (corresponding to a receiving device 110, a receiver 111, a file / segment unpackager 112, a dynamic mesh video decoder 113, a renderer 114, a parameter of a mesh, and a decoding of a mesh based on tiles according to an embodiment of the disclosure)

[0208] of the disclosure) decode mesh data based on a parameter of a mesh Figure 15 and included in a bitstream. Figure 16 and Figure 17 A decoder of the disclosure, Figure 18 a receiving device of the disclosure, as Figure 19 a decoding of a process inverse to a tile-based encoding, and Figure 20 a decoding of the disclosure) decode mesh data based on a parameter of a mesh Figure 21 and included in a bitstream. Figure 1 num_info_sets_minus1 + 1 indicates the number of extraction information sets contained in the MCTS extraction information set SEI message.

[0209] num_mcts_sets_minus1[ i ] + 1 indicates the number of MCTS sets that share the i-th extraction information set.

[0210] num_mcts_in_set_minus1[ i ][ j ] + 1 indicates the number of MCTSs in the j-th MCTS set associated with the i-th extraction information set.

[0211] num_mcts_sets_minus1[ i ] + 1 indicates the number of MCTS sets that share the i-th extraction information set.

[0212] idx_of_mcts_in_set[ i ][ j ][ k ] indicates the MCTS index of the k-th MCTS in the j-th MCTS set associated with the i-th extraction information set.

[0213] Slice_reordering_enabled_flag[ i ] equal to 1 indicates that extraction of the MCTS sub-bitstream using the i-th extraction information set includes reordering of extracted slice segments, and the Slice_segment_address of the j-th slice segment in bitstream order is associated with any extracted MCTS.

[0214] num_slice_segments_minus1[ i ] + 1 indicates the number of slice segments associated with the MCTS set of the i-th extraction information set when Slice_reordering_enabled_flag[ i ] is equal to 1.

[0215] output_slice_segment_address[ i ][ j ] indicates the slice segment address of the j-th slice segment in bitstream order associated with any MCTS set of the i-th extraction information set when Slice_reordering_enabled_flag[ i ] is equal to 1.

[0216] num_vps_in_info_set_minus1[ i ] + 1 indicates the number of alternative VPSs in the i-th extraction information set.

[0217] vps_rbsp_data_length[ i ][ j ] indicates the number of RBSP data bytes of the j-th alternative VPS in the i-th extraction information set.

[0218] num_sps_in_info_set_minus1[ i ] + 1 indicates the number of alternative SPSs in the i-th extraction information set.

[0219] num_sps_in_info_set_minus1[ i ] indicates the number of RBSP data bytes of the j-th alternative SPS in the i-th extraction information set.

[0220] num_pps_in_info_set_minus1[ i ] + 1 indicates the number of alternative PPSs in the i-th extraction information set.

[0221] num_pps_in_info_set_minus1[ i ] minus 1 indicates the temporal identifier of the jth alternative PPS in the ith extraction information set.

[0222] pps_rbsp_data_length[ i ][ j ] indicates the number of RBSP data bytes of the jth alternative PPS in the ith extraction information set.

[0223] The grid encoding method and encoder (corresponding to the transmission device 100, the dynamic grid video acquisition unit 101, the dynamic grid video encoder 102, the file / segment packager 103, the transmitter 104, the preprocessor 200, the encoder 201, the encoder of the tile-based grid data encoding, the transmission device, the decoding of the tile-based encoding as the inverse process, the tile-based parameter analysis, and the decoding of the Figure 11 Figure 12 Figure 14 Figures 15 to 17 the encoder of the tile-based grid data encoding, Figure 18 the transmission device of the tile-based grid data encoding, Figure 19 the tile-based encoding based on and Figure 21 the encoding) of the embodiments generate the tile-based parameters of the Figure 21 and Figure 20 and transmit the parameters in the bitstream.

[0224] The grid decoding method and decoder (corresponding to the reception device 110, the receiver 111, the file / segment unpackager 112, the dynamic grid video decoder 113, the renderer 114, the decoder of the tile-based grid data decoding, the reception device, the decoding of the tile-based encoding as the inverse process, the tile-based parameter analysis, and the decoding of the Figure 18 Figure 19 Figure 11 the decoder of the tile-based grid data decoding, Figure 12 the reception device of the tile-based grid data decoding, Figure 21 the decoding of the tile-based encoding as the inverse process, ​ and ​ the tile-based parameter analysis of the tile-based grid data decoding, and ​ the decoding of the tile-based grid data decoding) of the embodiments can identify the mapping relationship between the tiles and the displacement information in the ​ and ​ and decode the displacement information by tile based on the tile-based parameters of the ​ and ​ For example, the tile parameters of the ​ and ​ may include the number of tiles and the tile ID. In addition, they can further include the level value of the level of detail (LoD) of the displacement information of each tile ID. In addition, the tile parameters or the bitstream of the ​ and ​ may include the information about the tile configuration of the current frame and / or the attribute of mapping LoD and tile ID.

[0225] ​​​​​For each frame header in the bitstream, or as signaling data including the parameter set or frame characteristics of the frame, the following information can be further included: tile configuration (width, height of the tile, position of the tile within the frame) of the current frame, and properties (such as image characteristics, resolution, LoD indicator of the tile) mapping LoD and tile ID.

[0226] ​ A mesh data encoding method according to an embodiment is illustrated.

[0227] A mesh encoding method and encoder (corresponding to ​ a transmitting device 100, a dynamic mesh video acquisition unit 101, a dynamic mesh video encoder 102, a file / segment packager 103, a transmitter 104, ​ a preprocessor 200, an encoder 201, ​ and ​ an encoder, ​ a transmitting device, and ​ a tile-based mesh data encoding according to an embodiment) can encode and transmit mesh data as illustrated in ​ .

[0228] A mesh data encoding method according to an embodiment can include encoding mesh data (S2000).

[0229] A mesh data encoding method according to an embodiment can further include transmitting a bitstream containing the mesh data (S2001).

[0230] Referring to ​ and ​ , the operation S2000 of encoding the mesh data can include: generating a base mesh of the mesh data; encoding the base mesh; reconstructing the base mesh; generating displacement information based on the reconstructed base mesh; encoding the displacement information; reconstructing the displacement information; and encoding attribute information of the mesh data based on the reconstructed displacement information.

[0231] Referring to ​ , the encoding of the displacement information can include: encoding displacement information for a first level of a first layer and for a second level of a second layer for scalable processing. The encoding of the attribute information can include: encoding attribute information for a first level of a first layer and for a second level of a second layer.

[0232] Referring to ​ , the encoding method can further include: packing the displacement information and the attribute information into a frame, and the frame (image) can include one or more tiles.

[0233] Referring to ​The first level displacement information and the first level attribute information for the first layer can be included in the first tile of the frame, and the second level displacement information and the second level attribute information for the second layer can be included in the second tile of the frame.

[0234] In the operation S2001 of sending the bit stream, the bit stream may include parameters related to the plot ( ​ and ​ And additional parameters indicating tile-level data mapping.

[0235] ​ The method can be performed by a grid data transmitting device (encoder). The transmitting device includes a memory; and a processor configured to execute one or more instructions in the memory. The processor can perform operations including encoding the grid data and transmitting a bit stream containing the grid data.

[0236] ​ The illustration shows a grid data decoding method according to an embodiment.

[0237] According to the embodiments, the grid decoding method and decoder (corresponding to) ​ The receiver 110, receiver 111, file / fragment decapsulator 112, dynamic mesh video decoder 113, and renderer 114 are included. ​ and ​ decoder ​ The receiving device, as ​ Decoding based on the inverse process of tile encoding and ​ and ​ Patch-based parameter parsing can receive bitstreams and, as ​ The decoded grid data is shown in the figure.

[0238] ​ Decoding can be ​ The reverse process of encoding.

[0239] The grid data decoding method according to the embodiment may include receiving a bit stream containing grid data (S2100).

[0240] The grid data decoding method according to the embodiment may further include decoding grid data (S2101).

[0241] In the operation S2100 of receiving the bit stream, the bit stream may contain parameters related to the plot ( ​ and ​ And additional parameters indicating tile-level data mapping.

[0242] Reference ​ and ​The operation S2101 of decoding the mesh data can include decoding a base mesh in the bitstream, decoding displacement information in the bitstream, and decoding attribute information in the bitstream.

[0243] The decoding of the displacement information can include decoding a first level of displacement information for a first layer and a second level of displacement information for a second layer for scalable processing. The decoding of the attribute information can include decoding a first level of attribute information for the first layer and a second level of attribute information for the second layer.

[0244] The decoding method can further include deblocking the displacement information and the attribute information from the frame in the bitstream, wherein the frame can include one or more tiles.

[0245] The first level of displacement information for the first layer and the first level of attribute information for the first layer can be included in a first tile of the frame, and the second level of displacement information for the second layer and the second level of attribute information for the second layer can be included in a second tile of the frame.

[0246] ​ The method of claim 1 can be performed by a mesh data receiving device (decoder). The receiving device includes a memory; a processor configured to execute one or more instructions in the memory. The processor can perform operations including receiving a bitstream containing mesh data, and decoding the mesh data.

[0247] Embodiments can provide the following technical effects:

[0248] The conventional V-Mesh compression method is designed such that the displacement vector video and the texture map video generated in the encoding process are compressed and reconstructed using a video codec based on the overall resolution. Therefore, in a case where it is difficult to receive and utilize the entire mesh data according to the conditions of the receiver, the transmitted mesh data cannot be used. To solve this problem, embodiments include a scalable mesh decoding / partial decoding method that allows selective reconstruction and utilization of the mesh at the receiving side.

[0249] By adopting a method of integrating videos into a single frame structure based on tiles, instead of packing videos into corresponding video frames and compressing them using multiple video codecs as in the conventional method, scalable services can be supported in a simplified structure based on a single video codec.

[0250] By using a tile structure, appropriate videos can be extracted and provided according to the specifications of the receiver or the needs of the application (partial access, partial processing, and partial decoding).

[0251] Embodiments have been described in terms of methods and / or devices. The description of the methods and the description of the devices can complement each other.

[0252] Although the embodiments have been described with reference to each of the drawings, new embodiments can be designed by combining the embodiments shown in the drawings, if necessary. If a person skilled in the art designs a computer-readable recording medium in which a program for executing the embodiments mentioned in the above description is recorded, it can fall within the scope of the appended claims and equivalents thereof. The apparatus and method can not be limited by the configurations and methods of the above-described embodiments. The above-described embodiments can be configured by selectively fully or partially combining each of the above-described embodiments, so that various modifications can be made. Although the preferred embodiments have been described with reference to the drawings, it will be understood by those skilled in the art that various modifications and changes can be made to the embodiments without departing from the spirit or scope of the disclosure described in the appended claims. Such modifications should not be understood as departing from the technical idea or perspective of the embodiments.

[0253] The various elements of the apparatus of the embodiments can be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiments can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiments, components according to the embodiments can be implemented as separate chips, respectively. According to the embodiments, at least one or more components of the apparatus according to the embodiments can include one or more processors capable of executing one or more programs. The one or more programs can perform any one or more of the operations / methods according to the embodiments or include instructions for performing the same. Executable instructions for performing the methods / operations of the apparatus according to the embodiments can be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a transitory CRM or other computer program product configured to be executed by one or more processors. In addition, the memory according to the embodiments can be used as a concept that encompasses not only a volatile memory (e.g., a RAM) but also a non-volatile memory, a flash memory, and a PROM. In addition, it can also be implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, the processor-readable recording medium can be distributed to computer systems connected via a network, so that the processor-readable code can be stored and executed in a distributed manner.

[0254] In this document, the terms “ / ” and “,” should be interpreted to mean “and / or”. For example, the expression “A / B” can mean “A and / or B”. Also, “A, B” can mean “A and / or B”. Also, “A / B / C” can mean “at least one of A, B, and / or C”. “A, B, C” can also mean “at least one of A, B, and / or C”. Also, in this document, the term “or” should be interpreted as “and / or”. For example, the expression “A or B” can mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” in this document should be interpreted as “additionally or alternatively”.

[0255] In this document, the terms "and / or" and "or" are used to include any and all combinations of one or more of the associated items. For example, the expression "A / B" can mean "A and / or B." In addition, "A, B" can mean "A and / or B." In addition, "A / B / C" can mean "at least one of A, B, and / or C." "A, B, C" can also mean "at least one of A, B, and / or C." In addition, in this document, the term "or" is intended to mean "and / or". For example, the expression "A or B" can mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document is intended to mean "additionally or alternatively".

[0256] Various elements of the embodiments can be described using terms such as first and second. However, the various components according to the embodiments should not be limited by the above terms. The terms are used only to distinguish one element from another element. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of these terms should be interpreted not to depart from the scope of the various embodiments. The first user input signal and the second user input signal are both user input signals, but do not mean the same user input signal unless the context clearly specifies otherwise.

[0257] The terms used to describe the embodiments are used only for the purpose of describing particular embodiments and are not intended to limit the embodiments. As used in the description and the claims of the embodiments, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The expression "and / or" is used to include all possible combinations of the items. Terms such as "include" or "have" are intended to indicate the presence of the mentioned items, and should be understood as not excluding the possibility of the presence of additional items. As used herein, conditional language such as "if" and "when" is not limited to optional situations and is intended to be interpreted as executing an operation or interpreting a definition according to a particular condition.

[0258] The operations described in the specification according to the embodiments can be performed by a transmitting / receiving device including a memory and / or a processor according to the embodiments. The memory can store a program for processing / controlling operations according to the embodiments, and the processor can control various operations described in the specification. The processor can be referred to as a controller or the like. In the embodiments, the operations can be performed by firmware, software, and / or a combination thereof. The firmware, software, and / or a combination thereof can be stored in the processor or the memory.

[0259] Operations according to the above-described embodiments can be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device can include a transmitter / receiver configured to transmit and receive media data, a memory configured to store instructions (program code, algorithms, flowcharts, and / or data) for processes according to the embodiments, and a processor configured to control the transmitting / receiving device operations.

[0260] The processor can be referred to as a controller or the like, and can correspond to, for example, hardware, software, and / or a combination thereof. Operations according to the above-described embodiments can be performed by the processor. Furthermore, the processor can be implemented as an encoder / decoder for operations of the above-described embodiments.

[0261] Modes of the disclosure

[0262] Various embodiments have been described in the best mode for carrying out the disclosure.

[0263] Industrial applicability

[0264] As described above, embodiments can be applied, in whole or in part, to point cloud data transmitting / receiving devices and systems.

[0265] It will be apparent to those skilled in the art that various changes or modifications can be made to the embodiments within the scope of the embodiments.

[0266] Accordingly, it is intended that the embodiments encompass modifications / changes as long as they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for receiving grid data, the method comprising: Receive a bitstream containing grid data; as well as Decode the grid data.

2. The method according to claim 1, wherein, The decoding of the grid data includes: Decode the underlying grid in the bitstream; Decode the shift information in the bitstream; and Decode the attribute information in the bitstream.

3. The method according to claim 2, wherein, Decoding the displacement information includes decoding: Displacement information for the first level of the first layer; and Displacement information for the second level of the second layer. The decoding of the attribute information includes decoding: Attribute information for the first level of the first layer; and The attribute information for the second level of the second layer.

4. The method of claim 3, further comprising: The displacement information and attribute information are unpacked from the frames in the bitstream. The frame includes one or more tiles.

5. The method according to claim 4, wherein, The displacement information for the first level of the first layer and the attribute information for the first level of the first layer are included in the first tile of the frame. The displacement information for the second level of the second layer and the attribute information for the second level of the second layer are included in the second tile of the frame.

6. The method according to claim 1, wherein, The bitstream contains parameters related to the tiles.

7. A device for receiving grid data, comprising: Memory; as well as A processor, configured to execute one or more instructions in the memory. The processor is configured to execute: Receive a bitstream containing grid data; and Decode the grid data.

8. The device according to claim 7, wherein, The processor is configured to execute: Decode the underlying grid in the bitstream; Decode the shift information in the bitstream; and Decode the attribute information in the bitstream.

9. A method for transmitting grid data, the method comprising: Encode the grid data; as well as Send a bit stream containing the grid data.

10. The method according to claim 9, wherein, The encoding of the grid data includes: The base grid for generating the aforementioned grid data; The basic grid is encoded; Reconstruct the base mesh; Displacement information is generated based on the reconstructed base mesh; The displacement information is encoded; Reconstruct the displacement information; and Based on the reconstructed displacement information, the attribute information of the grid data is encoded.

11. The method according to claim 10, wherein, The encoding of the displacement information includes the following encoding: Displacement information for the first level of the first layer; and Displacement information for the second level of the second layer. The encoding of the attribute information includes the following encoding: Attribute information for the first level of the first layer; and The attribute information for the second level of the second layer.

12. The method of claim 11, further comprising: The displacement information and the attribute information are packaged into a frame. The frame includes one or more tiles.

13. The method according to claim 12, wherein, The displacement information for the first level of the first layer and the attribute information for the first level of the first layer are included in the first tile of the frame. The displacement information for the second level of the second layer and the attribute information for the second level of the second layer are included in the second tile of the frame.

14. The method according to claim 9, wherein, The bitstream contains parameters related to the tiles.

15. An apparatus for transmitting grid data, comprising: Memory; as well as A processor, configured to execute one or more instructions in the memory. The processor is configured to execute: Encoding grid data; and Send a bit stream containing the grid data.