3d data transmission device, 3d data transmission method, 3d data reception device, and 3d data reception method

By preprocessing, subgroup segmentation and motion vector encoding methods on 3D grid data, the problems of low transmission efficiency and high encoding complexity in the prior art are solved, and efficient data transmission and reduced delays are achieved.

CN120077645APending Publication Date: 2025-05-30LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073669.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-30
Filing Date
2023-08-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently transmit and receive 3D grid data of a large number of points, resulting in low data transmission efficiency and high encoding/decoding complexity.

Method used

By preprocessing the input grid data to generate the basic grid data, subgroup segmentation, motion vector calculation and encoding the basic grid data, and finally send a bit stream including the encoded grid data and signaling information.

Benefits of technology

It realizes efficient transmission and reception of 3D grid data, reduces the delay of data transmission and encoding/decoding complexity, and improves data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077645A_ABST
    Figure CN120077645A_ABST
Patent Text Reader

Abstract

A 3D data transmission method according to an embodiment may comprise the steps of: preprocessing input mesh data and outputting base mesh data; encoding the basic grid data; and transmitting a bitstream including the encoded mesh data and the signaling information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An embodiment provides a method for providing 3D content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and self-driving services to a user. Background Art

[0002] Point cloud data or mesh data in 3D content is a set of points in 3D space. However, due to a large number of points in 3D space, it is difficult to create point cloud data or mesh data.

[0003] In other words, high throughput is required to send and receive 3D data (e.g., point cloud or mesh data) having a significant number of points. Summary of the Invention

[0004] Technical Problem

[0005] An object of the present disclosure is to provide an apparatus and method for efficiently sending and receiving mesh data to solve the above problems.

[0006] Another object of the present disclosure is to provide an apparatus and method for solving latency and encoding / decoding complexity of mesh data.

[0007] The embodiments are not limited to the above objects, and the scope of the embodiments can be extended to other objects that can be inferred by those skilled in the art based on the entire content of the present disclosure.

[0008] Technical Solution

[0009] To achieve these objects and other advantages and in accordance with the purpose of the present invention, as specifically implemented and broadly described herein, a method for sending three-dimensional (3D) data may include the steps of: preprocessing input mesh data and outputting base mesh data; encoding the base mesh data; and sending a bitstream including the encoded mesh data and signaling information.

[0010] According to an embodiment, encoding of the base mesh data may include: dividing the reference base mesh data into subgroups; obtaining a motion vector between the base mesh data and the reference base mesh data for each subgroup; and encoding the obtained motion vector.

[0011] According to an embodiment, the motion vector is an average of motion vectors of vertices in a corresponding subgroup.

[0012] According to an embodiment, the signaling information may include information related to subgroup division.

[0013] According to an embodiment, the signaling information may further include motion vector related information for indicating whether to skip the motion vector.

[0014] According to an embodiment, the method may further include transmitting a motion vector or skipping the transmission of the motion vector based on motion vector related information.

[0015] According to an embodiment, the transmission of the motion vector is skipped, and a zero vector may be derived for the motion vector at the receiving side.

[0016] According to an embodiment, an apparatus for transmitting 3D data may include: a pre-processor configured to pre-process input mesh data and output base mesh data; an encoder configured to encode the base mesh data; and a transmitter configured to transmit a bitstream including the encoded mesh data and signaling information.

[0017] According to an embodiment, the encoder may include: a subgroup splitter configured to split reference base mesh data into subgroups; a motion vector calculator configured to obtain a motion vector between the base mesh data and the reference base mesh data for each subgroup; and an encoder configured to perform entropy coding on the obtained motion vectors.

[0018] According to an embodiment, the motion vector is an average of motion vectors corresponding to vertices in a subgroup.

[0019] According to an embodiment, the signaling information may include information related to subgroup splitting.

[0020] According to an embodiment, the signaling information may further include motion vector related information for indicating whether to skip the motion vector.

[0021] According to an embodiment, based on the motion vector related information, a motion vector may be transmitted, or the transmission of the motion vector may be skipped.

[0022] According to an embodiment, the transmission of the motion vector is skipped, and a zero vector may be derived for the motion vector at the receiving side.

[0023] According to an embodiment, a method for receiving 3D data may include the steps of: receiving a bitstream including encoded mesh data and signaling information; decoding the mesh data based on the motion vectors of each subgroup; and rendering the decoded mesh data.

[0024] Advantageous Effects

[0025] According to an embodiment, the 3D data transmission method, the 3D data transmission apparatus, the 3D data reception method, and the 3D data reception apparatus may provide high-quality 3D services.

[0026] According to an embodiment, the 3D data transmission method, the 3D data transmission apparatus, the 3D data reception method, and the 3D data reception apparatus may implement various video encoding and decoding schemes.

[0027] According to an embodiment, a 3D data transmission method, a 3D data transmission apparatus, a 3D data reception method, and a 3D data reception apparatus may support general 3D content, such as for autonomous driving services.

[0028] According to an embodiment, when encoding / decoding geometric information related to 3D dynamic mesh data through inter-frame prediction, the 3D data transmission method, the 3D data transmission apparatus, the 3D data reception method, and the 3D data reception apparatus may divide a reference base mesh into subgroups and calculate motion vectors based on each subgroup, such that (difference) motion vectors may be transmitted based on each subgroup. Accordingly, the amount of data to be transmitted may be reduced, and the compression efficiency of the geometric information may be increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application. The drawings illustrate embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar parts. In the drawings:

[0030] Figure 1 Illustrates a system for providing dynamic mesh content according to an embodiment.

[0031] Figure 2 Illustrates a V-MESH compression method according to an embodiment.

[0032] Figure 3 Illustrates preprocessing in V-MESH compression according to an embodiment.

[0033] Figure 4 Illustrates an edge midpoint subdivision method according to an embodiment.

[0034] Figure 5 Illustrates a displacement generation process according to an embodiment.

[0035] Figure 6 Illustrates an intra-frame encoding process of V-MESH data according to an embodiment.

[0036] Figure 7 Illustrates an inter-frame encoding process of V-MESH data according to an embodiment.

[0037] Figure 8 Illustrates a lifting transform process of displacement according to an embodiment.

[0038] Figure 9 Illustrates a process of packing transform coefficients into a 2D image according to an embodiment.

[0039] Figure 10Shows the attribute transfer process in the V-MESH compression method according to an embodiment.

[0040] Figure 11 Shows the intra-frame decoding process of V-MESH data according to an embodiment.

[0041] Figure 12 Shows the inter-frame decoding process of V-MESH data according to an embodiment.

[0042] Figure 13 Shows the mesh data transmitting device according to an embodiment.

[0043] Figure 14 Shows the mesh data receiving device according to an embodiment.

[0044] Figure 15 Shows the mesh data transmitting device according to an embodiment.

[0045] Figure 16 Is an exemplary detailed block diagram of a motion vector encoder according to an embodiment.

[0046] Figure 17 Is a diagram showing an example of a process of checking whether to skip a motion vector and calculating a motion vector according to an embodiment.

[0047] Figure 18 Shows an example of the Nth patch in the texture space according to an embodiment.

[0048] Figure 19 Shows an example of the motion vector resolution according to the motion vector resolution information according to an embodiment.

[0049] Figure 20 Is a diagram showing an example of the motion vector resolution based on each subgroup in the octree structure according to an embodiment.

[0050] Figure 21 Shows the mesh data receiving device according to an embodiment.

[0051] Figure 22 Is a diagram showing an example of a subgroup segmentation process according to an embodiment.

[0052] Figure 23 Shows an example of the subgroup segmentation method index according to an embodiment.

[0053] Figure 24 Is a diagram showing an example of deriving subgroup segmentation information from the octree structure according to an embodiment.

[0054] Figure 25 Is an exemplary detailed block diagram of a motion vector decoder according to an embodiment.

[0055] Figure 26 FIG. is a diagram showing an example of the detailed operation of a differential motion vector decoder according to an embodiment.

[0056] Figure 27 FIG. is a diagram showing an example of an exemplary motion vector estimation process according to an embodiment.

[0057] Figure 28 FIG. shows an exemplary syntax structure of subgroup segmentation information according to an embodiment.

[0058] Figure 29 FIG. shows an exemplary syntax structure of motion vector related information according to an embodiment.

[0059] Figure 30 FIG. is a flowchart showing an exemplary transmission method according to an embodiment.

[0060] Figure 31 FIG. is a flowchart showing an exemplary reception method according to an embodiment. DETAILED DESCRIPTION

[0061] Reference will now be made in detail to the preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description given with reference to the accompanying drawings is intended to illustrate the exemplary embodiments of the present disclosure, and not to show all the embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0062] Although most of the terms used in the present disclosure are selected from general terms widely used in the art, some terms are arbitrarily selected by the applicant, and their meanings are detailed in the following description as needed. Therefore, the present disclosure should be understood based on the intended meanings of the terms rather than their simple names or meanings.

[0063] With the recent progress of 3D data modeling and rendering technologies, active research has been conducted on generating and processing 3D data across various fields, including virtual reality (VR), augmented reality (AR), autonomous driving, computer-aided design (CAD) / computer-aided manufacturing (CAM), and geographic information system (GIS). 3D data can be represented as point clouds or meshes according to the representation format. A mesh consists of geographic information indicating the coordinates of individual vertices or points, connection information indicating the connections between vertices, a texture map representing color information about the mesh surface as 2D image data, and texture coordinates indicating the mapping information between the mesh surface and the texture map. In the present disclosure, when at least one element constituting the mesh changes over time, the mesh is defined as a dynamic mesh, and when it does not change, it is defined as a static mesh.

[0064] Compared with 2D image data, dynamic mesh data involves significantly larger amounts of element data to represent the mesh. As a result, techniques have been developed for efficiently compressing large amounts of mesh data for storage and transmission of the data.

[0065] Figure 1 A system for providing dynamic mesh content according to an embodiment is shown.

[0066] Figure 1 The system in includes a transmitting device 100 and a receiving device 110. The transmitting device 100 may include a mesh video acquisition unit (or portion) 101, a mesh video encoder 102, a file / fragment encapsulation module 103, and a transmitter 104. The receiving device 110 may include a receiver 111, a file / fragment decapsulation unit 112, a mesh video decoder 113, and a renderer 114. Figure 1 Each component in may correspond to hardware, software, a processor, and / or a combination thereof. In the following description, a mesh data transmitting device according to an embodiment may be interpreted to refer to a 3D data transmitting device or the transmitting device 100, or to a mesh video encoder (hereinafter referred to as an encoder) 102. A mesh data receiving device according to an embodiment may be interpreted to refer to a 3D data receiving device or the receiving device 110, or to a mesh video decoder (hereinafter referred to as a decoder) 113.

[0067] Figure 1 The system of may perform video-based dynamic mesh compression and decompression.

[0068] With the advancement of 3D capture, modeling, and rendering, users are allowed to access various forms of 3D content (e.g., AR, XR, the metaverse, and holograms) across multiple platforms and devices. 3D content is becoming increasingly complex and realistic in its object representation to provide users with an immersive experience. However, for the generation and use of 3D models, this requires significant amounts of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system using mesh content.

[0069] First, a method of compressing dynamic mesh data starts with a video-based point cloud compression (V-PCC) standard technique for point cloud data. Point cloud data is data having color information in the coordinates (X, Y, Z) of vertices (or points). In the present disclosure, vertex coordinates (i.e., position information) are referred to as geometric information, and color information about vertices is referred to as attribute information. Geometric information and attribute information together are referred to as vertex information or point cloud data. Mesh data refers to vertex information including connection information between vertices. Content may be initially created in the form of mesh data. Alternatively, connection information may be added to the point cloud data, and the point cloud data may be transformed into mesh data.

[0070] Currently, the MPEG standards group has defined two data types for dynamic mesh data: mesh data of category 1 with a texture map as color information and mesh data of category 2 with vertex colors as color information.

[0071] The mesh coding standard for category 1 data is currently in progress, and the standardization of category 2 data is expected to follow. As Figure 1 shown, the overall process for providing mesh content services may include an acquisition, encoding, transmission, decoding, rendering, and / or feedback process.

[0072] To provide mesh content services, 3D data acquired by multiple cameras or special cameras can be processed into a mesh data type through a series of steps to generate a video. The generated mesh video can be sent through a series of operations, and the receiving side can process the received data back into a mesh video for rendering. Through this process, the mesh video can be provided to the user, allowing the user to interactively utilize the mesh content according to their intentions.

[0073] As Figure 1 shown, the mesh compression system may include a transmitting device 100 and a receiving device 110. The transmitting device 100 can encode the mesh video to output a bitstream, which can be transmitted to the receiving device 110 in the form of a file or a stream (stream segment) via a digital storage medium or a network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0074] In the transmitting device 100, the encoder can be referred to as a mesh video / image / picture / frame encoding device. In the receiving device 110, the decoder can be referred to as a mesh video / image / picture / frame decoding device. The transmitter can be included in the mesh video encoder, and the receiver can be included in the mesh video decoder. The renderer 114 can include a display, and the renderer and / or the display can be configured as a separate device or an external component. The transmitting device 100 and the receiving device 110 can also include separate internal or external modules / units / components for the feedback process.

[0075] Mesh data uses multiple polygons to represent the surface of an object. Each polygon is defined by vertices in 3D space and connection information indicating how the vertices are connected. Additionally, vertex attributes such as color and normal vector can be included in the data. Mapping information that allows the mesh surface to be mapped onto a 2D plane can also be included in the mesh attributes. Mapping is typically described using a set of parametric coordinates related to the mesh vertices (referred to as UV coordinates or texture coordinates). The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as texture, normal, and displacement. Here, displacement can be used interchangeably with displacement information or displacement vector.

[0076] The mesh video acquisition unit 101 may include processing 3D object data acquired through a camera or the like into a mesh data type having the above-described attributes through a series of operations, and generating a video composed of the mesh data. In the mesh video, the attributes of the mesh (e.g., vertices, polygons, connections between vertices, colors, and normals) may change over time. A mesh video having attributes and connection information that change over time is referred to as a dynamic mesh video.

[0077] The mesh video encoder 102 may encode an input mesh video into one or more video streams. A video may include multiple frames, and each frame may correspond to a still image / frame. In the present disclosure, a mesh video may include mesh images / frames / frames. The term "mesh video" may be used interchangeably with mesh images / frames / frames. The mesh video encoder 102 may perform a video-based dynamic mesh (V-Mesh) compression process. For compression and encoding efficiency, the mesh video encoder 102 may perform a series of processes such as prediction, transformation, quantization, and entropy encoding. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0078] The file / fragment encapsulation module 103 may encapsulate the encoded mesh video data and / or mesh video-related metadata in the form of a file or the like. The mesh video-related metadata may be received from a metadata processor. The metadata processing unit may be included in the mesh video encoder 102 or may be configured as a separate component / module. The file / fragment encapsulation module 103 may encapsulate the data into a file format such as ISOBMFF or process it into a form such as DASH fragments. According to an embodiment, the file / fragment encapsulation module 103 may include mesh video-related metadata in the file format. For example, the mesh video metadata may be included in the boxes at various levels in the ISOBMFF file format or as data on a separate track in the file. In some embodiments, the file / fragment encapsulation module 103 may encapsulate the mesh video-related metadata into the file.

[0079] The sending processor may apply processing to the encapsulated mesh video data based on the file format for transmission. The sending processor may be included in the transmitter 104 or implemented as a separate component / module. The sending processor may process the mesh video data according to any transmission protocol. The processing for transmission may include transmission processing via a broadcast network and transmission processing via broadband. In some embodiments, the sending processor may receive the mesh video-related metadata and the mesh video data from the metadata processor and process them for transmission.

[0080] The transmitter 104 can send the encoded video / image information in the form of a file or a stream or the data output in the form of a bitstream to the receiver 111 of the receiving device 110 via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter 104 can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver 111 can extract the bitstream and transmit it to the decoding device.

[0081] The receiver 111 can receive the mesh video data sent by the mesh data transmission device. Depending on the channel used for transmission, the receiver 111 can receive the mesh video data via a broadcast network or a broadband network, or can receive the mesh video data via a digital storage medium.

[0082] The receiving processor can perform processing on the received mesh video data according to the transmission protocol. The receiving processor can be included in the receiver 111 or can be configured as a separate component / module. To correspond to the processing performed for transmission on the sending side, the receiving processor can perform the reverse process of the operations of the above-mentioned sending processor. The receiving processor can transmit the obtained mesh video data to the file / fragment de-packager 112 and transmit the obtained mesh video-related metadata to the metadata parser. The mesh video-related metadata obtained by the receiving processor can be in the form of a signaling table.

[0083] The file / fragment de-packager 112 can de-package the mesh video data in the form of a file received from the receiving processor. The file / fragment de-packager 112 can de-package the file, etc. according to ISOBMFF, etc. to obtain the mesh video bitstream or the mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to the mesh video decoder 113, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to the metadata processor. The mesh video bitstream can include metadata (metadata bitstream). The metadata processor can be included in the mesh video decoder 113 or can be configured as a separate component / module. The mesh video-related metadata obtained by the file / fragment de-packager 112 can be in the form of a box or a track in the file format. When necessary, the file / fragment de-packager 112 can receive the metadata required for de-packaging from the metadata processor. The mesh video-related metadata can be transmitted to the mesh video decoder 113 for use in the mesh video decoding process or transmitted to the renderer 114 for use in the mesh video rendering process.

[0084] The mesh video decoder 113 can receive an input bitstream and perform inverse operations corresponding to the operations of the mesh video encoder 102 to decode the video / image. The decoded mesh video / image can be displayed on the display of the renderer 114. The user can view all or part of the rendering result through a VR / AR display, a general display, etc.

[0085] The feedback process may include sending various types of feedback information that can be obtained during the rendering / display operation to the decoder on the sending side or the receiving side. The feedback process can provide interactivity when consuming mesh video. In some embodiments, the feedback process may include sending head orientation information, viewport information indicating the area that the user is currently viewing, etc. In some embodiments, the user can interact with objects implemented in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction during the feedback process can be transmitted to the sending side or the service provider. In some embodiments, the feedback process can be skipped.

[0086] The head orientation information may refer to information about the user's head position, angle, movement, etc. Based on this information, information about the area that the user is currently viewing within the mesh video (i.e., viewport information) can be calculated.

[0087] The viewport information can be information about the area that the user is currently viewing in the mesh video. Gaze analysis can be performed based on this information to determine how the user consumes the mesh video, how long the user looks at a specific area of the mesh video, etc. The gaze analysis can be performed on the receiving side, and the results can be transmitted to the sending side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0088] In some embodiments, the above feedback information can be not only transmitted to the transmitter, but also consumed on the receiving side. In other words, operations such as decoding and rendering can be performed on the receiving side based on the above feedback information. For example, based on the head orientation information and / or the viewport information, only the mesh video of the area that the user is currently viewing can be preferentially decoded and rendered.

[0089] The present disclosure relates to embodiments of dynamic mesh video compression as described above. The methods / embodiments disclosed herein can be applied to the video-based dynamic mesh compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or any next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time. It can perform lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, and AR / VR.

[0090] The dynamic mesh video compression method described below is based on the V-mesh method of MPEG.

[0091] In the present disclosure, a picture / frame generally may refer to a unit representing an image at a specific time.

[0092] A pixel or pel may refer to the smallest unit constituting a picture (or video). Additionally, the term "sample" may be used as a term corresponding to a pixel. A sample generally may indicate a pixel or a pixel value. It may indicate only the pixel / pixel value of a luminance component, or may indicate only the pixel / pixel value of a chrominance component, or may indicate only the pixel / pixel value of a depth component.

[0093] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. In some cases, the term unit may be used interchangeably with terms such as a block or a region. Generally, an M×N block may include a set (or array) of samples (or a sample array) or transform coefficients composed of M columns and N rows.

[0094] As described above, Figure 1 the encoding process is performed as follows.

[0095] In other words, a compression method based on video-based dynamic mesh compression (V-Mesh) may provide a method for compressing dynamic mesh video data based on 2D video coding and decoding (e.g., High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC)). During the V-Mesh compression process, the following data is received as input and compressed.

[0096] Input mesh: includes 3D coordinates of vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connections between vertices constituting the surface. The mesh surface may be represented by triangles or other polygons, and the connection information between vertices constituting the surface is stored according to a predetermined shape. The input mesh may be stored in the OBJ file format.

[0097] Attribute map (texture map may also be used interchangeably hereinafter): contains information about attributes (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the mesh surface onto a 2D image. The mapping indicating which part (surface or vertex) of the mesh corresponds to each piece of data in the attribute map is based on the mapping information included in the input mesh. Since the attribute map has data for each frame of the mesh video, it may also be referred to as an attribute map video. The attribute map in the V-Mesh compression method mainly contains color information about the mesh, and is stored in an image file format (PNG, BMP, etc.).

[0098] Material library file: contains information about material properties used in the mesh, especially information for linking the input mesh to the corresponding attribute map. It is stored in the Wavefront Material Template Library (MTL) file format.

[0099] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0100] Base mesh: The input mesh is represented by the minimum vertices determined according to user criteria by simplifying the input mesh through a preprocessing process.

[0101] Displacement: Displacement information used to represent the input mesh as similarly as possible using the base mesh, represented in 3D coordinates.

[0102] Atlas information: Metadata required to reconstruct the mesh using the base mesh, displacement, and attribute map information. It can be used to generate and use sub-units (sub-meshes, patches, etc.) of the mesh.

[0103] Reference Figures 2 to 7 Describes a method for encoding mesh position information (or vertex position information). Refer to Figures 6 to 10 And so on for describing a method for reconstructing mesh position information to encode attribute information (attribute map).

[0104] Figure 2 Shows the V-MESH compression method according to an embodiment.

[0105] Figure 2 Shows Figure 1 The encoding process, where the encoding process may include a preprocessing process and an encoding process. As Figure 2 Shown, Figure 1 The mesh video encoder 102 of Figure 1 The transmitting device of Figure 1 The mesh video encoder 102 of Figure 2 Shown, the V-Mesh compression method may include preprocessing 200 and encoding 201. Figure 2 The preprocessor 200 of Figure 2 May be located at the front end of the encoder 201 of Figure 2 The preprocessor 200 and the encoder 201 of

[0106] The preprocessor 200 may receive a static or dynamic mesh (M(i)) and / or an attribute map (A(i)). The preprocessor 200 may generate a base mesh m(i) and / or a displacement d(i) through preprocessing. The preprocessor 200 may receive feedback information from the encoder 201 and may generate a base mesh and / or a displacement based on the feedback information.

[0107] The encoder 201 may receive a base mesh m(i), a displacement d(i), a static or dynamic mesh M(i), and / or an attribute map A(i). In the present disclosure, at least one of the base mesh m(i), the displacement d(i), the static or dynamic mesh M(i), and / or the attribute map A(i) may be referred to herein as mesh-related data. The encoder 201 may encode the mesh-related data to generate a compressed bitstream.

[0108] Figure 3 Shows preprocessing in V-MESH compression according to an embodiment.

[0109] Figure 3 Shows Figure 2 the configuration and operation of the preprocessor. In Figure 3 it, the input mesh may include a static or dynamic mesh M(i) and / or an attribute map A(i). The input mesh may also include the 3D coordinates of the vertices that make up the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between the vertices that make up the surface.

[0110] Figure 3 Shows the process of performing preprocessing on the input mesh. The preprocessing 200 may include four operations: 1) Group of Frames (GoF) generation, 2) mesh simplification, 3) UV parameterization, and 4) fitting a subdivision surface (300). According to an embodiment, the GoF generation may be referred to as the GoF generation process or the GoF generator, the mesh simplification may be referred to as the mesh simplification process or the mesh simplification section, the UV parameterization may be referred to as the UV parameterization process or the UV parameterization section, and the fitting of the subdivision surface may be referred to as the fitting of the subdivision surface process or the fitting of the subdivision surface section. The preprocessor 200 may generate a displacement and / or a base mesh from the received input mesh and transmit it to the encoder 201. The preprocessor 200 may transmit GoF information related to the GoF generation to the encoder 201.

[0111] Next, each operation of Figure 3 is described.

[0112] GoF generation: The process of generating a reference structure for the mesh data. When the mesh of the previous frame and the current mesh have the same number of vertices, the same number of texture coordinates, the same vertex connection information, and the same texture coordinate connection information, the previous frame may be set as the reference frame. In other words, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, the encoder 201 may perform inter-frame encoding. Otherwise, it performs intra-frame encoding for the frame.

[0113] Mesh simplification: The process of simplifying the input mesh to create a simplified mesh (referred to as the base mesh). Vertices to be removed may be selected from the original mesh based on user-defined criteria, and then the selected vertices and the triangles connected to the selected vertices may be removed.

[0114] During the process of performing mesh simplification, the voxelized input mesh, the target triangle ratio (TTR), and the minimum triangle component (CCCount) information can be transmitted as inputs, and a simplified mesh can be obtained as an output. During this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.

[0115] UV parameterization: The process of mapping a 3D surface to the texture domain of the simplified mesh. The UVAtlas tool can be used to perform the parameterization. This process generates mapping information indicating where each vertex of the simplified mesh can be mapped onto a 2D image. The mapping information is represented as texture coordinates and stored, and a final base mesh is generated through this process.

[0116] Fitting a subdivision surface (300): The process of performing subdivision on the simplified mesh (i.e., the simplified mesh with texture coordinates). The displacement and the base mesh generated through this process are output to the encoder 201. A user-defined method (e.g., the edge midpoint method) can be applied as the subdivision method. The fitting process is performed such that the input mesh and the subdivision mesh become similar to each other. The mesh on which the fitting process is performed will be referred to as the fitting subdivision mesh in this article.

[0117] Figure 4 Shows the edge midpoint subdivision method according to an embodiment.

[0118] Figure 4 Shows with reference to Figure 3 the edge midpoint subdivision method of the fitting subdivision surface described. With reference to Figure 4 , the original mesh containing four vertices is subdivided to create a submesh. The submesh can be created by creating new vertices at the midpoints of the edges between the vertices. Then, the fitting process is performed to make the input mesh and the submesh similar to each other, thereby obtaining the fitting subdivision mesh.

[0119] Once the fitting subdivision mesh is generated, the displacement is calculated based on this result and the previously compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). In other words, the reconstructed base mesh is subdivided in the same manner as the fitting subdivision surface. The position difference between each vertex in the result and the fitting subdivision mesh is the displacement of each vertex. Since the displacement represents the position difference in 3D space, it is represented as a value in the (x, y, z) space in the Cartesian coordinate system. According to the user input parameters, the (x, y, z) coordinate values can be converted into (normal, tangent, binormal) coordinate values in the local coordinate system.

[0120] Figure 5 Shows the displacement generation process according to an embodiment. Figure 5 The displacement generation process of

[0121] Figure 5 Shows in detail how the displacement is calculated for the fitted subdivision surface 300 as described with reference to Figure 4 .

[0122] An encoder and / or pre-processor according to an embodiment may include 1) a subdivider, 2) a local coordinate system calculator, and 3) a displacement vector calculator. The subdivider may perform subdivision on the reconstructed base mesh to generate a subdivided reconstructed base mesh. Here, the reconstruction of the base mesh may be performed by the pre-processor 200 or may be performed by the encoder 201. The local coordinate system calculator may receive the fitted subdivision mesh and the subdivided reconstructed base mesh, and may transform the coordinate system related to the mesh into a local coordinate system based on the received meshes. The local coordinate system calculation may be optional. The displacement calculator calculates the position difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, it may generate the position difference between the vertices in the two input meshes. The position difference between the vertices is the displacement.

[0123] A mesh data transmission method and apparatus according to an embodiment may encode mesh data as follows. Mesh data is a term including point cloud data. Point cloud data according to an embodiment (which may be abbreviated as point cloud) may refer to data including vertex coordinates (also referred to as geometric information) and color information (also referred to as attribute information). Additionally, geometric images, attribute images, occupancy maps, and auxiliary information (also referred to as patch information) generated by patch generation and packing based on vertex coordinates and color information may also be referred to as point cloud data. Thus, point cloud data including connection information may be referred to as mesh data. The terms point cloud and mesh data may be used interchangeably herein.

[0124] According to an embodiment, the V-Mesh compression (reconstruction) method may include intra-frame encoding ( Figure 6 ) and inter-frame encoding ( Figure 7 ).

[0125] Based on the results generated by the above GoF, intra-frame encoding or inter-frame encoding is performed. In intra-frame encoding, the data to be compressed may be the base mesh, displacement, attribute map, etc. In inter-frame encoding, the data to be compressed may be the displacement, attribute map, and the motion field between the reference base mesh and the current base mesh.

[0126] Figure 6 Shows the intra-frame encoding process in the V-MESH compression method according to an embodiment. Figure 6 The respective components of the intra-frame encoding process correspond to hardware, software, processors, and / or combinations thereof.

[0127] Figure 6 The encoding process of Figure 1 details the encoding of the mesh video encoder 102 of Figure 1The encoding is a configuration of the grid video encoder 102 when intra-frame encoding is performed. Figure 6 The encoder may include a pre-processor 200 and / or an encoder 201. Figure 6 The pre-processor 200 and the encoder 201 may correspond to Figure 3 the pre-processor 200 and the encoder 201.

[0128] The pre-processor 200 may receive an input grid and perform the above-mentioned pre-processing. A base grid and / or a fitted subdivision grid may be generated through the pre-processing.

[0129] The quantizer 411 of the encoder 201 may quantize the base grid and / or the fitted subdivision grid. The static grid encoder 412 may encode the static grid (i.e., the quantized base grid) and generate a bitstream containing the encoded base grid (i.e., the compressed base grid bitstream). The static grid decoder 413 may decode the encoded static grid (i.e., the encoded base grid). The inverse quantizer 414 may inverse-quantize the quantized static grid (i.e., the base grid) and output the reconstructed (restored) base grid. The displacement calculator 415 may generate a displacement based on the reconstructed static grid (i.e., the base grid) and the fitted subdivision grid. According to an embodiment, the displacement calculator 415 subdivides the reconstructed base grid and then calculates the displacement, i.e., the position difference between the respective vertices of the subdivided base grid and the fitted subdivision grid. In other words, when the fitted subdivision grid is similar to the original grid, the displacement is a displacement vector that is the position difference between the vertices in the two grids. The forward linear lifting unit 416 may perform a lifting transformation on the input displacement to generate lifting coefficients (also referred to as transform coefficients). The quantizer 417 may quantize the lifting coefficients. The image packer 418 may pack an image based on the quantized lifting coefficients. The video encoder 419 may encode the packed image. That is, the quantized lifting coefficients are packed by the image packer 418 into a frame as a 2D image, compressed by the video encoder 419, and output as a displacement bitstream (i.e., the compressed displacement bitstream).

[0130] The video decoder 420 decodes the compressed displacement bitstream. The image unpacker 421 may unpack the decoded displacement frame to output the quantized lifting coefficients. The inverse quantizer 422 may inverse-quantize the quantized lifting coefficients. The inverse linear lifting unit 423 applies an inverse lift to the inverse-quantized lifting coefficients to generate a reconstructed displacement. The grid reconstructor 424 restores the reconstructed and deformed grid based on the reconstructed displacement output from the inverse linear lifting unit 423 and the reconstructed base grid (also referred to as the subdivided reconstructed base grid) output from the inverse quantizer 414. The reconstructed and deformed grid is referred to herein as the reconstructed deformed grid.

[0131] The attribute transfer 425 receives an input mesh and / or an input attribute map, and regenerates an attribute map based on the reconstructed deformed mesh. The attribute map refers to a texture map corresponding to the attribute information in the mesh data component. In the present disclosure, the terms attribute map and texture map may be used interchangeably. The push-pull filler 426 can fill data into the attribute map based on the push-pull method. The color space converter 427 can convert the space of the color components of the attribute map. For example, the attribute map can be converted from the RGB color space to the YUV color space. The video encoder 428 can encode the attribute map to output a compressed attribute bitstream.

[0132] The multiplexer 430 can multiplex the compressed base mesh bitstream, the compressed displacement bitstream, and the compressed attribute bitstream to generate a compressed bitstream.

[0133] In Figure 6 the displacement calculator 415 can be included in the pre-processor 200. Additionally, at least one of the quantizer 411, the static mesh encoder 412, the static mesh decoder 413, or the inverse quantizer 414 can be included in the pre-processor 200.

[0134] As Figure 6 described, the intra-frame encoding method includes base mesh encoding (also known as static mesh encoding). That is, when performing intra-frame encoding on the current input mesh frame, the base mesh generated during the preprocessing of the pre-processor 200 can be quantized by the quantizer 411 and then encoded by the static mesh encoder 412 using static mesh compression techniques. For example, in the V-Mesh compression method, the Draco technique is applied to encode the base mesh, and the vertex position information, mapping information (texture coordinates), vertex connection information, etc. related to the base mesh are subject to compression.

[0135] Figure 6 The encoder in Figure 7 compresses the base mesh, displacement, and attributes in the frame to generate a bitstream, while

[0136] Figure 7 The encoder in Figure 7 compresses the motion, displacement, and attributes between the current frame and the reference frame to generate a bitstream.

[0137] Figure 7 The encoding process in Figure 1 details the encoding in Figure 1 That is, it represents the configuration of the encoder when the encoding in Figure 7 is inter-frame encoding. Figure 7The preprocessor 200 and the encoder 201 may correspond to Figure 3 the preprocessor 200 and the encoder 201.

[0138] For the components corresponding to Figure 6 the encoding operation of Figure 7 the encoding operation, refer to Figure 6 the description of Figure 7 That is, the operations of the quantizer 511, displacement calculator 515, wavelet transformer 516, quantizer 517, image packer 518, video encoder 519, video decoder 520, image unpacker 521, inverse quantizer 522, and inverse wavelet transformer 523, mesh reconstructor 524, attribute transfer 525, push-pull padding 526, color space converter 527, video encoder 528, and multiplexer 530 in Figure 6 are the same as or similar to the operations of the quantizer 411, static mesh encoder 412, static mesh decoder 413, and inverse quantizer 414, displacement calculator 415, forward linear lifting unit 416, quantizer 417, image packer 418, video encoder 419, video decoder 420, image unpacker 421, inverse quantizer 422, inverse linear lifting unit 423, and mesh reconstructor 424, attribute transfer 425, push-pull filler 426, color space converter 427, video encoder 428, and multiplexer 430 in Figure 7 and thus will not be described in detail herein to avoid redundancy.

[0139] In Figure 7 for inter-frame based encoding, the motion encoder 512 may obtain and encode the motion vector between the reconstructed quantized reference base mesh and the quantized current base mesh, and output a compressed motion bitstream. The motion encoder 512 may be referred to as a motion vector encoder. The base mesh reconstructor 513 may reconstruct the base mesh based on the reconstructed quantized reference base mesh and the encoded motion vector. The reconstructed base mesh is inverse quantized by the inverse quantizer 514 and output to the displacement calculator 515.

[0140] In Figure 7 the displacement calculator 515 may be included in the preprocessor 200. Additionally, at least one of the quantizer 511, motion encoder 512, base mesh reconstructor 513, or inverse quantizer 514 may be included in the preprocessor 200.

[0141] As referred to Figure 7As described, the inter-frame encoding method may include motion field encoding (also known as motion vector encoding). When the reference grid and the current input grid have a one-to-one vertex correspondence and they differ only in the position information of the vertices, inter-frame encoding can be performed. When performing inter-frame encoding, the base grid may not be compressed. Instead, the difference between the vertices of the reference base grid and the current base grid, that is, the motion field (or motion vector), can be calculated and encoded. The reference base grid is the result of quantifying and decoding the base grid data and is determined by the reference frame index determined in the generation of GoF. The motion field can be encoded as it is. Alternatively, the predicted motion field can be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and as the difference between the value of the predicted motion field and the value of the motion field of the current vertex, the residual motion field can be encoded. The value of the residual motion field can be encoded using entropy encoding. In addition to the motion field encoding in inter-frame encoding, the process of encoding displacement and attribute maps is the same as the structure of the intra-frame encoding method except for the base grid encoding.

[0142] Figure 8 Shows the lifting transformation process of displacement according to an embodiment.

[0143] Figure 9 Shows the process of packing transformation coefficients (also known as lifting coefficients) into a 2D image according to an embodiment.

[0144] Figure 8 and Figure 9 respectively show Figure 6 and Figure 7 the processes of transforming displacement and packing transformation coefficients in the encoding process of.

[0145] The encoding method according to an embodiment includes displacement encoding.

[0146] After base grid encoding and / or motion field encoding, a reconstructed base grid can be generated through reconstruction and inverse quantization, and the displacement can be calculated between the subdivision result of the reconstructed base grid and the fitted subdivision grid generated by fitting the subdivision surface (see 415 in Figure 6 or 515 in Figure 7 ). The data transformation process (e.g., wavelet transformation) can be applied to the displacement information for efficient encoding (see 416 in Figure 6 or 516 in Figure 7 ).

[0147] Figure 8 Figure 6 is shown by Figure 7 the forward linear lifting unit 416 of Figure 7The process of the wavelet transformer 516 using the lifting transform to transform displacement information. For example, a lifting transform based on linear wavelets can be performed. The transform coefficients generated by the transform process are quantized by the quantizer 417 (or 517), and then packed into a 2D image by the image packer 418 (or 518), as Figure 9 shown. The transform coefficients can be organized into blocks, with one block for every 256 (= 16×16) units. Each block can be packed in a z-scan order. The number of rows in a block is fixed at 16, but the number of columns in a block can be determined by the number of vertices in the subdivided base mesh. Inside a block, the transform coefficients can be sorted and packed according to the Morton code. For the packed image, a displacement video can be generated per GoF. The displacement video can be encoded by the video encoder 419 (or 519) using conventional video compression encoding and decoding.

[0148] Referring to Figure 8 , the base mesh (original) can include the vertices and edges of LoD0. The first subdivided mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) the edges of the base mesh. The first subdivided mesh contains the vertices of LoD0 and the vertices of LoD1. LoD1 includes the subdivided vertices and the vertices from the base mesh (LoD0). The first subdivided mesh can be divided (or subdivided) to generate the second subdivided mesh. The second subdivided mesh contains LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 contains the vertices further divided (or subdivided) from LoD0, and LoD2 contains the vertices further divided (or subdivided) from LoD1. LoD is a level of detail indicating how detailed the mesh data content is. As the index of the level increases, the distance between vertices decreases and the level of detail increases. In other words, as the value of LoD decreases, the detail of the mesh data content deteriorates. As the value of LoD decreases, the detail of the mesh data content enhances. LoD N contains the vertices included in LoD N-1. In the case of further dividing the mesh (or vertices) by subdivision, the mesh can be encoded by considering the previous vertices v1 and v2 and the subdivided vertex v based on a prediction and / or update method. Instead of encoding the information of the current LoD N as it is, a residual relative to the previous LoD N-1 can be generated. Therefore, the residual can be used to encode the mesh to reduce the size of the bitstream. The prediction process refers to the operation of predicting the current vertex v from the previous vertices v1 and v2. Since adjacent subdivided meshes have similar data, this property can be utilized for efficient encoding. The current vertex position information is predicted from the residual of the previous vertex position information, and the previous vertex position information is updated by the residual. In the present disclosure, vertices and points can be used interchangeably. LoD can be defined in the subdivision of the base mesh. According to an embodiment, the subdivision of the base mesh can be performed by the preprocessor 200, or can be performed by a separate component / module.

[0149] Referring to Figure 9, vertices have transformation coefficients (also referred to as lifting coefficients) generated by a lifting transformation. The transformation coefficients of vertices related to the lifting transformation can be packed into an image by an image packer 418 (or 518) and then encoded by a video encoder 419 (or 519).

[0150] Figure 10 Illustrates an attribute transfer process in a V-MESH compression method according to an embodiment.

[0151] According to an embodiment, Figure 10 Illustrates Figure 6 、 Figure 7 and other detailed operations of attribute transfer 425 (or 525) in encoding.

[0152] Encoding according to an embodiment includes attribute graph encoding. According to an embodiment, attribute graph encoding can be performed by Figure 6 video encoder 428 of Figure 7 or video encoder 528 of

[0153] According to an embodiment, in the present disclosure, the encoder compresses information about an input mesh through base mesh encoding (i.e., intra-frame encoding), motion field encoding (i.e., inter-frame encoding), and displacement encoding. The input mesh compressed during the encoding process is reconstructed through base mesh decoding (intra-frame), motion field decoding (inter-frame), and displacement video decoding, and as a reconstruction result, the reconstructed deformed mesh (hereinafter referred to as the reconstructed deformed mesh) is used to compress the input attribute graph, as Figure 6 and Figure 7 shown. The reconstructed deformed mesh has position information about vertices, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as Figure 10 shown, in the V-Mesh compression method, a new attribute graph with color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of attribute transfer 425 (or 525).

[0154] According to an embodiment, the attribute transfer 425 (or 525) first checks for each point P(u,v) in the 2D texture domain whether the corresponding vertex is within the texture triangle of the reconstructed deformed mesh. When the corresponding vertex is within the texture triangle T, the attribute transfer calculates the barycentric coordinates (α,β,γ) of P(u,v) based on the triangle T. Then, it calculates the 3D coordinates M(x,y,z) of P(u,v) based on the 3D vertex positions of the triangle T and (α,β,γ). The vertex coordinates M’(x’,y’,z’) corresponding to the position closest to the calculated M(x,y,z) and the triangle T’ containing the vertex are searched for in the input mesh domain. Then, the barycentric coordinates (α’,β’,γ’) of M’(x’,y’,z’) in the triangle T’ are calculated. The texture coordinates (u’,v’) are calculated based on the texture coordinates corresponding to the three vertices of the triangle T’ and (α’,β’,γ’), and the color information corresponding to the coordinates is searched for in the input attribute map. The color information found in this way is then assigned to the (u,v) pixel position in the new input attribute map. If P(u,v) does not belong to any triangle, a filling algorithm (e.g., the push-pull algorithm of the push-pull filling 426 (or 526)) is used to fill the pixel at that position in the new input attribute map with a color value.

[0155] The new attribute map generated by the attribute transfer 425 (or 525) is bundled into the GoF to construct an attribute map video, which is compressed using the video encoding and decoding of the video encoder 428 (or 528).

[0156] It can be seen from Figure 10 the reference relationships among the input mesh, the input attribute map, the reconstructed deformed mesh, and the reconstructed attribute map.

[0157] Figure 1 The decoding process of Figure 1 can execute the inverse process of the encoding process of

[0158] Figure 11 FIG. shows the intra-frame decoding process of the V-Mesh technology according to an embodiment.

[0159] Figure 11 FIG. shows Figure 1 the configuration and operation of the mesh video decoder 113 of the receiving device of Figure 11 FIG. also shows that the mesh data can be reconstructed by executing the inverse process of the intra-frame encoding process of Figure 6 FIG. Figure 11 The various components of the intra-frame decoding process of

[0160] First, the bitstream (i.e., the compressed bitstream) received and input to the demultiplexer 611 of the intra-frame decoder 610 can be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information about the mesh (e.g., V-PCC / V3C). The term V-PCC (Video-based Point Cloud Compression) used in this disclosure may have the same meaning as V3C (Vision Volume Video-based Coding). The two terms may be used interchangeably. Thus, in this disclosure, the term V-PCC may be interpreted as V3C.

[0161] According to an embodiment, the mesh substream can be input to the static mesh decoder 612 and decoded by it, the displacement substream can be input to the video decoder 613 and decoded by it, and the attribute map substream can be input to the video decoder 617 and decoded by it.

[0162] According to an embodiment, the mesh substream can be decoded by the decoder 612 of the static mesh coding decoding used in the encoding (e.g., GoogleDraco) to reconstruct connection information, vertex geometry information, vertex texture coordinates, etc. related to the decoding result of the reconstructed quantization base mesh (e.g., the reconstructed base mesh).

[0163] According to an embodiment, the displacement substream can be decoded into a displacement video by the decoder 613 of the video compression coding decoding used in the encoding. Then, image unpacking is performed by the image unpacker 614, inverse quantization is performed by the inverse quantizer 615, and inverse transformation is performed by the inverse linear lifting unit 616 to reconstruct the displacement information about each vertex (i.e., the reconstructed displacement).

[0164] According to an embodiment, the base mesh reconstructed by the static mesh decoder 612 is inverse quantized by the inverse quantizer 620 and output to the mesh reconstructor 630. The mesh reconstructor 630 reconstructs the reconstructed deformed mesh (i.e., the decoded mesh) based on the reconstructed displacement output from the inverse linear lifting unit 616 and the reconstructed base mesh output from the inverse quantizer 620. In other words, the inverse quantized reconstructed base mesh is combined with the reconstructed displacement information to generate the final decoded mesh. In this disclosure, the final decoded mesh is referred to as the reconstructed deformed mesh.

[0165] According to an embodiment, the attribute map substream is decoded by the decoder 617 corresponding to the video compression coding decoding used in the encoding, and then the final attribute map (i.e., the decoded attribute map) is reconstructed by the color converter 640 through color format transformation, color space conversion, etc.

[0166] According to an embodiment, the reconstructed decoded mesh and the decoded attribute map can be used as the final mesh data available to the user on the receiving side.

[0167] Refer to Figure 11, the received compressed bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term substream is interpreted to mean a partial bitstream included in the bitstream. The bitstream contains patch information (data), mesh information (data), displacement information (data), and attribute map information (data).

[0168] As described above, Figure 11 the decoder performs intra-frame decoding as follows. The static mesh decoder 612 decodes the mesh substream to generate a reconstructed quantized base mesh, and the inverse quantizer 620 inversely applies the quantization parameters of the quantizer to generate a reconstructed base mesh. The video decoder 613 decodes the displacement substream, the image unpacker 614 unpacks the image of the decoded displacement video, and the inverse quantizer 615 inversely quantizes the quantized image. The inverse linear lifting unit 616 applies a lifting transform in the inverse process of the encoder to generate a reconstructed displacement. The mesh reconstructor 630 generates a reconstructed deformed mesh based on the reconstructed base mesh and the reconstructed displacement. The video decoder 617 decodes the attribute map substream, and the color transformer 64 transforms the color format and / or space of the decoded attribute map to generate a decoded attribute map.

[0169] Figure 12 Shows the inter-frame decoding process of the V-Mesh technology.

[0170] Figure 12 Shows Figure 1 the configuration and operation of the mesh video decoder 113 of the receiving device. In Figure 12 it, the mesh data can be reconstructed by performing the inverse process of the inter-frame encoding process of Figure 7 . Figure 12 The various components of the intra-frame decoding process of

[0171] First, the bitstream received and input to the demultiplexer 711 of the intra-frame decoder 710 can be separated into a motion substream (also referred to as a motion vector substream), a displacement substream, an attribute map substream, and a substream containing patch information about the mesh (e.g., V3C / V-PCC).

[0172] According to an embodiment, the motion substream can be input to and decoded by the motion decoder 712, the displacement substream can be input to and decoded by the video decoder 713, and the attribute map substream can be input to and decoded by the video decoder 717.

[0173] According to an embodiment, the motion substream is decoded by the motion decoder 712 through entropy decoding and inverse prediction to reconstruct motion information (also called motion vector information). The base grid reconstructor 718 combines the reconstructed motion information with the pre-reconstructed and stored reference base grid to generate a reconstructed quantized base grid for the current frame. The inverse quantizer 720 applies inverse quantization to the reconstructed quantized base grid to generate a reconstructed base grid. The video decoder 713 decodes the displacement substream, the image unpacker 714 unpacks the image of the decoded displacement video, and the inverse quantizer 715 inverse quantizes the quantized image. The inverse linear lifting unit 716 applies the lifting transform in the inverse process of the encoder to generate a reconstructed displacement. The grid reconstructor 730 generates a reconstructed deformed grid (i.e., the final decoded grid) based on the reconstructed base grid and the reconstructed displacement.

[0174] According to an embodiment, the video decoder 717 decodes the attribute map substream in the same manner as intra-frame decoding, and the color converter 740 converts the color format and / or space of the decoded attribute map to generate a decoded attribute map. The decoded grid and the decoded attribute map can be used as final grid data available to the user at the receiving side.

[0175] Reference Figure 12 , the bitstream contains motion information (also called motion vectors), displacements, and attribute maps. Because inter-frame decoding is performed, Figure 12 The process also includes decoding inter-frame motion information. The reconstructed basic grid is generated by decoding the motion information and generating a reconstructed quantized basic grid for the motion information based on the reference basic grid. Figure 12 Zhongyu Figure 11 The same operation as in Figure 11 Description.

[0176] Figure 13 A mesh data transmitting device according to an embodiment is shown.

[0177] Figure 13 Corresponds to Figure 1 The transmitting device 100 or the grid video encoder 102, Figure 2 , Figure 6 or Figure 7 encoder (preprocessor and encoder) and / or corresponding sending encoding device. Figure 13 The various components correspond to hardware, software, processors and / or a combination thereof.

[0178] The operation process of using V-Mesh compression technology to compress and send dynamic mesh data at the sending end can be as follows: Figure 13 Configuration shown. Figure 13 The transmitting device may perform intra-frame coding (also known as intra-frame coding or intra-picture coding) and / or inter-frame coding (also known as inter-frame coding or inter-picture coding).

[0179] The preprocessor 811 receives the original mesh and generates a simplified mesh (or base mesh) and a fitted subdivision (or subdivision) mesh. Simplification can be performed based on the vertices of the target number of the constituent mesh or the target number of polygons. Parametrization can be performed on the simplified mesh to generate per-vertex texture coordinates and texture connection information. For example, parametrization is a process of mapping a 3D surface to the texture domain of the simplified mesh. When parametrization is performed using the UVAtlas tool, mapping information indicating where each vertex of the simplified mesh can be mapped onto a 2D image is generated. The mapping information is represented as texture coordinates and stored, and a final base mesh is generated through this process. The mesh information can be quantized from floating-point form to fixed-point form. The result is a base mesh, which can be output to the motion vector encoder 813 or the static mesh encoder 814 through the switch unit 812. The preprocessor 811 can perform mesh subdivision on the base mesh to generate additional vertices. According to the subdivision method, vertex connection information including additional vertices, texture coordinates, and connection information about the texture coordinates can be generated. The preprocessor 811 can generate a fitted subdivision mesh by adjusting the vertex positions so that the subdivision mesh becomes similar to the original mesh.

[0180] According to an embodiment, when performing inter-frame coding on a mesh frame, the base mesh is output to the motion vector encoder 813 through the switch unit 812. When performing intra-frame coding on a mesh frame, the base mesh is output to the static mesh encoder 814 through the switch unit 812. The motion vector encoder 813 can be referred to as a motion encoder.

[0181] For example, when performing intra-frame coding on a mesh frame, the base mesh can be compressed by the static mesh encoder 814. In this case, connection information, vertex geometry information, vertex texture information, normal information, etc. related to the base mesh can be encoded. The base mesh bitstream generated by encoding is sent to the multiplexer 823.

[0182] As another example, when performing inter-frame coding on a mesh frame, the motion vector encoder 813 can receive the base mesh and the reference reconstructed base mesh (or reconstructed quantized reference base mesh) as inputs, calculate the motion vector between the two meshes, and encode its value. In addition, the motion vector encoder 813 can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by encoding is sent to the multiplexer 823.

[0183] The base mesh reconstructor 815 can receive the base mesh encoded by the static mesh encoder 814 or the motion vectors encoded by the motion vector encoder 813, and generate a reconstructed base mesh. For example, the base mesh reconstructor 815 can perform static mesh decoding on the base mesh encoded by the static mesh encoder 814 to reconstruct the base mesh. In this case, quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding. In another example, the base mesh reconstructor 815 can reconstruct the base mesh based on the reconstructed quantization reference base mesh and the motion vectors encoded by the motion vector encoder 813. The reconstructed base mesh is output to the displacement calculator (or displacement vector calculator) 816 and the mesh reconstructor 820.

[0184] The displacement calculator 816 can perform mesh subdivision on the reconstructed base mesh. The displacement calculator 816 can calculate displacement vectors, which are the values of the vertex position differences between the subdivided reconstructed base mesh and the fitted subdivision (or subdivided) mesh generated by the preprocessor 811. In this case, as many displacement vectors can be calculated as there are vertices in the subdivision mesh. The displacement calculator 816 can transform the displacement vectors calculated in the 3D Cartesian coordinate system to the local coordinate system based on the normal vectors of the respective vertices.

[0185] The displacement vector video generator 817 can include a linear lifting section, a quantizer, and an image packer. That is, in the displacement vector video generator 817, the linear lifting unit can transform the displacement vectors for efficient encoding. According to an embodiment, the transformation can be a lifting transformation, a wavelet transformation, etc. Additionally, the quantizer can perform quantization on the transformed displacement vector values (i.e., transformation coefficients). In this case, different quantization parameters can be applied to the axes of the transformation coefficients respectively. The quantization parameters can be derived through an agreement between the encoder / decoder. After the transformation and quantization, the displacement vector information can be packed into a 2D image by the image packer. The displacement vector video generator 817 can generate a displacement vector video by grouping the packed 2D images for each frame. A displacement vector video can be generated for each group of frames (GoF) of the input mesh.

[0186] The displacement vector video encoder 818 can encode the generated displacement vector video using video compression coding and decoding. The generated displacement vector video bitstream is sent to the multiplexer 823.

[0187] The displacement vector reconstructor 819 may include a video decoder, an image unpacker, an inverse quantizer, and an inverse linear lifting section. That is, in the displacement vector reconstructor 819, the encoded displacement vector is decoded by the video decoder, image unpacking is performed by the image unpacker, inverse quantization is performed by the inverse quantizer, and inverse transformation is performed by the inverse linear lifting unit to reconstruct the displacement vector. The reconstructed displacement vector is output to the mesh reconstructor 820. The mesh reconstructor 820 reconstructs a deformed mesh based on the base mesh reconstructed by the base mesh reconstructor 815 and the displacement vector reconstructed by the displacement vector reconstructor 819. The reconstructed mesh (also referred to as the reconstructed deformed mesh) has reconstructed vertices, vertex - to - vertex connection information, texture coordinates, and texture - coordinate - to - texture - coordinate connection information.

[0188] The texture map video generator 821 may regenerate a texture map based on the texture map (or attribute map) of the original mesh and the reconstructed deformed mesh output from the mesh reconstructor 820. According to an embodiment, the texture map video generator 821 may assign per - vertex color information in the texture map of the original mesh to the texture coordinates of the reconstructed deformed mesh. According to an embodiment, the texture map video generator 821 may generate a texture map video by grouping the frame - level regenerated texture maps into a GoF.

[0189] The generated texture map video may be encoded by the texture map video encoder 822 using video compression coding and decoding. The texture map video bitstream generated by encoding is sent to the multiplexer 823.

[0190] The multiplexer 823 multiplexes the motion vector bitstream (e.g., in the case of inter - frame coding), the base mesh bitstream (e.g., in the case of intra - frame coding), the displacement vector bitstream, and the texture map bitstream into a single bitstream. This single bitstream may be sent to the receiving side through the transmitter 824. Alternatively, for the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream, a file having one or more track data may be generated, or the bitstreams may be encapsulated into segments and sent to the receiving side through the transmitter 824.

[0191] Refer to Figure 13, the transmitter (encoder) can encode the mesh in an intra-frame or inter-frame manner. According to intra-frame encoding, the transmitting device can generate a base mesh, displacement vectors (or displacements), and texture maps (or attribute maps). According to inter-frame encoding, the transmitting device can generate motion vectors (or motions), displacement vectors (or displacements), and texture maps (or attribute maps). The texture map obtained from the data input unit is generated and encoded based on the reconstructed mesh. The displacement is generated and encoded based on the vertex position difference between the base mesh and the segmented (or subdivided) mesh. More specifically, the displacement is the position difference between the fitted subdivided mesh and the subdivided reconstructed base mesh, i.e., the vertex position difference between the two meshes. The base mesh is generated by simplifying the original mesh through preprocessing and encoding the simplified mesh. For motion, motion vectors are generated for the mesh in the current frame based on the reference base mesh in the previous frame.

[0192] Figure 14 Shows a mesh data receiving device according to an embodiment.

[0193] Figure 14 Corresponding to Figure 1 the receiving device 110 or the mesh video decoder 113, Figure 11 or Figure 12 the decoder and / or the corresponding receiving and decoding device. Figure 14 The respective components of Figure 14 correspond to hardware, software, processors, and / or combinations thereof. Figure 13 The receiving (decoding) operation of

[0194] The bitstream of the mesh data received by the receiver 910 undergoes file / segment de-encapsulation, and then is demultiplexed by the demultiplexer 911 into a compressed motion vector bitstream (e.g., inter-frame decoding) or a base mesh bitstream (e.g., intra-frame decoding), a displacement vector bitstream, and a texture map bitstream. For example, when the current mesh is inter-frame encoded, the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder 913 through the switch unit 912. In another example, when the current mesh is intra-frame encoded, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder 914 through the switch unit 912. Here, the motion vector decoder 913 can be referred to as a motion decoder.

[0195] According to an embodiment, in the case of applying inter-frame encoding to the current mesh based on the frame header information, the motion vector decoder 913 can decode the motion vector bitstream. According to an embodiment, the motion vector decoder 913 can use the previously decoded motion vector as a predictor and add it to the residual motion vector decoded from the bitstream to reconstruct the final motion vector.

[0196] According to an embodiment, in the case of applying intra coding to the current mesh based on the frame header information, the static mesh decoder 914 may decode the base mesh bitstream to reconstruct connection information, vertex geometry information, texture coordinates, normal information, etc. related to the base mesh.

[0197] According to an embodiment, the base mesh reconstructor 915 may reconstruct the current base mesh based on the decoded motion vector or the decoded base mesh. For example, in the case of applying inter coding to the current mesh, the base mesh reconstructor 915 may add the decoded motion vector to the reference base mesh and perform inverse quantization to generate the reconstructed base mesh. In another example, in the case of applying intra coding to the current mesh, the base mesh reconstructor 915 may perform inverse quantization on the base mesh decoded by the static mesh decoder 914 to generate the reconstructed base mesh.

[0198] According to an embodiment, the displacement vector video decoder 917 may decode the displacement vector bitstream into a video bitstream using video coding decoding.

[0199] According to an embodiment, the displacement vector reconstructor 918 extracts displacement vector transform coefficients from the decoded displacement vector video, and applies inverse quantization and inverse transformation to the extracted displacement vector transform coefficients to reconstruct the displacement vector. To this end, the displacement vector reconstructor 918 may include an image unpacker, an inverse quantizer, and an inverse linear lifting part. If the reconstructed displacement vector is a value in the local coordinate system, an inverse transformation to the Cartesian coordinate system may be performed.

[0200] The mesh reconstructor 916 may subdivide the reconstructed base mesh to generate additional vertices. Through subdivision, vertex connection information including additional vertices, texture coordinates, and connection information about the texture coordinates may be generated. In this case, the mesh reconstructor 916 may combine the subdivided reconstructed base mesh with the reconstructed displacement vector to generate a final reconstructed mesh (also referred to as a reconstructed deformed mesh).

[0201] According to an embodiment, the texture map video decoder 919 may decode the texture map bitstream into a video bitstream using video coding decoding to reconstruct the texture map. The reconstructed texture map has color information about each vertex in the reconstructed mesh, and the texture coordinates of each vertex may be used to obtain the color value of the vertex from the texture map.

[0202] According to an embodiment, the mesh reconstructed from the mesh reconstructor 916 and the texture map reconstructed from the texture map video decoder 919 are presented to the user through the rendering process in the mesh data renderer 920.

[0203] Refer to Figure 14, the receiving device (decoder) can decode the mesh in an intra-frame or inter-frame manner. According to intra-frame decoding, the receiving device can receive a base mesh, displacement vectors (or displacements), and texture maps (or attribute maps), and render the mesh data based on the reconstructed mesh and the reconstructed texture map. According to inter-frame decoding, the receiving device can receive motion vectors (or motions), displacement vectors (or displacements), texture maps (or attribute maps), and render the mesh data based on the reconstructed mesh and the reconstructed texture map.

[0204] The mesh data transmitting device and method according to an embodiment can preprocess mesh data, encode the preprocessed mesh data, and transmit a bitstream including the encoded mesh data. The point mesh data receiving device and method according to an embodiment can receive a bitstream including mesh data and decode the mesh data. The mesh data transmitting / receiving method / device according to an embodiment can be referred to as the method / device according to an embodiment. The mesh data transmitting / receiving method / device can also be referred to as a 3D data transmitting / receiving method / device or a point cloud data transmitting / receiving method / device.

[0205] As described above, the transmitting device for V-Mesh regenerates the texture map of the reconstructed mesh from the texture map of the input original mesh during the encoding process, and then processes the regenerated texture map into a video stream for easy compression. In addition, the encoding process in the transmitting device supports an intra-frame (intra-frame prediction) mode and an inter-frame (inter-frame prediction) mode.

[0206] According to an embodiment, the transmitting device can perform inter-frame prediction on mesh data vertex by vertex. In particular, when encoding geometric information related to 3D dynamic mesh data using inter-frame prediction, per-vertex motion vectors can be calculated, and per-vertex motion vectors or differential motion vectors can be transmitted. In this case, since per-vertex (differential) motion vectors need to be transmitted, the amount of data to be transmitted / parsed is large.

[0207] To solve this problem, the present disclosure proposes a method for encoding and decoding motion vectors or differential motion vectors based on each subgroup, aiming to improve the inter-frame (inter-frame prediction) mode technology of V-Mesh.

[0208] According to an embodiment, when encoding / decoding geometric information related to 3D dynamic mesh data using inter-frame prediction, the motion vectors of similar vertices can be divided into subgroups, and calculated based on each subgroup. Then, per-subgroup motion vectors or differential motion vectors can be transmitted. In particular, when the motion vectors of the vertices within a subgroup are similar, only the (differential) motion vectors of each subgroup can be transmitted, and the transmission of per-vertex (differential) motion vectors can be skipped. Thus, the number of bits to be transmitted / parsed can be reduced. In addition, the present disclosure can allow determining the resolution of motion vectors based on each subgroup.

[0209] In the present disclosure, geometric information (or geometry or geometric data) refers to one of the elements that make up a mesh, including vertices (or points), edges, and polygons. Here, vertices define positions in 3D space, edges represent connection information between vertices, and polygons formed by the combination of edges and vertices define the surface of the mesh. In other words, each vertex that makes up the mesh represents a position in 3D space, for example, represented by X, Y, and Z coordinates. Polygons can be triangles or rectangles. Thus, the geometry forms the skeleton of the 3D model, defining the shape of the model that is visually represented when rendered.

[0210] Figure 15 A mesh data transmitting device according to an embodiment is shown. Figure 15 The transmitting device in may be referred to as an encoder.

[0211] Figure 15 Corresponding to Figure 1 the transmitting device 100 or the mesh video encoder 102, Figure 2 , Figure 6 or Figure 7 the encoder (pre-processor and encoder) of, Figure 13 the transmitting device of and / or the corresponding transmitting and encoding device. Figure 15 Each component in corresponds to hardware, software, a processor, and / or a combination thereof. In Figure 15 , the execution order of the blocks can be changed, some blocks can be omitted, and new blocks can be added.

[0212] In Figure 15 the transmitting device of, the dynamic mesh encoder simplifies the dynamic mesh to create a base mesh, and then uses an incremental encoding method to gradually form a complex mesh from the base mesh to encode the mesh.

[0213] That is, the operation process of compressing and transmitting dynamic mesh data using the V-Mesh compression technology on the transmitting side can be performed as in Figure 15 . Figure 15 The transmitting device in can support both an intra-coding (also referred to as intra-frame coding) process and / or an inter-coding (also referred to as inter-frame coding) process.

[0214] In Figure 15 , the mesh simplification unit 11011 simplifies the input original mesh to generate a base mesh. The mesh simplification can be performed based on the number of target vertices or the number of target polygons that make up the mesh. For example, methods such as simplification can be used to simplify the original mesh. Specifically, the simplification can be a process of selecting vertices to be removed from the original mesh based on a specific reference point, and then removing the selected vertices and the triangles connected to the selected vertices. The base mesh generated by the mesh simplification unit 11011 is input into the mesh quantizer 11012 to be quantized. According to an embodiment, the mesh quantizer 11012 can quantize the mesh information in floating-point form into fixed-point form.

[0215] In addition, the base mesh generated by the mesh simplification unit 11011 is input to the mesh subdivider 11017 and subdivided. That is, the mesh subdivider 11017 performs mesh subdivision on the base mesh to generate additional vertices. According to the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The mesh fitting unit 11018 performs fitting by adjusting the vertex positions so that the subdivided mesh from the mesh subdivider 11017 becomes similar to the original mesh, thereby generating a fitted subdivided mesh.

[0216] In the present disclosure, the combination of the mesh simplification unit 11011, the mesh subdivider 11017, and the mesh fitting unit 11018 may be referred to as a pre-processor. According to an embodiment, the pre-processor may further include a displacement vector calculator.

[0217] In addition, the pre-processor may perform parameterization to generate per-vertex texture coordinates and texture connection information of the simplified mesh (i.e., the base mesh). For example, parameterization is the process of mapping the 3D surface of the simplified mesh to the texture domain. If parameterization is performed using the UVAtlas tool, mapping information for identifying where each vertex of the simplified mesh can be mapped on the 2D image is generated. The mapping information is represented and stored as texture coordinates. Through this process, the final base mesh is generated. The final base mesh (i.e., the simplified mesh with texture coordinates). The final base mesh (with texture coordinates) is input to the mesh subdivider 11017 and can be subdivided.

[0218] According to an embodiment, the base mesh from the mesh quantizer 11012 may be output to the motion vector encoder 11014 or the static mesh encoder 11015 through the switching unit 11013.

[0219] According to an embodiment, when performing inter-frame coding on a mesh frame, the base mesh is output to the motion vector encoder 11014 through the switching unit 11013. When performing intra-frame coding on a mesh frame, the base mesh is output to the static mesh encoder 11015 through the switching unit 11013. The motion vector encoder 11014 may be referred to as a motion encoder.

[0220] For example, when performing intra-frame coding on a mesh frame, the base mesh may be compressed by the static mesh encoder 11015. In this case, connection information, vertex geometry information, vertex texture information, normal information, etc. related to the base mesh may be encoded. The base mesh bitstream generated by the encoding is sent to a multiplexer (not shown).

[0221] As another example, when performing inter-frame coding on a mesh frame, the motion vector encoder 11014 may receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as inputs, calculate the motion vector between the two meshes, and encode its value. In addition, the motion vector encoder 11014 may perform prediction based on connection information using previously encoded / decoded motion vectors as predictors, and encode the differential motion vector (also referred to as the residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by encoding is sent to a multiplexer (not shown) as the base mesh bitstream. That is, in the case of intra-frame coding, the static mesh bitstream is input to the multiplexer as the base mesh bitstream. In the case of inter-frame coding, the motion vector bitstream is input to the multiplexer as the base mesh bitstream.

[0222] According to an embodiment, the motion vector encoder 11014 may calculate a motion vector or a differential motion vector based on subgroups.

[0223] The operation of the motion vector encoder 11014 proposed in the present disclosure to divide subgroups based on each subgroup and obtain (differential) motion vectors will be described in detail later.

[0224] In this regard, the subgroup division information output from the motion vector encoder 11014 is encoded as auxiliary information by the auxiliary information encoder 11016. According to an embodiment, the subgroup division information may include a subgroup division method. According to the subgroup division method, it may also include information such as the initial number of clusters (in the case of K-means clustering), the clustering level (in the case of hierarchical clustering), or an octree structure (in the case of octree splitting).

[0225] According to an embodiment, the auxiliary information encoder 11016 may encode the auxiliary information and output an auxiliary information bitstream to a multiplexer (not shown).

[0226] According to an embodiment, the auxiliary information may further include auxiliary patch information. According to an embodiment, the auxiliary patch information may include an index (clustering index) for identifying a projection plane (normal), the 3D spatial position of the patch (e.g., the tangential minimum of the patch (patch 3d shift tangential axis), the double tangential minimum of the patch (patch 3d shift double tangential axis), the normal minimum of the patch (patch 3d shift normal axis)), the 2D spatial position and size of the patch (e.g., the horizontal size (patch 2d size u), the vertical size (patch 2d size v), the horizontal minimum (patch 2d shift u), the vertical minimum (patch 2d shift u)), mapping information regarding each block and patch (e.g., a candidate index (note that when patches are sorted based on the 2D spatial position and size information of the patch, multiple patches may be mapped to a block. In this case, the mapped patches form a candidate list, and the index indicates the patch whose data exists in the block)), and a local patch index (an index indicating one of the patches present in the frame). In other words, patches may be generated for 2D image mapping of the mesh data, and auxiliary patch information may be generated as a result of patch generation. The auxiliary patch information may be used in the geometric reconstruction process. The generated patches may be mapped onto a 2D image through a patch packing process.

[0227] In one embodiment of the present disclosure, patch generation and patch packing may be performed during the parameterization of the pre-processor.

[0228] In Figure 15 , the mesh reconstructor 11020 may receive the base mesh encoded by the static mesh encoder 11015 or the motion vectors encoded by the motion vector encoder 11014, and may generate a reconstructed base mesh. For example, the mesh reconstructor 11020 may reconstruct the base mesh by performing static mesh decoding on the base mesh encoded by the static mesh encoder 11015. In this case, quantization may be applied before static mesh decoding, and inverse quantization may be applied after static mesh decoding. Alternatively, the mesh reconstructor 11020 may reconstruct the base mesh based on the reconstructed quantization reference base mesh and the motion vectors encoded by the motion vector encoder 11014. The reconstructed base mesh is output to the displacement vector calculator 11019 and the texture map generator 11022.

[0229] According to an embodiment, the displacement vector calculator 11019 may perform mesh subdivision on the reconstructed base mesh. In addition, the displacement vector calculator 11019 may calculate displacement vectors, i.e., the values of the vertex position differences between the subdivided reconstructed base mesh and the fitted subdivided mesh generated by the mesh fitting unit 11018. In this case, as many displacement vectors as there are vertices in the subdivided mesh may be calculated. The displacement vector calculator 11019 may transform the displacement vectors calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vectors of the respective vertices.

[0230] According to an embodiment, the displacement vector calculator 11019 or a displacement vector video generator (not shown) disposed between the displacement vector calculator 11019 and the displacement vector video encoder 11021 may include a linear lifting part, a quantizer, and an image packing part. In this case, the linear lifting unit may transform the displacement vector for efficient encoding. According to an embodiment, the applied transformation may be a lifting transformation, a wavelet transformation, etc. Additionally, the quantizer may perform quantization on the transformed displacement vector values (i.e., transformation coefficients). In this case, different quantization parameters may be applied to the axes of the transformation coefficients. The quantization parameters may be derived through an agreement between the encoder / decoder. After transformation and quantization, the image packer may pack the displacement vector information into a 2D image. The displacement vector video generator may generate a displacement vector video by grouping the packed 2D images for each frame. A displacement vector video may be generated for each group of frames (GoF) of the input mesh.

[0231] According to an embodiment, the displacement vector video encoder 11021 may encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is sent to a multiplexer (not shown). According to an embodiment, regarding the selection of the displacement vector video encoder 11021, a displacement vector video encoder agreed upon by the encoder (i.e., the transmitting side) and the decoder (i.e., the receiving side) may be used, or the encoder on the transmitting side may analyze the characteristics of the displacement vector and send the type of the selected displacement vector encoder to the decoder on the receiving side. According to an embodiment, the displacement vector video encoder 11021 may transform and quantize the input displacement vector.

[0232] According to an embodiment, a displacement vector reconstructor may be further disposed between the displacement vector video encoder 11021 and the texture map generator 11022.

[0233] According to an embodiment, the displacement vector reconstructor may include a video decoder, an image unpacker, an inverse quantizer, and an inverse linear lifting unit. That is, in the displacement vector reconstructor, the encoded displacement vector is decoded by the video decoder, image unpacking is performed by the image unpacker, inverse quantization is performed by the inverse quantizer, and inverse transformation is performed by the inverse linear lifting unit to reconstruct the displacement vector. Additionally, the displacement vector reconstructor may reconstruct the deformed mesh based on the reconstructed displacement vector and the base mesh reconstructed by the mesh reconstructor 11020. The reconstructed mesh (also referred to as the reconstructed deformed mesh) has reconstructed vertices, vertex - to - vertex connection information, texture coordinates, and texture - coordinate - to - texture - coordinate connection information.

[0234] According to an embodiment, the texture map generator 11022 may regenerate a texture map based on the texture map (or attribute map) of the original mesh and the base mesh reconstructed by the mesh reconstructor 11020 (or the deformed mesh reconstructed by the displacement vector reconstructor). According to an embodiment, the texture map generator 11022 may assign vertex-specific color information in the texture map of the original mesh to the texture coordinates of the reconstructed base mesh (or the reconstructed deformed mesh). According to an embodiment, the texture map generator 11022 may generate a texture map video by grouping the frame-level regenerated texture maps into GoFs.

[0235] According to an embodiment, the texture map video generated by the texture map generator 11022 may be encoded using a video compression codec of the texture map video encoder 11023. The generated texture map video bitstream by encoding is sent to a multiplexer (not shown).

[0236] According to an embodiment, the type of the texture map video encoder may include a video encoder (e.g., VVC, HEVC, etc.) and an entropy coding-based encoder. Regarding the selection of the texture map video encoder, the texture map video encoder agreed upon by the encoder (i.e., the transmitting side) and the decoder (i.e., the receiving side) may be used, or the encoder on the transmitting side may analyze the characteristics of the texture map video encoder and send the type of the selected texture map video encoder to the decoder on the receiving side.

[0237] According to an embodiment, the multiplexer may multiplex the input auxiliary information bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream into a single bitstream, which may then be sent to the receiving side through a transmitter (not shown). Alternatively, the auxiliary information bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be encapsulated into a file / fragment and sent to the receiving side through a transmitter.

[0238] Hereinafter, a detailed description of calculating and transmitting a motion vector or a differential motion vector based on each subgroup will be given.

[0239] Figure 16 is an exemplary detailed block diagram of a motion vector encoder according to an embodiment. That is, Figure 16 is Figure 15 an example of the motion vector encoder 11014 shown. Figure 16 Each component in Figure 16 corresponds to hardware, software, a processor, and / or a combination thereof. In

[0240] In Figure 16In one embodiment, the vertex motion vector calculator 12012 may receive geometric information about a reference reconstructed base mesh from the reference reconstructed mesh buffer 12011 and geometric information about a quantized base mesh from the mesh quantizer 11012, calculate the motion vector for each vertex, and output it to the subgroup splitter 12013.

[0241] In this case, since the input base mesh (e.g., the quantized base mesh) and the reference base mesh (e.g., the reference reconstructed base mesh) are in one-to-one correspondence between vertices, the motion vector can be obtained based on the difference in geometric information between the input base mesh and the reference base mesh having the same index. In the present disclosure, the reference reconstructed base mesh may have the same meaning as the reconstructed quantized reference base mesh, the reconstructed reference mesh, or the reference base mesh and may be used interchangeably.

[0242] According to an embodiment, the subgroup splitter 12013 may split the reference base mesh (or geometric information about the reference base mesh) into subgroups of similar objects based on similarity or distance (e.g., similarity of the motion vector of a vertex and geometric information about the vertex). In this case, the subgroup splitting of the subgroup splitter 12013 is calculated based on the splitting method in the encoder (e.g., the motion vector encoder). And the corresponding subgroup splitting information may be signaled, or the subgroup splitting information may be derived in the same way by the encoder (i.e., the motion vector encoder on the transmitting side) / decoder (i.e., the motion vector decoder on the receiving side) based on the reconstructed geometric information or the reconstructed motion vector of the reconstructed reference mesh (i.e., the reference base mesh).

[0243] According to an embodiment, the subgroup splitter 12013 may determine subgroups based on the vertex order in the base mesh or the order in which the motion vectors are encoded, and split the reference base mesh into the determined subgroups. In this case, the size of the subgroups may be predefined according to an agreement between the encoder and the decoder. Alternatively, the size of the subgroups may be signaled and sent to the decoder of the receiving device.

[0244] Various subgroup splitting methods may be applied. According to an embodiment, the subgroup splitting methods include octree splitting, K-means clustering, hierarchical clustering, Kd-tree splitting, and patch segmentation.

[0245] According to an embodiment, for the subgroup splitting method, the same splitting method agreed upon by the encoder on the transmitting side and the decoder on the receiving side may be selected, and / or the encoder on the transmitting side may add index information (e.g., partition_type_idx) related to the selected splitting method to the auxiliary header and send it to the decoder on the receiving side. In the latter case, the decoder on the receiving side may identify the splitting method based on the index information and split the reference base mesh into subgroups based on the identified splitting method, and reconstruct the motion vectors.

[0246] According to an embodiment, when the subgroup splitter 12013 performs splitting, the splitting method and the basis for deriving the splitting may vary according to the frame type of the reference base grid.

[0247] According to an embodiment, the splitting method may be implicitly determined based on the frame type of the reference base grid (e.g., I-frame, P-frame, or B-frame). That is, different data may be used during the splitting process according to the frame type of the reference base grid.

[0248] For example, when the frame type of the reference base grid is a B-frame or a P-frame, the grid may be split based on the reconstructed motion vector of the reference base grid. When the reference base grid is an I-frame, the grid may be split based on the reconstructed vertex geometry information related to the reference base grid. In other words, when the encoder / decoder uses the same method to derive the same subgroup splitting information, when the reference base grid is a B-frame or a P-frame, the splitting may be performed based on the reconstructed motion vector. When the reference base grid is an I-frame, the splitting may be performed based on the reconstructed vertex geometry information related to the reference base grid.

[0249] In other words, the goal of splitting into subgroups depends on whether the reference base grid is an I-frame or a P- or B-frame. When the reference base grid is an I-frame, it has no motion vector information, so the splitting goal is the vertex geometry information. When the reference base grid is a P-frame or a B-frame, the splitting goal is the motion vector.

[0250] Next, each subgroup splitting method will be described below.

[0251] 1) Example of splitting method 1 (octree splitting)

[0252] When the subgroup splitting method is octree splitting, the subgroup splitter 12013 may recursively split the cube and determine whether to perform splitting based on the minimum number of vertices in the splitting region, the vertex distribution in the splitting region, etc.

[0253] Then, the rate distortion may be calculated based on the motion vector, and the octree splitting may be calculated based on the result of the rate distortion. Then, the subgroup splitting information may be signaled. In other words, during the octree splitting process, the distribution of the values of the motion vectors split into the same subgroup may be checked to determine whether they are similar. If the values are similar, further splitting to a lower level may not be performed; if they are not similar, a recursive splitting operation to a lower level may be performed. This process is repeated. Then, in the finally split octree structure, each cube will represent a subgroup.

[0254] According to an embodiment, when the splitting method is octree splitting, the octree splitting information (e.g., octree_partitioning_data) may be included in the auxiliary information header and sent to the receiving side.

[0255] According to an embodiment, signaling of segmentation information for a space of vertices that cannot be derived based on geometric information about a reference base grid may be omitted on the transmitting side. In other words, the segmentation information may not be sent to the receiving side.

[0256] In addition, the bounding box size of the octree may be derived in the same manner by the encoder / decoder based on geometric information about the reference base grid.

[0257] 2) Example of segmentation method 2 (K-means clustering)

[0258] According to an embodiment, when the subgroup segmentation method is K-means clustering, the process of calculating the distance between the centroid of each cluster and each vertex or motion vector and forming clusters of vertices or motion vectors that are close to each other may be iteratively performed.

[0259] In K-means clustering, when the reference base grid is a P-frame or a B-frame, the encoder on the transmitting side may calculate the initial number of clusters based on the motion vectors of the reference base grid. The initial number may be derived by the encoder / decoder. Additionally / Alternatively, after the encoder on the transmitting side sets the initial number of clusters, the set cluster number information (e.g., number_of_cluster) may be included in the auxiliary header and sent to the decoder on the receiving side.

[0260] If the reference base grid is an I-frame, the encoder on the transmitting side may calculate the initial number of clusters based on the vertex geometric information about the reference base grid, and send the calculated cluster number information (e.g., number_of_cluster) to the decoder on the receiving side in the auxiliary header. Additionally / Alternatively, the encoder / decoder may use a fixed number of clusters.

[0261] In one embodiment, the distance (e.g., geodesic distance, Euclidean distance, etc.) between the centroid of the initialized clusters and each vertex may be calculated to form clusters of vertices that are close to each other. In the present disclosure, the term "cluster" may have the meaning of a subgroup. Additionally, the initial number of clusters may have the meaning of the number of clusters.

[0262] Therefore, when the reference base grid is a P-frame or a B-frame, the motion vectors of the reference base grid may be used to derive the initial number of clusters. For example, the distribution of the motion vectors may be calculated, and the initial number of clusters may be derived based on the characteristics of the distribution. The initial number of clusters indicates how many clusters should be set in the K-means clustering process. Based on the initial number of clusters, the number of subgroups to be segmented may be determined.

[0263] In addition, when the reference base grid is an I-frame, the method for calculating the initial number of clusters may include calculating the distribution of vertex geometric information (i.e., the distribution of vertex coordinates related to the reference base grid), and deriving the number of clusters based on the characteristics of the distribution. Alternatively, the initial number of clusters may be determined by user parameters (encoder parameters).

[0264] 3) Example of segmentation method 3 (hierarchical clustering)

[0265] According to an embodiment, when the subgroup segmentation method is hierarchical clustering, all vertices or motion vectors may be set as a single cluster in a top-down manner, and then clusters with high similarity may be sequentially merged for clustering. Alternatively, each vertex or motion vector may be set as a single cluster in a bottom-up manner, and then clusters with high similarity may be sequentially merged for clustering. In this case, in the initial process of hierarchical clustering, each entity may initially be set as a single cluster, and then the clustering process may be sequentially performed. The clusters obtained through the final clustering may become subgroups.

[0266] For example, when the reference base grid is a P-frame or a B-frame, it may be initialized to form a group for each vertex, and then clustering may be performed by merging vertices with similar motion vectors. As another example, when the reference base grid is an I-frame, it may be initialized to form a group for each vertex, and then clustering may be performed by merging similar vertices.

[0267] In this case, information about the hierarchical level (or clustering level) may be derived by the encoder / decoder. Alternatively, the encoder may set the hierarchical level (or clustering level), and then send the clustering level information (e.g., level_of_cluster) in the auxiliary information header to the decoder on the receiving side.

[0268] Here, the hierarchical level may be used interchangeably with the clustering level. That is, the hierarchical level (e.g., level_of_cluster) refers to the final depth at which hierarchical clustering is performed. Additionally, if the encoder / decoder uses a fixed hierarchical level, the hierarchical level may not be sent.

[0269] 4) Example of segmentation method 4 (patch segmentation)

[0270] When the subgroup segmentation method is patch segmentation, one or more patches may be formed into a single subgroup. A patch may be a set of faces and may be a connected component in the texture coordinate space. In other words, one or more patches may represent a single subgroup.

[0271] According to an embodiment, the encoder may calculate the rate distortion based on the vertex geometry or motion vectors of the reference base grid, segment the reference base grid into patches based on the calculated rate distortion, and signal the patch segmentation information (e.g., patch parameter set).

[0272] Here, a patch may be a set of vertex geometric information about one or more vertices or a set of motion vectors of one or more vertices. Additionally, a patch may be a set of one or more texture coordinates or a set of faces.

[0273] According to an embodiment, the encoder may add patch segmentation information (e.g., the number of patches per patch group of one or more patches (e.g., number_of_patch)) to the patch parameter set to be sent to the decoder on the receiving side.

[0274] Figure 18 An example of the Nth patch in the texture space according to an embodiment is shown. In this case, the patch may be a connected component in the texture coordinate space or may be given as shown Figure 18 shown.

[0275] Here, patches may be packed / unpacked in various orders (e.g., raster scan order, zigzag scan order, etc.).

[0276] According to an embodiment, the subgroup splitter 12013 may determine subgroups based on the vertex order of the base mesh or the order in which motion vectors are encoded. In this case, the size of the subgroup may be defined according to an agreement between the encoder / decoder. Alternatively, the size of the subgroup may be included in the subgroup segmentation information and sent to the decoder of the receiving device. In the latter case, the decoder of the receiving device may determine the size of the subgroup by parsing the signaling information including the subgroup segmentation information.

[0277] As described above, the subgroup splitter 12013 divides the reference base mesh into one or more subgroups. Then, the subgroup segmentation information including the segmentation method is encoded by the auxiliary information encoder, and then an auxiliary information bitstream related thereto is output.

[0278] According to an embodiment, the subgroup motion vector calculator 12014 may determine whether to skip the (differential) motion vectors of respective subgroups and / or vertices. Based on this determination, it may calculate and send the (differential) motion vectors of respective subgroups and / or vertices or skip the transmission. For example, when it is determined to skip, the motion vectors of all vertices in the subgroup may be set to a zero vector. In other words, when it is determined to skip, the corresponding (differential) motion vector is not sent to the receiving side.

[0279] According to an embodiment, the encoder signals skip flag information (or a skip flag) indicating whether to skip a (differential) motion vector. In the present disclosure, two items of skip flag information may be signaled, which may be simply referred to as first skip flag information (e.g., mvd_skip_flag) and second skip flag information (e.g., vertex_mvd_skip_flag) for simplicity. This is only an embodiment, and only one of the two items of skip flag information may be signaled.

[0280] That is, based on the motion vectors of subgroups or vertices within a subgroup, rate distortion may be calculated. Based on the calculation result, a skip flag may be sent for each subgroup or vertex, and the (differential) motion vector may not be sent. Here, calculating rate distortion based on the motion vectors of subgroups or vertices within a subgroup aims to determine whether the motion vectors of the vertices within the subgroup are similar.

[0281] Various methods may be used to calculate rate distortion based on the motion vectors of subgroups or vertices within a subgroup. For example, the rate distortion cost may be calculated based on the difference in point cloud-based D1-PSNR, D2-PSNR, and the number of bits between the case where the motion vectors of the vertices within a subgroup are not sent and the case where the motion vectors are sent.

[0282] According to an embodiment, the distribution of the motion vectors within a subgroup may be calculated. When the variance is small, it may be determined that the motion vectors within the subgroup are similar.

[0283] According to an embodiment, the encoder signals the first skip flag information (e.g., mvd_skip_flag). The first skip flag information may be referred to as the (differential) motion vector skip flag for subgroups and vertices. According to the present disclosure, based on the first skip flag information (e.g., mvd_skip_flag), the transmission of the (differential) motion vector for each subgroup and the (differential) motion vector of the vertices may be skipped. In this case, the decoder of the receiving device may derive the (differential) motion vector for each subgroup and the (differential) motion vector of the vertices as zero vectors.

[0284] According to an embodiment, the encoder signals the second skip flag information (e.g., vertex_mvd_skip_flag). The second skip flag information may be referred to as the differential motion vector skip flag for vertices. Based on the second skip flag information (e.g., vertex_mvd_skip_flag), the transmission of the (differential) motion vector for each vertex may be skipped.

[0285] According to an embodiment, based on at least one of the first skip flag information (e.g., mvd_skip_flag) and the second skip flag information (e.g., vertex_mvd_skip_flag), the (differential) motion vector for each subgroup etc. may be sent to the receiving device.

[0286] According to an embodiment, encoder parameters may be sent based on each tile or slice.

[0287] In the present disclosure, a tile / slice may be a unit that divides dynamic mesh data and encodes / decodes the divided regions independently.

[0288] According to an embodiment, a tile / slice may be composed of one or more subgroups.

[0289] According to an embodiment, the encoder may send various encoding parameters to the decoder of the receiving device, such as motion vector resolution information (e.g., mvd_resolution_idx) and quantization parameters for each divided unit (e.g., patch group, subgroup, slice, tile, etc.).

[0290] Figure 17 is a diagram showing an example of a process of checking whether to skip a motion vector and calculating a motion vector according to an embodiment. That is, Figure 17 shows Figure 16 an example of the detailed operation of the subgroup motion vector calculator 12014 in. In Figure 17 it, the execution order of blocks may be changed, some blocks may be omitted, and new blocks may be added.

[0291] That is, the vertex motion vector calculated by the vertex motion vector calculator 12012 is input to the subgroup motion vector calculator 12014 based on each subgroup or each vertex.

[0292] According to an embodiment, the subgroup motion vector calculator 12014 may check whether to skip the encoding of the differential motion vector (13011), and based on the check result, calculate the differential motion vector of the subgroup (i.e., the difference between the motion vector of the subgroup and the predicted motion vector of the subgroup).

[0293] According to an embodiment, it may be determined whether to skip the encoding of the differential motion vector based on at least one of first skip flag information (e.g., mvd_skip_flag) and second skip flag information (e.g., vertex_mvd_skip_flag).

[0294] In Figure 17In one embodiment, when the value of the first skip flag information (e.g., mvd_skip_flag) is 1 (13012), the transmission of the differential motion vectors of the subgroup can be skipped. In this case, according to one embodiment, the motion vectors of all vertices in the subgroup can be determined as the zero vector (0, 0, 0). Additionally, in one embodiment, when the value of the first skip flag information (e.g., mvd_skip_flag) is 1 (13012), the transmission of the differential motion vector of the vertex can be skipped. In this case, according to one embodiment, the motion vector of the vertex can be determined as the zero vector (0, 0, 0).

[0295] When the value of the first skip flag information (e.g., mvd_skip_flag) is 0 (13012), the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is checked to determine when it is 1 or 0 (13013).

[0296] In one embodiment, when the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 1 (13015), the subgroup motion vector calculator 12014 can calculate the differential motion vector based on each subgroup. When the value is 0, the calculator can calculate the differential motion vector based on each vertex and output it (13014).

[0297] According to an embodiment, for the motion vector of the subgroup, the subgroup differential motion vector calculator 13015 can obtain the average of the motion vectors of the vertices in the subgroup as the representative value of the motion vectors of the vertices in the subgroup. In other words, the average of the motion vectors of the vertices in the current subgroup can be used as the motion vector of the current subgroup.

[0298] According to an embodiment, the subgroup differential motion vector calculator 13015 can calculate the predicted motion vector of the subgroup based on the motion vectors of the earlier decoded neighbor vertices in the current frame or the motion vectors of the reference frame.

[0299] For example, the predicted motion vector of the subgroup can be obtained by predicting based on the motion vectors of the subgroups that are spatially close to the current subgroup among the earlier reconstructed subgroups or the average motion vector of the N vertices closest to the vertices of the current subgroup. In other words, the motion vectors of the subgroups that are spatially close to the current subgroup among the earlier reconstructed subgroups can be used as the predicted motion vector of the current subgroup, or the average motion vector of the N vertices closest to the vertices of the current subgroup can be used as the predicted motion vector of the current subgroup.

[0300] Additionally, when the reference base grid is a P frame or a B frame, the average motion vector of the reference base grid vertices having the same index as the current base grid vertices can be used as the predicted motion vector of the current subgroup.

[0301] According to an embodiment, the subgroup differential motion vector calculator 13015 may calculate the differential motion vector of a subgroup as the difference between the motion vector of the subgroup and the predicted motion vector of the subgroup. That is, the differential motion vector of the current subgroup is obtained by subtracting the predicted motion vector of the current subgroup from the motion vector of the current subgroup.

[0302] According to an embodiment, the vertex differential motion vector calculator 13014 may calculate the differential motion vector of each vertex by obtaining the difference between the predicted motion vector of the vertex and the motion vector of the vertex in the subgroup. Here, the motion vector of the vertex in the subgroup is the motion vector calculated by the vertex motion vector calculator 12012.

[0303] According to an embodiment, when the motion vectors of the vertices in the subgroup are not similar to each other, the vertex differential motion vector calculator 13014 calculates the differential motion vector of each vertex. That is, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 0, it may be determined that the motion vectors of the vertices in the subgroup are not similar to each other.

[0304] The vertex differential motion vector calculator 13014 may calculate the predicted motion vector of the vertices in the subgroup based on the previously reconstructed vertex motion vectors. In this case, the prediction may be performed by averaging the motion vectors or parallelogram prediction or averaging of multiple parallelogram predictions of adjacent vertices based on the connection information.

[0305] According to an embodiment, the second skip flag information (e.g., vertex_mvd_skip_flag) may be omitted. In this case, based on the first skip flag information (e.g., mvd_skip_flag), the differential motion vector of the subgroup may be skipped or transmitted. For example, the first skip flag information (e.g., mvd_skip_flag) is signaled for each subgroup. When the value of the first skip flag information (e.g., mvd_skip_flag) is 1, the motion vectors of all vertices in the corresponding subgroup may be determined as zero vectors (0, 0, 0) (i.e., skipped). In other words, when the second skip flag information (e.g., vertex_mvd_skip_flag) is omitted, the decoder of the receiving device may receive the motion vectors parsed per vertex and perform inverse quantization on the vertex (differential) motion vectors.

[0306] According to an embodiment, the differential motion vector per vertex or per subgroup output from the subgroup motion vector calculator 12014 is quantized by the motion vector quantizer 12015 and then output to the geometric information entropy encoder 12016. In the present disclosure, the motion vector quantizer 12015 may be omitted.

[0307] According to an embodiment, the geometric information entropy encoder 12016 may perform entropy encoding on subgroup differential motion vector transmission skip flags (e.g., first skip flag information), vertex differential motion vector transmission skip flags (e.g., second skip flag information), differential motion vectors per subgroup, differential motion vectors per vertex, etc., and output a geometric information bitstream. In the present disclosure, the geometric information bitstream may be referred to as a base mesh bitstream or a motion vector bitstream.

[0308] According to an embodiment, the geometric information entropy encoder 12016 may use various coding methods to encode subgroup differential motion vector transmission skip flags (e.g., first skip flag information), vertex differential motion vector transmission skip flags (e.g., second skip flag information), differential motion vectors per subgroup, and differential motion vectors per vertex, such as context-adaptive binary arithmetic coding (CABAC), exponential Golomb, variable length coding (VLC), or context-adaptive variable length coding (CAVLC). In the present disclosure, the type of entropy encoder / decoder may be determined by an agreement between the encoder (i.e., the transmitting side) / decoder (i.e., the receiving side), or the receiving device may determine the encoder / decoder type determined by the encoder of the transmitting device by receiving and parsing the bitstream.

[0309] Next, a method of selecting the resolution of the differential motion vector will be described.

[0310] When performing differential motion vector coding, the motion vector encoder 11014 may select the resolution of the differential motion vector based on each subgroup.

[0311] Then, the index of the differential motion vector resolution (e.g., mvd_resolution_idx) may be signaled per subgroup. In the present disclosure, the index of the differential motion vector resolution (e.g., mvd_resolution_idx) may be referred to as motion vector resolution information.

[0312] According to an embodiment, the motion vector resolution index for each higher-level unit (e.g., slice, tile, patch group, etc.) may be signaled. Thus, for a subgroup with less motion of the unit, a higher resolution of the differential motion vector may be selected. For a subgroup with more motion, a lower resolution may be selected for the differential motion vector.

[0313] The motion vector resolution may be selected based on each subgroup or based on each vertex motion vector.

[0314] In addition, for the x, y, and z components of the motion vector, the resolution of the differential motion vector may be the same, or different resolutions may be selected for the components respectively.

[0315] Figure 19An example of the motion vector resolution according to the motion vector resolution information according to an embodiment is shown. That is, Figure 19 An example of the differential motion vector resolution index (i.e., motion vector resolution information) based on the resolution of the differential motion vector is shown. For example, the motion vector resolution information (mvd_resolution_idx) set to 0 may indicate that a differential motion vector resolution equal to 0.25 is selected. The motion vector resolution information (mvd_resolution_idx) set to 1 may indicate that a differential motion vector resolution equal to 0.5 is selected. The motion vector resolution information (mvd_resolution_idx) set to 2 may indicate that a differential motion vector resolution equal to 1 is selected. The motion vector resolution information (mvd_resolution_idx) set to 3 may indicate that a differential motion vector resolution equal to 2 is selected.

[0316] Figure 20 FIG. is a diagram showing an example of the motion vector resolution based on each subgroup in an octree structure according to an embodiment. That is, Figure 20 An example of configuring a motion vector resolution index based on each subgroup when the subgroup splitting method is octree splitting is shown.

[0317] In the present disclosure, the resolution of the motion vector refers to the accuracy of the motion vector.

[0318] For example, when the motion vector resolution information (mvd_resolution_idx) is 3 (i.e., a differential motion vector resolution equal to 2 is selected), it means that during the decoding process of the motion vector on the receiving device, decoding is performed with twice the value of the decoded motion vector. For example, when the motion is large (i.e., the motion vector has a large resolution), the motion vector encoder 1104 may reduce the motion vector resolution to encode with fewer bits.

[0319] On the other hand, when the motion vector resolution information (mvd_resolution_idx) is 1, that is, when a differential motion vector resolution equal to 1 / 2 is selected, it means that the receiving device performs decoding with half of that value.

[0320] In the present disclosure, it is possible to determine whether the motion is large or small based on the motion vector value of each subgroup.

[0321] In addition, for each subgroup generated using not only octree splitting but also various splitting methods, the (differential) motion vector resolution may be determined differently. The motion vector resolution may be set for each subgroup based on motion vector information (e.g., motion vector distribution, motion vector average, etc.).

[0322] The auxiliary information bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream generated through the above process can be multiplexed by a multiplexer into a single bitstream. Then, it can be sent via a network or stored on a digital storage medium. Here, the network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0323] Figure 21 Shows a mesh data receiving device according to an embodiment. In the present disclosure, Figure 21 the receiving device can be referred to as a decoder.

[0324] Figure 21 Corresponds to Figure 1 the receiving device 110 or the mesh video decoder 113, Figure 11 or Figure 12 the decoder, Figure 14 the receiving device and / or the corresponding receiving and decoding device. Figure 21 Each component of Figure 21 corresponds to hardware, software, a processor, and / or a combination thereof. Figure 15 The receiving (decoding) operation in Figure 21 can follow a process opposite to the corresponding process of the sending (encoding) operation in

[0325] According to an embodiment, the mesh data bitstream received by a receiver (not shown) undergoes file / fragment de-encapsulation and then is demultiplexed by a demultiplexer (not shown) into an auxiliary information bitstream, a base mesh bitstream (or geometric information bitstream), a displacement vector bitstream, and a texture map bitstream. In the case of applying inter-frame coding to the current mesh, the base mesh bitstream (or geometric information bitstream) can be a motion vector bitstream.

[0326] According to an embodiment, the auxiliary information decoder 15011 decodes the auxiliary information bitstream and outputs the auxiliary information.

[0327] The base mesh bitstream is output to the motion vector decoder 15014 or the static mesh decoder 15015 via the switching unit 15013.

[0328] For example, in the case of applying inter-frame coding to the current mesh, a base mesh bitstream (i.e., a motion vector bitstream) is received, demultiplexed, and output to a motion vector decoder 15014 via a switching unit 15013. In another example, in the case of applying intra-frame coding to the current mesh, a base mesh bitstream is received, demultiplexed, and output to a static mesh decoder 15015 via the switching unit 15013. Here, the motion vector decoder 15014 may be referred to as a motion decoder.

[0329] According to an embodiment, the motion vector decoder 15014 may decode the motion vector bitstream based on per-vertex or per-subgroup. To this end, the receiving device may further include a subgroup splitter 15012. The subgroup splitter 15012 divides a reference base mesh into subgroups based on auxiliary information, and subgroup division information included in the decoded auxiliary information is output to the motion vector decoder 15014. A detailed operation of the subgroup splitter 15012 will be described later.

[0330] According to an embodiment, the motion vector decoder 15014 may reconstruct a final motion vector by using a previously decoded motion vector as a predictor and adding a differential motion vector (i.e., a residual motion vector) decoded from the bitstream thereto.

[0331] According to an embodiment, the static mesh decoder 15015 may decode the base mesh bitstream to reconstruct connection information, vertex geometry information, texture coordinates (i.e., attribute geometry information), normal information, etc. related to the base mesh.

[0332] According to an embodiment, the base mesh reconstructor 15016 may reconstruct a current base mesh based on the decoded motion vector or the decoded base mesh. For example, in the case of applying inter-frame coding to the current mesh, the base mesh reconstructor 15016 may reconstruct the base mesh by adding the decoded (or reconstructed) motion vector to a reference base mesh and performing inverse quantization. In another example, in the case of applying intra-frame coding to the current mesh, the base mesh reconstructor 15016 may reconstruct the base mesh by performing inverse quantization on the base mesh decoded (or reconstructed) by the static mesh decoder 15015.

[0333] According to an embodiment, the displacement vector video decoder 15018 may decode a displacement vector bitstream into a video bitstream using a video codec.

[0334] According to an embodiment, the type of the displacement vector video decoder 15018 may include a video decoder (e.g., VVC, HEVC) and an entropy coding-based decoder. For the displacement vector video decoder 15018, the displacement vector video decoder agreed upon by the encoder (i.e., the transmitting side) / decoder (i.e., the receiving side) may be selected, or the type of the displacement vector video decoder determined by the encoder may be parsed to select the displacement vector video decoder.

[0335] According to an embodiment, a displacement vector reconstructor may be further provided between the displacement vector video decoder 15018 and the mesh reconstructor 15017. The displacement vector reconstructor may extract displacement vector transform coefficients from the decoded displacement vector video, and apply inverse quantization and inverse transformation (e.g., inverse wavelet transformation) to the extracted displacement vector transform coefficients to reconstruct the displacement vector. To this end, the displacement vector reconstructor may include an image unpacker, an inverse quantizer, and an inverse linear lifting section. If the reconstructed displacement vector is in a local coordinate system, an inverse transformation to a Cartesian coordinate system may be performed. Alternatively, the displacement vector video decoder 15018 may further perform the above displacement vector reconstruction process.

[0336] The mesh reconstructor 15017 may subdivide the base mesh reconstructed by the base mesh reconstructor 15016 based on the auxiliary information to generate additional vertices. Through the subdivision, vertex connection information including the added vertices, texture coordinates, and connection information about the texture coordinates may be generated. Then, the mesh reconstructor 15017 may combine the subdivided reconstructed base mesh with the reconstructed displacement vector to generate a final reconstructed mesh (also referred to as a reconstructed deformed mesh).

[0337] According to an embodiment, the texture map video decoder 15019 may decode a texture map bitstream into a video bitstream using a video codec to reconstruct the texture map. The reconstructed texture map may have color information about each vertex included in the reconstructed mesh, and the color value of each vertex may be obtained from the texture map based on the texture coordinates of the vertex.

[0338] According to an embodiment, the type of the texture map video decoder 15019 may include a video decoder (e.g., VVC, HEVC) and an entropy coding-based decoder. As a method of selecting the texture map video decoder, the texture map video decoder agreed upon between the encoder / decoder may be used, or the type of the decoder determined by the encoder may be parsed to determine the texture map video decoder.

[0339] According to an embodiment, the reconstructed mesh from the mesh reconstructor 15017 and the reconstructed texture map from the texture map video decoder 15019 are rendered by a mesh data renderer (not shown) and displayed to the user through a rendering process.

[0340] Figure 22is a diagram showing an example subgroup segmentation process according to an embodiment. In Figure 22 the execution order of blocks can be changed, some blocks can be omitted, and new blocks can be added.

[0341] In Figure 22 the subgroup segmentation information parser 16011 parses the subgroup segmentation information from the decoded auxiliary information.

[0342] According to an embodiment, the subgroup segmentation information may include a subgroup segmentation method (e.g., partition_type_idx), information about the subgroup segmentation method (e.g., the initial number of clusters for K-means clustering (e.g., number_of_cluster), the clustering level for hierarchical clustering (e.g., level_of_cluster), the octree structure for octree partitioning (e.g., octree_partitioning_data), and the patch size for patch segmentation (e.g., patch_parameter_set)).

[0343] According to an embodiment, when performing motion vector decoding on a grid for each frame, the subgroup splitter 15012 may split the grid of a frame into one subgroup.

[0344] According to an embodiment, the subgroup splitter 15012 may check the subgroup segmentation method index (i.e., partition_type_idx) (16012) included in the subgroup segmentation information. Based on the subgroup segmentation method index (i.e., partition_type_idx), an implicit segmentation information derivation process (16013) may be performed, and / or a subgroup derivation process (16014) may be performed.

[0345] Figure 23 Shows an example of the subgroup segmentation method index (i.e., partition_type_idx) according to an embodiment. In other words, Figure 23 shows an exemplary method of configuring the subgroup segmentation method index (i.e., partition_type_idx) according to the subgroup segmentation method. For example, when partition_type_idx is set to 0, it may indicate segmentation method 1 (i.e., octree partitioning); when set to 1, it may indicate segmentation method 2 (i.e., K-means clustering); when set to 2, it may indicate segmentation method 3 (i.e., hierarchical clustering); when set to 3, it may indicate segmentation method 4 (i.e., patch segmentation).

[0346] According to an embodiment, the subgroup splitter 15012 may use various subgroup splitting methods. For example, the subgroup splitter 15012 may perform splitting using a fixed method based on an agreement between an encoder (i.e., a transmitting side) / a decoder (i.e., a receiving side) without transmitting partition_type_idx, or may implicitly derive a specific splitting pattern according to the frame type (inter-frame or intra-frame) of a reference grid. Additionally, as a subgroup splitting method, only some of various splitting methods may be used to configure an index.

[0347] According to an embodiment, the implicit split information derivator 16013 may derive split information based on geometric information about a reference grid reconstructed by a decoder, and / or may skip transmitting split information about a region where vertices do not exist.

[0348] According to an embodiment, the subgroup splitter 15012 may perform an implicit split information derivation process only when partition_type_idx indicates octree splitting as a subgroup splitting method.

[0349] For example, when the codeword of a voxel having a vertex is 1 and the codeword of a voxel having no vertex is 0, the codeword of a region where vertices do not exist may be derived as 0. In the present disclosure, the codeword lacking a vertex is derived as 0 so that the corresponding region is no longer split. A subgroup may be configured using a final octree structure generated through this process. For example, in the case where vertices exist in a voxel, an octree splitting operation may be recursively performed until the number of vertices in the voxel is reduced to a specific number N or less. When no vertex exists in a voxel, no further splitting is performed on the voxel.

[0350] Figure 24 is a diagram showing an example of deriving subgroup split information from an octree structure according to an embodiment. Specifically, Figure 24 shows an embodiment in which, when deriving octree split depth information, for a region where no vertex exists, implicit split information is derived as 0.

[0351] Therefore, when the subgroup splitting method index (i.e., partition_type_idx) is 0 (i.e., octree splitting), the subgroup splitter 15012 may derive split information through the implicit split information derivator 16013. In the present disclosure, the term "implicit split derivation" means that split information (e.g., split octree depth information) may be derived by a decoder of a receiving device without receiving split information.

[0352] According to an embodiment, when the subgroup splitting method index (i.e., partition_type_idx) is not 0, i.e., when the method is one of K-means clustering, hierarchical clustering, or patch segmentation, the subgroup derivator 16014 may perform subgroup splitting.

[0353] For example, the subgroup deriver 16014 may perform subgroup segmentation based on subgroup segmentation information determined by the encoder of the transmitting device, signaled, and then parsed, or may perform subgroup segmentation based on per-subgroup segmentation information derived by the decoder of the receiving device from the geometric information about the reference grid. In other words, since the current base grid and the reference base grid have the same connection information, texture coordinates, and number of vertices, the segmentation structure of the current base grid can be derived based on the reference base grid information. Additionally, the subgroup segmentation information can be derived from the geometric information related to the reference grid. According to an embodiment, the subgroup segmentation method may be K-means clustering. When the reference base grid is a P-frame or a B-frame, the motion vectors of the reference base grid can be used to perform clustering. When the grid is an I-frame, the K-means clustering process can be performed based on the vertex geometric information related to the reconstructed reference base grid to derive the subgroup segmentation information.

[0354] As described above, in the present disclosure, after deriving the subgroup segmentation information related to the reference base grid, the subgroup deriver 16014 may apply the derived subgroup segmentation information to the current base grid. In this case, as the applied method, the vertex index information about the vertices of the reference base grid included in each subgroup can be mapped to the corresponding vertex indices of the current base grid per subgroup to segment the current base grid into subgroups.

[0355] According to an embodiment, the subgroup segmenter 15012 may obtain the subgroup segmentation method by parsing the subgroup segmentation method index (partition_type_idx) included in the auxiliary information header transmitted from the transmitting device.

[0356] According to an embodiment, when the subgroup segmentation method is segmentation method example 1 (octree segmentation), the subgroup deriver 16014 may perform subgroup segmentation based on the octree segmentation information (octree_partitioning_data) included in the auxiliary information header.

[0357] According to an embodiment, when the subgroup segmentation method is segmentation method example 2 (K-means clustering), the initial number of clusters may be set based on the initial number of clusters information (number_of_cluster) included in the auxiliary information header, or may be set to a fixed initial number of clusters by the encoder / decoder. Then, subgroup segmentation may be performed based on the set initial number of clusters.

[0358] When the reference grid is a P-frame or a B-frame, K-means clustering can be performed based on the motion vectors of the reconstructed reference base grid. When the reference grid is an I-frame, K-means clustering can be performed based on the geometric information of the reconstructed reference base grid.

[0359] According to an embodiment, when the subgroup segmentation method is segmentation method example 3 (hierarchical clustering), and the reference grid is a P-frame or a B-frame, the subgroup derivator 16014 may be initialized to form one group per vertex, and then perform clustering by merging the motion vectors of similar vertices. When the reference grid is an I-frame, the subgroup derivator 16014 may be initialized to form one group per vertex, and then perform clustering by merging similar vertices.

[0360] In addition, when the subgroup segmentation method is hierarchical clustering, the clustering level may be set based on the clustering level information (level_of_cluster) included in the auxiliary information header, or may be set by the encoder / decoder to a fixed clustering level. In other words, subgroup segmentation may be performed based on the set clustering level.

[0361] According to an embodiment, when the subgroup segmentation method is segmentation method example 4 (patch segmentation), patch segmentation information such as the number of patches (number_of_patch) included in the patch parameter set may be parsed per frame to perform patch segmentation.

[0362] Figure 25 is an exemplary detailed block diagram of a motion vector decoder according to an embodiment. Figure 25 Each component in corresponds to hardware, software, a processor, and / or a combination thereof. In Figure 25 it, the execution order of the blocks may be changed, some blocks may be omitted, and new blocks may be added.

[0363] According to an embodiment, in the case of inter-frame prediction, the motion vector decoder 15014 receives a geometric information bitstream (also referred to as a motion vector bitstream) transmitted from the encoder of the transmitting device via the switching unit 15013, decodes the bitstream based on per-subgroup and / or per-vertex, and obtains a differential motion vector. Then, it adds the decoded differential motion vector to the derived predicted motion vector to reconstruct the motion vector, and reconstructs the geometric information about the current base grid based on the reconstructed motion vector. Here, the geometric information bitstream may be used interchangeably with the base grid bitstream.

[0364] More specifically, the differential motion vector decoder 17011 decodes the geometric information bitstream to reconstruct the differential motion vector. In this case, the differential motion vector decoder 17011 may decode the differential motion vector based on per-subgroup and / or per-vertex for easier reconstruction. To this end, subgroup segmentation information is provided to the differential motion vector decoder 17011 in the motion vector decoder 15014 via the subgroup segmenter 15012.

[0365] For example, when the value of the first skip flag information (e.g., mvd_skip_flag) is 1, the differential motion vector decoder 17011 may perform differential motion vector decoding based on per-vertex.

[0366] In another example, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0, the second skip flag information (e.g., vertex_mvd_skip_flag) is parsed to determine whether to decode the differential motion vector on a per-subgroup or per-vertex basis to decode the differential motion vector. For example, when the value of the second skip flag information (vertex_mvd_skip_flag) is 1, the differential motion vector is decoded on a per-subgroup basis. When the value of the second skip flag information (vertex_mvd_skip_flag) is 0, the differential motion vector is decoded on a per-vertex basis.

[0367] According to an embodiment, the motion vector reconstructor 17012 reconstructs a motion vector by adding the differential motion vector reconstructed by the differential motion decoder 17011 to the predicted motion vector output by the motion vector predictor 17014. The detailed operation of the motion vector predictor 17014 that predicts the motion vector based on the motion vectors stored in the motion vector buffer 17015 will be described later.

[0368] According to an embodiment, the geometry information reconstructor 17013 may reconstruct geometry information about the current base mesh based on the reference reconstructed base mesh and the reconstructed motion vector.

[0369] For example, when the motion vector is decoded on a per-subgroup basis, the motion vector predictor 17014 may be used to obtain a predicted motion vector on a per-subgroup basis, and the motion vector reconstructor 17012 may reconstruct the motion vector by adding the reconstructed differential motion vector to the predicted motion vector on a per-subgroup basis. Then, the geometry information reconstructor 17013 may reconstruct the geometry information about the current base mesh by adding the reconstructed motion vector to the vertices of the reference base mesh whose indices are the same as the vertices of the current base mesh.

[0370] According to an embodiment, the motion vector buffer 17015 may store the motion vector reconstructed by the motion vector reconstructor 17012. When storing the motion vector in the motion vector buffer 17015 (i.e., the memory), the bit depth of the motion vector set to a value pre-fixed by the encoder / decoder may be used. Here, the bit depth of the motion vector refers to the number of bits used to store each of the x, y, and z components of the motion vector (x, y, z).

[0371] According to an embodiment, the bit depth pre-fixed by the encoder / decoder may be represented as an integer value. For example, during the motion vector encoding / decoding process, the bit depth of the motion vector may be calculated as 10 bits and stored in the motion vector buffer 17015 as 8 bits (i.e., less data).

[0372] According to an embodiment, the motion vector buffer 17015 may store motion vectors per vertex.

[0373] According to an embodiment, the motion vectors may be stored in the motion vector buffer 17015 after reducing the spatial resolution of the motion vectors. Additionally, the motion vector buffer 17015 may store the motion vectors in n×k×l cells.

[0374] In the present disclosure, whether to reduce the spatial resolution of the motion vectors may be determined based on the magnitude of the motion vectors, or the resolution may be reduced by calculating the rate distortion of the motion vectors. In the present disclosure, the spatial resolution of the motion vectors is independent of the resolution of the differential motion vectors. That is, when sending (differential) motion vectors to the decoder of the receiving device, the resolution of the differential motion vectors is changed to reduce the amount of bits. When storing the motion vectors in the motion vector buffer 17015 of the receiving device, the spatial resolution of the motion vectors may be changed to reduce the amount of bits.

[0375] According to an embodiment, the order in which the reconstructed motion vectors are stored in the motion vector buffer 17015 may be the reconstruction order of the motion vectors. In other words, the reconstruction order of the motion vectors may be the index order of the vertices of the current base mesh. Further, when the motion vectors are in floating-point form, the motion vectors may be quantized and converted to fixed-point form and stored in the motion vector buffer 17015.

[0376] Figure 26 is a diagram showing an example detailed operation of a differential motion vector decoder according to an embodiment. That is, Figure 26 is a diagram showing an exemplary process in which the differential motion vector decoder 17011 decodes differential motion vectors. In Figure 26 , the execution order of the blocks may be changed, some blocks may be omitted, and new blocks may be added.

[0377] In Figure 26 , the geometric information entropy decoder 18011 may perform entropy decoding on the geometric information bitstream. For example, the geometric information entropy decoder 18011 may decode the geometric information bitstream using various decoding methods such as CABAC, exponential Golomb, VLC, or CAVLC. According to an embodiment, the geometric information bitstream input to the geometric information entropy decoder 18011 may include a differential motion vector transmission skip flag for a subgroup (e.g., first skip flag information), a differential motion vector transmission flag for a vertex (e.g., second skip flag information), subgroup differential motion vectors, and vertex differential motion vectors.

[0378] In the present disclosure, the geometric information entropy decoder 18011 may be determined by a predefined convention between the encoder / decoder or by parsing an index of the entropy decoder type sent from the encoder.

[0379] According to an embodiment, based on at least one of the first skip flag information (e.g., mvd_skip_flag) (18012) and the second skip flag information (e.g., vertex_mvd_skip_flag) (18014) decoded by entropy decoding, the differential motion vector decoder 17011 may perform inverse quantization on the differential motion vectors for each subgroup and / or for each vertex pair, or perform vertex differential motion vector derivation.

[0380] For example, when the value of the first skip flag information (e.g., mvd_skip_flag) is 1, vertex differential motion vector derivation (18013) may be performed. That is, when the value of the first skip flag information (e.g., mvd_skip_flag) is 1, the transmitting device skips the transmission of the differential motion vectors for each subgroup and for each vertex, so the receiving device does not receive the subgroup differential motion vectors. Since the subgroup differential motion vectors are not received (i.e., parsed), the vertex differential motion vector derivator 18013 derives the differential motion vectors for the subgroup and the vertex as zero vectors (0, 0, 0).

[0381] If the value of the first skip flag information (e.g., mvd_skip_flag) is 0, inverse quantization may be performed on the differential motion vectors for each subgroup (18015) or for each vertex (18017) based on the second skip flag information (e.g., vertex_mvd_skip_flag).

[0382] For example, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 1, the subgroup differential motion vector inverse quantizer 18015 may perform inverse quantization on the differential motion vectors for each subgroup. In another example, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 0, the vertex differential motion vector inverse quantizer 18017 may perform inverse quantization on the differential motion vectors for each vertex.

[0383] In the present disclosure, the subgroup differential motion vector inverse quantizer 18015 may perform inverse quantization on the differential motion vectors of a subgroup. According to an embodiment, the subgroup differential motion vector inverse quantizer 18015 may be omitted. Additionally, the vertex differential motion vector inverse quantizer 18017 may perform inverse quantization on the differential motion vectors of the vertices in a subgroup. According to an embodiment, the vertex differential motion vector inverse quantizer 18017 may be omitted.

[0384] According to an embodiment, the vertex differential motion vector derivation unit 18016 derives a vertex differential motion vector based on the subgroup differential motion vectors inverse quantized subgroup by subgroup by the subgroup differential motion vectors inverse quantizer 18015, and then outputs the reconstructed vertex differential motion vector. In other words, when parsing the subgroup differential motion vectors, the vertex differential motion vector derivation unit 18016 can derive the vertex differential motion vector in the current subgroup from the subgroup differential motion vectors. When there is no parsed subgroup differential motion vector, the vertex differential motion vector derivation unit 18016 can derive the vertex differential motion vector as a zero vector (0, 0, 0). That is, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 1, the transmitting device transmits the differential motion vectors of the subgroup, but does not transmit the differential motion vectors per vertex. Therefore, the vertex differential motion vector derivation unit 18016 derives the differential motion vectors per vertex to have the same value as the differential motion vectors of the subgroup because the subgroup differential motion vectors are parsed without parsing the differential motion vectors per vertex.

[0385] Next, the parsing of signaling information (e.g., mvd_resolution_idx, mvd_skip_flag, vertex_mvd_skip_flag) performed by the differential motion vector decoder 17011 and / or the subgroup splitter 15012 will be described.

[0386] According to an embodiment, the motion vector resolution information (mvd_resolution_idx) can be parsed to select the resolution of the differential motion vector. The motion vector resolution information (mvd_resolution_idx) may be referred to as an index of the differential motion vector resolution. In the present disclosure, the resolution of the differential motion vector can be determined by parsing the differential motion vector resolution index subgroup by subgroup or per higher unit (e.g., slice or tile).

[0387] According to an embodiment, in order to check whether a (differential) motion vector should be skipped, the (differential) motion vector skip flag information (e.g., mvd_skip_flag, vertex_mvd_skip_flag) can be parsed.

[0388] According to an embodiment, the first skip flag information (e.g., mvd_skip_flag) is a flag for checking whether to skip the (differential) motion vectors of the subgroup and the vertex. For example, if the (differential) motion vector skip flag of the subgroup (e.g., mvd_skip_flag) is 1, both the (differential) motion vector of the subgroup and the (differential) motion vector of the vertex can be derived as zero vectors (0, 0, 0). In other words, deriving the (differential) motion vector as a zero vector (0, 0, 0) means that the transmitting device has not transmitted the (differential) motion vector.

[0389] According to an embodiment, the second skip flag information (e.g., vertex_mvd_skip_flag) is a flag for checking whether to skip the (differential) motion vector of a vertex. For example, when the skip flag (e.g., vertex_mvd_skip_flag) of the (differential) motion vector of a vertex in a subgroup is 1, the (differential) motion vector of the vertex can be derived as a zero vector (0, 0, 0). In other words, deriving the (differential) motion vector as a zero vector (0, 0, 0) means that the transmitting device does not transmit the (differential) motion vector.

[0390] Therefore, when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 1, the differential motion vectors of the subgroup are transmitted from the transmitting device, but the per-vertex differential motion vectors are not transmitted. Accordingly, the receiving device parses the differential motion vectors of the subgroup, but does not parse the per-vertex differential motion vectors. In this case, the per-vertex differential motion vectors are derived to have the same values as the differential motion vectors of the subgroup.

[0391] In other words, deriving the (differential) motion vector as a zero vector means that the transmitting device does not transmit the per-subgroup differential motion vectors and / or the per-vertex differential motion vectors. In this case, the receiving device can obtain a predicted motion vector as a reconstructed motion vector.

[0392] That is, when the (differential) motion vector is derived as a zero vector, the predicted motion vector becomes the reconstructed motion vector (predicted motion vector (x, y, z) + differential motion vector (0, 0, 0) = reconstructed motion vector (x, y, z)).

[0393] According to an embodiment, the decoder in the receiving device can decode one or more subgroups based on each tile or each slice. Additionally, the decoding can be processed in parallel based on each tile or each slice because the tiles do not depend on each other.

[0394] Furthermore, the decoder of the receiving device can support spatial random access based on each tile or each slice. Additionally, encoder parameters such as a motion vector resolution index (e.g., mvd_resolution_idx), quantization parameters, etc. for each division unit (patch group, subgroup, slice, tile, etc.) can be received from the transmitting device.

[0395] According to an embodiment, the motion vector decoder 15014 can receive a per-subgroup differential motion vector resolution index (mvd_resolution_idx) obtained by parsing the reconstructed per-vertex differential motion vectors and multiplying them by the differential motion vector resolutions mapped to respective indexes. Here, the differential motion vector resolutions mapped to the differential motion vector resolution indexes can be as Figure 19Tabulate as shown to obtain resolution values mapped to indices.

[0396] Figure 27 is a diagram showing an exemplary motion vector estimation process according to an embodiment. Specifically, Figure 27 is a diagram showing Figure 26 an exemplary process in which the motion vector predictor 17014 predicts a motion vector. In Figure 27 it, the execution order of blocks can be changed, some blocks can be omitted, and new blocks can be added.

[0397] According to an embodiment, the motion vector predictor 17014 can first generate a predicted motion vector based on the reconstructed motion vector and then output it to the motion vector reconstructor 17012. In one embodiment, the reconstructed motion vector (19011) can be provided from the motion vector buffer.

[0398] According to an embodiment, the subgroup motion vector predictor 19012 predicts each subgroup motion vector based on at least one of first skip flag information (e.g., mvd_skip_flag) and second skip flag information (e.g., vertex_mvd_skip_flag).

[0399] For example, when the value of the first skip flag information (e.g., vertex_mvd_skip_flag) is 1, or when the value of the first skip flag information (e.g., mvd_skip_flag) is 0 and the value of the second skip flag information (e.g., vertex_mvd_skip_flag) is 1, the subgroup motion vector predictor 19012 calculates the predicted motion vector based on each subgroup. That is, when (mvd_skip_flag == 1) or (mvd_skip_flag == 0 && vertex_mvd_skip_flag == 1), the subgroup motion vector predictor 19012 can operate. The predicted motion vector calculated (or obtained) by the subgroup motion vector predictor 19012 is provided to the motion vector reconstructor 17012.

[0400] According to an embodiment, the subgroup motion vector predictor 19012 can calculate the predicted motion vector of the subgroup based on the motion vector of a previously reconstructed neighboring vertex in the current frame or the motion vector in the reference frame.

[0401] According to an embodiment, the predicted motion vector of the subgroup can be obtained by predicting based on the motion vector of a subgroup that is spatially close to the current subgroup among the previously reconstructed subgroups or the average motion vector of N vertices closest to the vertices of the current subgroup.

[0402] According to an embodiment, when the reference base grid is a P-frame or a B-frame, the average motion vector of the vertices of the reference base grid having the same index as the vertices of the current base grid can be used as the predicted motion vector of the current subgroup.

[0403] When the values of the first skip flag information (e.g., mvd_skip_flag) and the second skip flag information (e.g., vertex_mvd_skip_flag) are both 0, the motion vector is predicted vertex-by-vertex by the vertex motion vector predictor 19014 and provided to the motion vector reconstructor 17012.

[0404] That is, the vertex motion vector predictor 19014 predicts the motion vector based on vertex-by-vertex calculation. In other words, when (mvd_skip_flag == 0 && vertex_mvd_skip_flag == 0), motion vector prediction can be performed vertex-by-vertex within the current subgroup. In the present disclosure, the predicted motion vector of the vertices in the subgroup can be calculated based on the previously decoded vertex motion vectors. According to an embodiment, the prediction can be performed by averaging the motion vectors or parallelogram prediction or averaging of multiple parallelogram predictions of adjacent vertices based on connection information.

[0405] The motion vector predictor 17014 in the present disclosure can be omitted. In this case, the motion vector decoder can reconstruct by parsing the motion vector instead of the differential motion vector.

[0406] In addition, in the present disclosure, when the differential motion vector is derived as a zero vector (0, 0, 0), that is, when no differential motion vector is sent, the motion vector predicted by the motion vector predictor 17014 is regarded as the reconstructed motion vector.

[0407] Furthermore, based on at least one of the first skip flag information and / or the second skip flag information, the transmission of the subgroup (differential) motion vector and / or the vertex (differential) motion vector can be skipped. The (differential) motion vector for which the transmission is skipped can be derived as a zero vector by the decoder of the receiving device. The first skip flag information can be signaled for each subgroup, and the second skip flag information can be signaled for each vertex in the subgroup. That is, the skip mode can be applied subgroup-by-subgroup and / or vertex-by-vertex.

[0408] For example, when the value of the first skip flag information is 1, the (differential) motion vectors of all vertices in the subgroup can be derived as zero vectors.

[0409] In addition, when the value of the first skip flag information is 0, the (differential) motion vector can be sent vertex-by-vertex, and the decoder of the receiving device can parse and decode the (differential) motion vector vertex-by-vertex. In this case, the second skip flag information can be omitted.

[0410] In the present disclosure, relevant information may be signaled to add / perform an embodiment. The signaling information according to the embodiment may be used on the transmitting side or the receiving side. For example, the signaling information according to the embodiment may be generated and transmitted by a metadata processor (which may be referred to as a metadata generator, not shown) of a transmitting device, and may be received and acquired by a metadata parser (not shown) of a receiving device. The operation of the receiving device may be performed based on the signaling information.

[0411] Figure 28 An exemplary syntax structure showing subgroup segmentation information according to an embodiment is shown.

[0412] In one embodiment, Figure 28 The subgroup segmentation information (decode_auxiliary_data()) in may be signaled in the auxiliary information (in particular, the auxiliary information header) for transmission and reception. In one embodiment, the auxiliary information header is based on a per-grid frame to transmit the syntax related to the auxiliary information.

[0413] According to an embodiment, according to the subgroup segmentation method (e.g., partition_type_idx), the subgroup segmentation information (decode_auxiliary_data()) may include information about the subgroup segmentation method.

[0414] In other words, partition_type_idx indicates the subgroup segmentation method index as Figure 23 shown.

[0415] For example, when the value of partition_type_idx is 0, it indicates that the subgroup segmentation method is octree segmentation, and the subgroup segmentation information (decode_auxiliary_data()) may include segmentation information related to the octree structure (e.g., octree_partitioning_data).

[0416] When the value of partition_type_idx is 1, it indicates that the subgroup segmentation method is K-means clustering, and the subgroup segmentation information (decode_auxiliary_data()) may include information related to the initial number of clusters (e.g., number_of_cluster).

[0417] When the value of partition_type_idx is 2, it indicates that the subgroup segmentation method is hierarchical clustering, and the subgroup segmentation information (decode_auxiliary_data()) may include information related to the clustering level (e.g., level_of_cluster).

[0418] When the value of partition_type_idx is 3, it indicates that the subgroup splitting method is patch segmentation, and the subgroup splitting information (decode_auxiliary_data()) may include information related to the patch size (e.g., patch_parameter_set).

[0419] Figure 29 An exemplary syntax structure of motion vector related information according to an embodiment is shown.

[0420] Figure 29 The motion vector related information (decode_motion_vector()) in can be signaled in the mvd_subgroup_header in the geometric information bitstream for transmission and reception. In one embodiment, the mvd_subgroup_header is based on subgroup-by-subgroup transmission of motion vector related syntax.

[0421] According to an embodiment, the motion vector related information may include mvd_resolution_idx and mvd_skip_flag.

[0422] mvd_resolution_idx indicates the index of the resolution of the differential motion vector as Figure 19 shown.

[0423] mvd_skip_flag is a differential motion vector skip flag for subgroups and vertices, and may be simply referred to as first skip flag information for simplicity.

[0424] According to an embodiment, the motion vector related information may include mvd_skip_flag, followed by a loop executed as many times as the number of subgroups to be processed.

[0425] For example, when the value of mvd_skip_flag of the i-th subgroup is 1, the differential motion vector of the i-th subgroup is derived as a zero vector (0, 0, 0). The differential motion vectors of the respective vertices included in the i-th subgroup are also derived as zero vectors (0, 0, 0). In other words, the differential motion vectors of all subgroups and the differential motion vectors of all vertices in each subgroup are derived as zero vectors (0, 0, 0).

[0426] On the other hand, when the value of mvd_skip_flag of the i-th subgroup is not 1, i.e., when it is 0, the motion vector related information may further include vertex_mvd_skip_flag of the i-th subgroup.

[0427] As a differential motion vector skip flag for vertices, vertex_mvd_skip_flag may be simply referred to as second skip flag information for simplicity.

[0428] For example, when the value of mvd_skip_flag of the i-th subgroup is 0 and the value of vertex_mvd_skip_flag is 1, the (differential) motion vector of the i-th subgroup is derived from the value obtained by inverse quantization of the differential motion vector of the i-th subgroup, and the (differential) motion vector of the j-th vertex in the i-th subgroup is derived from sub_group_mvd[i]*mvd_resolution[mvd_resolution_idx]. In other words, the (differential) motion vector of the j-th vertex in the i-th subgroup can be derived from the (differential) motion vector of the i-th subgroup. In this case, a process of parsing the motion vector resolution index determined for each subgroup and multiplying the value corresponding to the differential motion vector resolution index of the vertex in the subgroup by the differential motion vector per vertex can be added.

[0429] In other words, when the value of vertex_mvd_skip_flag is 1, the motion vector decoder in the receiving device can determine to skip the differential motion vector per vertex, and then perform decoding by treating the (differential) motion vector of the subgroup as the motion vector of the vertex.

[0430] Thereafter, the decoded subgroup differential motion vector is inverse quantized by the subgroup differential motion vector inverse quantizer 18015, and the per-vertex differential motion vector is derived from the inverse quantized subgroup differential motion vector by the vertex differential motion vector derivator 18016.

[0431] In another example, when the value of mvd_skip_flag of the i-th subgroup is 0 and the value of vertex_mvd_skip_flag is also 0, the (differential) motion vector of the j-th vertex in the i-th subgroup is derived as the motion vector of the vertex (InvQuant(Vertex)MVD)*mvd_resolution[mvd_resolution_idx].

[0432] In other words, the motion vector decoder in the receiving device can decode the differential motion vector per vertex obtained by parsing. Then, the decoded differential motion vector per vertex can be inverse quantized by the vertex differential motion vector inverse quantizer 18017. In this case, a process of parsing the motion vector resolution index determined for each subgroup and multiplying the differential motion vector of the vertex in the subgroup by the value corresponding to the motion vector resolution index can be performed.

[0433] Figure 30 is a flowchart showing an exemplary transmission method according to an embodiment. The transmission method according to the embodiment may include: encoding mesh data (21011); and transmitting a bitstream including the encoded mesh data (21012).

[0434] According to an embodiment, when encoding is performed by inter-frame prediction, the encoding (21011) of mesh data may include dividing a reference base mesh into subgroups and encoding (reference) motion vectors based on each subgroup and / or each vertex pair. In this case, the subgroup division information may be signaled in the auxiliary information and sent to the receiving side.

[0435] According to an embodiment, the subgroup division information may include a subgroup division method (e.g., partition_type_idx) and information about the subgroup division method (e.g., in the case of K-means clustering, the initial number of clusters (e.g., number_of_cluster), in the case of hierarchical clustering, the clustering level (e.g., level_of_cluster), in the case of octree partitioning, the octree structure (e.g., octree_partitioning_data), in the case of patch segmentation, the patch size (e.g., patch_parameter_set)). Additionally, the encoding (21011) of mesh data may include signaling skip flag information (or simply referred to as a skip flag) for indicating whether to skip (differential) motion vectors.

[0436] For details of the subgroup division method, the encoding of motion vectors based on each subgroup and / or each vertex, and the subgroup division information and motion vector-related information, refer to Figures 15 to 20 , Figure 28 and Figure 29 The descriptions of which will not be described below to avoid redundancy.

[0437] As described above, the reference base mesh may be divided into subgroups. When the (differential) motion vectors of the vertices in each subgroup are similar to each other, only the differential motion vector of the subgroup may be sent, and the transmission of the differential motion vectors of each vertex may be skipped. Thereby, the amount of data to be sent may be reduced, and thus the compression efficiency of geometric information may be further improved.

[0438] Figure 31 is a flowchart showing an exemplary receiving method according to an embodiment. The receiving method according to an embodiment may include: receiving a bitstream (22011) containing mesh data; and decoding the mesh data contained in the bitstream (22012).

[0439] For details of the reception (22011) of the bitstream containing mesh data and the decoding (22012) of the mesh data contained in the bitstream, refer to Figures 21 to 29 The detailed description of which will not be described below to avoid redundancy.

[0440] Therefore, the present disclosure can improve the inter-frame (inter-frame prediction) mode technology of V-mesh by allowing each subgroup with similar motion vectors to perform encoding and decoding when performing inter-frame prediction. In other words, when encoding / decoding geometric information related to 3D dynamic mesh data through inter-frame prediction, the reference base mesh can be divided into subgroups based on the motion vectors of similar vertices, and motion vectors can be calculated for each of the divided subgroups to transmit the (differential) motion vector of each subgroup. Thereby, the amount of data to be transmitted can be reduced, and the compression efficiency of geometric information can be improved. In particular, when the motion vectors of the vertices in a subgroup are similar, only the differential motion vector of each subgroup can be transmitted, and the transmission of the differential motion vector of the vertices can be skipped, thereby reducing the amount of bits to be transmitted or parsed. Additionally, according to the present disclosure, the resolution of the motion vector can be determined subgroup by subgroup.

[0441] Each of the above-mentioned parts, modules, or units can be software, a processor, or a hardware part that executes a continuous process stored in a memory (or storage unit). Each of the steps described in the above embodiments can be executed by a processor, software, or a hardware part. Each of the above-mentioned modules / blocks / units can operate as a processor, software, or hardware. Additionally, the method presented in the embodiments can be executed as code. The code can be written on a processor-readable storage medium and thus read by the processor provided by the device.

[0442] In this specification, when a part "includes" or "contains" an element, unless otherwise mentioned, it means that the part also includes or contains another element. Additionally, the term "... module (or unit)" disclosed in this specification means a unit for processing at least one function or operation, and can be implemented by hardware, software, or a combination of hardware and software.

[0443] Although the embodiments have been illustrated and described with reference to the respective drawings for simplicity, new embodiments can be designed by combining the embodiments shown in the drawings. If a person skilled in the art designs a computer-readable recording medium recording a program for executing the embodiments mentioned in the above description, it can fall within the scope of the appended claims and their equivalents.

[0444] The device and method can be unrestricted by the configurations and methods of the above embodiments. The above embodiments can be configured by selectively combining with each other completely or partially, so as to enable various modifications.

[0445] Although the preferred embodiments of the embodiments have been shown and described, the embodiments are not limited to the above specific embodiments. Without departing from the spirit of the embodiments claimed in the claims, various modifications can be made by those of ordinary skill in the art, and these modifications should not be understood separately from the technical idea or perspective of the embodiments.

[0446] The various elements of the device of the embodiment may be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiment may be implemented by a single chip (e.g., a single hardware circuit). According to an embodiment, the components according to the embodiment may be separately implemented as separate chips. According to an embodiment, at least one or more components of the device according to the embodiment may include one or more processors capable of executing one or more programs. The one or more programs may execute any one or more of the operations / methods according to the embodiment, or include instructions for executing them. The executable instructions for executing the method / operation of the device according to the embodiment may be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transitory CRM or other computer program product configured to be executed by one or more processors. Additionally, the memory according to the embodiment may be used as a concept that encompasses not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it may also be implemented in the form of a carrier wave (e.g., transmission via the Internet). Additionally, the processor-readable recording medium may be distributed among computer systems connected via a network such that the processor-readable code may be stored and executed in a distributed manner.

[0447] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". "A / B / C" may mean "at least one of A, B, and / or C". Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0448] The various elements of the embodiment may be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiment may be executed by a single chip (e.g., a single hardware circuit). According to an embodiment, the elements may be selectively executed separately by separate chips. According to an embodiment, at least one element of the embodiment may be executed in one or more processors including instructions for executing the operations according to the embodiment.

[0449] The operations according to the embodiments described in this specification can be performed by a transmitting / receiving device according to an embodiment including one or more memories and / or one or more processors. The one or more memories may store programs for processing / controlling the operations according to the embodiment, and the one or more processors may control the various operations described in this specification. The one or more processors may be referred to as a controller or the like. In an embodiment, the operations may be performed by firmware, software, and / or a combination thereof. The firmware, software, and / or a combination thereof may be stored in the processor or the memory.

[0450] Terms such as first and second may be used to describe various elements of an embodiment. However, the various components according to the embodiment should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be construed as not departing from the scope of the various embodiments. Both the first user input signal and the second user input signal are user input signals, but do not mean the same user input signal unless the context clearly dictates otherwise. The terms used to describe the embodiments are only used to describe specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and the claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. The expression "and / or" is used to include all possible combinations of items. Terms such as "including" or "having" are intended to indicate the presence of diagrams, numbers, steps, elements, and / or components, and should be understood as not excluding the possibility of the additional presence of diagrams, numbers, steps, elements, and / or components.

[0451] As used herein, conditional expressions such as "if" and "when" are not limited to optional cases and are intended to be interpreted as performing the relevant operations or interpreting the relevant definitions according to the specific conditions when the specific conditions are met. The embodiments may include variations / modifications within the scope of the claims and their equivalents. It will be apparent to those skilled in the art that various modifications and changes can be made in this disclosure without departing from the spirit and scope of the disclosure. Therefore, this disclosure is intended to cover the modifications and changes of this disclosure as long as they fall within the scope of the appended claims and their equivalents.

[0452] The mode of the present disclosure

[0453] As described above, the relevant content has been described in the best mode of implementing the embodiments.

[0454] Industrial applicability

[0455] As described above, the embodiments can be applied, in whole or in part, to 3D data transmission / reception apparatuses and systems. It will be apparent to those skilled in the art that various changes or modifications can be made to the embodiments within the scope of the embodiments. Accordingly, the embodiments are intended to cover modifications and variations as long as they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for transmitting 3D data, the method comprises the following steps: Preprocess the input mesh data and output the base mesh data; Encode the base mesh data; And Transmit a bitstream including the encoded mesh data and signaling information.

2. The method according to claim 1, wherein, The encoding of the base mesh data comprises the following steps: Divide the reference base mesh data into subgroups; Obtain the motion vectors between the base mesh data and the reference base mesh data for each of the subgroups; and Encode the obtained motion vectors.

3. The method according to claim 2, wherein, The motion vector is the average of the motion vectors of the vertices in a corresponding one of the subgroups.

4. The method according to claim 2, wherein, The signaling information includes information related to subgroup division.

5. The method according to claim 2, wherein, The signaling information further includes motion vector related information for indicating whether to skip the motion vector.

6. The method according to claim 5, the method further comprises the following steps: Transmit the motion vector or skip the transmission of the motion vector based on the motion vector related information.

7. The method according to claim 6, wherein, Based on skipping the transmission of the motion vector, derive a zero vector for the motion vector at the receiving side.

8. An apparatus for transmitting 3D data, the apparatus comprises: A preprocessor configured to preprocess the input mesh data and output the base mesh data; An encoder configured to encode the base mesh data; And A transmitter configured to transmit a bitstream including the encoded mesh data and signaling information.

9. The apparatus according to claim 8, wherein, The encoder comprises: A subgroup splitter configured to divide the reference base mesh data into subgroups; A motion vector calculator configured to obtain the motion vectors between the base mesh data and the reference base mesh data for each of the subgroups; and An encoding unit configured to perform entropy encoding on the obtained motion vectors.

10. The apparatus according to claim 9, wherein, The motion vector is the average of the motion vectors of the vertices in a corresponding one of the subgroups.

11. The apparatus according to claim 9, wherein, The signaling information includes information related to subgroup division.

12. The apparatus according to claim 11, wherein, The signaling information further includes motion vector related information for indicating whether to skip the motion vector.

13. The apparatus according to claim 11, wherein, Transmit the motion vector or skip the transmission of the motion vector based on the motion vector related information.

14. The apparatus according to claim 13, wherein, Based on skipping the transmission of the motion vector, derive a zero vector for the motion vector at the receiving side.

15. A method for receiving 3D data, the method comprises the following steps: Receive a bitstream including encoded mesh data and signaling information; The reference grid data is segmented into subgroups based on the signaling information, and the grid data is decoded based on the motion vectors for each of the subgroups; and render the decoded grid data.