3d data transmission device, 3d data transmission method, 3d data reception device, and 3d data reception method

By dividing the grid data into subgroups and preprocessing and encoding, the efficient and partial access problems of sending and receiving 3D grid data is solved, and high-quality 3D service and resource optimization is achieved.

CN120266481APending Publication Date: 2025-07-04LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083823.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2023-12-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently transmit and receive large amounts of 3D grid data, resulting in high latency and encoding/decoding complexity, and it is difficult to realize compression and reconstruction of partial access.

Method used

By dividing the grid data into subgroups, methods of preprocessing, encoding and signaling information, including extraction, texture coordinate generation, fitting subdivision, encoding and reconstruction, generate bitstreams to support partial access.

Benefits of technology

It realizes high-quality 3D services, supports general 3D content such as autonomous driving services, reduces processing time and resource use, and improves compression efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266481A_ABST
    Figure CN120266481A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D data transmitting method, a 3D data transmitting device, a 3D data receiving method and a 3D data receiving device. The 3D data receiving method may comprise the steps of: determining a target subset based on a target area selected by a user and signaling information; and extracting the bit stream of the target subset from the bit stream, and performing decoding based on the extracted bit stream of the target subset to reconstruct the grid data of the target subset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An embodiment provides a method for providing 3D content to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services to a user. Background Art

[0002] Point cloud data or mesh data in 3D content is a collection of points in 3D space. However, due to a large number of points in 3D space, it is difficult to create point cloud data or mesh data.

[0003] In other words, a large throughput is required to send and receive 3D data with a large number of points, such as point cloud or mesh data. Summary of the Invention

[0004] Technical Problem

[0005] An object of the present disclosure is to provide an apparatus and method for effectively sending and receiving mesh data to solve the above problems.

[0006] Another object of the present disclosure is to provide an apparatus and method for solving the latency and encoding / decoding complexity of mesh data.

[0007] An object of the present disclosure is to provide an apparatus and method for compressing and reconstructing mesh data for partial access.

[0008] Another object is to provide an apparatus and method for compressing and reconstructing mesh data for partial access by dividing the mesh data into subgroups for compression and reconstruction to achieve partial access while improving the compression efficiency of the mesh data.

[0009] The objects of the present disclosure are not limited to the above objects, and other objects not mentioned above in the present disclosure will become apparent to those of ordinary skill in the art after studying the following description.

[0010] Technical Solution

[0011] The object of the present disclosure can be achieved by providing a method for sending 3D data, the method including: preprocessing input mesh data and outputting base mesh data, encoding the base mesh data, and sending a bitstream including the encoded mesh data and signaling information.

[0012] According to an embodiment, the method may further include dividing the input mesh data into one or more subgroups.

[0013] According to an embodiment, the preprocessing may include: extracting a subgroup of the input mesh data and generating the extracted mesh data, generating texture coordinates for each vertex of the extracted mesh data based on the subgroup and outputting the base mesh data with the texture coordinates, and subdividing the base mesh data with the texture coordinates based on the subgroup, and generating a fitted subdivided mesh data by performing fitting so that the subdivided base mesh data becomes similar to the input mesh data.

[0014] According to an embodiment, based on that a vertex of one of the polygons configured by connecting vertices of the extracted mesh data is included in two or more of the subgroups, connectivity information related to the vertex constituting the polygon may be redundantly included in the two or more subgroups.

[0015] According to an embodiment, the encoding may include encoding the base mesh data of each subgroup in the subgroup and generating a bitstream for each subgroup in the subgroup, and inserting subgroup identification information before each bitstream in the bitstream to identify the corresponding one in the subgroup.

[0016] According to an embodiment, the encoding may further include: reconstructing the encoded base mesh data, generating displacement information based on the fitted subdivided mesh data and the reconstructed base mesh data, encoding the displacement information and generating a displacement information bitstream, reconstructing the encoded displacement information, reconstructing the mesh data based on the reconstructed base mesh data and the reconstructed displacement information, regenerating a texture map based on the texture map of the original mesh data and the reconstructed mesh data, and encoding the regenerated texture map and generating a texture map bitstream.

[0017] According to an embodiment, the signaling information may include subgroup-related signaling information, where the subgroup-related signaling information may include at least one of information for identifying the number of one or more subgroups, information for identifying the method of dividing the subgroup, or information for identifying the compilation type for the base mesh data.

[0018] According to an embodiment, an apparatus for transmitting 3D data may include: a preprocessor configured to preprocess input mesh data and output base mesh data; an encoder configured to encode the base mesh data; and a transmitter configured to transmit a bitstream including the encoded mesh data and the signaling information.

[0019] According to an embodiment, the apparatus may further include a subgroup divider configured to divide the input mesh data into one or more subgroups.

[0020] According to an embodiment, a pre-processor may include: a mesh extractor configured to extract input mesh data of a subgroup and generate the extracted mesh data; a parameterizer configured to generate texture coordinates for each vertex of the extracted mesh data based on the subgroup and output base mesh data having the texture coordinates; and a fitting subdivider configured to subdivide the base mesh data having the texture coordinates based on the subgroup and generate a fitted subdivided mesh data by performing fitting such that the subdivided base mesh data becomes similar to the input mesh data.

[0021] According to an embodiment, based on that a vertex of one polygon configured by connecting vertices of the extracted mesh data in the parameterizer is included in two or more subgroups among the subgroups, connectivity information related to the vertex constituting the polygon may be redundantly included in the two or more subgroups.

[0022] According to an embodiment, an encoder may include: a base mesh encoder configured to encode the base mesh data of each subgroup in the subgroup and generate a bitstream for each subgroup in the subgroup, and insert subgroup identification information before each bitstream in the bitstreams to identify the corresponding one in the subgroup; a base mesh reconstructor configured to reconstruct the encoded base mesh data; a displacement information generator configured to generate displacement information based on the fitted subdivided mesh data and the reconstructed base mesh data; a displacement information encoder configured to encode the displacement information and generate a displacement information bitstream; a displacement information reconstructor configured to reconstruct the encoded displacement information; a mesh reconstructor configured to reconstruct mesh data based on the reconstructed base mesh data and the reconstructed displacement information; a texture map generator configured to regenerate a texture map based on the texture map of the original mesh data and the reconstructed mesh data; and a texture map encoder configured to encode the regenerated texture map and generate a texture map bitstream.

[0023] According to an embodiment, signaling information may include subgroup-related signaling information, where the subgroup-related signaling information may include at least one of information for identifying the number of one or more subgroups, information for identifying a method of dividing subgroups, or information for identifying a compilation type for base mesh data.

[0024] According to an embodiment, a method of receiving 3D data may include: receiving a bitstream including encoded mesh data and signaling information; decoding the encoded mesh data in the bitstream based on the signaling information; and rendering the decoded mesh data.

[0025] According to an embodiment, decoding of mesh data may include determining a target subgroup based on a target region selected by a user and signaling information, extracting a bitstream of the target subgroup from a bitstream, and reconstructing mesh data of the target subgroup by performing decoding based on the extracted bitstream of the target subgroup.

[0026] Advantageous Effects

[0027] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may provide high-quality 3D services.

[0028] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may implement various video codec schemes.

[0029] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may support general 3D content, such as for autonomous driving services.

[0030] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may allow mesh data to be divided into one or more subgroups, and enable encoding and decoding to be performed on a per-subgroup basis, thereby allowing access to only the mesh data corresponding to a user-specified partial region.

[0031] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may allow the transmitting side to perform subgroup-based encoding after dividing mesh data into one or more subgroups, and allow the receiving side to receive mesh region information related to a necessary region from a user by a device or software configured to render mesh data, and only decode and display mesh data of a partial region, thereby providing a significantly more efficient service in terms of processing time and resource usage compared to processing the entire mesh data.

[0032] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data reception method, and a 3D data reception device may enable the utilization of mesh data in various network environments and applications, and extend the availability of mesh data by allowing flexible use of the resources of a receiver. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the disclosure and, together with the description, serve to explain the principles of the disclosure. To better understand the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to refer to the same or like parts. In the drawings:

[0034] Figure 1 illustrates a system for providing dynamic mesh content according to an embodiment;

[0035] Figure 2 illustrates a V-MESH compression method according to an embodiment;

[0036] Figure 3 illustrates preprocessing in V-MESH compression according to an embodiment;

[0037] Figure 4 illustrates an intermediate edge subdivision method according to an embodiment;

[0038] Figure 5 illustrates a displacement generation process according to an embodiment;

[0039] Figure 6 illustrates an intra-coding process for V-MESH data according to an embodiment;

[0040] Figure 7 illustrates an inter-coding process for V-MESH data according to an embodiment;

[0041] Figure 8 illustrates a lifting transform process for displacement according to an embodiment;

[0042] Fig. 9 illustrates a process of packing transform coefficients into a 2D image according to an embodiment;

[0043] Fig.10 illustrates an attribute transfer process in a V-MESH compression method according to an embodiment;

[0044] Fig.11 illustrates an intra-decoding process for V-MESH data according to an embodiment;

[0045] Fig.12 illustrates an inter-decoding process for V-MESH data according to an embodiment;

[0046] Fig.13 illustrates a mesh data sending device according to an embodiment;

[0047] Fig.14 illustrates a mesh data receiving device according to an embodiment;

[0048] Fig.15 (a) to Fig.15 (c) are diagrams showing examples of subgroup generation according to an embodiment;

[0049] Fig.16 shows an example of information for identifying a subgroup generation method according to an embodiment;

[0050] Fig.17 shows an example syntax structure of subgroup-related information including information for identifying a subgroup generation method according to an embodiment;

[0051] Fig.18 is a diagram showing an example method of determining a subgroup for a vertex located at a subgroup boundary according to an embodiment;

[0052] Fig.19 is a diagram showing an example result of subgroup-based atlas parameterization according to an embodiment;

[0053] Fig. 20 is a diagram showing a first embodiment of a base mesh bitstream generation method according to the present disclosure;

[0054] Fig.21 is a diagram showing a second embodiment of a base mesh bitstream generation method according to the present disclosure;

[0055] Fig. 22 shows an example syntax structure of subgroup index information for each base mesh vertex / texture coordinate according to an embodiment;

[0056] Fig.23 shows an example syntax structure of subgroup index information for each base mesh polygon / texture polygon according to an embodiment;

[0057] Fig.24 shows an example of type information for identifying a subgroup encoding method according to an embodiment;

[0058] Fig.25 shows an example syntax structure of subgroup base mesh compilation information according to an embodiment;

[0059] Fig.26 is a diagram showing an example of a subgroup-specific motion field bitstream according to an embodiment;

[0060] Fig. 27 is a diagram showing another example of a subgroup-specific motion field bitstream according to an embodiment;

[0061] Fig.28 shows an example syntax structure of subgroup index information for each mesh vertex / texture coordinate according to an embodiment;

[0062] Fig.29 Shows an example syntax structure of subgroup index information for each mesh polygon / texture polygon according to an embodiment;

[0063] Fig.30 Is a block diagram showing another example of a transmitting device according to an embodiment;

[0064] Fig.31 Is a block diagram showing another example of a receiving device according to an embodiment;

[0065] Fig.32 Is a flowchart showing an example of a transmitting method according to an embodiment; and

[0066] Fig.33 Is a flowchart showing an example of a receiving method according to an embodiment. Detailed Description of the Invention

[0067] Now, preferred embodiments of the present disclosure will be described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to illustrate the exemplary embodiments of the present disclosure, rather than showing all the embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0068] Although most of the terms used in the present disclosure are selected from general terms widely used in the art, some terms are arbitrarily selected by the applicant, and their meanings are described in detail in the following description as needed. Therefore, the present disclosure should be understood based on the intended meanings of the terms rather than their simple names or meanings.

[0069] With the recent progress of 3D data modeling and rendering technologies, active research has been conducted on generating and processing 3D data across various fields, including virtual reality (VR), augmented reality (AR), autonomous driving, computer-aided design (CAD) / computer-aided manufacturing (CAM), and geographic information system (GIS). 3D data can be represented as a point cloud or a mesh according to the representation format. A mesh consists of geographic information indicating the coordinates of individual vertices or points, connection information indicating the connections between vertices, a texture map representing color information about the mesh surface as 2D image data, and texture coordinates indicating the mapping information between the mesh surface and the texture map. In the present disclosure, when at least one element constituting the mesh changes over time, the mesh is defined as a dynamic mesh, and when it does not change, it is defined as a static mesh.

[0070] Compared with 2D image data, dynamic mesh data involves significantly more element data to represent the mesh. As a result, techniques for efficiently compressing large amounts of mesh data have been developed to store and transmit the data.

[0071] Figure 1 A system for providing dynamic mesh content according to an embodiment is shown.

[0072] Figure 1 The system in includes a transmitting device 100 and a receiving device 110. The transmitting device 100 may include a mesh video acquisition unit (or component) 101, a mesh video encoder 102, a file / fragment encapsulator 103, and a transmitter 104. The receiving device 110 may include a receiver 111, a file / fragment decapsulator 112, a mesh video decoder 113, and a renderer 114. Figure 1 Each component in may correspond to hardware, software, a processor, and / or a combination thereof. In the following description, a mesh data transmitting device according to an embodiment may be interpreted to refer to a 3D data transmitting device or the transmitting device 100, or to a mesh video encoder (hereinafter referred to as an encoder) 102. A mesh data receiving device according to an embodiment may be interpreted to refer to a 3D data receiving device or the receiving device 110, or to a mesh video decoder (hereinafter referred to as a decoder) 113.

[0073] Figure 1 The system of may perform video-based dynamic mesh compression and decompression.

[0074] With the advancement of 3D capture, modeling, and rendering, users are allowed to access various forms of 3D content (e.g., AR, XR, the metaverse, and holograms) across multiple platforms and devices. 3D content is becoming increasingly complex and realistic in its object representation to provide users with an immersive experience. However, for the generation and use of 3D models, this requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system using mesh content.

[0075] First, a method of compressing dynamic mesh data starts with a video-based point cloud compression (V-PCC) standard technique for point cloud data. Point cloud data is data having color information in the coordinates (X, Y, Z) of vertices (or points). In the present disclosure, vertex coordinates (i.e., position information) are referred to as geometric information, and color information about vertices is referred to as attribute information. Geometric information and attribute information together are referred to as vertex information or point cloud data. Mesh data refers to vertex information including connection information between vertices. Content may be initially created in the form of mesh data. Alternatively, connection information may be added to the point cloud data, and the point cloud data may be transformed into mesh data.

[0076] Currently, the MPEG standards group has defined two data types for dynamic mesh data: mesh data of category 1 with a texture map as color information and mesh data of category 2 with vertex colors as color information.

[0077] The mesh compilation standard for category 1 data is currently in progress, and the standardization of category 2 data is expected to follow. As Figure 1 shown, the overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback processes.

[0078] To provide mesh content services, 3D data acquired through multiple cameras or special cameras can be processed into a mesh data type through a series of steps to generate a mesh video. The generated mesh video can be sent through a series of operations, and the receiving side can process the received data back into a mesh video for rendering. Through this process, the mesh video can be provided to the user, allowing the user to interactively utilize the mesh content according to their intentions.

[0079] As Figure 1 shown, the mesh compression system may include a transmitting device 100 and a receiving device 110. The transmitting device 100 can encode the mesh video to output a bitstream, which can be transmitted to the receiving device 110 in the form of a file or a stream (stream segment) via a digital storage medium or a network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0080] In the transmitting device 100, the encoder can be referred to as a mesh video / image / picture / frame encoding device. In the receiving device 110, the decoder can be referred to as a mesh video / image / picture / frame decoding device. The transmitter can be included in the mesh video encoder, and the receiver can be included in the mesh video decoder. The renderer 114 can include a display, and the renderer and / or the display can be configured as a separate device or an external component. The transmitting device 100 and the receiving device 110 can also include separate internal or external modules / units / components for the feedback process.

[0081] Mesh data uses multiple polygons to represent the surface of an object. Each polygon is defined by vertices in 3D space and connection information indicating how the vertices are connected. Additionally, vertex attributes such as color and normal vector can be included in the data. Mapping information that allows the mesh surface to be mapped onto a 2D plane can also be included in the mesh attributes. Mapping is typically described using a set of parametric coordinates (referred to as UV coordinates or texture coordinates) related to the mesh vertices. The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as texture, normal, and displacement. Here, displacement can be used interchangeably with displacement information or displacement vector.

[0082] The mesh video acquisition unit 101 may include processing 3D object data acquired through a camera or the like into a mesh data type with the above-mentioned attributes through a series of operations, and generating a video composed of the mesh data. In the mesh video, the attributes of the mesh (e.g., vertices, polygons, connections between vertices, colors, and normals) may change over time. A mesh video with attributes and connection information that change over time is called a dynamic mesh video.

[0083] The mesh video encoder 102 may encode the input mesh video into one or more video streams. A video may include multiple frames, and each frame may correspond to a still image / frame. In the present disclosure, the mesh video may include mesh images / frames / frames. The term "mesh video" may be used interchangeably with mesh images / frames / frames. The mesh video encoder 102 may perform a video-based dynamic mesh (V-Mesh) compression process. For compression and compilation efficiency, the mesh video encoder 102 may perform a series of processes such as prediction, transformation, quantization, and entropy compilation. The encoded data (encoded video / image information) may be output in the form of a bitstream.

[0084] The file / fragment encapsulation module 103 may encapsulate the encoded mesh video data and / or mesh video-related metadata in the form of a file or the like. The mesh video-related metadata may be received from a metadata processor. The metadata processing unit may be included in the mesh video encoder 102 or may be configured as a separate component / module. The file / fragment encapsulation module 103 may encapsulate the data into a file format such as ISOBMFF or process it into a form such as DASH fragments. According to an embodiment, the file / fragment encapsulator 103 may include mesh video-related metadata in the file format. For example, the mesh video metadata may be included in the boxes at various levels in the ISOBMFF file format or as data on a separate track in the file. In some embodiments, the file / fragment encapsulator 103 may encapsulate the mesh video-related metadata into the file.

[0085] The sending processor may apply processing to the encapsulated mesh video data based on the file format for transmission. The sending processor may be included in the transmitter 104 or implemented as a separate component / module. The sending processor may process the mesh video data according to any transmission protocol. The processing for transmission may include transmission processing via a broadcast network and transmission processing via broadband. In some embodiments, the sending processor may receive the mesh video-related metadata and the mesh video data from the metadata processor and process them for transmission.

[0086] The transmitter 104 can send the encoded video / image information in the form of a file or a stream or the data output in the form of a bitstream to the receiver 111 of the receiving device 110 via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter 104 can include elements for generating a media file in a predetermined file format and can include elements for transmitting via a broadcast / communication network. The receiver 111 can extract the bitstream and transfer it to a decoding device.

[0087] The receiver 111 can receive the mesh video data sent by the mesh data transmission device. According to the channel for transmission, the receiver 111 can receive the mesh video data via a broadcast network or a broadband network, or can receive the mesh video data via a digital storage medium.

[0088] The receiving processor can perform processing on the received mesh video data according to the transmission protocol. The receiving processor can be included in the receiver 111 or can be configured as a separate component / module. In order to correspond to the processing performed for transmission on the sending side, the receiving processor can perform the reverse process of the operations of the above-mentioned sending processor. The receiving processor can transfer the obtained mesh video data to the file / fragment de-packager 112 and transfer the obtained mesh video-related metadata to the metadata parser. The mesh video-related metadata obtained by the receiving processor can be in the form of a signaling table.

[0089] The file / fragment de-packager 112 can de-package the mesh video data in the form of a file received from the receiving processor. The file / fragment de-packager 112 can de-package a file, etc. according to ISOBMFF, etc. to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transferred to the mesh video decoder 113, and the obtained mesh video-related metadata (metadata bitstream) can be transferred to the metadata processor. The mesh video bitstream can include metadata (metadata bitstream). The metadata processor can be included in the mesh video decoder 113 or can be configured as a separate component / module. The mesh video-related metadata obtained by the file / fragment de-packager 112 can be in the form of a box or a track in a file format. When necessary, the file / fragment de-packager 112 can receive the metadata required for de-packaging from the metadata processor. The mesh video-related metadata can be transferred to the mesh video decoder 113 for use in the mesh video decoding process or transferred to the renderer 114 for use in the mesh video rendering process.

[0090] The mesh video decoder 113 can receive an input bitstream and perform inverse operations corresponding to the operations of the mesh video encoder 102 to decode the video / image. The decoded mesh video / image can be displayed on the display of the renderer 114. The user can view all or part of the rendering result through a VR / AR display, a general display, etc.

[0091] The feedback process can include sending various types of feedback information that can be obtained during the rendering / display operation to the decoder on the sending side or the receiving side. The feedback process can provide interactivity when consuming mesh video. In some embodiments, the feedback process can include sending head orientation information, viewport information indicating the area that the user is currently viewing, etc. In some embodiments, the user can interact with objects implemented in a VR / AR / MR / autonomous driving environment. In this case, information related to the interaction during the feedback process can be transmitted to the sending side or the service provider. In some embodiments, the feedback process can be skipped.

[0092] The head orientation information can refer to information about the user's head position, angle, movement, etc. Based on this information, information about the area that the user is currently viewing within the mesh video (i.e., viewport information) can be calculated.

[0093] The viewport information can be information about the area that the user is currently viewing in the mesh video. Gaze analysis can be performed based on this information to determine how the user consumes the mesh video, how long the user looks at a specific area of the mesh video, etc. Gaze analysis can be performed on the receiving side, and the results can be transmitted to the sending side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0094] In some embodiments, the above feedback information can be transmitted not only to the transmitter, but also consumed on the receiving side. In other words, operations such as decoding and rendering can be performed on the receiving side based on the above feedback information. For example, based on the head orientation information and / or the viewport information, only the mesh video of the area that the user is currently viewing can be preferentially decoded and rendered.

[0095] This disclosure relates to embodiments of dynamic mesh video compression as described above. The methods / embodiments disclosed herein can be applied to the video-based dynamic mesh compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or any next-generation video / image compilation standard. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time. It can perform lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, and AR / VR.

[0096] The dynamic mesh video compression method described below is based on the V-mesh method of MPEG.

[0097] In the present disclosure, a picture / frame generally may refer to a unit representing an image at a specific time.

[0098] A pixel or pel may refer to the smallest unit constituting a picture (or video). Additionally, the term "sample" may be used as a term corresponding to a pixel. A sample generally may indicate a pixel or a pixel value. It may indicate only a pixel / pixel value of a luminance component, or may indicate only a pixel / pixel value of a chrominance component, or may indicate only a pixel / pixel value of a depth component.

[0099] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. In some cases, the term unit may be used interchangeably with terms such as block or region. Generally, an M×N block may include a set (or array) of samples (or a sample array) or transform coefficients composed of M columns and N rows.

[0100] As described above, Figure 1 the encoding process is performed as follows.

[0101] In other words, a compression method based on video-based dynamic mesh compression (V-Mesh) may provide a method for compressing dynamic mesh video data based on a 2D video codec (e.g., High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC)). During the V-Mesh compression process, the following data is received as input and compressed.

[0102] Input mesh: Includes 3D coordinates of vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connections between vertices constituting the surface. The mesh surface may be represented by triangles or other polygons, and connection information between vertices constituting the surface is stored according to a predetermined shape. The input mesh may be stored in the OBJ file format.

[0103] Attribute map (texture map may also be used interchangeably hereinafter): Contains information about attributes (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the mesh surface onto a 2D image. The mapping indicating which part of the mesh (surface or vertex) corresponds to each piece of data in the attribute map is based on the mapping information included in the input mesh. Since the attribute map has data for each frame of the mesh video, it may also be referred to as an attribute map video. The attribute map in the V-Mesh compression method mainly contains color information about the mesh, and is stored in an image file format (PNG, BMP, etc.).

[0104] Material library file: Contains information about material properties used in the mesh, especially information for linking the input mesh to the corresponding attribute map. It is stored in the Wavefront Material Template Library (MTL) file format.

[0105] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0106] Base mesh: The input mesh is represented by using the minimum vertices determined according to user criteria through extraction of the input mesh via a preprocessing process.

[0107] Displacement: Displacement information used to represent the input mesh as similarly as possible using the base mesh, represented in 3D coordinates.

[0108] Atlas information: Metadata required to reconstruct the mesh using the base mesh, displacement, and attribute map information. It can be used to generate and use sub-units (sub-meshes, patches, etc.) of the mesh.

[0109] Refer to Figures 2 to 7 Describes a method for encoding mesh position information (or vertex position information). Refer to Figures 6 to 10 etc. for methods of reconstructing mesh position information to encode attribute information (attribute map).

[0110] Figure 2 Shows the V-MESH compression method according to an embodiment.

[0111] Figure 2 Shows Figure 1 The encoding process, where the encoding process may include a preprocessing process and an encoding process. As Figure 2 shown, Figure 1 The mesh video encoder 102 of Figure 1 can include a preprocessor 200 and an encoder 201. Additionally, Figure 1 The sending device of Figure 2 can be generally referred to as an encoder, and Figure 2 The mesh video encoder 102 of Figure 2 can be referred to as an encoder. As Figure 2 shown, the V-Mesh compression method may include preprocessing 200 and encoding 201.

[0112] The preprocessor 200 can receive a static or dynamic mesh (M(i)) and / or an attribute map (A(i)). The preprocessor 200 can generate a base mesh m(i) and / or a displacement d(i) through preprocessing. The preprocessor 200 can receive feedback information from the encoder 201 and can generate a base mesh and / or a displacement based on the feedback information.

[0113] The encoder 201 may receive a base mesh m(i), a displacement d(i), a static or dynamic mesh M(i), and / or an attribute map A(i). In the present disclosure, at least one of the base mesh m(i), the displacement d(i), the static or dynamic mesh M(i), and / or the attribute map A(i) may be referred to herein as mesh-related data. The encoder 201 may encode the mesh-related data to generate a compressed bitstream.

[0114] Figure 3 Illustrates preprocessing in V-MESH compression according to an embodiment.

[0115] Figure 3 Illustrates Figure 2 the configuration and operation of the preprocessor. In Figure 3 it, the input mesh may include a static or dynamic mesh M(i) and / or an attribute map A(i). The input mesh may also include 3D coordinates of vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between vertices constituting the surface.

[0116] Figure 3 Illustrates the process of performing preprocessing on the input mesh. The preprocessing 200 may include four operations: 1) Group of Frames (GoF) generation, 2) mesh decimation, 3) UV parameterization, and 4) fitting a subdivision surface (300). According to an embodiment, GoF generation may be referred to as the GoF generation process or the GoF generator, mesh decimation may be referred to as the mesh simplification process or the mesh decimation section, UV parameterization may be referred to as the UV parameterization process or the UV parameterization section, and fitting a subdivision surface may be referred to as the fitting subdivision surface process or the fitting subdivision surface section. The preprocessor 200 may generate a displacement and / or a base mesh from the received input mesh and transmit it to the encoder 201. The preprocessor 200 may transmit GoF information related to GoF generation to the encoder 201.

[0117] Next, describe Figure 3 each operation of

[0118] GoF generation: A process of generating a reference structure of mesh data. When the mesh of the previous frame and the current mesh have the same number of vertices, the same number of texture coordinates, the same vertex connection information, and the same texture coordinate connection information, the previous frame may be set as the reference frame. In other words, if only the vertex coordinate values differ between the current input mesh and the reference input mesh, the encoder 201 may perform inter-frame coding. Otherwise, it performs intra-frame coding for the frame.

[0119] Mesh decimation: A process of simplifying the input mesh to create a simplified mesh (referred to as the base mesh). Vertices to be removed may be selected from the original mesh based on user-defined criteria, and then the selected vertices and the triangles connected to the selected vertices may be removed.

[0120] During the process of performing mesh extraction, the voxelized input mesh, the target triangle ratio (TTR), and the minimum triangle component (CCCount) information can be transmitted as inputs, and the extracted mesh can be obtained as an output. During this process, the connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.

[0121] UV parameterization: The process of mapping a 3D surface to a texture domain for the extracted mesh. The parameterization can be performed using the UVAtlas tool. This process generates mapping information indicating where each vertex of the extracted mesh can be mapped onto a 2D image. The mapping information is represented as texture coordinates and stored, and the final base mesh is generated through this process.

[0122] Fitting a subdivision surface (300): The process of performing subdivision on the extracted mesh (i.e., the extracted mesh with texture coordinates). The displacement and the base mesh generated through this process are output to the encoder 201. A user-defined method (e.g., the edge midpoint method) can be applied as the subdivision method. The fitting process is performed such that the input mesh and the subdivided mesh become similar to each other. The mesh on which the fitting process is performed will be referred to as the fitted subdivision mesh in this article.

[0123] Figure 4 Shows the edge midpoint subdivision method according to an embodiment.

[0124] Figure 4 Shows with reference to Figure 3 the edge midpoint subdivision method of the fitted subdivision surface described. With reference to Figure 4 , the original mesh containing four vertices is subdivided to create a submesh. The submesh can be created by creating new vertices at the midpoints of the edges between the vertices. Then, the fitting process is performed to make the input mesh and the submesh similar to each other, thereby obtaining the fitted subdivision mesh.

[0125] Once the fitted subdivision mesh is generated, the displacement is calculated based on this result and the previously compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). In other words, the reconstructed base mesh is subdivided in the same manner as the fitted subdivision surface. The position difference between each vertex in the result and the fitted subdivision mesh is the displacement of each vertex. Since the displacement represents the position difference in 3D space, it is represented as a value in the (x, y, z) space in the Cartesian coordinate system. According to the user input parameters, the (x, y, z) coordinate values can be converted to the (normal, tangent, binormal) coordinate values in the local coordinate system.

[0126] Figure 5 Shows the displacement generation process according to an embodiment. Figure 5 The displacement generation process of can be performed by the preprocessor 200, or can be performed by the encoder 201.

[0127] Figure 5 Shows in detail how the displacement is calculated for the fitted subdivision surface 300 as described with reference to Figure 4 description.

[0128] The encoder and / or preprocessor according to an embodiment may include 1) a subdivider, 2) a local coordinate system calculator, and 3) a displacement vector calculator. The subdivider may perform subdivision on the reconstructed base mesh to generate a subdivided reconstructed base mesh. Here, the reconstruction of the base mesh may be performed by the preprocessor 200 or may be performed by the encoder 201. The local coordinate system calculator may receive the fitted subdivision mesh and the subdivided reconstructed base mesh, and may transform the coordinate system related to the mesh into a local coordinate system based on the received meshes. The local coordinate system calculation may be optional. The displacement calculator calculates the position difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, it may generate the position difference between the vertices in the two input meshes. The position difference between the vertices is the displacement.

[0129] The mesh data sending method and apparatus according to an embodiment may encode the mesh data as follows. Mesh data is a term including point cloud data. The point cloud data according to an embodiment (which may be abbreviated as point cloud) may refer to data including vertex coordinates (also referred to as geometric information) and color information (also referred to as attribute information). Additionally, geometric images, attribute images, occupancy maps, and auxiliary information (also referred to as patch information) generated by patch generation and packing based on vertex coordinates and color information may also be referred to as point cloud data. Therefore, point cloud data including connection information may be referred to as mesh data. The terms point cloud and mesh data may be used interchangeably herein.

[0130] According to an embodiment, the V-Mesh compression (reconstruction) method may include intra-frame coding ( Figure 6 ) and inter-frame coding ( Figure 7 ).

[0131] Based on the result generated by the above-mentioned GoF, intra-frame coding or inter-frame coding is performed. In intra-frame coding, the data to be compressed may be the base mesh, displacement, attribute map, etc. In inter-frame coding, the data to be compressed may be the displacement, attribute map, and the motion field between the reference base mesh and the current base mesh.

[0132] Figure 6 Shows the intra-frame coding process in the V-MESH compression method according to an embodiment. Figure 6 Each component of the intra-frame coding process corresponds to hardware, software, a processor, and / or a combination thereof.

[0133] Figure 6 The encoding process of Figure 1 details the encoding of the mesh video encoder 102 of Figure 1The encoding is a configuration of the grid video encoder 102 when intra-frame encoding is performed. Figure 6 The encoder may include a pre-processor 200 and / or an encoder 201. Figure 6 The pre-processor 200 and the encoder 201 may correspond to Figure 3 the pre-processor 200 and the encoder 201.

[0134] The pre-processor 200 may receive an input grid and perform the above-mentioned pre-processing. A base grid and / or a fitted subdivision grid may be generated through the pre-processing.

[0135] The quantizer 411 of the encoder 201 may quantize the base grid and / or the fitted subdivision grid. The static grid encoder 412 may encode the static grid (i.e., the quantized base grid) and generate a bitstream containing the encoded base grid (i.e., the compressed base grid bitstream). The static grid decoder 413 may decode the encoded static grid (i.e., the encoded base grid). The inverse quantizer 414 may inverse-quantize the quantized static grid (i.e., the base grid) and output the reconstructed (restored) base grid. The displacement calculator 415 may generate a displacement based on the reconstructed static grid (i.e., the base grid) and the fitted subdivision grid. According to an embodiment, the displacement calculator 415 subdivides the reconstructed base grid and then calculates the displacement, i.e., the position difference of each vertex between the subdivided base grid and the fitted subdivision grid. In other words, when the fitted subdivision grid is similar to the original grid, the displacement is a displacement vector that is the position difference between the vertices in the two grids. The forward linear lifting unit 416 may perform a lifting transformation on the input displacement to generate lifting coefficients (also referred to as transform coefficients). The quantizer 417 may quantize the lifting coefficients. The image packer 418 may pack the image based on the quantized lifting coefficients. The video encoder 419 may encode the packed image. That is, the quantized lifting coefficients are packed by the image packer 418 into frames as 2D images, compressed by the video encoder 419, and output as a displacement bitstream (i.e., the compressed displacement bitstream).

[0136] The video decoder 420 decodes the compressed displacement bitstream. The image unpacker 421 may unpack the decoded displacement frame to output the quantized lifting coefficients. The inverse quantizer 422 may inverse-quantize the quantized lifting coefficients. The inverse linear lifting unit 423 applies an inverse lift to the inverse-quantized lifting coefficients to generate a reconstructed displacement. The grid reconstructor 424 restores the reconstructed and deformed grid based on the reconstructed displacement output from the inverse linear lifting unit 423 and the reconstructed base grid (also referred to as the subdivided reconstructed base grid) output from the inverse quantizer 414. The reconstructed and deformed grid is referred to as the reconstructed deformed grid herein.

[0137] The attribute transfer 425 receives an input mesh and / or an input attribute map, and regenerates an attribute map based on the reconstructed deformed mesh. The attribute map refers to a texture map corresponding to the attribute information among the mesh data components. In the present disclosure, the terms attribute map and texture map may be used interchangeably. The push-pull filling unit 426 may fill data into the attribute map based on the push-pull method. The color space converter 427 may convert the space of the color components of the attribute map. For example, the attribute map may be converted from the RGB color space to the YUV color space. The video encoder 428 may encode the attribute map to output a compressed attribute bitstream.

[0138] The multiplexer 430 may multiplex the compressed base mesh bitstream, the compressed displacement bitstream, and the compressed attribute bitstream to generate a compressed bitstream.

[0139] In Figure 6 the displacement calculator 415 may be included in the preprocessor 200. Additionally, at least one of the quantizer 411, the static mesh encoder 412, the static mesh decoder 413, or the inverse quantizer 414 may be included in the preprocessor 200.

[0140] As Figure 6 described, the intra-frame encoding method includes base mesh encoding (also known as static mesh encoding). That is, when performing intra-frame encoding on the current input mesh frame, the base mesh generated during the preprocessing of the preprocessor 200 may be quantized by the quantizer 411 and then encoded by the static mesh encoder 412 using static mesh compression techniques. For example, in the V-Mesh compression method, the Draco technique is applied to encode the base mesh, and vertex position information, mapping information (texture coordinates), vertex connection information, etc. related to the base mesh are subject to compression.

[0141] Figure 6 The encoder in Figure 7 compresses the base mesh, displacement, and attributes in the frame to generate a bitstream, while

[0142] Figure 7 shows the inter-frame encoding process in the V-MESH compression method according to an embodiment. Figure 7 The respective components of the inter-frame encoding process of

[0143] Figure 7 correspond to hardware, software, processors, and / or combinations thereof. Figure 1 The encoding process of Figure 1 details the encoding of Figure 7 That is, it represents the configuration of the encoder when the encoding of Figure 7The pre-processor 200 and the encoder 201 may correspond to Figure 3 the pre-processor 200 and the encoder 201.

[0144] For the components corresponding to Figure 6 the encoding operation of Figure 7 the encoding operation, refer to Figure 6 the description of Figure 7 That is, the operations of the quantizer 511, displacement calculator 515, wavelet transformer 516, quantizer 517, image packer 518, video encoder 519, video decoder 520, image unpacker 521, inverse quantizer 522, inverse wavelet transformer 523, mesh reconstructor 524, attribute transfer 525, push-pull padding 526, color space converter 527, video encoder 528, and multiplexer 530 in Figure 6 are the same as or similar to the operations of the quantizer 411, static mesh encoder 412, static mesh decoder 413, and inverse quantizer 414, displacement calculator 415, forward linear lifting unit 416, quantizer 417, image packer 418, video encoder 419, video decoder 420, image unpacker 421, inverse quantizer 422, inverse linear lifting unit 423, and mesh reconstructor 424, attribute transfer 425, push-pull padding 426, color space converter 427, video encoder 428, and multiplexer 430 in Figure 7 and thus will not be described in detail herein to avoid redundancy.

[0145] In Figure 7 , for inter-frame based encoding, the motion encoder 512 may obtain and encode the motion vector between the reconstructed quantized reference base mesh and the quantized current base mesh, and output a compressed motion bitstream. The motion encoder 512 may be referred to as a motion vector encoder. The base mesh reconstructor 513 may reconstruct the base mesh based on the reconstructed quantized reference base mesh and the encoded motion vector. The reconstructed base mesh is inverse quantized by the inverse quantizer 514 and output to the displacement calculator 515.

[0146] In Figure 7 , the displacement calculator 515 may be included in the pre-processor 200. Additionally, at least one of the quantizer 511, motion encoder 512, base mesh reconstructor 513, or inverse quantizer 514 may be included in the pre-processor 200.

[0147] As referred to Figure 7As described, the inter-frame encoding method may include motion field encoding (also known as motion vector encoding). When the reference mesh and the current input mesh have a one-to-one vertex correspondence and differ only in the position information of the vertices, inter-frame encoding can be performed. When performing inter-frame encoding, the base mesh may not be compressed. Instead, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field (or motion vector), can be calculated and encoded. The reference base mesh is the result of quantifying the decoded base mesh data and is determined by the reference frame index determined in the generation of GoF. The motion field can be encoded as it is. Alternatively, a predicted motion field can be calculated by taking the average of the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and as the difference between the value of the predicted motion field and the value of the motion field of the current vertex, the residual motion field can be encoded. The value of the residual motion field can be encoded using entropy coding. In addition to the motion field encoding in inter-frame encoding, the process of encoding displacement and attribute maps is the same as the structure of the intra-frame encoding method except for the base mesh encoding.

[0148] Figure 8 Shows the lifting transformation process of displacement according to an embodiment.

[0149] Fig. 9 Shows the process of packing transformation coefficients (also known as lifting coefficients) into a 2D image according to an embodiment.

[0150] Figure 8 and Fig. 9 respectively show Figure 6 and Figure 7 the processes of transforming displacement and packing transformation coefficients in the encoding process.

[0151] The encoding method according to an embodiment includes displacement encoding.

[0152] After base mesh encoding and / or motion field encoding, a reconstructed base mesh can be generated through reconstruction and inverse quantization, and the displacement can be calculated between the subdivision result of the reconstructed base mesh and the fitted subdivision mesh generated by fitting the subdivision surface (see 415 in Figure 6 or 515 in Figure 7 ). The data transformation process (e.g., wavelet transformation) can be applied to the displacement information for efficient encoding (see 416 in Figure 6 or 516 in Figure 7 Figure 7 ).

[0153] Figure 8 Shown by Figure 6 the forward linear lifting unit 416 of Figure 7The wavelet transformer 516 uses a lifting transform to transform the displacement information. For example, a linear wavelet-based lifting transform can be performed. The transform coefficients generated by the transform process are quantized by a quantizer 417 (or 517), and then packed into a 2D image by an image packer 418 (or 518), as shown in Fig. 9 shown. The transform coefficients can be organized into blocks, with one block for every 256 (=16×16) units. Each block can be packed in a z-scan order. The number of rows in a block is fixed at 16, but the number of columns in a block can be determined by the number of vertices in the subdivided base grid. Inside a block, the transform coefficients can be sorted and packed according to the Morton code. For the packed image, a displacement video can be generated every GoF. The displacement video can be encoded by a video encoder 419 (or 519) using a conventional video compression codec.

[0154] Referring to Figure 8 , the base grid (original) can include the vertices and edges of LoD0. The first subdivided grid generated by dividing (or subdividing) the base grid includes vertices generated by further dividing (or subdividing) the edges of the base grid. The first subdivided grid contains the vertices of LoD0 and the vertices of LoD1. LoD1 includes the subdivided vertices and the vertices from the base grid (LoD0). The first subdivided grid can be divided (or subdivided) to generate a second subdivided grid. The second subdivided grid contains LoD2. LoD2 includes the base grid vertices (LoD0), LoD1 contains the vertices further divided (or subdivided) from LoD0, and LoD2 contains the vertices further divided (or subdivided) from LoD1. LoD is a level of detail indicating how detailed the mesh data content is. As the index of the level increases, the distance between vertices decreases and the level of detail increases. In other words, as the value of LoD decreases, the detail of the mesh data content deteriorates. As the value of LoD decreases, the detail of the mesh data content enhances. LoD N contains the vertices included in LoD N-1. In the case of further dividing the grid (or vertices) by subdivision, the grid can be encoded by considering the previous vertices v1 and v2 and the subdivided vertex v based on a prediction and / or update method. Instead of encoding the information of the current LoD N as it is, a residual with respect to the previous LoD N-1 can be generated. Therefore, the residual can be used to encode the grid to reduce the size of the bitstream. The prediction process refers to the operation of predicting the current vertex v from the previous vertices v1 and v2. Since adjacent subdivided meshes have similar data, this property can be utilized for efficient encoding. The current vertex position information is predicted from the residual of the previous vertex position information, and the previous vertex position information is updated by the residual. In the present disclosure, vertices and points can be used interchangeably. LoD can be defined in the subdivision of the base grid. According to an embodiment, the subdivision of the base grid can be performed by the preprocessor 200, or can be performed by a separate component / module.

[0155] Referring to Fig. 9, vertices have transformation coefficients (also referred to as lifting coefficients) generated by a lifting transformation. The transformation coefficients of vertices related to the lifting transformation can be packed into an image by an image packer 418 (or 518), and then encoded by a video encoder 419 (or 519).

[0156] Fig.10 Shows the attribute transfer process in the V-MESH compression method according to an embodiment.

[0157] According to an embodiment, Fig.10 Shows Figure 6 、 Figure 7 The detailed operations of the attribute transfer 425 (or 525) in the encoding of etc.

[0158] The encoding according to an embodiment includes attribute map encoding. According to an embodiment, the attribute map encoding can be performed by Figure 6 the video encoder 428 of Figure 7 or the video encoder 528 of

[0159] According to an embodiment, in the present disclosure, the encoder compresses the information about the input mesh through base mesh encoding (i.e., intra-frame encoding), motion field encoding (i.e., inter-frame encoding), and displacement encoding. The input mesh compressed during the encoding process is reconstructed through base mesh decoding (intra-frame), motion field decoding (inter-frame), and displacement video decoding, and as a reconstruction result, the reconstructed deformed mesh (hereinafter referred to as the reconstructed deformed mesh) is used to compress the input attribute map, as Figure 6 and Figure 7 shown. The reconstructed deformed mesh has position information about vertices, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as Fig.10 shown, in the V-Mesh compression method, a new attribute map with color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of the attribute transfer 425 (or 525).

[0160] According to an embodiment, the attribute transfer 425 (or 525) first checks for each point P(u, v) in the 2D texture domain whether the corresponding vertex is within the texture triangle of the reconstructed deformed mesh. When the corresponding vertex is within the texture triangle T, the attribute transfer calculates the barycentric coordinates (α, β, γ) of P(u, v) based on the triangle T. Then, it calculates the 3D coordinates M(x, y, z) of P(u, v) based on the 3D vertex positions of the triangle T and (α, β, γ). The vertex coordinates M'(x', y', z') corresponding to the position closest to the calculated M(x, y, z) and the triangle T' containing this vertex are searched for in the input mesh domain. Then, the barycentric coordinates (α', β', γ') of M'(x', y', z') in the triangle T' are calculated. The texture coordinates (u', v') are calculated based on the texture coordinates corresponding to the three vertices of the triangle T' and (α', β', γ'), and the color information corresponding to the coordinates is searched for in the input attribute map. The color information found in this way is then assigned to the (u, v) pixel position in the new input attribute map. If P(u, v) does not belong to any triangle, the pixel at this position in the new input attribute map is filled with a color value using a filling algorithm (e.g., the push-pull algorithm of the push-pull filling 426 (or 526)).

[0161] The new attribute map generated by the attribute transfer 425 (or 525) is bundled into the GoF to construct an attribute map video, which is compressed using the video codec of the video encoder 428 (or 528).

[0162] It can be seen from Fig.10 the reference relationships among the input mesh, the input attribute map, the reconstructed deformed mesh, and the reconstructed attribute map.

[0163] Figure 1 The decoding process of Figure 1 can execute the inverse process of the encoding process of

[0164] Fig.11 shows the intra-frame decoding process of the V-Mesh technology according to an embodiment.

[0165] Fig.11 shows Figure 1 the configuration and operation of the mesh video decoder 113 of the receiving device of Fig.11 In addition, Figure 6 shows that the mesh data can be reconstructed by executing the inverse process of the intra-frame encoding process of Fig.11 The various components of the intra-frame decoding process of

[0166] ​First, the bitstream (i.e., the compressed bitstream) received and input to the demultiplexer 611 of the intra-frame decoder 610 can be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information about the mesh (e.g., V-PCC / V3C). The term V-PCC (Video-based Point Cloud Compression) used in this disclosure may have the same meaning as V3C (Visual Volume Video-based Coding). The two terms may be used interchangeably. Therefore, in this disclosure, the term V-PCC may be interpreted as V3C.

[0167] According to an embodiment, the mesh substream can be input to the static mesh decoder 612 and decoded by it, the displacement substream can be input to the video decoder 613 and decoded by it, and the attribute map substream can be input to the video decoder 617 and decoded by it.

[0168] According to an embodiment, the mesh substream can be decoded by the decoder 612 of the static mesh codec (e.g., GoogleDraco) used in the encoding to reconstruct connection information, vertex geometry information, vertex texture coordinates, etc. related to the decoding result of the reconstructed quantization base mesh (e.g., the reconstructed base mesh).

[0169] According to an embodiment, the displacement substream can be decoded into a displacement video by the decoder 613 of the video compression codec used in the encoding. Then, image unpacking is performed by the image unpacker 614, inverse quantization is performed by the inverse quantizer 615, and inverse transformation is performed by the inverse linear lifting unit 616 to reconstruct the displacement information about each vertex (i.e., the reconstructed displacement).

[0170] According to an embodiment, the base mesh reconstructed by the static mesh decoder 612 is inverse quantized by the inverse quantizer 620 and output to the mesh reconstructor 630. The mesh reconstructor 630 reconstructs the reconstructed deformed mesh (i.e., the decoded mesh) based on the reconstructed displacement output from the inverse linear lifting unit 616 and the reconstructed base mesh output from the inverse quantizer 620. In other words, the inverse quantized reconstructed base mesh is combined with the reconstructed displacement information to generate the final decoded mesh. In this disclosure, the final decoded mesh is referred to as the reconstructed deformed mesh.

[0171] According to an embodiment, the attribute map substream is decoded by the decoder 617 corresponding to the video compression codec used in the encoding, and then the final attribute map (i.e., the decoded attribute map) is reconstructed by the color converter 640 through color format transformation, color space conversion, etc.

[0172] According to an embodiment, the reconstructed decoded mesh and the decoded attribute map can be used as the final mesh data available to the user on the receiving side.

[0173] Refer to Fig.11, the received compressed bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term substream is interpreted to mean a partial bitstream included in the bitstream. The bitstream contains patch information (data), mesh information (data), displacement information (data), and attribute map information (data).

[0174] As described above, Fig.11 The decoder of performs intra-frame decoding as follows. The static mesh decoder 612 decodes the mesh substream to generate a reconstructed quantized base mesh, and the inverse quantizer 620 inversely applies the quantization parameters of the quantizer to generate a reconstructed base mesh. The video decoder 613 decodes the displacement substream, the image unpacker 614 unpacks the image of the decoded displacement video, and the inverse quantizer 615 inversely quantizes the quantized image. The inverse linear lifting unit 616 applies a lifting transform in the inverse process of the encoder to generate a reconstructed displacement. The mesh reconstructor 630 generates a reconstructed deformed mesh based on the reconstructed base mesh and the reconstructed displacement. The video decoder 617 decodes the attribute map substream, and the color transformer 64 transforms the color format and / or space of the decoded attribute map to generate a decoded attribute map.

[0175] Fig.12 Shows the inter-frame decoding process of the V-Mesh technology.

[0176] Fig.12 Shows Figure 1 The configuration and operation of the mesh video decoder 113 of the receiving device of. In Fig.12 It is possible to reconstruct the mesh data by performing the inverse process of the inter-frame encoding process of Figure 7 . Fig.12 The various components of the intra-frame decoding process of correspond to hardware, software, and / or a combination thereof.

[0177] First, the bitstream received and input to the demultiplexer 711 of the intra-frame decoder 710 can be separated into a motion substream (also referred to as a motion vector substream), a displacement substream, an attribute map substream, and a substream containing patch information about the mesh (e.g., V3C / V-PCC).

[0178] According to an embodiment, the motion substream can be input to and decoded by the motion decoder 712, the displacement substream can be input to and decoded by the video decoder 713, and the attribute map substream can be input to and decoded by the video decoder 717.

[0179] According to an embodiment, the motion sub-stream is entropy decoded and inverse prediction decoded by the motion decoder 712 to reconstruct motion information (also referred to as motion vector information). The base mesh reconstructor 718 combines the reconstructed motion information with a pre-reconstructed and stored reference base mesh to generate a reconstructed quantized base mesh for the current frame. The inverse quantizer 720 applies inverse quantization to the reconstructed quantized base mesh to generate a reconstructed base mesh. The video decoder 713 decodes the displacement sub-stream, the image unpacker 714 unpacks the images of the decoded displacement video, and the inverse quantizer 715 applies inverse quantization to the quantized images. The inverse linear lifting unit 716 applies a lifting transform in the inverse process of the encoder to generate a reconstructed displacement. The mesh reconstructor 730 generates a reconstructed deformed mesh (i.e., the final decoded mesh) based on the reconstructed base mesh and the reconstructed displacement.

[0180] According to an embodiment, the video decoder 717 decodes the attribute map sub-stream in the same manner as intra-frame decoding, and the color transformer 740 transforms the color format and / or space of the decoded attribute map to generate a decoded attribute map. The decoded mesh and the decoded attribute map can be used as the final mesh data available to the user on the receiving side.

[0181] Refer to Fig.12 , the bitstream contains motion information (also referred to as motion vectors), displacement, and attribute maps. Since inter-frame decoding is performed, Fig.12 the process also includes decoding inter-frame motion information. A reconstructed base mesh is generated by decoding the motion information and generating a reconstructed quantized base mesh for the motion information based on the reference base mesh. For Fig.12 the operations in Fig.11 that are the same as the operations in Fig.11 , refer to the description in

[0182] Fig.13 Fig. shows a mesh data transmitting device according to an embodiment.

[0183] Fig.13 Corresponds to Figure 1 the transmitting device 100 or the mesh video encoder 102 corresponding to Figure 2 , Figure 6 or Figure 7 the encoder (pre-processor and encoder) of Fig.13 and / or the corresponding transmitting and encoding device.

[0184] The operation process of compressing and transmitting dynamic mesh data using the V-Mesh compression technology at the transmitting end can be configured as shown in Fig.13 . Fig.13 The transmitting device of

[0185] The pre-processor 811 receives the original mesh and generates an extracted mesh (or base mesh) and a fitted subdivision (or subdivision) mesh. Extraction can be performed based on the vertices of the target number constituting the mesh or the target number of polygons. Parameterization can be performed on the extracted mesh to generate per-vertex texture coordinates and texture connection information. For example, parameterization is the process of mapping a 3D surface onto a texture domain for the extracted mesh. When parameterization is performed using the UVAtlas tool, mapping information indicating where each vertex of the extracted mesh can be mapped onto a 2D image is generated. The mapping information is represented as texture coordinates and stored, and a final base mesh is generated through this process. The mesh information can be quantized from floating-point form to fixed-point form. The result is a base mesh, which can be output to the motion vector encoder 813 or the static mesh encoder 814 through the switch unit 812. The pre-processor 811 can perform mesh subdivision on the base mesh to generate additional vertices. According to the subdivision method, vertex connection information including additional vertices, texture coordinates, and connection information regarding the texture coordinates can be generated. The pre-processor 811 can generate a fitted subdivision mesh by adjusting the vertex positions so that the subdivision mesh becomes similar to the original mesh.

[0186] According to an embodiment, when performing inter-frame encoding on a mesh frame, the base mesh is output to the motion vector encoder 813 through the switch unit 812. When performing intra-frame encoding on a mesh frame, the base mesh is output to the static mesh encoder 814 through the switch unit 812. The motion vector encoder 813 can be referred to as a motion encoder.

[0187] For example, when performing intra-frame encoding on a mesh frame, the base mesh can be compressed by the static mesh encoder 814. In this case, connection information, vertex geometry information, vertex texture information, normal information, etc. related to the base mesh can be encoded. The base mesh bitstream generated by encoding is sent to the multiplexer 823.

[0188] As another example, when performing inter-frame encoding on a mesh frame, the motion vector encoder 813 can receive the base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as inputs, calculate the motion vector between the two meshes, and encode its value. In addition, the motion vector encoder 813 can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by encoding is sent to the multiplexer 823.

[0189] The base mesh reconstructor 815 can receive the base mesh encoded by the static mesh encoder 814 or the motion vectors encoded by the motion vector encoder 813, and generate a reconstructed base mesh. For example, the base mesh reconstructor 815 can perform static mesh decoding on the base mesh encoded by the static mesh encoder 814 to reconstruct the base mesh. In this case, quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding. In another example, the base mesh reconstructor 815 can reconstruct the base mesh based on the reconstructed quantized reference base mesh and the motion vectors encoded by the motion vector encoder 813. The reconstructed base mesh is output to the displacement calculator (or displacement vector calculator) 816 and the mesh reconstructor 820.

[0190] The displacement calculator 816 can perform mesh subdivision on the reconstructed base mesh. The displacement calculator 816 can calculate displacement vectors, which are the values of the vertex position differences between the subdivided reconstructed base mesh and the fitted subdivision (or subdivided) mesh generated by the preprocessor 811. In this case, as many displacement vectors as there are vertices in the subdivision mesh can be calculated. The displacement calculator 816 can transform the displacement vectors calculated in the 3D Cartesian coordinate system to the local coordinate system based on the normal vectors of the respective vertices.

[0191] The displacement vector video generator 817 can include a linear lifting section, a quantizer, and an image packer. That is, in the displacement vector video generator 817, the linear lifting unit can transform the displacement vectors for efficient encoding. According to an embodiment, the transformation can be a lifting transformation, a wavelet transformation, etc. Additionally, the quantizer can perform quantization on the transformed displacement vector values (i.e., the transformation coefficients). In this case, different quantization parameters can be applied to the axes of the transformation coefficients respectively. The quantization parameters can be derived through an agreement between the encoder / decoder. After the transformation and quantization, the displacement vector information can be packed into a 2D image by the image packer. The displacement vector video generator 817 can generate a displacement vector video by grouping the packed 2D images for each frame. A displacement vector video can be generated for each group of frames (GoF) of the input mesh.

[0192] The displacement vector video encoder 818 can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is sent to the multiplexer 823.

[0193] The displacement vector reconstructor 819 may include a video decoder, an image unpacker, an inverse quantizer, and an inverse linear lifting section. That is, in the displacement vector reconstructor 819, the encoded displacement vector is decoded by the video decoder, image unpacking is performed by the image unpacker, inverse quantization is performed by the inverse quantizer, and inverse transformation is performed by the inverse linear lifting unit to reconstruct the displacement vector. The reconstructed displacement vector is output to the mesh reconstructor 820. The mesh reconstructor 820 reconstructs a deformed mesh based on the base mesh reconstructed by the base mesh reconstructor 815 and the displacement vector reconstructed by the displacement vector reconstructor 819. The reconstructed mesh (also referred to as the reconstructed deformed mesh) has reconstructed vertices, vertex - to - vertex connection information, texture coordinates, and texture - coordinate - to - texture - coordinate connection information.

[0194] The texture map video generator 821 may regenerate a texture map based on the texture map (or attribute map) of the original mesh and the reconstructed deformed mesh output from the mesh reconstructor 820. According to an embodiment, the texture map video generator 821 may assign per - vertex color information in the texture map of the original mesh to the texture coordinates of the reconstructed deformed mesh. According to an embodiment, the texture map video generator 821 may generate a texture map video by grouping the frame - level regenerated texture maps into a GoF.

[0195] The generated texture map video may be encoded by the texture map video encoder 822 using a video compression codec. The texture map video bitstream generated by encoding is sent to the multiplexer 823.

[0196] The multiplexer 823 multiplexes a motion vector bitstream (e.g., in the case of inter - frame coding), a base mesh bitstream (e.g., in the case of intra - frame coding), a displacement vector bitstream, and a texture map bitstream into a single bitstream. This single bitstream may be sent to the receiving side through the transmitter 824. Alternatively, for the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream, a file having one or more track data may be generated, or the bitstreams may be encapsulated into segments and sent to the receiving side through the transmitter 824.

[0197] Refer to Fig.13, the transmitter (encoder) can encode the mesh in an intra-frame or inter-frame manner. According to intra-frame encoding, the transmitting device can generate a base mesh, displacement vectors (or displacements), and a texture map (or attribute map). According to inter-frame encoding, the transmitting device can generate motion vectors (or motions), displacement vectors (or displacements), and a texture map (or attribute map). Generate and encode the texture map obtained from the data input unit based on the reconstructed mesh. Generate and encode the displacement based on the vertex position difference between the base mesh and the segmented (or subdivided) mesh. More specifically, the displacement is the position difference between the fitted subdivided mesh and the subdivided reconstructed base mesh, that is, the vertex position difference between the two meshes. Generate the base mesh by decimating the original mesh via preprocessing and encoding the decimated mesh. For motion, generate motion vectors for the mesh in the current frame based on the reference base mesh in the previous frame.

[0198] Fig.14 Shows a mesh data receiving device according to an embodiment.

[0199] Fig.14 Corresponding to Figure 1 the receiving device 110 or the mesh video decoder 113, Fig.11 or Fig.12 the decoder and / or the corresponding receiving and decoding device. Fig.14 The respective components of Fig.14 correspond to hardware, software, processors, and / or combinations thereof. Fig.13 The receiving (decoding) operation of

[0200] The bitstream of the mesh data received by the receiver 910 undergoes file / fragment de-encapsulation and then is demultiplexed by the demultiplexer 911 into a compressed motion vector bitstream (e.g., inter-frame decoding) or a base mesh bitstream (e.g., intra-frame decoding), a displacement vector bitstream, and a texture map bitstream. For example, when the current mesh is inter-frame encoded, the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder 913 through the switch unit 912. In another example, when the current mesh is intra-frame encoded, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder 914 through the switch unit 912. Here, the motion vector decoder 913 can be referred to as a motion decoder.

[0201] According to an embodiment, in the case of applying inter-frame encoding to the current mesh based on the frame header information, the motion vector decoder 913 can decode the motion vector bitstream. According to an embodiment, the motion vector decoder 913 can use the previously decoded motion vector as a predictor and add it to the residual motion vector decoded from the bitstream to reconstruct the final motion vector.

[0202] According to an embodiment, in a case where intra coding is applied to a current mesh based on header information, the static mesh decoder 914 may decode the base mesh bitstream to reconstruct connection information, vertex geometry information, texture coordinates, normal information, etc. related to the base mesh.

[0203] According to an embodiment, the base mesh reconstructor 915 may reconstruct the current base mesh based on the decoded motion vector or the decoded base mesh. For example, in a case where inter coding is applied to the current mesh, the base mesh reconstructor 915 may add the decoded motion vector to a reference base mesh and perform inverse quantization to generate a reconstructed base mesh. In another example, in a case where intra coding is applied to the current mesh, the base mesh reconstructor 915 may perform inverse quantization on the base mesh decoded by the static mesh decoder 914 to generate a reconstructed base mesh.

[0204] According to an embodiment, the displacement vector video decoder 917 may decode a displacement vector bitstream into a video bitstream using a video codec.

[0205] According to an embodiment, the displacement vector reconstructor 918 extracts displacement vector transform coefficients from the decoded displacement vector video, and applies inverse quantization and inverse transformation to the extracted displacement vector transform coefficients to reconstruct a displacement vector. To this end, the displacement vector reconstructor 918 may include an image unpacker, an inverse quantizer, and an inverse linear lifting section. If the reconstructed displacement vector is a value in a local coordinate system, an inverse transformation to a Cartesian coordinate system may be performed.

[0206] The mesh reconstructor 916 may subdivide the reconstructed base mesh to generate additional vertices. Through subdivision, vertex connection information including additional vertices, texture coordinates, and connection information about the texture coordinates may be generated. In this case, the mesh reconstructor 916 may combine the subdivided reconstructed base mesh with the reconstructed displacement vector to generate a final reconstructed mesh (also referred to as a reconstructed deformed mesh).

[0207] According to an embodiment, the texture map video decoder 919 may decode a texture map bitstream into a video bitstream using a video codec to reconstruct a texture map. The reconstructed texture map has color information about each vertex in the reconstructed mesh, and the texture coordinates of each vertex may be used to obtain the color value of the vertex from the texture map.

[0208] According to an embodiment, the mesh reconstructed by the mesh reconstructor 916 and the texture map reconstructed by the texture map video decoder 919 are presented to a user through a rendering process in the mesh data renderer 920.

[0209] Refer to Fig.14, the receiving device (decoder) can decode the mesh in an intra-frame or inter-frame manner. According to intra-frame decoding, the receiving device can receive a base mesh, displacement vectors (or displacements), and texture maps (or attribute maps), and render the mesh data based on the reconstructed mesh and reconstructed texture maps. According to inter-frame decoding, the receiving device can receive motion vectors (or motions), displacement vectors (or displacements), texture maps (or attribute maps), and render the mesh data based on the reconstructed mesh and reconstructed texture maps.

[0210] The mesh data transmitting device and method according to an embodiment can preprocess mesh data, encode the preprocessed mesh data, and transmit a bitstream including the encoded mesh data. The point mesh data receiving device and method according to an embodiment can receive a bitstream including mesh data and decode the mesh data. The mesh data transmitting / receiving method / device according to an embodiment can be referred to as the method / device according to an embodiment. The mesh data transmitting / receiving method / device according to an embodiment can also be referred to as a 3D data transmitting / receiving method / device or a point cloud data transmitting / receiving method / device.

[0211] In the present disclosure, geometric information (or geometry or geometric data) refers to one of the elements constituting a mesh, including vertices (or points), edges, and polygons. Here, a vertex defines a position in 3D space, an edge represents connectivity information between vertices, and a polygon formed by a combination of edges and vertices defines the surface of the mesh. In other words, each vertex constituting the mesh represents a position in 3D space, represented by, for example, X, Y, and Z coordinates. A polygon can be a triangle or a rectangle. Thus, the geometric structure forms the skeleton of a 3D model, defining the shape of the model, which is visually represented when rendered.

[0212] According to an embodiment, the device and method can process mesh data (also referred to as partial decoding) considering partial access. More specifically, the present disclosure describes a method for selectively decoding a part of data effectively when necessary during the transmission and reception of mesh data due to factors such as receiver performance or transmission speed.

[0213] That is to say, in the video-based dynamic mesh compilation (V-DMC) technology, which is a method of compressing 3D dynamic mesh data using a traditional 2D video codec, the transmitter is designed to compress and transmit the entire input mesh data (i.e., the entire mesh content), and the receiver is designed to decode the entire bitstream to reconstruct the entire mesh area. In particular, current V-DMC technology does not employ any compilation method or signaling for partial access. However, when a user utilizes the content at the device, application, or software level, there may be cases where only a specific region rather than the entire content is required. In such cases, it is more efficient for the receiving side to receive information about the desired region in terms of time and resource utilization, extract only the bitstream corresponding to that region, and perform decoding and reconstruction on that region.

[0214] Therefore, the present disclosure proposes an apparatus and method for supporting partial access (or partial reconstruction) and reconstruction of mesh data. Specifically, the present disclosure proposes an encoding apparatus and method and a decoding apparatus and method for supporting partial access, which only allow selective decoding and reconstruction of a specific region of the mesh data (or mesh content) in a V-DMC manner on the receiving side.

[0215] To this end, in the present disclosure, partial units for partial region access are defined, and preprocessing (e.g., atlas parameterization, base mesh generation) and compression of the mesh data are performed based on the partial units. Thereafter, the present disclosure describes the encoding, decoding, and signaling methods for the base mesh data (including the motion field) among the elements (i.e., the base mesh (including the motion field), displacement, and attribute map (also referred to as the texture map)) that make up the compressed bitstream of the mesh data.

[0216] In other words, to allow the receiving side (e.g., the decoder) to independently decode and reconstruct only a specific mesh data region specified by the user, the present disclosure describes methods for generating and signaling a mesh set in the partial units, and methods for atlas parameterization, base mesh generation, encoding, decoding, and signaling based on the mesh set in the partial units. In addition, the operations of the transmitting and receiving apparatuses applying these methods are described. In the present disclosure, the partial units are referred to as subgroups or submeshes.

[0217] Therefore, the present disclosure implements the partial access (or partial decoding) and reconstruction functions to enable the reconstruction and utilization of a specific region of the dynamic mesh.

[0218] Therefore, the present disclosure describes the definition, generation method, and signaling method for subgroups in the mesh, the atlas parameterization method based on subgroups, the base mesh generation method based on subgroups, the base mesh encoding and decoding methods based on subgroups, the new bitstream structure and signaling method, the motion field encoding and decoding methods based on subgroups, and the new bitstream structure and / or signaling method for subgroup-based extraction of displacement information.

[0219] First, a method of encoding and signaling grid data by a grid data transmission device / method based on each subgroup will be described.

[0220] The grid data transmission device according to an embodiment may be Figure 1 the transmission device in Figure 2 the transmission device in Figure 6 the transmission device in Figure 7 the transmission device in Fig.13 the transmission device in Fig.30 or one of the transmission devices in

[0221] Fig.30 is a block diagram showing another example of a transmission device according to an embodiment.

[0222] Fig.30 The transmission device in Fig.30 corresponds to Figure 1 the transmission device 100 in Figure 2 , 6 , 7 or Fig.13 the transmission device in

[0223] Specifically, Fig.30 the transmission device in Fig.13 is configured by adding a subgroup generator 1110 and a subgroup-related auxiliary information bitstream to the transmission device in Fig.30 . Therefore, for details not described with respect to Fig.13 , refer to the description of the transmission device in Fig.30 . The execution order of the blocks in the transmission device of Fig.30 can be changed, some blocks can be omitted, and some blocks can be newly added.

[0224] a) Method of defining subgroups of a grid and generating subgroups

[0225] Hereinafter, a method of defining subgroups of a grid and generating subgroups will be described.

[0226] The present disclosure defines a subgroup as a partial unit for accessing a partial area of a grid. In the present disclosure, a subgroup may also be referred to as a sub-grid or a sub-block. In the present disclosure, depending on user settings, a single subgroup may be a unit including the entire grid or a unit including only a partial area of the grid. Here, the partial area of the grid may correspond to a part of an object, a part of a bounding box, or a part within a grid frame.

[0227] In the present disclosure, the entire mesh can be defined as a single frame or one or more objects.

[0228] That is, the mesh data (or mesh frame or mesh content) can include one or more mesh objects. Each mesh object (hereinafter simply referred to as an object) occupies a specific region in 3D space. The bounding box represents the smallest bounding box that encloses (or surrounds) a specific object. The box is typically a rectangular cuboid in 3D space and is adjusted to the smallest size that completely contains the object.

[0229] For example, when a single object constitutes the entire mesh, the bounding box can enclose the entire mesh. In another example, when multiple objects constitute the entire mesh, a bounding box can be generated for each object. That is, the bounding box can enclose the entire mesh, or can enclose each object individually.

[0230] Fig.15 (a) to Fig.15 (c) are diagrams showing examples of subgroup generation according to an embodiment.

[0231] Fig.15 (a) shows an example in which a single bounding box consists of a single subgroup, and Fig.15 (b) shows an example in which a single bounding box consists of two subgroups. Fig.15 (c) shows an example in which one bounding box consists of two subgroups and another bounding box consists of one subgroup.

[0232] According to an embodiment of the present disclosure, there can be one or more methods for generating subgroups as disclosed below. As Fig.16 shown, information for identifying the subgroup generation method (e.g., subgroup_decision_method) can be signaled in signaling information (e.g., mesh_subgroup_info_set()). The signaling information can also include information for identifying the number of subgroups (e.g., num_total_subgroup).

[0233] According to an embodiment, the signaling information (e.g., mesh_subgroup_info_set()) can be referred to as subgroup-related information or subgroup information. Alternatively, the information included in the signaling information (e.g., mesh_subgroup_info_set()) can be referred to as subgroup-related information or subgroup information.

[0234] Fig.16 shows an example of information for identifying the subgroup generation method (e.g., subgroup_decision_method) according to an embodiment.

[0235] In Fig.16 In Fig.16 , in the value of subgroup_decision_method, 0 may indicate user-guided (manual), 1 may indicate uniform partitioning, and 2 may indicate object-based separation. In addition, three or more values of subgroup_decision_method may indicate that one or more of user-guided (manual), uniform partitioning, or object-based separation are combined to generate subgroups, or may be reserved for future use.

[0236] A first embodiment of the subgroup generation method according to the present disclosure corresponds to subgroup_decision_method set to 0. That is, in encoding, the mesh object is divided into K parts (where K is greater than or equal to 1) based on user-specified 3D (x, y, z) coordinates, and the divided parts are respectively designated as subgroup 0 to subgroup K-1. In this case, the region information of each user-specified subgroup (e.g., subgroup_boundingbox_position_x, subgroup_boundingbox_position_y, subgroup_boundingbox_position_z, subgroup_boundingbox_size_x_minus1, subgroup_boundingbox_size_y_minus1, subgroup_boundingbox_size_z_minus1) can be signaled and sent in signaling information (e.g., mesh_subgroup_info_set()).

[0237] A second embodiment of the subgroup generation method according to the present disclosure corresponds to subgroup_decision_method set to 1. That is, in encoding, based on the number of subgroups specified by the user, the bounding box enclosing the mesh object is evenly divided into K rectangular cuboids. The partitioned parts are respectively designated as subgroup 0 to subgroup K-1. In this case, the partitioning information (e.g., num_subgroup_x_minus1, num_subgroup_y_minus1, num_subgroup_z_minus1) for identifying the number of subgroups into which the bounding box is divided can be signaled and sent in signaling information (e.g., mesh_subgroup_info_set()).

[0238] The third embodiment of the subgroup generation method according to the present disclosure corresponds to subgroup_decision_method set to 2. For example, when there are K objects in the content, the corresponding different or independent and unconnected objects are designated as subgroup 0 to subgroup K-1.

[0239] The fourth embodiment of the subgroup generation method according to the present disclosure is a method of generating one or more subgroups by combining one or more of the above methods. In this case, information according to the combination method can be added to the signaling information (e.g., mesh_subgroup_info_set()) and sent.

[0240] Fig.15 (a) and Fig.15 (b) may correspond to the first or second embodiment, and Fig.15 (c) may correspond to the fourth embodiment, which is a combination of the second and third embodiments.

[0241] According to an embodiment, in one embodiment of the present disclosure, the subgroup generator 1110 may generate subgroups. That is, the subgroup generator 1110 divides the original mesh into subgroups based on at least one of the first to fourth embodiments.

[0242] In the present disclosure, subgroups may be configured by a user or a system for each frame. Alternatively, the user (or system) may configure new subgroups every specific number of frames. In this case, the pre-configured subgroups are maintained for each frame until new subgroups are configured. Alternatively, when the objects included in each subgroup move out of the corresponding area due to movement, the user (or system) may reconfigure the subgroups. Then, the configured subgroups are maintained for each frame until new subgroups are configured.

[0243] According to an embodiment of the present disclosure, subgroup-related information (also referred to as subgroup information) may be signaled for the entire sequence, at a constant frame interval, for each frame, or whenever the subgroups are updated.

[0244] Fig.17 Shows an example syntax structure of the signaling information (e.g., mesh_subgroup_info_set()) of the subgroup generation method including Fig.16 .

[0245] In Fig.17 , the num_total_subgroup field indicates the number of subgroups. For example, it may indicate the total number of subgroups generated by user settings. The value of the num_total_subgroup field is based on the frame and may change per sequence / frame / fixed number of frames, or whenever the number changes.

[0246] The subgroup_decision_method field indicates the subgroup generation method used in the present disclosure. For example, it can indicate the subgroup generation method used by the user. In the present invention, as defined in Fig.16 and signals the corresponding value according to the generation method.

[0247] In the present disclosure, when subgroup_decision_method is set to 0, which indicates user guidance, the signaling information (e.g., mesh_subgroup_info_set()) can include a loop that iterates as many times as the value of num_total_subgroup. In this case, in one embodiment, k can be initialized to 0 and incremented by 1 with each iteration of the loop, which can iterate until k reaches the value of the num_total_subgroup field. The loop can include the subgroup_boundingbox_position_x[k] field, subgroup_boundingbox_position_y[k] field, subgroup_boundingbox_position_z[k] field, subgroup_boundingbox_size_x_minus1[k] field, subgroup_boundingbox_size_y_minus1[k] field, and subgroup_boundingbox_size_z_minus1[k] field.

[0248] The subgroup_boundingbox_position_x[k] field indicates the x coordinate of the starting point in the 3D space of the bounding box that encloses the mesh region of the k-th subgroup.

[0249] The subgroup_boundingbox_position_y[k] field indicates the y coordinate of the starting point in the 3D space of the bounding box that encloses the mesh region of the k-th subgroup.

[0250] The subgroup_boundingbox_position_z[k] field indicates the z coordinate of the starting point in the 3D space of the bounding box that encloses the mesh region of the k-th subgroup.

[0251] Adding 1 to subgroup_boundingbox_size_x_minus1[k] represents the size of the bounding box that encloses the mesh region of the k-th subgroup along the x-axis direction.

[0252] subgroup_boundingbox_size_y_minus1[k] plus 1 represents the size of the bounding box enclosing the grid area of ​​the kth subgroup along the y-axis.

[0253] subgroup_boundingbox_size_z_minus1[k] plus 1 represents the size of the bounding box enclosing the grid area of ​​the kth subgroup along the z-axis.

[0254] When subgroup_decision_method is set to 1, that is, indicating uniform partitioning, signaling information (eg, mesh_subgroup_info_set()) may further include a num_subgroup_x_minus1 field, a num_subgroup_y_minus1 field, and a num_subgroup_z_minus1 field.

[0255] num_subgroup_x_minus1 plus 1 indicates the nubmer of parts into which the bounding box enclosing the entire mesh is divided along the x-axis.

[0256] num_subgroup_y_minus1 plus 1 indicates the nubmer of parts into which the bounding box enclosing the entire mesh is divided along the y-axis.

[0257] num_subgroup_z_minus1 plus 1 indicates the number of parts into which the bounding box enclosing the entire mesh is divided along the z-axis.

[0258] b) Subgroup-based atlas parameterization (preprocessing) methods

[0259] Next, a method for performing atlas parameterization on a subgroup basis by a preprocessor is described below.

[0260] According to an embodiment, the preprocessor 1111 performs preprocessing before compressing the input mesh data. That is, the preprocessor 1111 receives the input mesh (or original mesh) of each subgroup generated by the subgroup generator 1110 and performs preprocessing.

[0261] According to an embodiment of the present disclosure, the preprocessing process can be Figure 3 The preprocessing process is the same as the preprocessor in, or may additionally include subgroup processing.

[0262] That is to say, Figure 3 The application Figure 3 Alternatively, atlas parameterization can be performed on a per-subgroup basis. The latter case considers technical extensions for attribute map video compression efficiency and future support for partial access to attribute maps.

[0263] Next, a subgroup-based atlas parameterization method is described.

[0264] First, the input mesh (also referred to as the original mesh) is simplified by the mesh extraction part of the preprocessor 1111. The extracted input mesh is input to the atlas parameterization part of the preprocessor 1111. That is, the mesh extraction part can select vertices to be removed from the original mesh according to user-defined criteria, and then remove the selected vertices and the triangles connected to the selected vertices.

[0265] In the present disclosure, the input mesh includes 3D coordinates of vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connectivity information between vertices constituting the surface.

[0266] According to an embodiment, the atlas parameterization part configures vertex and vertex connectivity information for each subgroup based on the vertices of the simplified mesh obtained as a result of mesh extraction. That is, based on the position values of the vertices of the extracted mesh and the position information for each user-specified subgroup, the vertices corresponding to each subgroup are determined, and the determined vertices together with their associated connectivity information are included in the components of the corresponding subgroup.

[0267] As Fig.18 shown, if vertices belonging to different subgroups are connected to form a single mesh polygon (face), the connected vertices can be redundantly included in the subgroup to which the connected vertices belong even when they belong to different subgroups. That is, when a polygon overlaps at the boundaries of multiple different subgroups, the vertices constituting the polygon can be included as constituent vertices of all subgroups forming the boundary.

[0268] Fig.18 is a diagram showing an example method for determining the subgroup of vertices located at the subgroup boundary according to an embodiment.

[0269] In Fig.18 , each of polygons 1102-1 to 1102-7 is a polygon formed by connecting vertices belonging to subgroup 1 and subgroup 2. That is, at least one of the vertices constituting each polygon is included in subgroup 1, and at least one is included in subgroup 2. For example, among the three vertices constituting polygon 1102-1, one vertex belongs to subgroup 1, and the other two vertices belong to subgroup 2. In contrast, in Fig.18 , all three vertices constituting polygon 1101 belong to subgroup 1, and all three vertices constituting polygon 1103 belong to subgroup 2.

[0270] In one embodiment of the present disclosure, all vertices constituting polygons 1102-1 to 1102-7 may be included in both subgroup 1 and subgroup 2.

[0271] Then, for the vertices classified according to each subgroup as Fig.18 shown, the atlas parameterization unit performs atlas parameterization on each subgroup.

[0272] According to an embodiment, the atlas parameterization unit maps the 3D surface of the extracted mesh onto a texture domain on the basis of each subgroup. Through this operation, mapping information indicating positions in the 2D image is generated, to which each vertex of the extracted mesh of each subgroup can be mapped. In other words, mapping information is generated for each subgroup. The mapping information is represented and stored as texture coordinates. Through this operation, a final base mesh is generated for each subgroup (i.e., for each subgroup). That is, the atlas parameterization unit performs parameterization on each subgroup, generating texture coordinates (UV coordinates) and texture connectivity information for each vertex of the extracted mesh.

[0273] Then, the texture coordinates (also referred to as UV coordinates) obtained as a result of atlas parameterization are set to have different UV coordinate ranges for each subgroup, such that the coordinates do not overlap within one 2D image.

[0274] For example, as Fig.19 shown, the UV coordinates of subgroup 1 may be set to exist within the range of 0 label (u, v) < 0.03, the UV coordinates of subgroup 2 may be set to exist within the range of 0.0 to be set to exist and 0.3 to be set to exist, and the UV coordinates of subgroup 3 may be set to exist within the range of 0.3 to be set to exist and 0.3 to be set to exist. In this way, the method of arranging the UV coordinates of each subgroup within one atlas image such that they do not overlap with each other can be configured by the user in the encoder based on the number of vertices in each subgroup.

[0275] Fig.19 is a diagram showing an example result of subgroup-based atlas parameterization according to an embodiment.

[0276] As Fig.19 shown, the UV coordinates of each subgroup are arranged within one atlas image such that they do not overlap with each other, such that when a new attribute map is generated through attribute transfer (or the texture map video generator 1121), the attribute information is presented on the image according to each subgroup. Thereby, the texture map video encoder 1122 can improve the efficiency of temporal compilation between consecutive attribute map frames during attribute map video compression, and can also help facilitate partial region bitstream access within the video codec.

[0277] c) Subgroup-based base mesh generation method

[0278] Next, a subgroup-based base mesh generation method is described below. In an embodiment of the present disclosure, the atlas parameterization unit of the preprocessor 1111 may generate a base mesh based on subgroups.

[0279] As described above, each subgroup can be defined based on the 3D vertex coordinates that make up the mesh extracted by preprocessing. The vertex coordinates of the extracted mesh within the 3D region where the subgroup is defined, the connectivity information between the vertex coordinates, the texture coordinates (also referred to as UV coordinates) obtained by atlas parameterization of the vertices, the connectivity information between the texture coordinates, and the additional attribute information become the components of the base mesh that defines the subgroup.

[0280] In addition, as described with reference to Fig.18 When vertices belonging to different subgroups are connected to form a mesh polygon, the connected vertices can be redundantly included in the subgroups to which they belong, even if they are from different subgroups. That is, when the polygon overlaps at the boundaries of multiple different subgroups, the vertices that make up the polygon can be included as constituent vertices of all subgroups, thus forming the base mesh.

[0281] d) Subgroup-based fitting subdivision surface method

[0282] Next, a subgroup-based fitting subdivision surface method is described below. In an embodiment of the present disclosure, the fitting subdivision surface part of the preprocessor 1111 performs a subdivision operation on the base mesh based on subgroups. That is, the base mesh generated for each subgroup is subdivided on a subgroup basis. For details of the subdivision operation, refer to the description in Figure 4 . In other words, fitting is performed so that the input mesh and the subdivided mesh become similar to each other on a subgroup basis. In the present disclosure, the mesh obtained through the fitting operation is referred to as a fitting subdivision mesh (or a fitted subdivision mesh). The subdivided vertices, the connectivity information thereon, and the texture information are also included in the corresponding subgroups.

[0283] Based on the fitting subdivision meshes for each subgroup obtained through this process and the previously compressed and decoded base meshes for each subgroup, the displacement component is calculated by the displacement vector calculator 1116. That is, the displacement component (or displacement vector) is calculated based on the fitted subdivision mesh and the base mesh that has been previously compressed and decoded (i.e., reconstructed) for the same subgroup.

[0284] According to an embodiment, the displacement vector calculator 1116 may perform mesh subdivision on the reconstructed base mesh on a subgroup basis and calculate displacement vectors, which are the vertex position differences between the subdivided reconstructed base mesh and the fitted subdivided mesh generated by the fitted subdivision surface portion of the preprocessor 1111 on a subgroup basis. In this operation, as many displacement vectors as the number of vertices in the subdivided mesh may be calculated.

[0285] e) Subgroup-based base mesh encoding method

[0286] Next, a method of encoding a base mesh on a subgroup basis is described below. In one embodiment of the present disclosure, the static mesh encoder 1114 may encode the base mesh on a subgroup basis and generate a base mesh bitstream containing the encoded base mesh.

[0287] For the method of encoding a base mesh on a subgroup basis according to the present disclosure and generating a base mesh bitstream, at least one of the first to third embodiments disclosed below may be applied.

[0288] The first embodiment is a method in which the base meshes of subgroups are compressed individually, and a single base mesh bitstream is generated and transmitted.

[0289] Fig. 20 FIG. is a diagram showing a first embodiment of a base mesh bitstream generation method. According to an embodiment, the base mesh of each subgroup may be compressed by a static mesh encoder (i.e., a static base mesh compilation tool (Draco, TFan, etc.)) 1114. That is, a base mesh bitstream is generated on a subgroup-by-subgroup basis. Then, the base mesh bitstreams of the corresponding subgroups may be concatenated (or multiplexed) into a single base mesh bitstream for transmission here, and the multiplexing may be performed by at least one of the static mesh encoder 1114, the multiplexer 1123, or the transmitter 1130, or by a separate block.

[0290] In other words, all the base mesh bitstreams of the corresponding subgroups may be concatenated into a single bitstream to be transmitted. In this case, subgroup identification information (e.g., subgroup index information, subgroup base mesh bitstream size information) may be signaled before each subgroup's base mesh bitstream and transmitted together with the bitstream so that the decoder on the receiving side can distinguish the subgroups and decode the base mesh bitstream on a subgroup basis.

[0291] Reference Fig. 20, for example, there are K base mesh bitstreams, and subgroup index information (e.g., 0 to K - 1) is signaled before each base mesh bitstream to identify the corresponding subgroup, and subgroup base mesh bitstream size information (e.g., N[0] to N[K - 1]) is signaled to identify the size of the base mesh bitstream of the corresponding subgroup.

[0292] The second embodiment is a method in which the base mesh of each subgroup is compressed separately, and each base mesh bitstream is generated and transmitted.

[0293] Fig.21 is a diagram showing a second embodiment of a method for generating a base mesh bitstream. According to the embodiment, the base mesh of each subgroup can be compressed by a static mesh encoder (i.e., a static base mesh compilation tool (Draco, TFan, etc.)). That is, the base mesh bitstream can be generated on a subgroup-by-subgroup basis. Then, the base mesh bitstream of the subgroup can be transmitted.

[0294] In other words, when transmitting the base mesh bitstream of a subgroup, subgroup identification information (e.g., subgroup index information) can be signaled before each subgroup's base mesh bitstream and transmitted together with the bitstream so that the decoder on the receiving side can distinguish the subgroups and decode the base mesh bitstream on a subgroup-by-subgroup basis.

[0295] Reference Fig.21 , for example, there are K base mesh bitstreams, and subgroup index information (e.g., 0 to K - 1) is signaled before each base mesh bitstream to identify the corresponding subgroup.

[0296] The third embodiment is a method in which the entire base mesh is compressed without distinguishing subgroups, and a base mesh bitstream is generated and transmitted. The compression and bitstream generation for the base mesh according to the third embodiment can be performed by a static mesh encoder (i.e., a static base mesh compilation tool (Draco, TFan, etc.)) 1114. In the third embodiment, the same base mesh compression method described in reference Figure 6 , 7 , 13 or 30 can be applied. In the third embodiment, the entire base mesh can be compressed by the static mesh encoder 1114 without considering the base mesh generated on each subgroup, but additional information of the components for accessing the base mesh of a specified subgroup during reconstruction can be signaled and transmitted in the signaling information.

[0297] That is, for all vertices or polygons of the base mesh decoded after compression, the encoder of the transmitting device can signal the subgroup index to which each vertex or polygon belongs, as Fig. 22 and Fig.23As shown. Then, the decoder of the receiving device can reconstruct the entire base mesh from the base mesh bitstream, and then can access the vertex or polygon index information belonging to the subgroup index specified by the user, access the corresponding vertex or polygon and use it for partial reconstruction.

[0298] Fig. 22 Shows an example syntax structure of subgroup index information for each base mesh vertex / texture coordinate (basemesh_vertex_to_subgroup_info()) according to an embodiment.

[0299] Fig. 22 The subgroup index information for each base mesh vertex / texture coordinate (basemesh_vertex_to_subgroup_info()) in may include a first loop that iterates as many times as the value of the num_basemesh_vertices_in_current_frame field.

[0300] The num_basemesh_vertices_in_current_frame field indicates the number of vertices that make up the base mesh in the current frame.

[0301] The first loop includes the subgroup_count_per_basesh_vertex_minus1[k] field. The value of subgroup_count_per_basemesh_vertex_minus1[k] plus 1 can indicate the number of subgroups to which the k-th vertex that makes up the base mesh in the current frame belongs.

[0302] According to an embodiment, the subgroup index information for each base mesh vertex / texture coordinate may include a second loop that iterates as many times as the value of subgroup_count_per_basesh_vertex_minus1[k]. The second loop may include the subgroup_id_per_basesh_vertex[k][n] field. The subgroup_id_per_basemesh_vertex[k][n] field indicates the index of the n-th subgroup to which the k-th vertex that makes up the base mesh in the current frame belongs.

[0303] According to an embodiment, the subgroup index information for each base mesh vertex / texture coordinate may further include a third loop that iterates as many times as the value of the num_texture_coords_in_current_frame field.

[0304] The num_texture_coords_in_current_frame field indicates the number of texture coordinates that make up the base mesh in the current frame.

[0305] The third loop includes the subgroup_count_per_basesh_texture_coord_minus1[k] field. The value of subgroup_count_per_basemesh_texture_coord_minus1[k] plus 1 can indicate the number of subgroups to which the k-th texture coordinate of the base mesh in the current frame belongs.

[0306] According to an embodiment, the subgroup index information for each base mesh vertex / texture coordinate may include a fourth loop that iterates as many times as the value of the subgroup_count_per_basesh_texture_coord_minus1[k] field. The fourth loop may include the subgroup_id_per_basemesh_texture_coord[k][n] field. The subgroup_id_per_basemesh_texture_coord[k][n] field indicates the index of the n-th subgroup to which the k-th texture coordinate of the base mesh in the current frame belongs.

[0307] Fig.23 An example syntax structure of the subgroup index information for each base mesh polygon / texture polygon (basemesh_polygon_to_subgroup_info()) according to an embodiment is shown.

[0308] Fig.23 The subgroup index information for each base mesh polygon / texture polygon (basemesh_polygon_to_subgroup_info()) in may include a first loop that iterates as many times as the value of the num_basemesh_polygons_in_current_frame field.

[0309] The num_basemesh_polygons_in_current_frame field indicates the number of polygons that make up the base mesh in the current frame.

[0310] The first loop includes the subgroup_count_per_basesh_polygon_minus1[k] field. The value of subgroup_count_per_basemesh_polygon_minus1[k] plus 1 can indicate the number of subgroups to which the k-th polygon of the base mesh in the current frame belongs.

[0311] According to an embodiment, the subgroup index information for each base mesh polygon / texture polygon may include a second loop that iterates the number of times of the value of the subgroup_count_per_basemesh_polygon_minus1[k] field. The second loop may include the subgroup_id_per_basesh_polygon[k][n] field. The subgroup_id_per_basemesh_polygon[k][n] field indicates the index of the n-th subgroup to which the k-th polygon of the base mesh in the current frame belongs.

[0312] According to an embodiment, the subgroup index information for each base mesh polygon / texture polygon may further include a third loop that iterates the same number of times as the value of the num_basesh_texture_polygons_in_current_frame field.

[0313] The num_basemesh_texture_polygons_in_current_frame field indicates the number of texture polygons of the base mesh in the current frame.

[0314] The third loop may include the subgroup_count_per_basesh_texture_polygon_minus1[k] field. The value of subgroup_count_per_basemesh_texture_polygon_minus1[k] plus 1 can indicate the number of subgroups to which the k-th texture polygon of the base mesh in the current frame belongs.

[0315] According to an embodiment, the subgroup index information for each base mesh polygon / texture polygon may include a fourth loop that iterates as many times as the value of the subgroup_count_per_basesh_texture_polygon_minus1[k] field. The fourth loop may include the subgroup_id_per_basesh_texture_polygon[k][n] field. The subgroup_id_per_basemesh_texture_polygon[k][n] field indicates the index of the n-th subgroup to which the k-th texture polygon of the base mesh in the current frame belongs.

[0316] That is, the static mesh encoder 1114 of the present disclosure may apply at least one of the first to third embodiments to compress the base mesh and generate a base mesh bitstream for each corresponding subgroup or a single base mesh bitstream.

[0317] According to an embodiment of the present disclosure, inter-frame encoding or intra-frame encoding may be performed. In one embodiment, intra-frame encoding may be performed by the static mesh encoder 1114, and inter-frame encoding may be performed by the motion vector encoder 1113. In intra-frame encoding, the data to be compressed may include the base mesh, displacement, and attribute map. That is, vertex position information, mapping information (texture coordinates), vertex connectivity information, etc. related to the base mesh may be the compression targets. In inter-frame encoding, the data to be compressed may include the displacement between the reference base mesh and the current base mesh, the attribute map, and the motion field (motion vector). Here, the motion field refers to the vertex difference between the reference base mesh and the current base mesh and is used interchangeably with the motion vector. In the present disclosure, the value of the motion field may be encoded. Alternatively, a predicted motion field may be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between the value of the predicted motion field and the value of the motion field of the current vertex, may be encoded.

[0318] In the present disclosure, intra-frame encoding or inter-frame encoding may be performed on a per-subgroup basis.

[0319] In the present disclosure, information for identifying the coding method applied to each subgroup (e.g., subgroup_basesh_coding_type) may be signaled and transmitted in signaling information (e.g., subgroup_basesh_coding_info()). In the present disclosure, information for identifying the coding method applied to each subgroup (e.g., subgroup_basesh_coding_type) may be referred to as subgroup base grid compilation type information. In addition, signaling information (e.g., subgroup_basesh_coding_info()) may be referred to as subgroup base grid compilation information.

[0320] Fig.24 An example of type information for identifying a subgroup coding method (e.g., subgroup_basesh_coding_type) according to an embodiment is shown.

[0321] In Fig.24 , subgroup_basesh_coding_type set to 0 indicates an intra mode, while subgroup_basesh_coding_type set to 1 may indicate an inter mode. That is, in the present disclosure, subgroup_basesh_coding_type can be used to indicate the base grid compilation type for each subgroup.

[0322] Fig.25 An example syntax structure of subgroup base grid compilation information (subgroup_basesh_coding_info()) according to an embodiment is shown.

[0323] That is, Fig.25 the subgroup base grid compilation information (subgroup_basesh_coding_info()) in

[0324] may include a loop iterated as many times as the value of the num_total_subgroup field. The loop may include the subgroup_basesh_coding_type[k] field. Fig.17 The num_total_subgroup field indicates the number of subgroups. According to an embodiment, as an example, the num_total_subgroup field may be included in

[0325] The subgroup_basesh_coding_type[k] field indicates the coding type of the k-th subgroup. For example, subgroup_basesh_coding_type[k] set to 0 may indicate that the intra mode is applied to the k-th subgroup. Subgroup_basesh_coding_type[k] set to 1 indicates that the inter mode is applied to the k-th subgroup.

[0326] Thus, in the present disclosure, the base mesh coding type of each subgroup can be defined as shown in Fig.24 and the base mesh coding type of each subgroup can be signaled as shown in Fig.25 When at least one of the above first to third embodiments is applied to generate the base mesh bitstream, subgroup_basesh_coding_type is set to the intra mode and signaled in the subgroup base mesh coding information (subgroup_basesh_coding_info()).

[0327] Fig.17 The signaling information (e.g., mesh_subgroup_info_set()) of Fig.25 and / or the signaling information (subgroup_basemesh_coding_info()) of

[0328] f) Subgroup-based motion field generation and coding method

[0329] Next, a method for generating and coding a motion field (also referred to as a motion vector) based on subgroups is described below.

[0330] As described above, for a specific subgroup, the inter mode can be applied for compression. In other words, when inter-frame coding (also referred to as inter-coding) is allowed, inter-frame coding can be performed on each subgroup. In other words, under the same conditions as V-DMC, when the number of vertices, connectivity information, the number of texture coordinates, and the connectivity information related to the base mesh are the same for the same subgroup from consecutive frames, inter-frame coding can be performed on the base mesh component of the subgroup, and a base mesh motion field can be obtained based on the subgroup.

[0331] In the present disclosure, the motion field components that can be generated for each subgroup as described above can be encoded. For example, the values of the motion field can be encoded. Alternatively, a predicted motion field can be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and a residual motion field that is the difference between the value of the predicted motion field and the value of the motion field of the current vertex can be encoded. For example, the motion vector encoder 1113 can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as inputs on a per-subgroup basis, calculate the motion field (or motion vector) between the two meshes, and encode the value. Alternatively, the motion vector encoder 1113 can perform prediction based on connectivity information using a previously encoded / decoded motion field as a predictor on a per-subgroup basis, and encode the residual motion field obtained by subtracting the predicted motion field from the current motion field.

[0332] Through the inter-frame encoding operation, a motion field bitstream (or a motion vector bitstream) is generated for each subgroup.

[0333] According to an embodiment, once the motion field bitstream is generated by applying the above inter-frame encoding method, subgroup_basesh_coding_type is set to the inter-frame mode and signaled in the subgroup base mesh compilation information (subgroup_basesh_coding_info()). In other words, for the subgroup to which the inter-frame encoding is applied, Fig.24 the subgroup_basesh_coding_type in Fig.25 can be set to the inter-frame mode and signaled in the subgroup base mesh compilation information (subgroup_basesh_coding_info()) in

[0334] At this time, the motion field bitstreams of the corresponding subgroups can be concatenated (or multiplexed) into a single motion field bitstream to be transmitted. Here, the multiplexing can be performed by at least one of the motion vector encoder 1113, the multiplexer 1123, or the transmitter 1130 or by a separate block.

[0335] In other words, all the motion field bitstreams of the corresponding subgroups can be concatenated into a single bitstream to be transmitted. In this case, subgroup identification information (e.g., at least one of subgroup index information, bitstream size information, or compilation type information) can be signaled before the motion field bitstream of each subgroup and transmitted together with the bitstream, so that the decoder on the receiving side can distinguish the subgroups and decode the motion field bitstream on a per-subgroup basis. Here, the compilation type information is the subgroup base mesh compilation type information of each subgroup and corresponds to Fig.24the subgroup_basemesh_coding_type in. Since the coding type information has been signaled via Fig.24 and Fig.25 it can be omitted before the motion field bitstream of each subgroup.

[0336] Fig.26 is a diagram showing an example of a subgroup-specific motion field bitstream according to an embodiment. In the example of Fig.26 before the motion field bitstream of each subgroup, subgroup index information (e.g., 0 to K-1), motion field bitstream size information related to the corresponding subgroup (e.g., N[0] to N[K-1]), and coding type information are included.

[0337] Fig. 27 is a diagram showing another example of a subgroup-specific motion field bitstream according to an embodiment. In the example of Fig. 27 before the motion field bitstream of each subgroup, subgroup index information (e.g., 0 to K-1) related to the corresponding subgroup and motion field bitstream size information (e.g., N[0] to N[K-1]) are included, but the coding type information is omitted.

[0338] As described above, inter-frame coding or intra-frame coding can be performed based on subgroups.

[0339] For example, intra-frame coding can be performed on subgroup 1 to generate a base mesh bitstream, while inter-frame coding can be performed on subgroup 2 to generate a motion field bitstream. In this case, the subgroup_basesh_coding_type of subgroup 1 is set to the intra-frame mode, and the subgroup_basesh_coding_type of subgroup 2 is set to the inter-frame mode, which is signaled in the subgroup base mesh coding information (subgroup_basesh_coding_info()).

[0340] Then, the bitstreams of the corresponding subgroups can be concatenated (or multiplexed) into a single bitstream, as shown in Fig.26 or 27. In this case, the bitstream compressed using intra-frame coding and the bitstream compressed using inter-frame coding can coexist. As shown in Fig.26 and Fig. 27 depending on whether the bitstream is a base mesh bitstream or a motion field bitstream, the bitstream size information can indicate the size of the base mesh bitstream or the motion field bitstream. In addition, Fig.26 the coding type information in can indicate the intra-frame mode or the inter-frame mode according to whether the bitstream is a base mesh bitstream or a motion field bitstream.

[0341] Therefore, when there are subgroups in the current frame for which inter-frame coding is not applicable, as shown in Fig. 20 or Fig.21 The intra coding shown can be used to generate a base mesh bitstream and can be combined with other subgroups of motion field bitstreams generated by inter coding into a single bitstream to be transmitted, as Fig.26 or Fig. 27 shown in

[0342] Alternatively, the motion field bitstream and the base mesh bitstream can be transmitted separately as independent bitstreams. That is, the motion field bitstream and the base mesh bitstream can be transmitted separately without being combined into a single bitstream. In this case, since both the motion field bitstream and the base mesh bitstream can be distinguished by subgroup indices, partial decoding can be performed by the decoder of the receiving device.

[0343] In the present disclosure, the switching section 1112, the motion vector encoder 1113, and the static mesh encoder 1114 can be grouped and referred to as a base mesh encoder.

[0344] g) Displacement and attribute mapping coding method

[0345] Next, a method for coding the displacement and attribute maps is described below.

[0346] In the present invention, the same compression method as described with reference to Figure 6 , 7 or 13 can be used to compress the components of the displacement component (also referred to as the displacement vector) and the attribute map (also referred to as the texture map) into a video bitstream.

[0347] According to an embodiment, although the displacement component is calculated for each subgroup, the processes of quantization, transformation, packing into a 2D image, video generation, and compressing data using a video codec can be performed in the same manner as the compression process described with reference to Figure 6 , 7 or Fig.13 described.

[0348] According to an embodiment, for the attribute map, attribute transfer and video compression can be performed on the input original attribute map image in the same manner as described with reference to Figure 6 , 7 or Fig.13 described.

[0349] Specifically, the base mesh reconstructor 1115 can receive the encoded base mesh from the static mesh encoder 1114 or the encoded motion field from the motion vector encoder 1113 and generate a reconstructed base mesh. For example, the base mesh reconstructor 1115 can reconstruct the base mesh by performing static mesh decoding on the base mesh encoded by the static mesh encoder 1114. In this case, quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding. In another example, the base mesh reconstructor 1115 can reconstruct the base mesh based on the reconstructed quantized reference base mesh and the motion field encoded by the motion vector encoder 1113. The reconstructed base mesh is output to the displacement vector calculator 1116 and the mesh reconstructor 1120.

[0350] In the present disclosure, the displacement vector calculator 1116 can perform mesh subdivision on the reconstructed base mesh. The displacement vector calculator 1116 can calculate displacement vectors (also referred to as displacement components), which are the vertex position differences between the subdivided reconstructed base mesh and the fitted subdivision mesh generated by the preprocessor 1111. At this time, as many displacement vectors as the number of vertices in the subdivision mesh can be calculated. The displacement vector calculator 1116 can transform the displacement vectors calculated in the 3D Cartesian coordinate system to the local coordinate system based on the normal vector of each vertex.

[0351] In the present disclosure, the displacement vector video generator 1117 can include a linear lifting section, a quantizer, and an image packer. That is, the linear lifting section in the displacement vector video generator 1117 can transform the displacement vectors for efficient encoding. Depending on the embodiment, the transformation can be a lifting transformation, a wavelet transformation, etc. The quantizer can quantize the transformed displacement vector values (i.e., transformation coefficients). Different quantization parameters can be applied to the corresponding axes of the transformation coefficients. The parameters can be derived through an agreement between the encoder and the decoder. Then, the transformed and quantized displacement vector information can be packed into a 2D image by the image packer. The displacement vector video generator 1117 can generate a displacement vector video by grouping the packed 2D images for each frame. A displacement vector video can be generated for each group of frames (GoF) of the input mesh.

[0352] In the present disclosure, the displacement vector video encoder 1118 can encode the generated displacement vector video using a video compression codec.

[0353] The displacement vector reconstructor 1119 may include a video decoder, an image unpacker, an inverse quantizer, and an inverse linear lifting section. That is, the displacement vector reconstructor 1119 decodes the encoded displacement vector using the video decoder, unpacks the image using the image unpacker, performs inverse quantization using the inverse quantizer, and then performs an inverse transform using the inverse linear lifting section to reconstruct the displacement vector. The reconstructed displacement vector is output to the mesh reconstructor 1120. The mesh reconstructor 1120 reconstructs a deformed mesh based on the base mesh reconstructed by the base mesh reconstructor 1115 and the displacement vector reconstructed by the displacement vector reconstructor 1119. The reconstructed mesh (also referred to as the reconstructed deformed mesh) has reconstructed vertices, vertex connectivity information, texture coordinates, texture coordinate connectivity information, etc.

[0354] In the present disclosure, the texture map video generator 1121 may regenerate a texture map based on the texture map (or attribute map) of the original mesh and the reconstructed deformed mesh output from the mesh reconstructor 1120. According to an embodiment, the texture mapping video generator 1121 may assign vertex-specific color information in the texture mapping of the original mesh to the texture coordinates of the reconstructed deformed mesh.

[0355] The texture map video generated by the texture map video generator 1121 may be encoded using the video compression codec of the texture map video encoder 1122. The encoded texture map video bitstream is sent to the multiplexer 1123.

[0356] In the present disclosure, for the displacement component, it is necessary to extract only the displacement information regarding the vertices belonging to the target subgroup from the reconstructed displacement data for all vertices in the current frame. To this end, according to the present disclosure, subgroup index information related to all vertices or each polygon may be signaled, thereby allowing the user to select only the displacement information related to the desired subgroup region.

[0357] According to an embodiment, for all vertices or polygons of the mesh, the encoder of the transmitting device may signal the subgroup index to which each vertex or polygon belongs, as Fig.28 and Fig.29 shown, and send it to the receiving device.

[0358] According to an embodiment, the decoder of the receiving device may retrieve only the necessary information from the reconstructed displacement data based on the indices of the vertices or polygons belonging to the target subgroup. In the case where the base mesh encoding is performed according to the third embodiment of the subgroup-based base mesh encoding method in e) above, the necessary information may be inferred from the signaling information in Fig. 22 and Fig.23 sent in this case. Therefore, Fig.28 and Fig.29The signaling information in Fig.28 and Fig.29 does not need to be sent separately. In other words, when the base grid coding is performed according to the third embodiment of the above-mentioned subgroup-based base grid coding method in e), the transmission of the signaling information in

[0359] Fig.28 can be skipped.

[0360] Fig.28 The subgroup index information for each mesh vertex / texture coordinate (vertex_to_subgroup_info()) in

[0361] can include a first loop that iterates as many times as the value of the num_vertices_in_current_frame field. The num_vertices_in_current_frame field indicates the number of vertices in the current frame.

[0362] The first loop can include the subgroup_count_per_vertex_minus1[k] field. The value of subgroup_count_per_vertex_minus1[k] plus 1 can indicate the number of subgroups to which the k-th vertex in the current frame belongs.

[0363] According to an embodiment, the subgroup index information for each mesh vertex / texture coordinate can include a second loop that iterates as many times as the value of the subgroup_count_per_vertex_minus1[k] field. The second loop can include the subgroup_id_per_vertex[k][n] field. The subgroup_id_per_vertex[k][n] field indicates the index of the n-th subgroup to which the k-th vertex in the current frame belongs.

[0364] According to an embodiment, the subgroup index information for each mesh vertex / texture coordinate can further include a third loop that iterates as many times as the value of the num_texture_coords_in_current_frame field.

[0365] The num_texture_coords_in_current_frame field indicates the number of texture coordinates in the current frame.

[0366] The third loop includes the subgroup_count_per_texture_coord_minus1[k] field. Incrementing the value of subgroup_count_per_texture_coord_minus1[k] by 1 can indicate the number of subgroups to which the k-th texture coordinate in the current frame belongs.

[0367] According to an embodiment, the subgroup index information for each mesh vertex / texture coordinate may include a fourth loop that iterates as many times as the value of the subgroup_count_per_texture_coord_minus1[k] field. The fourth loop may include the subgroup_id_per_texture_coord[k][n] field. The subgroup_id_per_texture_coord[k][n] field indicates the index of the n-th subgroup to which the k-th texture coordinate in the current frame belongs.

[0368] Fig.29 An example syntax structure (polygon_to_subgroup_info()) of the subgroup index information for each mesh polygon / texture polygon according to an embodiment is shown.

[0369] Fig.29 The subgroup index information (polygon_to_subgroup_info()) for each mesh polygon / texture polygon in may include a first loop that iterates as many times as the value of the num_polygons_in_current_frame field.

[0370] The num_polygons_in_current_frame field indicates the number of polygons in the current frame.

[0371] The first loop may include the subgroup_count_per_polygon_minus1[k] field. The value of subgroup_count_per_polygon_minus1[k] incremented by 1 can indicate the number of subgroups to which the k-th polygon in the current frame belongs.

[0372] According to an embodiment, the subgroup index information for each mesh polygon / texture polygon may include a second loop that iterates as many times as the value of the subgroup_count_per_polygon_minus1[k] field. The second loop may include the subgroup_id_per_polygon[k][n] field. The subgroup_id_per_polygon[k][n] field indicates the index of the nth subgroup to which the kth polygon in the current frame belongs.

[0373] According to an embodiment, the subgroup index information for each mesh polygon / texture polygon may further include a third loop that iterates as many times as the value of the num_texture_polygons_in_current_frame field.

[0374] The num_texture_polygons_in_current_frame field indicates the number of texture polygons in the current frame.

[0375] The third loop may include the subgroup_count_per_texture_polygon_minus1[k] field. The value of subgroup_count_per_texture_polygon_minus1[k] plus 1 may indicate the number of subgroups to which the kth texture polygon in the current frame belongs.

[0376] According to an embodiment, the subgroup index information for each mesh polygon / texture polygon may include a fourth loop that iterates as many times as the value of the subgroup_count_per_texture_polygon_minus1[k] field. The fourth loop may include the subgroup_id_per_texture_polygon[k][n] field. The subgroup_id_per_texture_polygon[k][n] field indicates the index of the nth subgroup to which the kth texture polygon in the current frame belongs.

[0377] As described above, Fig.30 the sending device in is a device configured to perform the definition of subgroups in the above-mentioned a) mesh and subgroup generation method, b) subgroup-based atlas parameterization (preprocessing) method, c) subgroup-based base mesh generation method, d) subgroup-based fitting subdivision surface method, e) subgroup-based base mesh encoding method, f) subgroup-based motion field generation and encoding method, and g) displacement and attribute mapping encoding method.

[0378] In one embodiment, in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 shown and referred to in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 The subgroup-related signaling information described can be encoded as a subgroup-related auxiliary information bitstream and then transmitted to the receiving device via the multiplexer 1123 and the transmitter 1130. Additionally, in one embodiment, in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 shown and referred to in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 Some or all of the signaling information described can be sent as header information or sent together with the base grid bitstream of each subgroup. Alternatively, it can be sent as header information before all the base grid bitstreams are sent.

[0379] That is, when the subgroup generator 1110 generates subgroups according to the method specified by the user for the input original grid, subsequent encoding operations can be performed based on this. Subgroup-based atlas parameterization can be applied in the preprocessor 1111. Additionally, the base grid generated by the subgroup-based base grid generation method can be compressed by the static grid encoder 1114 (i.e., subgroup-based base grid encoding operation) using one of the encoding methods proposed in the first to third embodiments. Moreover, when encoding the subgroup base grid in an inter-frame mode is allowed, the subgroup-based motion field generation and encoding method can be performed by the motion vector encoder 1113. The process of encoding the displacement vector and texture map has been described in detail above, so its detailed description is omitted. Additionally, according to the present disclosure, the auxiliary information as shown in Fig.28 and Fig.29 can be signaled on the receiving side for subgroup-based grid reconstruction.

[0380] So far, the encoding of the transmitting device or the encoder of the transmitting device has been described.

[0381] Hereinafter, a method for decoding and rendering grid data on a subgroup basis by a grid data receiving device / method based on signaling information will be given.

[0382] According to an embodiment, the grid data receiving device can be the receiving device 110 in Figure 1 , the receiving device in Fig.11 , the receiving device in Fig.12 , the receiving device in Fig.14 , or one of the receiving devices in Fig.31 .

[0383] A method for a mesh data receiving device / method to decode and render mesh data on a subgroup basis based on signaling information is described below.

[0384] According to an embodiment, the mesh data receiving device may be Figure 1 the receiving device 110 in Fig.11 the receiving device in Fig.12 the receiving device in Fig.14 the receiving device in Fig.31 or the receiving device in

[0385] Fig.31 is a block diagram showing another example of a receiving device according to an embodiment.

[0386] Fig.31 The receiving device in Fig.31 corresponds to Figure 1 the receiving device 110 in Fig.11 , 12 or Fig.14 the receiving device in

[0387] Specifically, Fig.31 the receiving device in Fig.14 is configured by adding a user-specified area receiver 1311, a target subgroup determiner 1312, a target subgroup displacement vector extractor 1315, and subgroup-related auxiliary information bitstream to Fig.31 the receiving device in Fig.11 and modifying the target subgroup-based mesh reconstructor 1313 and the target subgroup-based mesh reconstructor 1314. Therefore, for parts not described with respect to Fig.31 reference is made to the description of the receiving device in Fig.31 The execution order of the blocks in the receiving device of

[0388] In one embodiment, with reference to Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 the subgroup-related signaling information shown and described may be encoded into a subgroup-related auxiliary information bitstream and then received by the receiver 1210 of the receiving device. The subgroup-related signaling information included in the subgroup-related auxiliary information bitstream received by the receiver 1210 is demultiplexed by the demultiplexer 1211 and provided to the target subgroup determiner 1312 and the target subgroup displacement vector extractor 1315. In addition, in one embodiment, in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 shown and referenced in Fig.16 , 17 , 22, 23, 24, 25, 28, and Fig.29 Some or all of the subgroup-related signaling information described can be received as header information or received together with the base grid bitstream for each subgroup. Alternatively, it can be received as header information before all of the base grid bitstreams are received.

[0389] h) Target subgroup determination method

[0390] Next, a method for determining a target subgroup is described below.

[0391] The user-specified area receiver 1311 of the present disclosure can receive a specified desired partial area within a desired grid through a device, application, or software that uses and displays grid content. The specified area is referred to herein as the target area.

[0392] Based on the 3D position of the specified target area, the target subgroup determiner 1312 of the present disclosure determines the corresponding target subgroup. The target subgroup or target subgroup information determined by the target subgroup determiner 1312 is provided to the motion vector decoder 1213, the static grid decoder 1214, and the target subgroup displacement vector extractor 1315.

[0393] According to an embodiment, the target subgroup determiner 1312 can infer the subgroup determination method, the number of subgroups, and the area information of each subgroup based on subgroup-related information (e.g., mesh_sub-group_info_set()) sent from the encoder (such as Fig.17 shown). Then, based on the inferred area information, the index of the target subgroup including the target area can be determined. It can include one or more subgroup indices. For details of the information included in the Fig.17 subgroup-related information (also referred to as subgroup information), refer to the description in Fig.17 . Its further description is omitted to avoid redundancy. The target subgroup index determined by the target subgroup determiner 1312 is provided to the motion vector decoder 1213, the static grid decoder 1214, and the target subgroup displacement vector extractor 1315.

[0394] According to an embodiment, the target subgroup can be updated by decoding the input subgroup-related information for each sequence, at fixed frame intervals, or for each frame. Then, the grid area of the determined target subgroup is reconstructed through subsequent decoding operations.

[0395] i) Decoding method for the base grid of the target subgroup

[0396] Next, a method for decoding the base mesh of the target subgroup by the static mesh decoder 1214 will be described.

[0397] According to an embodiment, the decoder can Fig.25 identify the coding type of the target subgroup through the subgroup_basesh_coding_type field included in the subgroup_basesh_coding_info() of the subgroup base mesh compilation information, and can control the switching section 1212 to perform base mesh decoding or motion field decoding according to the coding type, as described below. For example, when the value of the subgroup_basesh_coding_type field indicates the intra mode, the static mesh decoder 1214 can perform base mesh decoding on the target subgroup to reconstruct the base mesh. When the value indicates the inter mode, the motion vector decoder 1213 can perform motion field decoding on the target subgroup to reconstruct the motion field.

[0398] j) Method used when the base mesh coding type (subgroup_basesh_coding_type) of the target subgroup is the intra mode

[0399] Next, when the base mesh coding type (subgroup_basemesh_coding_type) of the target subgroup is the intra mode, a method for the static mesh decoder 1214 to perform base mesh decoding by applying at least one of the first to third embodiments disclosed below will be given.

[0400] The decoding methods of the following first to third embodiments correspond to the reverse process of the e) sub-block-based base mesh coding method described for the encoder of the transmission device. In the e) subgroup-based base mesh coding method, the first embodiment (see Fig. 20 ), the second embodiment (see Fig.21 ), or at least one of the third embodiments is applied to code the base mesh for each subgroup.

[0401] The first embodiment performed by the decoder of the present disclosure is the following case: as Fig. 20 shown, the base mesh bitstreams of each subgroup are concatenated and sent as a single bitstream from the transmitting device and input to the decoder of the receiving device.

[0402] That is to say, the first embodiment is a method for the decoder to reconstruct the received bitstream when the encoder of the transmitting device compresses and transmits the base grid of each subgroup based on the first embodiment of the subgroup-based base grid coding method in e). In this case, subgroup identification information (e.g., subgroup index information and subgroup base grid bitstream size information) signaled before each subgroup's bitstream in the input bitstream can be used to extract the base grid bitstream region corresponding to the target subgroup. Then, the extracted target base grid bitstream can be decoded and reconstructed by the static grid decoder 1214 using the static base grid compilation tools (Draco, TFan, etc.) used during encoding.

[0403] A second embodiment performed by the decoder of the present disclosure is the following case: as Fig.21 shown, the base grid bitstream of each subgroup is sent separately and input separately into the decoder of the receiving device.

[0404] That is to say, the second embodiment is a method for the decoder to reconstruct the received bitstream when the encoder of the transmitting device compresses and transmits the base grid of each subgroup based on the second embodiment of the subgroup-based base grid coding method in e). In this case, subgroup identification information (e.g., subgroup index information) signaled before the input single bitstream can be used to select the base grid bitstream corresponding to the target subgroup. The selected target base grid bitstream can be decoded by the static grid decoder 1214 using the static base grid compilation tools (Draco, TFan, etc.) used by the encoder.

[0405] A third embodiment performed by the decoder of the present disclosure is the case where the encoder of the transmitting device transmits the compressed bitstream of the entire base grid without subgroup differentiation, and this bitstream is input into the decoder of the receiving device.

[0406] That is to say, the third embodiment describes a method for the decoder to reconstruct the received bitstream when the encoder of the transmitting device has compressed and transmitted the base grid of each subgroup based on the third embodiment of the subgroup-based base grid coding method in e). In this case, the static grid decoder 1214 can decode the bitstream by using the static base grid compilation tools to reconstruct the entire base grid in the same manner as Figure 1 、 11 、12、14 or Fig.31 described. Then, the subgroup base grid reconstructor 1313 can extract the components of the base grid corresponding to the target subgroup index and based on those for Fig. 22Subgroup index information for each base mesh vertex / texture coordinate (basemesh_vertex_to_subgroup_info()) and / or Fig.23 Subgroup index information for each base mesh polygon / texture polygon (basemesh_polygon_to_subgroup_info()) in Fig.23 to reconstruct the base mesh of the target subgroup by signaling and receiving subgroup index information for each vertex or polygon.

[0407] k) Method used when the base mesh compilation type (subgroup_basesh_coding_type) of the target subgroup is the inter-frame mode

[0408] Next, when the base mesh compilation type (subgroup_basemesh_coding_type) of the target subgroup is the inter-frame mode, a description of the method for performing motion field decoding by the motion field decoder 1213 will be given.

[0409] When the base mesh compilation type (subgroup_basesh_coding_type) is the inter-frame mode, send a bitstream in the form of Fig.26 or Fig. 27 and receive it by the receiving device. That is, instead of the base mesh, a motion field bitstream is sent from the transmitter because the inter-frame mode is applied to the base mesh of each subgroup. Then, the decoder of the receiving device extracts subgroup identification information (e.g., target subgroup index and compilation type (if signaled), and bitstream size information) signaled before the bitstream of each subgroup, and based on the extracted subgroup identification information, extracts and decodes the motion field bitstream of the corresponding size. In one embodiment, the motion vector decoder 1213 can decode the motion field bitstream. Then, the subgroup base mesh reconstructor 1313 can restore the base mesh of the target subgroup in the current frame by adding the reconstructed motion field information related to the target subgroup to the base mesh of the same subgroup index reconstructed from the previous frame.

[0410] l) Method for displacement decoding and extracting displacement information related to the target subgroup

[0411] Next, a description of the method for decoding displacement information by the decoder and the method for extracting displacement information related to the target subgroup will be given.

[0412] In the present invention, the displacement components are compressed into a video bitstream using the same compression method as described in reference Figure 6 , 7 , 13 or 30. The displacement information related to the entire mesh can be sent without subgroup discrimination. Therefore, as described in reference Fig.11, 12 As described in 14 or 31, all displacement components can be restored using a video codec, and information corresponding only to the target subgroup can be selectively extracted based on the restored components.

[0413] According to an embodiment, the displacement vector video decoder 1217 can use a video codec to decode a displacement vector bitstream into a video bitstream.

[0414] According to an embodiment, the displacement vector reconstructor 1218 can extract displacement vector transform coefficients from the decoded displacement vector video and restore the displacement vector by applying inverse quantization and inverse transform to the extracted coefficients. To this end, the displacement vector reconstructor 1218 can include an image unpacker, an inverse quantizer, and an inverse linear lifting section. If the reconstructed displacement vector is in a local coordinate system, it can be inversely transformed to a Cartesian coordinate system.

[0415] To allow selection of only displacement information related to a subgroup region desired by the user from the reconstruction result, the decoder of the present disclosure uses subgroup index information (vertex_to_subgroup_info()) for each vertex / texture coordinate in Fig.28 and / or subgroup index information (polygon_to_subgroup_info()) for each mesh polygon / texture polygon in Fig.29 According to an embodiment, vertices corresponding to the target subgroup can be extracted based on the subgroup index information for each vertex or polygon, and displacement information corresponding to the target subgroup can be extracted based on the indices of the extracted vertices.

[0416] According to an embodiment, instead of Fig.28 subgroup index information (vertex_to_subgroup_info()) for each vertex / texture coordinate in Fig.29 and / or subgroup index information (polygon_to_subgroup_info()) for each mesh polygon / texture polygon in Fig. 22 subgroup index information (basesh_vertex_to_subgroup_info()) for each base mesh vertex / texture coordinate in Fig.23 and subgroup index information (basesh_polygon_to_subgroup_info()) for each base mesh polygon / texture polygon in Fig. 22 can be sent from the transmitting device. In this case, the subgroup index of the subdivided vertices generated by the subdivision of the base mesh can be obtained from the subgroup index information (basesh_vertex_to_subgroup_info()) for each base mesh vertex / texture coordinate in Fig.23 Inference of subgroup index information for each base mesh polygon / texture polygon (basesh_polygon_to_subgroup_info()). As a result, the subgroup indices of all vertices or polygons can be identified. Based on the identified indices, the target subgroup displacement vector extractor 1315 can extract the vertices corresponding to the target subgroup and extract the displacement vector (or displacement information) corresponding to the target subgroup based on the indices of the extracted vertices.

[0417] According to an embodiment, the target subgroup mesh reconstructor 1314 can finally recover the mesh information related to the target subgroup based on the displacement information extracted through the above process. In other words, the target subgroup mesh reconstructor 1314 can generate the mesh of the finally recovered target subgroup (also referred to as the reconstructed deformed mesh) by combining the base mesh of the target subgroup reconstructed by the target subgroup base mesh reconstructor 1313 with the displacement vector of the target subgroup extracted by the target subgroup displacement vector extractor 1315.

[0418] m) Method for decoding the attribute map and extracting attribute information related to the target subgroup

[0419] Next, a description of the method for decoding the attribute map and extracting attribute information related to the target subgroup will be given. In one embodiment, the decoding of the attribute map and the extraction of the attribute information related to the target subgroup are performed by the texture map video decoder 1219.

[0420] In the present disclosure, the attribute map component is compressed into a video bitstream in the same manner as the compression method performed in Figure 1 , 6 , 7, 13 or Fig.30 , and the attribute map information related to the entire mesh is compressed without subgroup distinction. In this case, the texture map video decoder 1219 can use the video codec to decode the texture map bitstream into a video bitstream to reconstruct the texture map. The reconstructed texture map can have color information about each vertex included in the reconstructed mesh, and the color value of each vertex can be obtained from the texture map based on the texture coordinates of each vertex. In other words, the reconstructed attribute map contains the attribute information about all vertices, and this information can be accessed using the texture coordinates of the reconstructed mesh. In the present disclosure, the vertex-specific attribute information corresponding to the target subgroup can also be obtained from the corresponding pixels in the reconstructed attribute map image based on the texture coordinates of the vertices of the reconstructed mesh of the target subgroup. The attribute information related to the target subgroup extracted in this way is used to finally recover the mesh color of the target subgroup.

[0421] According to an embodiment, the mesh reconstructed by the target subgroup mesh reconstructor 1314 and the texture map reconstructed by the texture map video decoder 1219 are rendered by the mesh data renderer 1220 and displayed to the user.

[0422] In this way, Fig.31 the receiving apparatus is configured to perform the above method, including h) the target subgroup determination method, i) the decoding method for the base mesh of the target subgroup, j) the base mesh decoding method used when the base mesh compilation type for the target subgroup is the intra mode, k) the motion field decoding method used when the base mesh compilation type of the target subgroup is the inter mode, l) the method for displacement decoding and extracting displacement information related to the target subgroup, and m) the method for decoding the attribute map and extracting attribute information related to the target subgroup.

[0423] That is, based on the received subgroup-related auxiliary information bitstream and the target area information received by the user-specified area receiver 1311, the subgroup determiner 1312 determines the target subgroup. Then, in order to reconstruct a partial area of the mesh for the determined target subgroup, the corresponding bitstream is decoded. A bitstream area corresponding to the target subgroup index is extracted from the base mesh bitstream or the motion vector bitstream, and the extracted bitstream area is decoded based on the subgroup-related signaling information included in the subgroup-related auxiliary information bitstream. As a result, the base mesh of the target subgroup is reconstructed. In addition, when all displacement components (or displacement vectors) are reconstructed from the displacement vector bitstream by the displacement vector video decoder 1217 and the displacement vector reconstructor 1218, the target subgroup displacement vector extractor 1315 extracts only the information corresponding to the target subgroup from the reconstructed displacement vectors. Then, the extracted displacement vectors of the target subgroup and the reconstructed base mesh of the target subgroup are combined by the target subgroup mesh reconstructor 1314 to reconstruct the mesh of the target subgroup. Based on the texture coordinates of the reconstructed mesh of the target subgroup, the mesh color of the target subgroup can be extracted from the texture map reconstructed by the texture map video decoder 1219. Therefore, the mesh data renderer 1220 can display and utilize the finally reconstructed mesh of the target local area.

[0424] Fig.32 is a flowchart showing an example of the transmission method according to an embodiment. The transmission method according to an embodiment may include preprocessing and encoding mesh data (21011) and transmitting a bitstream containing the encoded mesh data (21012).

[0425] According to an embodiment, the operation 21011 of preprocessing and encoding mesh data may include the above methods, including a) a method of defining subgroups of a mesh and generating the subgroups, b) a method of atlas parameterization based on the subgroups, c) a method of generating a base mesh based on the subgroups, d) a method of fitting a subdivision surface based on the subgroups, e) a method of encoding the base mesh based on the subgroups, f) a method of generating and encoding a motion field based on the subgroups, and g) a method of encoding displacement and attribute maps.

[0426] In other words, in the encoding operation 21011, subgroups are generated for the input original mesh according to a user-specified method, and atlas parameterization is performed based on the subgroups. Further, depending on whether the encoding type is an intra mode or an inter mode, the base mesh is encoded based on the subgroups to generate a base mesh bitstream, or the motion vectors (also referred to as a motion field) are encoded to generate a motion vector bitstream. Additionally, the displacement vectors are encoded to generate a displacement vector bitstream, and the texture maps are encoded to generate a texture map bitstream. Then, the generated bitstreams are multiplexed and sent to a receiving device. Subgroup-related signaling information is included in the subgroup-related auxiliary information bitstream to be sent. At this time, all or part of the subgroup-related signaling information may be sent as header information, or sent together with the base mesh bitstream of each subgroup. Alternatively, it may be sent as header information before all the base mesh bitstreams are sent. For a more specific and detailed description of the subgrouping method, subgroup-based encoding, and subgroup-related signaling information for the subgrouping method, it will not be provided below to avoid redundancy. Refer to Figures 15 to 30 the description of.

[0427] Fig.33 is a flowchart showing an example of a receiving method according to an embodiment. The receiving method according to an embodiment may include receiving a bitstream containing mesh data (22011) and decoding the mesh data contained in the bitstream (22012).

[0428] According to an embodiment, the operation 22012 of decoding mesh data includes the above methods, including h) a method of determining a target subgroup, i) a method of decoding a base mesh for the target subgroup, j) a base mesh decoding method used when the encoding type of the base mesh for the target subgroup is an intra mode, k) a motion field decoding method used when the encoding type of the base mesh of the target subgroup is an inter mode, l) a method of decoding displacement and extracting displacement information related to the target subgroup, and m) a method of decoding an attribute map and extracting attribute information related to the target subgroup.

[0429] That is to say, in operation 22012 of decoding mesh data, a target subgroup is determined based on the received subgroup-related signaling information and target region information, basic mesh decoding or motion vector decoding is performed based on the determined target subgroup, and the basic mesh of the target subgroup is decoded and reconstructed. In addition, only the information corresponding to the target subgroup is extracted from the reconstructed displacement vectors, and then combined with the reconstructed basic mesh of the target subgroup to reconstruct the mesh of the target subgroup. Then, based on the texture coordinates of the reconstructed mesh of the target subgroup, the mesh color of the target subgroup can be extracted from the reconstructed texture map, and the final reconstructed mesh of the target local region can be displayed and utilized by the renderer. For a more specific and detailed description of operation 22012 of decoding mesh data, it will not be provided below to avoid redundancy. Refer to Figures 15 to 31 for a detailed description.

[0430] In other words, in various usage scenarios, devices, and applications that utilize dynamic meshes, there may be cases where only information about a specific region is needed, even though the entire received dynamic mesh data can be used. Therefore, in such cases, the present disclosure may allow a device or software configured to render mesh data to receive mesh region information related to the necessary region from a user, decode and display only that region, and also make the information available for other regions. In terms of the associated costs of processing time and resources being compared, this method may be more efficient than processing the entire mesh data.

[0431] In summary, the present disclosure defines the concept of a subgroup for accessing a partial region and proposes a method for generating a subgroup. Based on this proposal, it presents atlas parameterization, basic mesh compilation, and motion field compilation methods as well as the necessary signaling. Therefore, independent local region reconstruction can be performed on the basic mesh, and the reconstruction result can be combined with the reconstruction results of other mesh components to reconstruct the mesh of the local region desired by the user.

[0432] Each of the above parts, modules, or units may be software, a processor, or a hardware part that executes a continuous process stored in a memory (or storage unit). Each of the steps described in the above embodiments may be executed by a processor, software, or a hardware part. Each of the above modules / blocks / units described in the above embodiments may operate as a processor, software, or hardware. Additionally, the method presented in the embodiments may be executed as code. The code may be written on a processor-readable storage medium and thus read by the processor provided by the device.

[0433] In this specification, when a part "includes" or "contains" an element, unless otherwise mentioned, it means that the part also includes or contains another element. Additionally, the term "... module (or unit)" disclosed in this specification refers to a unit for processing at least one function or operation and may be implemented by hardware, software, or a combination of hardware and software.

[0434] Although, for simplicity, the embodiments have been described with reference to the respective drawings, new embodiments can be designed by combining the embodiments shown in the drawings. If a person skilled in the art designs a computer-readable recording medium recording a program for executing the embodiments mentioned in the above description, it may fall within the scope of the appended claims and their equivalents.

[0435] The devices and methods may not be limited by the configurations and methods of the above embodiments. The above embodiments can be configured by selectively combining each other completely or partially, so that various modifications can be made.

[0436] Although the preferred embodiments of the embodiments have been shown and described, the embodiments are not limited to the above specific embodiments. Without departing from the spirit of the embodiments claimed in the claims, various modifications can be made by those of ordinary skill in the art, and these modifications should not be understood separately from the technical idea or vision of the embodiments.

[0437] Various elements of the devices of the embodiments can be implemented by hardware, software, firmware, or a combination thereof. Various elements in the embodiments can be implemented by a single chip (e.g., a single hardware circuit). According to an embodiment, components according to the embodiment can be implemented separately as separate chips. According to an embodiment, at least one or more components of the device according to the embodiment can include one or more processors capable of executing one or more programs. The one or more programs can execute any one or more of the operations / methods according to the embodiment, or include instructions for executing it. Executable instructions for executing the methods / operations of the device according to the embodiment can be stored in a non-transitory CRM or other computer program products configured to be executed by one or more processors, or can be stored in a transitory CRM or other computer program products configured to be executed by one or more processors. Additionally, the memory according to the embodiment can be used as a concept that covers not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it can also be implemented in the form of a carrier wave (e.g., transmission via the Internet). Additionally, the processor-readable recording medium can be distributed in computer systems connected via a network, so that the processor-readable code can be stored and executed in a distributed manner.

[0438] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Additionally, "A, B" may mean "A and / or B". Additionally, "A / B / C" may mean "at least one of A, B, and / or C". "A / B / C" may mean "at least one of A, B, and / or C". Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0439] The various elements of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various elements in the embodiments may be executed by a single chip (e.g., a single hardware circuit). According to an embodiment, the elements may be selectively executed by separate chips respectively. According to an embodiment, at least one element of the embodiments may be executed in one or more processors including instructions for performing the operations according to the embodiments.

[0440] The operations according to the embodiments described in this specification may be executed by a transmitting / receiving device according to the embodiments including one or more memories and / or one or more processors. One or more memories may store programs for processing / controlling the operations according to the embodiments, and one or more processors may control the various operations described in this specification. One or more processors may be referred to as a controller or the like. In an embodiment, the operations may be executed by firmware, software, and / or a combination thereof. The firmware, software, and / or a combination thereof may be stored in a processor or a memory.

[0441] Terms such as first and second may be used to describe various elements of the embodiments. However, the various components according to the embodiments should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, a first user input signal may be referred to as a second user input signal. Similarly, a second user input signal may be referred to as a first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Both the first user input signal and the second user input signal are user input signals, but do not mean the same user input signal unless the context clearly states otherwise. The terms used to describe the embodiments are only used to describe specific embodiments and are not intended to limit the embodiments. As used in the description and claims of the embodiments, unless the context clearly states otherwise, the singular forms "a", "an", and "the" include plural referents. The expression "and / or" is used to include all possible item combinations. Terms such as "including" or "having" are intended to indicate the existence of diagrams, numbers, steps, elements, and / or components and should be understood as not excluding the possibility of the additional existence of diagrams, numbers, steps, elements, and / or components.

[0442] As used herein, conditional statements such as "if" and "when" are not limited to optional cases and are intended to be interpreted as performing related operations or interpreting related definitions according to specific conditions when the specific conditions are met. Embodiments may include variations / modifications within the scope of the claims and their equivalents. It will be apparent to those skilled in the art that various modifications and changes can be made to the present disclosure without departing from the spirit and scope of the present disclosure. Accordingly, the present disclosure is intended to cover modifications and variations of the present disclosure as long as they fall within the scope of the appended claims and their equivalents.

[0443] [Publication Mode]

[0444] As described above, the relevant content has been described in the best mode of carrying out the embodiments.

[0445] [Industrial Applicability]

[0446] As described above, the embodiments can be applied, in whole or in part, to 3D data sending / receiving devices and systems. It will be apparent to those skilled in the art that various changes or modifications can be made to the embodiments within the scope of the embodiments. Accordingly, the embodiments are intended to cover modifications and variations as long as they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for transmitting 3D data, the method comprising: Preprocessing input mesh data and outputting base mesh data; Encoding the base mesh data; And Transmitting a bitstream containing the encoded mesh data and signaling information.

2. The method according to claim 1, further comprising: Dividing the input mesh data into one or more subgroups.

3. The method according to claim 2, wherein The preprocessing includes: Extracting the input mesh data of the subgroup and generating extracted mesh data; Generating texture coordinates for each vertex of the extracted mesh data based on the subgroup, and outputting base mesh data having the texture coordinates; and Subdividing the base mesh data having the texture coordinates based on the subgroup, and generating a fitted subdivided mesh data by performing fitting such that the subdivided base mesh data becomes similar to the input mesh data.

4. The method according to claim 3, wherein Based on the fact that vertices constituting one of the polygons configured by connecting the vertices of the extracted mesh data in the generation of the texture coordinates are included in two or more of the subgroups, connectivity information related to the vertices constituting the polygon is redundantly included in the two or more subgroups.

5. The method according to claim 4, wherein, The encoding includes: Encoding the base mesh data of each of the subgroups and generating a bitstream for each subgroup in the subgroup, and Inserting subgroup identification information before each bitstream in the bitstream to identify the corresponding one in the subgroup.

6. The method according to claim 5, wherein, The encoding further includes: Reconstructing the encoded base mesh data; Generating displacement information based on the fitted subdivided mesh data and the reconstructed base mesh data; Encoding the displacement information and generating a displacement information bitstream; Reconstructing the encoded displacement information; Reconstructing mesh data based on the reconstructed base mesh data and the reconstructed displacement information; Regenerating a texture map based on the texture map of the input mesh data and the reconstructed mesh data; and Encoding the regenerated texture map and generating a texture map bitstream.

7. The method according to claim 6, wherein The signaling information includes subgroup-related signaling information. Wherein, the subgroup-related signaling information includes at least one of the following: Information for identifying the number of the one or more subgroups; Information for identifying the method of dividing the subgroup; or Information for identifying the compilation type for the base mesh data.

8. An apparatus for transmitting 3D data, comprising: A preprocessor configured to preprocess input mesh data and output base mesh data; An encoder configured to encode the base mesh data; And A transmitter configured to transmit a bitstream containing the encoded mesh data and signaling information.

9. The apparatus according to claim 8, further comprising: A subgroup divider configured to divide the input mesh data into one or more subgroups.

10. The apparatus according to claim 9, wherein, The preprocessor includes: A mesh extractor configured to extract the input mesh data of the subgroup and generate extracted mesh data; A parameterization unit configured to generate texture coordinates for each vertex of the meshed data for the extraction based on the subgroup and output base meshed data having the texture coordinates; and A fitting subdivider configured to subdivide the base meshed data having the texture coordinates based on the subgroup and generate a fitted subdivided meshed data by performing fitting such that the subdivided base meshed data becomes similar to the input meshed data.

11. The device according to claim 10, wherein, Based on two or more subgroups in which vertices of one polygon among the polygons configured by connecting vertices of the extracted meshed data are included in the subgroup, connectivity information related to the vertices constituting the polygon is redundantly included in the two or more subgroups.

12. The apparatus according to claim 11, wherein, The encoder includes: A base mesh encoder configured to: Encode the base meshed data for each subgroup in the subgroup and generate a bitstream for each subgroup in the subgroup; and Insert subgroup identification information before each bitstream in the bitstream to identify the corresponding one in the subgroup; A base mesh reconstructor configured to reconstruct the encoded base meshed data; A displacement information generator configured to generate displacement information based on the fitted subdivided meshed data and the reconstructed base meshed data; A displacement information encoder configured to encode the displacement information and generate a displacement information bitstream; A displacement information reconstructor configured to reconstruct the encoded displacement information; A mesh reconstructor configured to reconstruct meshed data based on the reconstructed base meshed data and the reconstructed displacement information; A texture map generator configured to regenerate a texture map based on the texture map of the input meshed data and the reconstructed meshed data; and A texture map encoder configured to encode the regenerated texture map and generate a texture map bitstream.

13. The device according to claim 12, wherein, The signaling information includes subgroup-related signaling information. Wherein, the subgroup-related signaling information includes at least one of the following: Information for identifying the number of the one or more subgroups; Information for identifying the method of dividing the subgroup; or Information for identifying the compilation type for the base meshed data.

14. A method for receiving 3D data, the method comprising: Receiving a bitstream including encoded meshed data and signaling information; Decoding the encoded meshed data in the bitstream based on the signaling information; And Rendering the decoded meshed data.

15. The method according to claim 14, wherein, The decoding of the meshed data includes: Determining a target subgroup based on a target area selected by a user and the signaling information; and Extracting the bitstream of the target subgroup from the bitstream and reconstructing the meshed data of the target subgroup by performing decoding based on the extracted bitstream of the target subgroup.