Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
Patent Information
- Application Number
- PCT/KR2026/003766
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-12
- Filing Date
- 2026-03-09
- Publication Date
- 2026-09-17
Smart Images

Figure KR2026003766_17092026_PF_FP_ABST
Abstract
Description
Mesh data encoding device, mesh data encoding method, mesh data decoding device and mesh data decoding method
[0001] The embodiments provide a method for providing Point Cloud content to provide various services to users, such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.
[0002] A point cloud is a collection of points in 3D space. There is a problem in that it is difficult to generate point cloud data because there are many points in 3D space.
[0003] Mesh data refers to a form of data in which connectivity information between the vertices of a mesh is added to point cloud data.
[0004] There is a problem in that a large amount of throughput is required to transmit and receive dynamic mesh data.
[0005] The technical problem according to the embodiments is to provide a point mesh data transmission device, a transmission method, a mesh data reception device, and a reception method for efficiently transmitting and receiving mesh data in order to solve the aforementioned problems, etc.
[0006] The technical problem according to the embodiments is to provide a point mesh transmission device, a transmission method, a mesh data receiving device, and a receiving method for solving latency and encoding / decoding complexity.
[0007] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that can be inferred by a person skilled in the art based on the entire content of this document.
[0008] To achieve the above-described purpose and other advantages, the decoding method according to the embodiments may include the step of decoding base mesh data in a bitstream; the step of decoding displacement data in a bitstream; and the step of decoding attribute data in a bitstream. The encoding method according to the embodiments may include the step of encoding base mesh data of mesh data; the step of encoding displacement data of mesh data; and the step of encoding attribute data of mesh data.
[0009] A mesh data transmission method, a transmission device, a mesh data reception method, and a reception device according to the embodiments can provide a high-quality mesh service.
[0010] A mesh data transmission method, a transmission device, a mesh data reception method, and a reception device according to the embodiments can achieve various video codec methods.
[0011] The mesh transmission method, transmission device, mesh data reception method, and reception device according to the embodiments can provide general-purpose mesh content such as autonomous driving services.
[0012] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments. For a better understanding of the various embodiments described below, one must refer to the description of the embodiments below in relation to the following drawings, which include parts corresponding to similar reference numerals throughout the drawings.
[0013] FIG. 1 shows a V-DMC-based encoder and decoder according to embodiments.
[0014] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.
[0015] FIG. 3 illustrates a V-MESH compression method according to embodiments.
[0016] FIG. 4 shows the pre-processing of V-MESH compression according to the embodiments.
[0017] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.
[0018] FIG. 6 illustrates a displacement generation process according to embodiments.
[0019] FIG. 7 illustrates the V-DMC encoding process according to the embodiments.
[0020] FIG. 8 illustrates a lifting conversion process for displacement according to embodiments.
[0021] FIG. 9 illustrates the process of packing conversion coefficients according to embodiments into a 2D image.
[0022] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.
[0023] FIG. 11 illustrates a V-DMC decoding process according to embodiments.
[0024] FIG. 12 illustrates a V-DMC encoding process according to embodiments.
[0025] FIG. 13 illustrates a V-DMC decoding process according to embodiments.
[0026] Figure 14 shows the conventional LoD extraction information SEI payload syntax.
[0027] Figures 15a and 15b show a conventional tile submesh mapping SEI payload syntax.
[0028] Figure 16 shows the conventional attribute extraction information SEI payload syntax.
[0029] FIG. 17 shows an encoded dynamic bitstream structure according to embodiments.
[0030] FIG. 18 shows the V3C unit payload syntax according to the embodiments.
[0031] FIG. 19 shows a dynamic mesh encoder configuration according to embodiments.
[0032] FIGS. 20 to 22 show the configuration of a displacement vector encoding unit according to embodiments.
[0033] FIGS. 23 and 24 show examples of designating the quantized displacement vector transformation coefficient packing region per LoD according to embodiments as MCTS.
[0034] FIG. 25 shows an example of specifying a packing area for quantized displacement vector transformation coefficients per LoD according to embodiments.
[0035] FIG. 26 shows an example of specifying a packing area for quantized displacement vector transformation coefficients per LoD according to embodiments.
[0036] FIGS. 27 and 28 illustrate a method for predicting the current displacement vector transformation coefficient based on the restored displacement vector transformation coefficient of a reference mesh according to embodiments.
[0037] FIG. 29 illustrates a lifting conversion process according to embodiments.
[0038] FIG. 30 shows a dynamic mesh decoder according to embodiments.
[0039] FIGS. 31 and 32 illustrate the displacement vector coordinate system inverse transformation process of the displacement vector decoder according to the embodiments.
[0040] FIG. 33 illustrates a displacement vector decoding process according to embodiments.
[0041] FIG. 34 illustrates the lifting inverse conversion process according to the embodiments.
[0042] FIG. 35 shows the extraction information SEI payload syntax according to the embodiments.
[0043] FIG. 36 shows the extraction information SEI payload syntax according to the embodiments.
[0044] FIGS. 37a, FIGS. 37b, and FIGS. 37c show mapping information SEI payload syntax according to embodiments.
[0045] FIGS. 38a and FIGS. 38b show mapping information SEI payload syntax according to embodiments.
[0046] FIGS. 39a and FIGS. 39b show mapping information SEI payload syntax according to embodiments.
[0047] FIG. 40 shows a configuration diagram of a dynamic mesh content receiver according to embodiments.
[0048] FIG. 41 illustrates the process of LoD extraction and attribute extraction being performed at a receiver when parsing an extraction information SEI message according to the embodiments.
[0049] FIGS. 42a and FIGS. 42b illustrate the process of LoD extraction, attribute extraction, and target atlas tile decoding being performed at a receiver when parsing a mapping information SEI message according to the embodiments.
[0050] FIGS. 43a and FIGS. 43b illustrate the process of decoding video tiles of target atlas tiles and sub-meshes at a receiver when parsing mapping information SEI messages according to the embodiments.
[0051] FIG. 44 illustrates a encoding method according to embodiments.
[0052] FIG. 45 illustrates a decoding method according to embodiments.
[0053]
[0054] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments may be practiced without these details.
[0055] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.
[0056] FIG. 1 shows a V-DMC-based encoder and decoder according to embodiments.
[0057] The basic structure of the currently ongoing V-DMC (v-mesh) is shown in Fig. 1. The encoder and decoder according to Fig. 1 perform the encoding and decoding processes of media representing a dynamic mesh using V3C technology. The preprocessor converts the input dynamic mesh representation into several V3C components (base mesh, displacement set, 2D representation of attributes, and atlas). The original mesh is simplified into a base mesh. The base mesh can be encoded using any mesh codec. Displacement vectors can be represented by a profile or encoded into V3C geometric video components using any video codec via SEI messages. For example, depending on the profile, displacement vectors (displacement data) can be encoded using arithmetic coding. Attribute data may include additional attributes. For example, texture or material information may be included as additional attributes and can be encoded using any video codec. Atlas data contains information on how to perform inverse reconstruction and is provided to the V3C decoding and / or rendering system. For example, atlas data may include methods for performing subdivision of the base mesh, methods for applying displacement vectors to the vertices of the subdivided mesh, and methods for applying attributes to the reconstructed mesh.
[0058] The encoder may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a preprocessor, an atlas encoder, a basemesh encoder, a displacement vector encoder, a video encoder, and a multiplexer.
[0059] The atlas encoding unit encodes the atlas of the mesh data to generate an atlas bitstream. The basemesh encoding unit encodes the basemesh of the mesh data to generate a basemesh bitstream. The displacement vector encoding unit encodes the displacement vector of the mesh data to generate a displacement vector bitstream. The video encoding unit encodes the attributes of the mesh data to generate an attribute bitstream. The encoder generates parameter information (which may be referred to as signaling information, metadata, etc.) related to each encoding. The encoder can generate a bitstream containing parameter information, the atlas, the basemesh, the displacement vector, and / or attributes.
[0060] The decoder may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a demultiplexer, an atlas decoder, a basemesh decoder, a displacement vector decoder, and a video decoder.
[0061] The atlas decoder decodes the atlas within the bitstream. The basemesh decoder decodes the basemesh within the bitstream. The displacement vector decoder decodes the displacement vector within the bitstream. The video decoder decodes the attributes within the bitstream. The decoder can perform each decoding operation based on parameter information within the bitstream. The decoder can reconstruct dynamic mesh data based on the atlas, displacement vector, attributes, and basemesh.
[0062] Below, the operation of the V-DMC encoder and decoder of FIG. 1 is explained in more detail.
[0063] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.
[0064] The system of FIG. 2 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a receiver (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or the dynamic mesh video encoder (hereinafter, encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to the receiving device (110) or the dynamic mesh video decoder (hereinafter, decoder) (113).
[0065] The system of Fig. 2 can perform video-based dynamic mesh compression and decompression.
[0066] With advancements in 3D capture, modeling, and rendering, users can access various forms of 3D content, such as AR, XR, the metaverse, and holograms, across multiple platforms and devices. 3D content represents objects more sophisticatedly and realistically to enable users to enjoy immersive experiences, and for this purpose, the creation and use of 3D models require a large amount of data. Among the various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. The embodiments include a series of processing steps in a system that uses such mesh content.
[0067] First, the method for compressing dynamic mesh data originates from the V-PCC (Video-based point cloud compression) standard technology. Point cloud data consists of data containing color information along with vertex coordinates (X, Y, Z). Mesh data refers to data where connectivity information between vertices is added to this vertex data. Content can be created in a mesh data format from the outset when generating content. Point cloud data can be converted into mesh data and used by adding connectivity information.
[0068] Currently, the MPEG standards organization defines the data types for dynamic mesh data as the following two types: Category 1: Mesh data containing texture maps as color information. Category 2: Mesh data containing vertex colors as color information.
[0069] Mesh coding standards for Category 1 data are currently being developed, and standardization work for Category 2 data is also scheduled to proceed in the future. As shown in Fig. 1, the entire process for providing mesh content services may include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.
[0070] To provide mesh content services, 3D data acquired through multiple cameras or special cameras can be processed into a mesh data type through a series of processes and then generated as a video. The generated mesh video is transmitted after undergoing a series of processes, and at the receiving end, the received data can be processed back into a mesh video and rendered. Through this, the mesh video is provided to the user, and the user can use the mesh content according to their intention through interaction.
[0071] A mesh compression system may include a transmission device and a reception device. The transmission device can encode mesh video to output a bitstream and transmit it to the reception device via a digital storage medium or network in the form of a file or streaming (streaming segment). The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0072] The transmission device may schematically include a mesh video acquisition unit, a mesh video encoder, and a transmission unit. The receiving device may schematically include a receiver, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / image / picture / frame encoding device, and the decoder may be referred to as a mesh video / image / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be composed of separate devices or external components. The transmission device and the receiving device may further include separate internal or external modules / units / components for a feedback process.
[0073] Mesh data represents the surface of an object using multiple polygons. Each polygon is defined by a vertex in 3D space and connectivity information indicating how those vertices are connected. It may also include vertex attributes such as color and normals. Mapping information, which enables the surface of the mesh to be mapped onto a 2D planar area, can also be included as an attribute of the mesh. Mapping can generally be described by a set of parametric coordinates, referred to as UV coordinates or texture coordinates, associated with the mesh vertices. The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as textures, normals, and displacements.
[0074] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc., into a mesh data type having the attributes described above through a series of processes, and generating a video composed of such mesh data. In a mesh video, the attributes of the mesh, namely vertices, polygons, connectivity information between vertices, color, normals, etc., may change over time. A mesh video having attributes and connectivity information that change over time in this way can be described as a dynamic mesh video.
[0075] A mesh video encoder can encode input mesh video into one or more video streams. A single video may contain multiple frames, and a single frame may correspond to a still image or picture. In this document, the term "mesh video" may include mesh images, frames, or pictures, and the terms mesh video and mesh images, frames, or pictures may be used interchangeably. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. To improve compression and coding efficiency, a mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0076] The file / segment encapsulation module can encapsulate encoded mesh video data and / or mesh video-related metadata into a file or the like. Here, the mesh video-related metadata may be received from a metadata processing module or the like. The metadata processing module may be included in the mesh video encoder or may be configured as a separate component / module. The encapsulation module can encapsulate the data into a file format such as ISOBMFF or process it into other forms such as DASH segments. Depending on the embodiment, the encapsulation module may include mesh video-related metadata in the file format. Mesh video metadata may be included, for example, in various levels of boxes in the ISOBMFF file format or as data within a separate track in the file. Depending on the embodiment, the encapsulation module may encapsulate the mesh video-related metadata itself into a file.
[0077] The transmission processing unit can apply transmission processing to mesh video data encapsulated according to the file format. The transmission processing unit may be included in the transmission unit or may be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. Processing for transmission may include processing for delivery via a broadcast network and processing for delivery via broadband. According to an embodiment, the transmission processing unit may receive mesh video-related metadata from the metadata processing unit in addition to mesh video data, and apply transmission processing to it.
[0078] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.
[0079] The receiver can receive mesh video data transmitted by a mesh video transmission device. Depending on the transmission channel, the receiver may receive mesh video data via a broadcast network or via broadband. Alternatively, it may receive mesh video data via a digital storage medium.
[0080] The receiving processing unit can perform processing on the received mesh video data according to the transmission protocol. The receiving processing unit may be included in the receiving unit or may be configured as a separate component or module. Corresponding to the processing for transmission performed on the transmitting side, the receiving processing unit may perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit may transmit the acquired mesh video data to the decapsulation processing unit and the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the receiving processing unit may be in the form of a signaling table.
[0081] The file / segment decapsulation module can decapsulate mesh video data in file format received from the receiving module. The decapsulation module can decapsulate files based on ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing module. The mesh video bitstream may also contain metadata (metadata bitstream). The metadata processing module may be included in the mesh video decoder or configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation module may be in the form of boxes or tracks within the file format. If necessary, the decapsulation module may receive metadata required for decapsulation from the metadata processing module. Mesh video-related metadata may be passed to a Mesh video decoder and used in the Mesh video decoding process, or passed to a renderer and used in the Mesh video rendering process.
[0082] The Mesh Video Decoder can decode video by receiving a bitstream as input and performing operations corresponding to those of the Mesh Video Encoder. The decoded Mesh Video can be displayed through a display unit. The user can view all or part of the rendered result through a VR / AR display or a standard display.
[0083] The feedback process may include the process of transmitting various feedback information obtainable during the rendering / display process to the transmitting side or to the decoder of the receiving side. Interactivity in mesh video consumption may be provided through the feedback process. According to an embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted during the feedback process. According to an embodiment, the user may interact with elements implemented in a VR / AR / MR / autonomous driving environment, and in this case, information related to such interaction may be transmitted to the transmitting side or the service provider side during the feedback process. According to an embodiment, the feedback process may not be performed.
[0084] Head orientation information can refer to information regarding the user's head position, angle, movement, etc. Based on this information, viewport information—that is, information about the area the user is currently viewing within the mesh video—can be calculated.
[0085] Viewport information may be information about the area currently being viewed by the user within a mesh video. Through this, gaze analysis can be performed to determine how the user consumes the mesh video and to what extent they gaze at specific areas of the mesh video. Gaze analysis may be performed at the receiving end and transmitted to the transmitting end via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0086] According to an embodiment, the aforementioned feedback information may not only be transmitted to the transmitting side but may also be consumed at the receiving side. That is, decoding and rendering processes at the receiving side may be performed using the aforementioned feedback information. For example, using head orientation information and / or viewport information, only the mesh video for the area currently viewed by the user may be preferentially decoded and rendered.
[0087] This document relates to dynamic mesh video compression as described above. The methods / embodiments disclosed in this document may be applied to the MPEG (Moving Picture Experts Group) Video-based Dynamic Mesh Compression Method (V-Mesh) standard or next-generation video / image coding standards. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.
[0088] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0089] In this document, "picture" or "frame" generally refers to a unit representing a single image of a specific time period.
[0090] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the lumina component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0091] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0092] The encoding process of Fig. 2 is as follows.
[0093] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. In the V-Mesh compression process, compression is performed by receiving the following data as input.
[0094] Input mesh: It includes the 3D coordinates (geometry) of the vertices constituting the mesh, normal information for each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between the vertices constituting the surface. The surface of the mesh can be represented by triangles or polygons of greater size, and connection information between the vertices constituting each surface is stored according to a defined shape. The input mesh can be saved in the OBJ file format.
[0095] Attribute map: (Hereafter, Texture map is used with the same meaning): It contains information on the attributes (color, normals, displacement, etc.) of a mesh and stores data in the form of mapping the mesh surface onto a 2D image. Mapping which part of the mesh (surface or vertex) corresponds to each data point in this attribute map is based on the mapping information contained in the input mesh. Since the attribute map holds data for each frame of the mesh video, it can also be referred to as an attribute map video (or simply "attribute"). In the V-Mesh compression method, the attribute map primarily contains the mesh's color information and is stored in image file formats (PNG, BMP, etc.).
[0096] Material Library File: Contains material attribute information used in the mesh, specifically information that links the input mesh with its corresponding attribute map. This is saved in the Wavefront Material Template Library (MTL) file format.
[0097] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0098] Base Mesh: By simplifying (decimating) the input mesh through a preprocessing stage, objects from the input mesh are represented using the minimum number of vertices determined according to user criteria.
[0099] Displacement: Displacement information used to represent the input mesh as similarly as possible using the base mesh, expressed in the form of 3D coordinates.
[0100] Atlas information: This is metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be generated and utilized as sub-units (sub-mesh, patch, etc.) that constitute the mesh.
[0101] Referring to FIGS. 3 to 7, a method for encoding mesh position information (vertices) is described, and referring to FIGS. 7-10 and others, a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0102] FIG. 3 illustrates a V-MESH compression method according to embodiments.
[0103] FIG. 3 illustrates the encoding process of FIG. 2, and the encoding process may include pre-processing and encoding processes. The encoder of FIG. 2 may include a pre-processor (200) and an encoder (201) as in FIG. 3. The transmitting device of FIG. 2 may be broadly referred to as an encoder, and the dynamic mesh video encoder of FIG. 2 may be referred to as an encoder. The V-Mesh compression method may include pre-processing (200) and encoding (201) processes as in FIG. 3. The pre-processor of FIG. 3 may be located in front of the encoder of FIG. 3. The pre-processor and encoder of FIG. 3 may be referred to as a single encoder.
[0104] The pre-processor can receive a static dynamic mesh and / or attribute map. The pre-processor can generate a base mesh and / or displacement through preprocessing. The pre-processor can receive feedback information from an encoder and generate a base mesh and / or displacement based on the feedback information.
[0105] The encoder can receive a base mesh, displacement, static dynamic mesh, and / or attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.
[0106] FIG. 4 shows the pre-processing of V-MESH compression according to the embodiments.
[0107] Figure 4 shows the configuration and operation of the pre-processor of Figure 3.
[0108] FIG. 3 illustrates a process of performing preprocessing on an input mesh. The preprocessing process (200) may include four main steps: 1) generation of a Group of Frame (GoF), 2) mesh decimation, 3) UV parameterization, and 4) fitting subdivision surface (300). The pre-processor (200) may receive the input mesh, generate displacement and / or a base mesh, and transmit it to the encoder (201). The pre-processor (200) may transmit GoF information associated with GoF generation to the encoder (201).
[0109] Below, each step of Fig. 4 is explained.
[0110] GoF Generation: This is the process of creating a reference structure for mesh data. If the number of vertices, texture coordinates, vertex connectivity, and texture coordinate connectivity of the mesh in the previous frame and the current mesh are all identical, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values differ between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame undergoes intra-frame encoding.
[0111] Mesh Decimation: This is the process of simplifying the input mesh to generate a simplified mesh, or base mesh. After selecting vertices to remove from the original mesh based on user-defined criteria, the selected vertices and the triangles connected to them can be removed.
[0112] In the process of performing mesh decimation, information regarding the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) is passed as input, and a decimated mesh can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0113] UV Parameterization: This is the process of mapping 3D surfaces to a texture domain for a decimated mesh. Parameterization can be performed using a UV Atlas tool. Through this process, mapping information is generated regarding where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process.
[0114] OrthoAtlas technology is a technique that generates texture coordinates using orthographic projection. In orthoAtlas technology, the processes of patch generation and patch packing are performed sequentially. First, Connected Components (CCs) are generated by splitting adjacent triangles, and then the optimal CCs are merged using a cost function to generate a patch. The cost function can measure the cost based on the degree of distortion that occurs when orthogonally projecting the patch in each direction. Finally, texture coordinates can be calculated by packing the patch that minimizes the cost function into the texture domain. In the case of orthoAtlas technology, texture coordinates can be derived in the base mesh decoder without compressing texture coordinate and texture connection information during the base mesh encoding process.
[0115] Fitting subdivision surface: This is the process of performing subdivision on a decimated mesh. User-defined methods, such as the mid-edge method, can be applied as the subdivision method. A fitting process is performed to make the input mesh and the subdivision mesh similar to each other.
[0116] This is a process of performing fitting so that the mesh obtained by subdividing the base mesh becomes similar to the surface of the input mesh. As for the subdivision method, a user-defined method such as the mid-edge method (Fig. 5), loop method, or LS3 method may be applied.
[0117] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.
[0118] Figure 5 illustrates the mid-edge method of the fitting subdivision surface described in Figure 4. Referring to Figure 5, an original mesh containing 4 vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex at the midpoint of the edge between vertices.
[0119] When a fitted subdivided mesh (hereinafter referred to as the fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The positional difference between this result and the fitted subdivided mesh for each vertex becomes the displacement for each vertex. Since displacement represents a positional difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of the Cartesian coordinate system. Depending on user input parameters, (x, y, z) coordinate values can be converted into (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0120] FIG. 6 illustrates a displacement generation process according to embodiments.
[0121] FIG. 6 illustrates in detail the displacement calculation method of a fitting subdivision surface (300) as described in FIG. 5.
[0122] An encoder and / or pre-processor according to the embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a restored base mesh and generate a subdivided restored base mesh. The local coordinate system calculation unit may receive a fitted subdivided mesh and a subdivided restored base mesh and convert a coordinate system relating to the mesh to a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit calculates the positional difference between the fitted subdivision mesh and the subdivided restored base mesh. For example, it may generate a positional difference value between the vertices of the two input meshes. The vertex positional difference value becomes the displacement.
[0123] The method and apparatus for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data according to the embodiments (which may be referred to simply as point cloud) may refer to data including vertex coordinates and color information. Point cloud is a term that includes mesh data, and in this document, point cloud and mesh data may be used interchangeably.
[0124] The V-Mesh compression (restoration) method according to the embodiments may include intra-frame encoding (Fig. 6) and inter-frame encoding (Fig. 7).
[0125] Intra-frame encoding or inter-frame encoding is performed based on the results of the aforementioned GoF generation. In the case of intra-frame encoding, the data to be compressed may include a base mesh, displacement, and attribute map. In the case of inter-frame encoding, the data to be compressed may include displacement, attribute map, and the motion field between the reference base mesh and the current base mesh.
[0126] FIG. 7 illustrates the V-DMC encoding process according to the embodiments.
[0127] The encoding process of Fig. 7 illustrates the encoding of Figs. 1 and 2 in detail.
[0128] A pre-processor receives an input mesh and can perform the aforementioned preprocessing. Through preprocessing, a base mesh and / or a fitted subdivided mesh can be generated. A quantizer can quantize the base mesh and / or the fitted subdivided mesh. A static mesh encoder can encode a static mesh. The static mesh encoder can generate a bitstream containing the encoded base mesh. A motion encoder can encode motion vectors for the base mesh based on inter-frame motion estimation and motion compensation for inter-prediction. An atlas encoder can encode an atlas for the vertices of the base mesh. The encoded base mesh can be reconstructed and inversely quantized through an inverse quantizer. A displacement computer receives the reconstructed mesh and, based on the fitted subdivided mesh, can generate displacement, which is the position difference. A lifting transform can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. Depending on the encoding method, the image packing unit can pack the image based on the quantized lifting coefficients. The video encoder can encode the packed image. Depending on the encoding method, it can apply interprediction to the quantized lifting coefficients and encode the predicted lifting coefficients according to an arithmetic encoding method. The mesh restoration unit restores the deformed mesh through the restored displacement and the restored base mesh. The displacement data is restored, and the deformed mesh is restored based on the restored displacement data and the restored base mesh and provided to the attribute transfer. The attribute transfer receives the input mesh and / or input attribute map and generates an attribute map based on the restored deformed mesh. Push-pull padding can pad data into the attribute map based on a push-pull method. The color space converter can convert the space of the color component that is an attribute. The video encoder can encode the attributes.A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.
[0129] Base Mesh Encoding: Base mesh compression methods can be divided into INTRA, INTER, and SKIP types depending on the base mesh type, and encoding can be performed in different ways for each. If the base mesh is of the INTRA type, it can be encoded using a static mesh encoding method. If the base mesh is of the INTER type, the motion field between the reference base mesh and the current base mesh can be encoded. If the current base mesh is of the SKIP type, the reference base mesh can be derived into the current base mesh.
[0130] After being encoded in the encoder, the base mesh can be subdivided into a subdivided mesh through a subdivision process. Subdivision algorithms such as mid-point subdivision and loop subdivision can be used.
[0131] Static Basemesh Encoding (Intra Basemesh Encoding): When performing intra encoding on the current basemesh, the base mesh generated during the preprocessing stage can be encoded using static mesh compression technology after undergoing a quantization process. Static mesh compression utilizes MPEG EdgeBreaker (MEB) technology, and the base mesh's vertex position information, mapping information (texture coordinates), vertex connectivity information, and normals are subject to compression.
[0132] The technology for compressing connection information can be encoded based on the edgebreaker algorithm. The edgebreaker algorithm is a technique that sequentially traverses triangles according to rules, maps symbols based on the characteristics of each triangle, and then encodes the corresponding symbols.
[0133] Techniques for compressing vertex location information can calculate predicted values based on prediction techniques such as multiple parallelogram prediction, and then encode the residual value, which is the difference between the current vertex and the predicted value.
[0134] A technique for compressing mapping information (texture coordinates) can calculate a predicted value based on a prediction technique such as stretching, and then encode the residual value, which is the difference between the current mapping information (texture coordinates) and the predicted value.
[0135] Normal compression techniques can obtain predicted values based on prediction techniques such as delta coding, multiple parallelogram prediction, and cross product-based prediction, and then encode the residual value, which is the difference between the current normal and the predicted value.
[0136] Motion Field Encoding (Inter Basemesh Encoding): Inter basemesh encoding can be performed when a one-to-one correspondence exists between the reference mesh and the current input mesh, differing only in vertex position information. When performing inter encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh—that is, the motion field—is calculated and this information is encoded. The reference base mesh is the result of quantizing already decoded base mesh data and is determined by the reference frame index.
[0137] The motion field can be encoded as is, or the predicted motion field can be calculated by averaging the motion fields of the restored vertices among those connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the current vertex's motion field value, can be encoded. This value can be encoded using entropy coding.
[0138] Displacement Encoding: After encoding the base mesh, reconstruction and inverse quantization are performed to reconstruct it. Once the base mesh is generated and subdivision is performed on it, the displacement between the result and the fitted subdivided mesh can be calculated. For effective encoding, data transform processes such as wavelet transform can be applied to the displacement information, and Figure 7 shows the process of transforming displacement information using the lifting transform in V-Mesh. The transform coefficients generated through the transformation process are quantized, and the quantized transform coefficients can be compressed through a video codec or through arithmetic coding, depending on the compression method.
[0139] When compressed through a video codec, the data is packed into a 2D image as shown in Figure 8. Transform coefficients are organized into one block for every N^2 (N*N) units, and each block can be packed in z-scan order. The number of horizontal blocks is fixed at N, while the number of vertical blocks can be determined by the number of vertices in the subdivided base mesh. Within a single block, transform coefficients can be packed by aligning them using Morton code. The packed images generate a displacement video for every GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0140] When compressed via arithmetic coding, cross-frame prediction can be performed on the quantized displacement vector transformation coefficients. When cross-frame prediction is performed on the current quantized displacement vector transformation coefficients, the residual value, which is the difference between the current displacement vector transformation coefficient and the reference displacement vector transformation coefficient, can be encoded, and information about the reference target can be encoded. Depending on the displacement vector type, the quantized displacement vector transformation coefficients can be arithmetic encoded if it is of the INTRA type, and the residual value if it is of the INTER type. Arithmetic coding can be performed based on Context Adaptive Binary Arithmetic Coding (CABAC). The CABAC process can first binarize the displacement vector data and map it to a bin string. The bin string can be an output binarized into 0s and 1s, where each 0 or 1 can be a bin. Each bin can be arithmetic encoded using context information selected from the context model, and a process of updating probabilities can be performed.
[0141]
[0142] FIG. 8 illustrates a lifting conversion process for displacement according to embodiments.
[0143] FIG. 9 illustrates the process of packing conversion coefficients according to embodiments into a 2D image.
[0144] FIGS. 8-9 respectively illustrate the process of converting the displacement of the encoding process of FIG. 7 and the process of packing the conversion coefficients.
[0145] The encoding method according to the embodiments includes displacement encoding.
[0146] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through reconstruction and inverse quantization. Subdivision is then performed on this reconstructed base mesh, and the displacement between the result and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as a wavelet transform can be applied to the displacement information.
[0147] Figure 8 illustrates the process of transforming displacement information using a lifting transform in V-Mesh. The transformation coefficients generated through the transformation process are quantized and then packed into a 2D image as shown in Figure 9. The transformation coefficients are 256 (=16 16) Each unit is composed of one block, and each block can be packed in z-scan order. The number of blocks is fixed at 16, but the number of blocks can be determined by the number of vertices of the subdivided base mesh. Transform coefficients can be packed by aligning them with a Morton code within a single block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using a conventional video compression codec.
[0148] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. A first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes subdivided vertices and vertices of the base mesh (LoD0). A second subdivision mesh may be generated by dividing the first subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 containing vertices additionally generated from LoD0, and vertices additionally divided from LoD1. LoD is a Level of Detail indicating the degree of detail; as the index of the level increases, the distance between vertices becomes closer, and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When vertices are further subdivided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1 and v2 and the subdivided vertex v. Instead of encoding the information for the current LoD N as is, the size of the bitstream can be reduced by generating residuals between the previous LoD N-1 and encoding the mesh using these residuals. The prediction process refers to the operation of predicting the current vertex v based on the previous vertices v1 and v2. Since adjacent subdivided meshes possess similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted using the residuals from the previous vertex position information, and the previous vertex position information is updated using these residuals.
[0149] Referring to FIG. 9, the vertices have coefficients generated through a lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.
[0150] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.
[0151] Figure 10 shows the detailed operation of the attribute transfer of the encoding of Figure 7.
[0152] The encoding according to the embodiments includes attribute map encoding.
[0153] Information regarding the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. The input mesh compressed during the encoding process is restored through base mesh decoding (intra frame), motion field encoding (inter frame), and displacement video decoding. The resulting restored deformed mesh (hereinafter referred to as Recon. deformed mesh) is used to compress the input attribute map, as shown in FIGS. 6 and 7. The Recon. deformed mesh possesses vertex position information, texture coordinates, and corresponding connection information, but lacks color information corresponding to the texture coordinates. Accordingly, as shown in FIG. 10, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is generated through the attribute transfer process in the V-Mesh compression method.
[0154] Attribute transfer first checks for every point P(u, v) in the 2D texture domain whether the point belongs to a texture triangle of the reconstructed deformed mesh, and if it does, determines the barycentric coordinates (α, of P(u, v) according to that triangle T). Calculate , γ). Then, the 3D vertex positions of triangle T and (α, Calculate the 3D coordinates M(x, y, z) of P(u, v) using , γ). Find the vertex coordinates M'(x', y', z') corresponding to the location most similar to the calculated M(x, y, z) in the input mesh domain, and find triangle T' containing this point. Then, in triangle T', find the coordinates of the centroid of M'(x', y', z') (α', Calculate , γ'). Texture coordinates corresponding to the three vertices of Triangle T' and (α', Texture coordinates (u', v') are calculated using , γ'), and color information corresponding to these coordinates is found in the input attribute map. The found color information is then assigned to the (u, v) pixel location in the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.
[0155] The new attribute map generated through attribute transfer is bundled in GoF units to form an attribute map video, which is then compressed using a video codec.
[0156] Atlas Encoding: Atlas information may be transmitted during the aforementioned process. The Atlas consists of information required during the mesh decoding and / or rendering process, and may include information required during the process of performing subdivision, displacement decoding, base mesh decoding, etc., as well as tile information, patch information, etc. Atlas data may be encoded using Exp-Golomb coding, etc.
[0157] Referring to FIG. 10, the reference relationships between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.
[0158] The decoding process of Fig. 1 can perform the reverse process of the corresponding process of the encoding process of Fig. 1. The specific decoding process is as follows.
[0159] FIG. 11 illustrates a V-DMC decoding process according to embodiments.
[0160] Figure 11 shows the configuration and operation of the decoder of the receiving device of Figure 1.
[0161] The input bitstream can be separated into a Basemesh sub-stream, a Displacement sub-stream, an Attribute map sub-stream, and an Atlas sub-stream.
[0162] The Atlas sub-stream can be decoded through Exp-Golomb coding, and as a result, information necessary for performing decoding, tile information, patch information, etc. can be obtained.
[0163] Depending on the basemesh type, if the basemesh sub-stream is of the INTRA type, it can be decoded through a static mesh decoder based on MEB (MPEG EdgeBreaker) technology, and as a result, the connectivity information, vertex geometry information, and vertex mapping information (texture coordinates) of the base mesh can be restored.
[0164] When the texture parameterization method in the encoder is orthoAtlas, the decoder can derive mapping information (texture coordinates) and attribute information (texture) connection information using vertex coordinates. The process of deriving mapping information (texture coordinates) and connection information can generate mapping information (texture coordinates) and attribute information (texture) connection information by calculating the homography transform of each face and then projecting the vertex based on it.
[0165] If the Basemesh type is INTER type, motion information can be decoded through entropy decoding and inverse prediction processes. The restored motion information is combined with the reference Basemesh that has already been restored and stored in the buffer to generate a Reconstructed quantized basemesh for the current frame. An inverse quantization process can be performed on the restored Basemesh.
[0166] If the displacement sub-stream is compressed through a video codec according to the compression method used in encoding, it is decoded into a displacement video through the video compression codec's decoder, and then the image unpacking process is performed.
[0167] When compressed through arithmetic coding, the displacement vector bitstream can decode binarized syntax elements through arithmetic decoding, and a Contextual Probability Model (CPM) can be adaptively determined according to each bin of the syntax elements, and arithmetic decoding can be performed by predicting the probability of occurrence of the bin through the CPM. The binarized syntax elements can be decoded through inverse binarization. Quantized displacement vector transformation coefficients can be derived from the decoded syntax elements. Depending on the displacement information type, if it is INTER (where inter prediction is performed), an inverse inter prediction process is performed using reference information for the quantized displacement coefficients.
[0168] The quantized displacement coefficient is restored as displacement information for each vertex through inverse quantization, inverse transform, and coordinate system transformation processes.
[0169] The restored Base mesh and the restored Displacement information are combined to generate the final Decoded mesh. The Attribute map sub-stream is decoded through the decoder of the video compression codec used in Encoding, and then restored to the final Attribute map through processes such as color format conversion.
[0170] The restored Decoded mesh and Decoded attribute map can be utilized at the receiving end as final mesh data available to the user.
[0171] The atlas decoder decodes atlas data within the bitstream.
[0172] When mesh data within the bitstream is encoded based on inter-prediction, the motion decoder derives the motion field of the current frame's basemesh through motion estimation and compensation, based on the basemesh within the reference frame. When mesh data within the bitstream is encoded based on intra-prediction, the sectic decoder decodes the basemesh. Depending on the encoding method, it decodes displacement data by applying either arithmetic coding or video decoding. The video decoder decodes attribute data within the bitstream.
[0173] The decoding method of FIG. 11 can follow the inverse process of the encoding method according to the embodiments.
[0174] FIG. 12 illustrates a V-DMC encoding process according to embodiments.
[0175] FIG. 12 illustrates the configuration and operation of an encoder of a transmitting device such as FIG. 1 or FIG. 2. Each component of FIG. 12 corresponds to hardware, software, a processor, and / or a combination thereof.
[0176] Figure 12 shows the encoding process of V-Mesh technology.
[0177] The mesh preprocessing unit receives the original mesh as input and generates a decimated mesh. Decimation can be performed based on the number of target vertices or polygons constituting the mesh. For the decimated mesh, parameterization can be performed to generate mapping information (texture coordinates) and attribute information (texture) connection information per vertex. Additionally, floating-point mesh information can be quantized into fixed-point form. This result can be encoded as a base mesh through the static mesh encoding unit. The mesh preprocessing unit can generate additional vertices by performing mesh subdivision on the base mesh. Depending on the subdivision method, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information can be generated. The subdivided mesh can be fitted by adjusting vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.
[0178] The base mesh generated through the mesh preprocessing unit can perform intra-encoding or inter-encoding depending on the base mesh type. If the base mesh frame undergoes intra-encoding, it can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the base mesh's connectivity information, vertex geometry information, vertex texture information, normal information, etc. If the base mesh frame undergoes inter-encoding, the motion vector encoding unit is executed; using the base mesh and the reference-restored base mesh as inputs, the motion vector between the two meshes is calculated and its value encoded. The motion vector encoding unit performs connectivity-based prediction using previously encoded / decoded motion vectors as predictors and can encode the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The base mesh bitstream generated through the base mesh encoding unit is transmitted to the multiplexer.
[0179] The encoded base mesh bitstream can generate a restored base mesh through the base mesh restoration unit.
[0180] The displacement vector calculation unit can perform mesh subdivision on the reconstructed base mesh. The displacement vector can be calculated as the difference in vertex positions between the subdivided reconstructed base mesh and the fitted subdivision mesh generated by the preprocessing unit. As a result, displacement vectors can be calculated for each vertex of the subdivision mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0181] The displacement vector processing unit can transform the displacement vector for effective encoding. Depending on the embodiment, the transformation may be performed using a lifting transformation, a wavelet transformation, etc. Additionally, quantization can be performed on the transformed displacement vector values, i.e., the transformation coefficients. Different quantization parameters can be applied to each axis of the transformation coefficients, and the quantization parameters can be derived by the agreement between the encoder and decoder. The quantized displacement vector transformation coefficients calculated by the displacement vector processing unit can be encoded through the displacement vector video encoding unit or the displacement vector arithmetic encoding unit, depending on the compression method.
[0182] The displacement vector video encoding unit can pack displacement vector information that has undergone transformation and quantization into 2D images. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh. The generated displacement vector video can be encoded using a video compression codec. The generated displacement vector video bitstream is transmitted to the multiplexer.
[0183] In the displacement vector arithmetic encoding unit, for the quantized displacement vector transformation coefficients, if the displacement vector type is INTER type, inter-frame prediction can be performed. The inter-frame prediction process may be a process of calculating the residual value, which is the difference between the current transformation coefficient and the reference transformation coefficient. The displacement vector transformation coefficient or the residual value can be encoded through the arithmetic encoding process.
[0184] The displacement vector restored through the displacement vector restoration unit and the base mesh restored and subdivided through the base mesh restoration unit are restored through the mesh restoration unit, and the restored mesh contains restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0185] The attribute information (texture map) of the original mesh can be regenerated into the attribute information (texture map) for the restored mesh through the attribute information (texture map) video generation unit. Vertex-specific color information contained in the original mesh's texture map can be assigned to the texture coordinates of the restored mesh. The texture maps regenerated for each frame can be grouped by GoF unit to generate a texture map video.
[0186] The generated texture map video can be encoded using a video compression codec through the texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to the multiplexer.
[0187] The atlas encoding unit can encode the atlas, which is additional information required for the mesh decoding and rendering processes. The generated atlas bitstream is transmitted to the multiplexer.
[0188] The generated base mesh bitstream, displacement vector bitstream, texture map bitstream, and atlas bitstream can be multiplexed into a single bitstream and transmitted to the receiver via the transmitter. Alternatively, the generated base mesh bitstream, displacement vector bitstream, texture map bitstream, and atlas bitstream can be generated into a file with one or more track data or encapsulated into segments and transmitted to the receiver (decoder) via the transmitter.
[0189] The data input unit can receive the original mesh and / or original texture map ('attribute'). The mesh preprocessing unit can simplify the original mesh to generate a base mesh and fit it to generate a refined mesh. If the mesh encoding method is inter-prediction, the motion vector encoding unit can generate motion vectors (motion fields) by referencing the restored base mesh within a previously processed reference frame and encode them based on a motion estimation and compensation method. If the mesh encoding method is intra-prediction, the static mesh encoding unit can encode the base mesh within the frame. The displacement vector calculation unit can calculate displacement vectors for vertices from the fitted refined mesh based on the restored base mesh. The displacement vector processing unit can process the displacement vectors into a form suitable for encoding. Depending on the encoding method for the displacement vectors, the displacement vectors can be encoded based on a video method or an arithmetic encoding method. The displacement vectors are restored and can be provided to the mesh restoration unit along with the restored base mesh. Based on the restored mesh, the attribute (texture map) video generation unit can generate a video for encoding the texture map using the original mesh and the texture map for the original mesh. The attributes are encoded based on the video method. The atlas is encoded by the atlas encoding unit.
[0190] FIG. 13 illustrates a V-DMC decoding process according to embodiments.
[0191] FIG. 13 corresponds to a decoder such as FIG. 1 to FIG. 3. Each component of FIG. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0192] The received Mesh bitstream is demultiplexed into a compressed base mesh bitstream, displacement vector bitstream, attribute information (texture map) bitstream, and atlas bitstream after file / segment decapsulation.
[0193] If the current mesh has inter-frame encoding applied based on the frame header information, decoding can be performed on the base mesh bitstream in the motion vector decoder. The final motion vector can be restored by using the previously decoded motion vector as a predictor and adding it to the residual motion vector decoded from the bitstream. The current base mesh can be restored by adding the decoded motion vector to the reference base mesh.
[0194] If the current mesh has in-frame encoding applied based on the frame header information, the base mesh bitstream can restore the base mesh's connectivity information, vertex geometry, texture coordinates, normal information, etc., through the static mesh decoding unit.
[0195] In the base mesh restoration unit, a restored base mesh can be generated by performing inverse quantization on the decoded base mesh.
[0196] Depending on the encoding codec type, if the displacement vector bitstream is encoded through a video codec, it can be decoded using the video codec and then subjected to a reverse packing process. If encoded through arithmetic coding, arithmetic decoding can be performed through the displacement vector arithmetic decoding unit, and if inter-frame prediction is performed, the current displacement vector transformation coefficient can be generated by adding the residual value to the reference displacement vector transformation coefficient through inter-frame prediction.
[0197] The displacement vector restoration unit restores the displacement vector by applying the decoded displacement vector transformation coefficients to inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, an inverse transformation to the Cartesian coordinate system can be performed.
[0198] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and connection information of the texture coordinates can be generated. The subdivided restored base mesh can be combined with the restored displacement vectors to generate the final restored mesh.
[0199] The texture map bitstream can be decoded as a video bitstream using a video codec in the texture map video decoder. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of the corresponding vertex can be retrieved from the texture map using the texture coordinates of each vertex.
[0200] The atlas bitstream can be decoded through the atlas decoding unit.
[0201] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.
[0202] The decoder receives the encoded bitstream and decodes the base mesh, displacement vector, attributes, and atlas within the bitstream based on the parameter information (which may be referred to as signaling information, metadata, etc.) contained within the bitstream. The decoding process may follow the inverse of the encoding process. Based on the decoded atlas, a mesh is reconstructed from the reconstructed base mesh and the reconstructed displacement mesh. Based on the reconstructed mesh and the reconstructed attributes, the mesh can be rendered.
[0203] A point cloud data encoding device and method according to the embodiments can encode mesh data and transmit a bitstream containing the encoded mesh data. A point cloud data decoding device and method according to the embodiments can receive a bitstream containing mesh data and decode the mesh data. A point cloud data encoding / decoding method / device according to the embodiments may be referred to simply as a method / device according to the embodiments. A point cloud data encoding / decoding method / device according to the embodiments may also be referred to as a mesh data encoding / decoding method / device according to the embodiments. Additionally, it may be used in this document simply as an encoding / decoding method / device.
[0204] The encoding method / device (or transmission method / device) according to the embodiments may include, and can perform, an encoder of FIG. 1, a transmission device (100) of FIG. 2, an acquisition unit (101), an encoder (102), an encapsulator (103), a transmitter (104), a pre-processor (200) of FIG. 3 to 4, an encoder (201), an encoding process of FIG. 7 and FIG. 12, generation of an encoded dynamic bitstream of FIG. 17, structuring of a bitstream according to the V3C unit payload syntax of FIG. 18, displacement vector encoding and transform coefficient processing of FIG. 19 to 29 (including MCTS / subpicture-based packing area designation), generation of an extraction information SEI payload and / or mapping information SEI payload (including tile-submesh mapping and video codec tile index signaling) of FIG. 35 to 39, and an encoding method of FIG. 44.
[0205] The decoding method / device (or receiving method / device) according to the embodiments may include and perform the decoder of FIG. 1, the receiving device (110) of FIG. 2, the receiving unit (111), the decapsulator (112), the decoder (113), the renderer (114), the decoding process of FIG. 11 and FIG. 13, the dynamic bitstream acquisition and parsing of FIG. 17, the dynamic mesh decoding and displacement vector coordinate system inverse transformation / decoding / lifting inverse transformation of FIG. 30 to FIG. 34, the extraction information SEI message and / or mapping information SEI message parsing of FIG. 35 to FIG. 39, the receiver operation (LoD extraction, attribute extraction, target atlas tile decoding, video tile decoding) of FIG. 40 to FIG. 43, and the decoding method of FIG. 45.
[0206] The contents of this specification may be understood based on V-DMC standard specification documents, including ISO / IEC 23090-29, disclosed at the time of the priority date or filing date of this application.
[0207] The encoding and encoder (or encoding unit) according to the embodiments may be interpreted as having the same meaning as encoding and encoder, respectively. The decoding and decoder (or decoder) according to the embodiments may be interpreted as having the same meaning as decoding and decoder, respectively.
[0208] Embodiments according to the present invention relate to Video-based Dynamic Mesh Compression (V-DMC), a method for compressing three-dimensional dynamic mesh data using an existing 2D video codec, and the embodiments may signal information related thereto as a Supplemental Enhancement Information (SEI) message to support the function of partially decoding a mesh to a sub-mesh or a specific Level of Detail (LoD) in a V-DMC encoder / decoder.
[0209] However, according to the conventional V-DMC standard document (based on N01099), SEI messages supporting the function of partially decoding attributes of a specific sub-mesh, SEI messages supporting the function of partially decoding a specific sub-mesh, and SEI messages supporting the function of partially decoding a mesh up to a specific LoD are classified as independent SEI messages, but there may be information that is redundantly signaled between SEI messages. The embodiments include or perform a method of signaling SEI messages by taking into account the redundant information between SEI messages, thereby reducing redundant signaling and decreasing the number of bits.
[0210] The embodiments relate to Video-based Dynamic Mesh Compression (V-DMC), a method for compressing three-dimensional dynamic mesh data using an existing 2D video codec, and the embodiments can generate or obtain SEI message syntax and semantics information that consider redundancy information of SEI messages used independently to partially decode attribute information of a specific sub-mesh (e.g., texture map, etc.) or a mesh corresponding to a specific LoD, or to partially decode a video tile mapped to a specific atlas tile.
[0211] Recently, V-DMC technology has been undergoing standardization since the CfP Response in April 2022 and is currently in the DIS phase as of February 2025. In the V-DMC standard document (N01099, Technologies for Video-based mesh coding), LoD extraction information SEI payload syntax, Tile submesh mapping SEI payload syntax, and Attribute extraction information SEI payload syntax are defined as SEI messages related to mapping information, so that mapping information necessary to support partial decoding at the submesh, LoD, or atlas tile level can be signaled in the form of SEI messages.
[0212] According to the above V-DMC standard document (based on N1099), the LoD extraction information SEI message is configured to support LoD-based geometry video extraction based on Motion-constrained Tile Sets (MCTS) index and / or subpicture index information mapped by refinement level among the packed regions of the displacement video.
[0213] According to the above V-DMC standard document (based on N1099), the Attribute extraction information SEI message is configured to support submesh-based attribute video extraction based on MCTS index and / or subpicture index information mapped per attribute patch or submesh.
[0214] According to the above V-DMC standard document (based on N1099), the Tile submesh mapping SEI message is configured to support atlas tile-based decoding and / or parallel processing based on submesh information mapped per atlas tile and video tile information mapped per atlas tile.
[0215] Table 1 below is a table comparing and analyzing the commonalities and differences of SEI messages in V-DMC standard documents (based on N1099) related to data matching information.
[0216] SEI message Condition LoD extraction information Attribute extraction information Tile submesh mapping Relevance to Geometry (displacement video) X Relevance to Attribute video X Whether displacement rectangular packing is performed O Functionality Submesh-based LoD partial extraction Submesh-based attribute partial extraction Atlas tile-based partial decoding Mapping information Mapping information between MCTS or subpicture index and LoD level Mapping information between MCTS or subpicture index and submesh ID Submesh ID information within atlas tile Mapping information between video tile and atlas tile 1 atlas tile - 1 submesh Submesh ID information mapped to video tile is known 1 atlas tile - Multiple submesh Submesh ID information mapped to video tile is unknown Multiple LoD required--
[0217] Referring to Table 1, there may be similarities and differences in the information handled among the SEI messages related to mapping information.
[0218] Information related to geometry video (or displacement video) can be handled in LoD extraction information SEI messages and Tile submesh mapping SEI messages.
[0219] Information related to attribute video can be handled in the Attribute extraction information SEI message and the Tile submesh mapping SEI message.
[0220] In the displacement video packing process, whether rectangular packing is performed may be irrelevant to other SEI messages, whereas rectangular packing may be required to extract rectangular regions per LoD in LoD extraction information SEI messages.
[0221] The LoD extraction information SEI message can be configured to support the function of decoding a mesh up to a specific LoD at the sub-mesh level. The Attribute extraction information SEI message can be configured to support the function of decoding attributes at the sub-mesh level. The Tile submesh mapping SEI message can be configured to support the function of decoding video tiles at the atlas tile level.
[0222] LoD extraction information may correspond to an MCTS index or subpicture index for the LoD region of the displacement video mapped to the submesh. Attribute extraction information may correspond to an MCTS index or subpicture index for the attribute video region mapped to the submesh. Tile submesh mapping information may correspond to submesh IDs existing within an atlas tile and a video tile index corresponding to that atlas tile.
[0223] LoD extraction information and Attribute extraction information can signal information corresponding to a sub-mesh ID. On the other hand, Tile submesh mapping information can only signal atlas tile information mapped to a video tile when there are two or more sub-meshes within an atlas tile, and as a result, it is difficult to know the video tile information corresponding to each sub-mesh, which may limit sub-mesh-based decoding. However, if there is a 1:1 correspondence between an atlas tile and a sub-mesh, sub-mesh unit decoding may be possible using the video tile information corresponding to the atlas tile.
[0224] LoD extraction information can provide the ability to decode a mesh into a specific LoD. On the other hand, Attribute extraction information and Tile submesh mapping information SEI messages can be configured independently of LoDs, which may limit LoD-based decoding.
[0225] Additionally, since the Tile submesh mapping SEI message signals the mapping relationship between atlas tiles and video tiles, if multiple submeshes exist within an atlas tile, it becomes difficult to map the video tile corresponding to each submesh, and a problem may arise where partial decoding for specific submeshes is limited.
[0226] SEI messages related to mapping information can be defined by including them in the Supplemental enhancement information of Annex F of the V-DMC standard document.
[0227] Figure 14 shows the conventional LoD extraction information SEI payload syntax.
[0228] The LoD extraction information of Fig. 14 may be defined in the V-DMC standard document (based on N1099) F.2.6 LoD extraction information SEI payload syntax.
[0229] Referring to FIG. 14, LoD_extraction_information(payloadSize) can be configured to signal information for partial extraction and partial decoding of packing regions corresponding to Level of Detail (LoD) levels per specific sub-mesh in the displacement vector bitstream.
[0230] The extractable unit type index (lei_extractable_unit_type_idx) indicates the type of extractable unit within the video bitstream. If the value of lei_extractable_unit_type_idx is 0, it indicates that the displacement video is encoded as Motion-Constrained Tile Sets (MCTS) as specified in ISO / IEC 23008-2, and if the value is 1, it indicates that the displacement video is encoded as a subpicture as specified in ISO / IEC 23090-3.
[0231] The value obtained by adding 1 to the number of submesh (lei_number_of_submesh_minus1) represents the number of submesh.
[0232] The submesh ID (lei_submesh_id[i]) represents the submesh ID of the i-th submesh.
[0233] The subdivision iteration count (lei_subdivision_iteration_count[i]) represents the number of subdivision iterations applied to the i-th submesh.
[0234] The MCTS index (lei_mcts_idx[i][j]) represents the identifier of the MCTS corresponding to the region within the video bitstream where the displacement data of the j-th refinement level of the i-th sub-mesh is located. lei_mcts_idx[i][j] may be identical to idx_of_mcts_in_set[m][n][l] in ISO / IEC 23008-2:2023:D.3.43. In this case, idx_of_mcts_in_set[m][n][l] represents the MCTS index of the l-th MCTS within the n-th MCTS set associated with the m-th extraction information set.
[0235] The subpicture index (lei_subpic_idx[i][j]) represents the identifier of the subpicture corresponding to the region where the j-th refinement level displacement data of the i-th submesh is located within the video bitstream.
[0236] Referring to FIG. 14, the LoD extraction information may include lei_extractable_unit_type_idx and lei_number_of_submesh_minus1, and may include lei_submesh_id[i] and lei_subdivision_iteration_count[i] of each submesh while iterating over the number of submeshes. Additionally, while iterating over each submesh to correspond to the number of LoD (or refinement level) levels, it may be structured to include either lei_mcts_idx[i][j] or lei_subpic_idx[i][j] depending on the value of lei_extractable_unit_type_idx.
[0237] The LoD extraction information SEI message can support the function of partially extracting and partially decoding only a portion of the bitstream corresponding to the LoD to be decoded among the regions packed at the submesh unit and LoD level unit within the displacement vector bitstream in order to restore the mesh up to a specific LoD.
[0238] In addition, the LoD extraction information SEI message can enable partial bitstream extraction and partial decoding up to a specific LoD by signaling the MCTS index or subpicture index corresponding to the packed area for each LoD within the displacement video per submesh.
[0239] A receiver (or decoder) checks for the existence of an LoD extraction information SEI message within an atlas bitstream, and if the SEI message exists, sequentially parses lei_extractable_unit_type_idx, lei_number_of_submesh_minus1, and lei_submesh_id[i] and lei_subdivision_iteration_count[i] corresponding to the number of submeshes, and then parses lei_mcts_idx[i][j] or lei_subpic_idx[i][j] according to lei_extractable_unit_type_idx, thereby obtaining mapping information for partial extraction / partial decoding at the submesh unit and LoD unit levels.
[0240] Figure 15 (Figures 15a and 15b) shows a conventional tile submesh mapping SEI payload syntax.
[0241] The tile submesh mapping of Fig. 15 may be defined in V-DMC standard document (based on N1099) F.2.7 Tile submesh mapping SEI payload syntax.
[0242] Referring to FIG. 15, tile_submesh_mapping(payloadSize) includes a descriptor and may include tmsm_persistance_mapping_flag, tmsm_num_tiles_minus1, tmsm_tile_id_length_minus1, and tmsm_codec_tile_signal_flag. Additionally, when tmsm_codec_tile_signal_flag is 1, it may further include tmsm_geo_codec_tile_alignment_flag and tmsm_attr_codec_tile_alignment_flag and be configured to initialize currCount[0] and currCount[1] to 0.
[0243] The tile persistence flag (tmsm_persistance_mapping_flag) can indicate whether the tile-submesh mapping is persistent, and if the value is 1, the mapping persists, and if the value is 0, it is valid only for the current frame.
[0244] The persistence range of a tile submesh mapping SEI message can be the remainder of the bitstream (i.e., until the end of the stream) or until a new tile submesh mapping SEI message appears, and if tmsm_persistance_mapping_flag is 1, the mapping defined in the previous SEI message can be maintained.
[0245] Additionally, if a tile submesh mapping SEI message exists in any access unit within a Coded Atlas Sequence (CAS), the above tile submesh mapping SEI message must also exist in the first access unit of the CAS, and the above tile submesh mapping SEI message may persist from the current access unit to the end of the CAS in the decoding order.
[0246] The value obtained by adding 1 to the tile count (tmsm_num_tiles_minus1) represents the number of tiles within CAS. tmsm_num_tiles_minus1 can be equal to TotalTileCount - 1. Here, TotalTileCount can be calculated by adding afti_num_tiles_in_atlas_frame_minus1 + 1 and accumulating aftai_num_tiles_in_atlas_frame_minus1[i] + 1 while iterating over AspsAttributeNominalFrameSizeCount. tmsm_num_tiles_minus1 + 1 can represent the total number of atlas tiles and attribute tiles.
[0247] The value obtained by adding 1 to the tile ID length (tmsm_tile_id_length_minus1) represents the number of bits used to represent tmsm_tile_id[i]. tmsm_tile_id_length_minus1 can range from 0 to 15.
[0248] The codec tile signal flag (tmsm_codec_tile_signal_flag) can indicate whether to signal the tile index information of the video codec corresponding to the atlas tile or attribute tile in this SEI message, and if the value is 1, it indicates that it is signaling, and if the value is 0, it indicates that it is not signaling.
[0249] The geometry codec tile alignment flag (tmsm_geo_codec_tile_alignment_flag) indicates whether the tile division structure in the geometry video codec is the same as the atlas tile division structure in V-DMC. If the value of tmsm_geo_codec_tile_alignment_flag is 1, it means that the division structures are the same in that each i-th tile contains only the area of the i-th codec tile, and when i and j are not the same, the i-th tile does not contain any area of the j-th codec tile. If the value of tmsm_geo_codec_tile_alignment_flag is 0, the tile division of the frame may not be the same as the tile division of the geometry sub-bitstream, and if omitted, it can be inferred as 0.
[0250] The attribute codec tile alignment flag (tmsm_attr_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the attribute video codec is the same as the attribute tile partitioning structure in V-DMC. If the value of tmsm_attr_codec_tile_alignment_flag is 1, it means that the partitioning structure is the same in that each i-th attribute tile contains only the area of the i-th codec tile, and when i and j are not the same, the i-th attribute tile does not contain any area of the j-th codec tile. If the value of tmsm_attr_codec_tile_alignment_flag is 0, the attribute tile partitioning of the frame may not be the same as the tile partitioning of the attribute sub-bitstream, and if omitted, it can be inferred as 0.
[0251] Tile ID (tmsm_tile_id[i]) represents the tile ID of the i-th tile. If omitted, it can be inferred that tmsm_tile_id[i]=i for each i from 0 to tmsm_num_tiles_minus1, and the bitstream conformance requirement that tmsm_tile_id[i] and tmsm_tile_id[k] must be different when i and k are not the same may apply.
[0252] The tile type flag (tmsm_tile_type_flag[i]) indicates whether the i-th tile is a geometry-related tile or an attribute-related tile. If the value of tmsm_tile_type_flag[i] is 0, it indicates that the tile is P_TILE or I_TILE and conveys geometry-related information, and if the value of tmsm_tile_type_flag[i] is 1, it indicates that the tile is P_TILE_ATTR or I_TILE_ATTR and conveys attribute-related information.
[0253] The value obtained by adding 1 to the number of submeshes (tmsm_num_submeshes_minus1[i]) represents the number of submeshes included in the i-th tile (e.g., the atlas tile corresponding to tile ID i). tmsm_num_submeshes_minus1[i] can range from 0 to 63.
[0254] The value obtained by adding 1 to the submesh ID length (tmsm_submesh_id_length_minus1[i]) represents the number of bits used to represent tmsm_submesh_id[i][j]. The value of tmsm_submesh_id_length_minus1[i] can range from 0 to 15, and if omitted, it can be inferred as the same value as Ceil(Log2(NumSubMeshes) - 1).
[0255] The submesh ID (tmsm_submesh_id[i][j]) may represent the submesh ID of the j-th submesh associated with tile index i, and if omitted, the value of tmsm_submesh_id[i][j] may be inferred as j for each j from 0 to NumSubMeshes-1, and a bitstream conformity requirement may apply that tmsm_submesh_id[i][j] and tmsm_submesh_id[i][k] must be different when j and k are not the same within the same tile i. The length of tmsm_submesh_id[i][j] may be tmsm_submesh_id_length_minus1[i] + 1 bit.
[0256] The value obtained by adding 1 to the number of codec tiles (tmsm_num_codec_tiles_in_tile_minus1[i]) represents the number of codec tiles included in the i-th tile. If omitted, it can be inferred to be 0 for each i from 0 to tmsm_num_tiles_minus1.
[0257] The codec tile index (tmsm_codec_tile_idx[i][j]) can represent the index of the j-th codec tile included in the i-th tile, and if the value does not exist in the syntax, tmsm_codec_tile_idx[i][0] can be inferred according to the rules listed in the syntax table.
[0258] The codec tile index reset flag (tmsm_codec_tile_idx_reset_flag[i]) indicates whether to initialize the variable currCount[1], which is used for inferring tmsm_codec_tile_idx[i][j], to 0. If the value of tmsm_codec_tile_idx_reset_flag[i] is 1, currCount[1] is reset to 0, and if the value of tmsm_codec_tile_idx_reset_flag[i] is 0, it may not be reset.
[0259] Referring to FIG. 15, tile_submesh_mapping(payloadSize) may include tmsm_tile_id[i], tmsm_tile_type_flag[i], tmsm_num_submeshes_minus1[i], tmsm_submesh_id_length_minus1[i], and tmsm_submesh_id[i][j] for each tile while iterating over the number of tiles (tmsm_num_tiles_minus1+1), and TileIdxToID[i] and SubmeshIdxToID[i][j] may be set to tmsm_tile_id[i] and tmsm_submesh_id[i][j], respectively.
[0260] In addition, when tmsm_codec_tile_signal_flag is 1, (1) it is a geometry-related tile and tmsm_geo_codec_tile_alignment_flag is 0, or (2) it is an attribute-related tile and tmsm_attr_codec_tile_alignment_flag is 0, tmsm_num_codec_tiles_in_tile_minus1[i] and a plurality of tmsm_codec_tile_idx[i][j] can be signaled.
[0261] In addition, if at least one of tmsm_geo_codec_tile_alignment_flag or tmsm_attr_codec_tile_alignment_flag is 1, tmsm_codec_tile_idx[i][0] is inferred to currCount[tmsm_tile_type_flag[i]] according to the tile type, and then the inference can be performed in such a way that currCount[tmsm_tile_type_flag[i]] is increased by 1, and in particular, if it is an attribute-related tile and currCount[1] is not 0, the inference can be performed after initializing currCount[1] to 0 according to tmsm_codec_tile_idx_reset_flag[i].
[0262] The tile submesh mapping SEI message defines the mapping between the atlas tiles signaled in the AFPS and the submeshes of the basemesh sub-bitstream, and can be configured to transmit a list of associated submesh IDs for each atlas tile ID, and the submesh IDs of the basemesh sub-bitstream components may be unique.
[0263] Additionally, the tile submesh mapping SEI message can be used to signal an initial mapping to the decoder at the beginning of the V-DMC bitstream and to signal a subsequent mapping update when the association between the atlas tile and the submesh ID changes.
[0264] The tile submesh mapping SEI message can represent not only atlas tiles and submesh mapping information within those tiles, but also mapping information between atlas tiles and video tiles, and can support atlas tile-unit partial decoding functions through video tile information corresponding to a specific atlas tile.
[0265] The receiver (or decoder) can check for the existence of a tile-submesh mapping SEI message within the atlas sub-bitstream, and if the SEI message exists, sequentially parse tmsm_persistance_mapping_flag, tmsm_num_tiles_minus1, tmsm_tile_id_length_minus1, tmsm_codec_tile_signal_flag (and alignment flags if necessary), and construct a tile-submesh mapping table by parsing (or inferring) tmsm_tile_id[i] and tmsm_submesh_id[i][j] from the tile iteration phrase.
[0266] In addition, when tmsm_codec_tile_signal_flag is 1, tile-codec tile mapping information that can be used for atlas tile unit part decoding, etc. can be obtained by parsing tmsm_codec_tile_idx[i][j] according to tile type and alignment flag or by inferring tmsm_codec_tile_idx[i][0] based on currCount.
[0267] Figure 16 shows the conventional attribute extraction information SEI payload syntax.
[0268] The attribute extraction information of Fig. 16 may be defined in V-DMC standard document (based on N1099) F.2.8 Attribute extraction information SEI payload syntax.
[0269] Referring to FIG. 16, attribute_extraction_information(payloadSize) can be configured to signal information for partial extraction and partial decoding of only a portion of the bitstream corresponding to a specific submesh among the regions packed in submesh units in the attribute bitstream.
[0270] In addition, the attribute extraction information SEI message can support submesh-based attribute video extraction based on MCTS or subpicture index information mapped to each attribute patch (submesh).
[0271] Referring to FIG. 16, attribute_extraction_information(payloadSize) may include a cancellation flag (aei_cancel_flag), and may be configured to be followed by attribute extraction information when the value of the cancellation flag (aei_cancel_flag) is 0.
[0272] If the value of the cancellation flag (aei_cancel_flag) is 1, the SEI message may indicate that the persistence of a previously existing attribute extraction information SEI message in the output order is canceled.
[0273] Referring to FIG. 16, when the value of the cancellation flag (aei_cancel_flag) is 0, the attribute extraction information SEI message may include an extraction unit type index (aei_extractable_unit_type_idx), a number of submesh (aei_number_of_submesh_minus1), a number of attributes (aei_attribute_count), a submesh ID (aei_submesh_id[i]), an extraction information present flag (aei_extraction_info_present_flag[j]), and an MCTS index (aei_mcts_idx[i][j]) or a subpicture index (aei_subpic_idx[i][j]).
[0274] The extractable unit type index (aei_extractable_unit_type_idx) indicates the type of extractable unit within the video bitstream. If the value of the extractable unit type index (aei_extractable_unit_type_idx) is 0, it indicates that the attribute video is encoded as Motion-Constrained Tile Sets (MCTS) as specified in ISO / IEC 23008-2, and if the value is 1, it indicates that the attribute video is encoded as a subpicture as specified in ISO / IEC 23090-3.
[0275] The value obtained by adding 1 to the number of submesh (aei_number_of_submesh_minus1) represents the number of submesh.
[0276] The submesh ID (aei_submesh_id[i]) represents the submesh ID of the i-th submesh.
[0277] The attribute count (aei_attribute_count) represents the number of attributes associated with the basemesh, and the value of the attribute count (aei_attribute_count) can be in the range of 0 to 127.
[0278] The extraction info present flag (aei_extraction_info_present_flag[j]) indicates whether extraction info exists for an attribute video corresponding to a submesh for an attribute signaled in the atlas attribute nominal frame at index j in the meshpatch data unit. If the value of the extraction info present flag (aei_extraction_info_present_flag[j]) is 0, it indicates that extraction info does not exist, and if the value is 1, it indicates that extraction info exists.
[0279] When the value of the extraction information present flag (aei_extraction_info_present_flag[j]) is 1, the MCTS index (aei_mcts_idx[i][j]) or subpic index (aei_subpic_idx[i][j]) may be signaled according to the value of the extraction unit type index (aei_extractable_unit_type_idx).
[0280] The MCTS index (aei_mcts_idx[i][j]) represents the identifier of the MCTS corresponding to the region where the j-th attribute video data of the i-th submesh in the attribute bitstream is located. The MCTS index (aei_mcts_idx[i][j]) may be identical to idx_of_mcts_in_set[m][n][l] of ISO / IEC 23008-2:2023:D.3.43, and idx_of_mcts_in_set[m][n][l] may represent the MCTS index of the l-th MCTS in the n-th MCTS set associated with the m-th extraction information set.
[0281] The subpicture index (aei_subpic_idx[i][j]) represents the identifier of the subpicture corresponding to the region where the video data of the j-th attribute of the i-th submesh is located within the attribute bitstream.
[0282] Referring to FIG. 16, attribute_extraction_information(payloadSize) may include the submesh ID (aei_submesh_id[i]) of each submesh while iterating for the number of submeshes (aei_number_of_submesh_minus1+1), and may include the extraction information present flag (aei_extraction_info_present_flag[j]) while iterating for the number of attributes (aei_attribute_count) for each submesh.
[0283] In addition, if the value of the extraction information present flag (aei_extraction_info_present_flag[j]) is 1, if the value of the extraction unit type index (aei_extractable_unit_type_idx) is 0, the MCTS index (aei_mcts_idx[i][j]) is included, and if the value of the extraction unit type index (aei_extractable_unit_type_idx) is 1, the subpicture index (aei_subpic_idx[i][j]) is included.
[0284] When the associated atlas frame is defined as aFrmA, the attribute extraction information SEI message can retain its meaning in the output order of the current layer.
[0285] The persistence may be maintained until at least one of the following occurs: when a new CAS (Coded Atlas Sequence) of the current layer starts, when the bitstream ends, or when an atlas frame aFrmB output after aFrmA in the current layer satisfies the condition that AtlasFrmOrderCnt(aFrmB) is greater than AtlasFrmOrderCnt(aFrmA), and a frame within a coded atlas access unit containing an attribute extraction information SEI message applicable to the current layer is output.
[0286] The receiver (or decoder) checks for the existence of an attribute extraction information SEI message within the atlas bitstream, and if such an SEI message exists, it can parse the cancellation flag (aei_cancel_flag).
[0287] If the value of the cancellation flag (aei_cancel_flag) is 0, the receiver (or decoder) sequentially parses the extractable unit type index (aei_extractable_unit_type_idx), the number of submeshes (aei_number_of_submesh_minus1), the number of attributes (aei_attribute_count), the submesh ID (aei_submesh_id[i]), and the extract information present flag (aei_extraction_info_present_flag[j]), and if the value of the extract information present flag (aei_extraction_info_present_flag[j]) is 1, the MCTS index (aei_mcts_idx[i][j]) or the subpicture index (aei_subpic_idx[i][j]) is additionally parsed according to the extractable unit type index (aei_extractable_unit_type_idx), thereby obtaining mapping information for partial extraction and partial decoding of attribute bitstreams at the submesh level. there is.
[0288] In addition, the receiver (or decoder) can derive an MCTS index or subpicture index corresponding to a specific submesh and a specific attribute using the acquired mapping information, and recover the attribute of the specific submesh by extracting only a portion of the bitstream of the region corresponding to the derived index and then decoding it.
[0289] FIG. 17 shows an encoded dynamic bitstream structure according to embodiments.
[0290] The encoding device / method illustrated in FIGS. 1 to 4, FIGS. 7, FIGS. 12, FIGS. 17, FIGS. 18, FIGS. 19, FIGS. 44, etc., according to embodiments can encode mesh data and generate a bitstream containing the encoded mesh data according to the bitstream structure illustrated in FIG. 17.
[0291] The decoding device / method illustrated in FIG. 1, FIG. 2, FIG. 11, FIG. 13, FIG. 30, FIG. 40, FIG. 45, etc., according to embodiments can acquire a bitstream and decode mesh data based on the bitstream structure and parameter information within the bitstream illustrated in FIG. 17.
[0292] The bitstream of FIG. 17 according to the embodiments may include the V3C unit payload syntax of FIG. 18 and the mapping information SEI payload syntax of FIG. 35 to FIG. 39.
[0293] The embodiments may include or perform a method of aggregating SEI messages into one by considering redundant information in the syntax of independent SEI messages related to mapping information among the SEI messages of V-DMC. The embodiments may reduce the number of signaled bits by omitting redundant syntax between SEI messages through such aggregation.
[0294] Specifically, the technical problem according to the embodiments may be that, as the LoD extraction information SEI message (F.2.6, FIG. 14), tile submesh mapping SEI message (F.2.7, FIG. 15), and attribute extraction information SEI message (F.2.8, FIG. 16) are defined independently of each other in the V-DMC standard document (based on N1099), the same or similar common information is repeatedly included in multiple SEI messages, thereby increasing signaling overhead.
[0295] As described above, the LoD extraction information of FIG. 14 includes the number of submesh (lei_number_of_submesh_minus1) and the submesh ID (lei_submesh_id[i]), and the attribute extraction information of FIG. 16 also includes the number of submesh (aei_number_of_submesh_minus1) and the submesh ID (aei_submesh_id[i]), and the tile submesh mapping of FIG. 15 can be configured to signal the number of submesh (tmsm_num_submeshes_minus1[i]) and the submesh ID (tmsm_submesh_id[i][j]) within the tile iteration syntax. Additionally, FIGS. 14 and FIGS. 16 may be configured to signal MCTS or subpicture-based extraction units through extraction unit type indices (lei_extractable_unit_type_idx, aei_extractable_unit_type_idx), respectively, and to repeatedly signal MCTS indices or subpicture indices corresponding to LoD levels or attribute indices in submesh units.
[0296] As such, since identification information such as the number of submeshes and submeshes ID, the submeshes repetition phrase range, and the representation methods of the extraction unit type and extraction area index are duplicated across multiple SEI messages, a technical problem may arise in which a receiver or decoder must parse and interconnect multiple SEI messages respectively to perform LoD-based partial extraction, submeshes-based attribute partial extraction, and tile-submeshes mapping.
[0297] As a technical means to solve the above technical problem, the mapping information SEI message according to the embodiments can be configured to aggregate common information that was repeatedly signaled in the multiple SEI messages of FIGS. 14 to 16 into a single payload structure and to selectively include only the extension information required for each function through flag-based conditional branching. Specifically, the mapping information SEI message can provide the basic mapping relationship between tiles and submeshes with a single signaling by including a tile persistence flag, the number of tiles and tile ID length, tile ID, tile type, and the number of submeshes and submeshe IDs associated with the tile.
[0298] In addition, according to the embodiments, when supporting partial extraction of LoD-based displacement video, the inclusion of LoD-related syntax is controlled by a LoD mapping signal flag, and the information that was distributed and transmitted as lei_number_of_submesh_minus1, lei_submesh_id[i], lei_subdivision_iteration_count[i], and lei_mcts_idx[i][j] / lei_subpic_idx[i][j] of FIG. 14, including the number of subdivision iterations per submesh and the extraction area index per LoD level, can be provided in a form that matches the tile-submesh iteration syntax.
[0299] In addition, according to the embodiments, when supporting partial extraction of submesh-based attribute video, the inclusion of attribute-related syntax is controlled by an attribute mapping signal flag, and the information that was distributed and transmitted as aei_number_of_submesh_minus1, aei_submesh_id[i], aei_attribute_count, and aei_mcts_idx[i][j] / aei_subpic_idx[i][j] of FIG. 16, including the number of attributes and an attribute extraction area index per submesh, can be provided within the same mapping framework.
[0300] Furthermore, according to the embodiments, if video tile unit decoding corresponding to an atlas tile or submesh is supported, the codec tile signal flag and codec tile alignment flag are included, and if necessary, the codec tile index is included, thereby allowing information related to tmsm_codec_tile_signal_flag and tmsm_codec_tile_idx of FIG. 15 to be provided within the same SEI message.
[0301] According to the embodiments, the aggregation structure described above can provide a technical solution in which syntaxes previously defined separately are rearranged into a common index system (an iterative syntax of tile index and submesh index), common information is signaled only once at a higher level, and purpose-specific information is included by conditional branching based on flags. Accordingly, the redundant signaling information shown in FIGS. 14 to 16 is integrated into a single mapping information SEI message or extraction information SEI message, thereby reducing the number of signaling bits and reducing processing complexity caused by redundant parsing and cross-referencing between multiple SEI messages. The mapping information SEI message or extraction information SEI message structure according to specific embodiments, and the operation of the encoder / decoder accordingly, etc., are described below.
[0302] According to the embodiments, V-DMC may also be referred to as V-Mesh below, and this may be used with an equivalent meaning.
[0303] According to the embodiments, dynamic mesh content can be encoded into a bitstream structure such as FIG. 17.
[0304] Referring to FIG. 17, a sample stream data unit used when encoding V3C content of the V3C codec standard (ISO / IEC 23090-5) can be used to configure the bitstream of dynamic mesh content.
[0305] Each abbreviation used in Fig. 17 may mean the following.
[0306] VPS can refer to the V3C / V-DMC parameter set.
[0307] AD can stand for atlas data.
[0308] BMD can stand for base mesh data.
[0309] DD can mean displacement data and can correspond to cases where displacement data is encoded by arithmetic coding.
[0310] GVD can mean geometry video data and can correspond to cases where displacement data is encoded using a video codec.
[0311] AVD can mean attribute video data and can be encoded by video coding.
[0312] PVD can mean packing video data and can be encoded by video coding.
[0313] The bitstream of FIG. 17 may be composed of or divided into an atlas sub-bitstream containing atlas data, a basemesh sub-bitstream containing basemesh data, a displacement sub-bitstream containing displacement data, and an attribute sub-bitstream containing attribute data, so as to correspond to each component. The atlas sub-bitstream, basemesh sub-bitstream, displacement sub-bitstream, and attribute sub-bitstream may correspond to the atlas bitstream, basemesh bitstream, displacement vector bitstream, and attribute information bitstream of FIG. 1, respectively.
[0314] The bitstream of FIG. 17 may include a video sub-bitstream. The video sub-bitstream may be configured to transmit at least one of geometry video data (GVD), attribute video data (AVD), and packing video data (PVD).
[0315] FIG. 18 shows the V3C unit payload syntax according to the embodiments.
[0316] Referring to FIG. 18, v3c_unit_payload(numBytesInV3CPayload) can be defined as a syntax structure that constructs a payload by receiving the length of the V3C unit payload as numBytesInV3CPayload.
[0317] vuh_unit_type indicates the unit type signaled in the V3C unit header. Conditional branching can be performed so that the syntax structure included in the payload varies depending on the unit type.
[0318] If vuh_unit_type is V3C_VPS, the payload can be configured to include v3c_parameter_set(numBytesInV3CPayload).
[0319] If vuh_unit_type is V3C_AD or V3C_CAD, the payload can be configured to include atlas_sub_bitstream(numBytesInV3CPayload).
[0320] If vuh_unit_type is any one of V3C_OVD, V3C_GVD, V3C_AVD, or V3C_PVD, the payload can be configured to include video_sub_bitstream(numBytesInV3CPayload).
[0321] V3C_VPS indicates that the V3C unit type is a V3C parameter set, V3C_AD indicates that the V3C unit type is Atlas data, V3C_OVD indicates that the V3C unit type is Occupancy video data, V3C_GVD indicates that the V3C unit type is Geometry video data, V3C_AVD indicates that the V3C unit type is Attribute video data, V3C_PVD indicates that the V3C unit type is Packed video data, and V3C_CAD indicates that the V3C unit type is Common atlas data.
[0322] FIG. 19 shows a dynamic mesh encoder configuration according to embodiments.
[0323] According to the embodiments, the dynamic mesh encoder of FIG. 19 receives the original mesh as input, generates and encodes base mesh data, displacement data, and attribute data, and can configure the encoded result into a bitstream.
[0324] The mesh simplification unit can take the original mesh as input and generate a simplified base mesh.
[0325] The mesh simplification unit can perform simplification of the input mesh based on the number of target vertices or the number of target faces.
[0326] The mesh parameterization unit can perform parameterization to generate vertex-specific texture coordinates (UV coordinates) and texture connection information of the input mesh.
[0327] The mesh subdivision unit can generate additional vertices by performing subdivision on the base mesh. According to the embodiments, geometric connectivity information, texture coordinate connectivity information, and texture coordinates can be implicitly derived and generated depending on the subdivision method.
[0328] Mesh segmentation can be performed n times by user parameters or encoder / decoder promises, and the vertices of the base mesh , the newly generated vertices by performing subdivision 1 , vertices generated by performing subdivision n times When defined as It can be defined as follows.
[0329]
[0330] The mesh fitting unit can perform vertex position adjustments so that the subdivided mesh becomes similar to the original mesh. According to the embodiments, the mesh simplification unit, mesh parameterization unit, mesh subdivision unit, and mesh fitting unit may be omitted. When these processes are omitted, the original mesh may be applied as input to the mesh quantization unit, and the displacement vector encoding process (displacement vector calculation unit, displacement vector coordinate system transformation unit, displacement vector encoding unit) may be omitted.
[0331] The mesh quantization unit can quantize at least one of geometric information (x, y, z), texture coordinates (u, v), and normal information (nx, ny, nz) in floating-point form into a fixed-point form. According to the embodiments, quantization for specific components may be omitted.
[0332] The static mesh encoding unit can perform encoding on at least one of the base mesh's connectivity information, vertex geometry information, vertex texture coordinates, and normal information.
[0333] The motion vector encoding unit can perform motion vector encoding by calculating motion vectors using a reference restored base mesh and a current base mesh as inputs. The motion vector encoding unit can perform connectivity-based prediction using previously encoded / decoded motion vectors as predictors, and perform entropy encoding on the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. According to the embodiments, motion vector encoding can be performed at the vertex level or subgroup level.
[0334] The displacement vector calculation unit can calculate the displacement vector between the fitted subdivided mesh and the mesh obtained by performing subdivision on the restored current base mesh. Through the displacement vector calculation unit, a displacement vector corresponding to the number of vertices of the subdivided mesh can be calculated.
[0335] The displacement vector coordinate system transformation unit can transform a vertex displacement vector calculated in (x, y, z) space into a (normal, tangential, bi-tangential) coordinate system based on the normal vector of each vertex. According to the embodiments, only the normal component of the (normal, tangential, bi-tangential) coordinate system may be encoded, or it may be determined that only the normal component is encoded by an encoder / decoder agreement, or the encoder may determine this by signaling a 1-bit flag (onlyNormFlag).
[0336] According to the embodiments, the normal vector can be calculated for each subdivided vertex based on the geometric information or connectivity information of the surrounding vertices.
[0337] Whether to apply a displacement vector coordinate system transformation can be determined by an encoder / decoder agreement, or by transmitting a coordinate system transformation flag (applyLocalCoord) in units such as a sequence, GOF (Group of Frame), frame, or submesh.
[0338] The base mesh restoration unit can restore the base mesh according to the encoding type of the current mesh (inter-frame encoding or intra-frame encoding). When inter-frame encoding is performed, the current base mesh can be generated by adding the restored motion vector to the reference restored base mesh. According to the embodiments, if the motion vector is not quantized, the motion vector restoration process is omitted, and the current base mesh can be restored using the motion vector calculated by the motion vector encoding unit. According to the embodiments, when intra-frame encoding is performed, the current base mesh can be restored by performing inverse quantization on the base mesh quantized through the mesh quantization unit.
[0339] The base mesh inverse quantization unit can perform inverse quantization on at least one of the reconstructed geometric information (x, y, z), texture coordinates (u, v), and normal information (nx, ny, nz) of the reconstructed base mesh. According to the embodiments, inverse quantization for a specific component may be omitted.
[0340] The displacement vector recovery unit can decode a bitstream that has been packed into a 2D image or video and encoded through a 2D video encoder using a 2D video decoder, and perform inverse packing. The displacement vector recovery unit can calculate a recovered displacement vector by performing inverse quantization and inverse transform on the quantization transform coefficients obtained as a result of inverse packing.
[0341] The mesh restoration unit can generate subdivided vertex position information, texture coordinates, and connectivity information by performing subdivision on the restored base mesh dequantized by the base mesh dequantization unit. Additionally, the mesh restoration unit can generate restored vertex position information by adding a restoration displacement vector to the subdivided vertex position information.
[0342] The texture map generation unit can generate a texture map of the restored mesh based on the texture coordinates and connection information of the restored mesh and the relationship between the original mesh and the texture map of the original mesh.
[0343] The texture map encoding unit can construct a texture map video by stacking the texture maps generated by the texture map generation unit in the order of mesh frames, and perform encoding through a 2D video encoder.
[0344] According to the embodiments, if the color space of the texture map is RGB444, encoding can be performed after converting to the YUV420 or YUV444 color space.
[0345] FIGS. 20 to 22 show the configuration of a displacement vector encoding unit according to embodiments.
[0346] According to the embodiments, the displacement vector encoding unit of FIGS. 20 to 22 may be configured as the dynamic mesh encoder of FIG. 19.
[0347] The displacement vector encoding unit can encode the displacement vector using a 2D video encoder such as H.264, HEVC, or VVC, or can encode the displacement vector using a zero run-length encoder or an arithmetic encoder.
[0348] According to the embodiments, the displacement vector encoding method may be determined as a specific encoding method by an agreement between the encoder and the decoder, or the encoding method determined by the encoder may be transmitted to the decoder by signaling it as a flag or index (dispEncType).
[0349] According to the embodiments, the displacement vector encoding method may be one of a method combining a video codec-based encoding method and a zero-run length encoding method, a method combining a video codec-based encoding method and an arithmetic encoding method, a video codec-based encoding method, a zero-run length encoding method, and an arithmetic encoding method.
[0350] According to the embodiments, a displacement vector encoding method can be determined according to a profile defined in an encoder and a decoder, and according to the embodiments, the encoder can signal an index (profileToolsetIdx) representing profile information, and the decoder can determine a displacement vector decoding method according to profileToolsetIdx.
[0351] Referring to FIG. 20, when the displacement vector encoding method is a video codec-based encoding method, the displacement vector transformation coefficient packing unit can perform the process of packing the quantized displacement vector transformation coefficients into a single frame according to the Level of Detail (LoD) level through the displacement vector transformation coefficient quantization unit. The areas packed according to the LoD level can be arranged in a raster-scan order or in a rectangular shape. The packing method can be determined by deriving it from the same rule in the encoder and decoder, or the encoder can signal the packing method so that the decoder determines the packing method.
[0352] Referring to FIG. 21, when the displacement vector encoding method is a zero-run length encoding method, the displacement vector transformer can transform the displacement vector to generate displacement vector transform coefficients. The displacement vector transform coefficient quantizer can perform quantization on the displacement vector transform coefficients to generate quantized displacement vector transform coefficient levels. The restored displacement vector transform coefficient level buffer can store previously encoded displacement vector transform coefficient levels. The displacement vector transform coefficient level prediction unit can predict the current displacement vector transform coefficient level using level information stored in the restored displacement vector transform coefficient level buffer. The displacement vector transform coefficient zero-run length encoding unit can represent consecutive zero-level intervals among the residual levels based on the prediction result as run information, and generate an encoded bit sequence by combining the run information and non-zero level values. According to the embodiments, the zero-run length based encoding method can be selected by dispEncType or profileToolsetIdx.
[0353] Referring to FIG. 22, when the displacement vector encoding method is an arithmetic encoding method, the displacement vector transformation unit can generate displacement vector transformation coefficients by transforming the displacement vector. The displacement vector transformation coefficient quantization unit can generate quantized displacement vector transformation coefficient levels by performing quantization on the displacement vector transformation coefficients. The restored displacement vector transformation coefficient level buffer can store previously encoded displacement vector transformation coefficient levels. The displacement vector transformation coefficient level prediction unit can predict the current displacement vector transformation coefficient level using the level information stored in the restored displacement vector transformation coefficient level buffer. The displacement vector transformation coefficient arithmetic encoding unit can perform probabilistic modeling on residual levels based on the prediction result and perform arithmetic encoding based on the corresponding probabilistic model. According to the embodiments, the arithmetic encoding-based encoding method can be selected by dispEncType or profileToolsetIdx.
[0354] According to the embodiments, when video coding is applied as in FIG. 20, the packing area by LoD level may have different division units depending on the type of video codec. Specifically, when the video codec is HEVC, the packing area by LoD level may be divided into MCTS (Motion-Constrained Tile Sets), and when the video codec is VVC, the packing area by LoD level may be divided into subpictures.
[0355] FIGS. 23 and 24 show examples of designating the quantized displacement vector transformation coefficient packing region per LoD according to embodiments as MCTS.
[0356] Referring to FIGS. 23 and 24, when video encoding is performed on frames in which displacement vector transformation coefficients are packed in rectangular areas for each refinement level, the areas can be separated into MCTS.
[0357] According to the embodiments, the refinement level may refer to vertices in the LoD N mesh excluding vertices corresponding to the LoD N-1 mesh. For example, LoD 0 may be vertices composed of refinement level 0, LoD 1 may be vertices composed of refinement level 0 and refinement level 1, and LoD 2 may be vertices composed of refinement level 0, refinement level 1, and refinement level 2.
[0358] As shown in Fig. 23, when MCTS areas are designated by LoD level, LoD 0 can be designated as MCTS 2, LoD 1 as MCTS 1, and LoD 2 as MCTS 0.
[0359] As shown in Fig. 24, when MCTS areas are designated by refinement level, refinement level 0 may be designated as MCTS 2, refinement level 1 may be designated as MCTS 1, and refinement level 2 may be designated as MCTS 0.
[0360] FIG. 25 shows an example of specifying a packing area for quantized displacement vector transformation coefficients per LoD according to embodiments.
[0361] According to the embodiments, the MCTS or subpicture and the LoD level may be set to a one-to-one correspondence, or a one-to-one correspondence may not be established.
[0362] FIG. 25 is an example of a case where MCTS or a subpicture and a LoD level are set in a one-to-one correspondence relationship. Referring to FIG. 25, refinement level 3 can be designated as subpicture 0, refinement level 2 can be designated as subpicture 1, refinement level 1 can be designated as subpicture 2, and refinement level 0 can be designated as subpicture 3.
[0363] FIG. 26 shows an example of specifying a packing area for quantized displacement vector transformation coefficients per LoD according to embodiments.
[0364] FIG. 26 is an example of a case where the MCTS or subpicture and the LoD level are not set to a one-to-one correspondence. Referring to FIG. 26, refinement level 3 can be designated as subpicture 0, refinement level 2 can be designated as subpicture 1, and refinement levels 1 and 0 can be designated as subpicture 2.
[0365] FIGS. 27 and 28 illustrate a method for predicting the current displacement vector transformation coefficient based on the restored displacement vector transformation coefficient of a reference mesh according to embodiments.
[0366] Referring to FIGS. 27 and 28, according to the embodiments, the restored displacement vector transformation coefficient can be stored in a buffer according to the reference structure, and the restored displacement vector transformation coefficient (refDispCoeff) of the reference mesh mapped to the current mesh vertex can be used as a predictor of the current displacement vector transformation coefficient to perform a prediction as shown in Table 2 below.
[0367] for (size_t v = 0; v < N; v++) {for (size_t d = 0; d < dim; d++) {dispCoeff[v][d] = curdispCoeff[v][d] - refDispCoeff[v][d]}}
[0368] According to the embodiments, the displacement vector transformation coefficient quantization unit can perform quantization on the difference displacement vector transformation coefficient (dispCoeff) obtained by subtracting the predicted displacement vector transformation coefficient (refDispCoeff) from the current displacement vector transformation coefficient (curdispCoeff).
[0369] FIG. 29 illustrates a lifting conversion process according to embodiments.
[0370] According to the embodiments, the displacement vector transformation unit can perform a transformation on the displacement vector, and the displacement vector can be processed in the (x, y, z) coordinate system or the (n, t, bt) coordinate system depending on the coordinate system. According to the embodiments, when a coordinate system transformation to the (n, t, bt) coordinate system is performed, a one-dimensional scalar displacement vector of the normal (n) component can be applied as an input to the displacement vector transformation unit, and transformation, quantization, and encoding can be performed on the displacement value of the normal component. According to the embodiments, the displacement vector transformation may include a lifting transformation or a wavelet transformation, etc.
[0371] According to the embodiments, when a lifting transformation is performed, a displacement vector transformation process such as that shown in FIG. 29 may be applied. Referring to FIG. 29, the number of lifting transformations may be determined using the number of subdivision levels (lodCount) of the mesh, and the lifting transformation process may be performed in units of the mesh subdivision levels. Additionally, lifting transformation prediction and lifting transformation updates may be performed during the lifting transformation process.
[0372] According to the embodiments, when the lifting transformation prediction unit predicts the displacement vector of a vertex of a specific subdivision level, it may perform the prediction by using the displacement vector of a vertex with the same or lower subdivision level as a predictor. Specifically, the lifting transformation prediction unit [predicts] a vertex of the k-th subdivision level When performing displacement vector prediction, t( or The vertex of the )th level of subdivision The displacement vector of can be used as a predictor to predict the displacement vector of the k-th level of refinement.
[0373] According to the embodiments, the prediction of the displacement vector can be performed by selecting n points close based on connectivity information among vertices with a lower level of refinement than the current vertex, and then using an average or distance-based weighted average prediction method.
[0374] According to the embodiments, a prediction can be performed based on the displacement vectors of n vertices used to generate the current vertex in the mesh subdivision step.
[0375] According to the embodiments, a residual signal can be generated based on the difference between the displacement vector of the segmentation level and the predicted displacement vector.
[0376] According to the embodiments, the lifting transformation update unit may perform the process of updating the displacement vector of the vertex used for prediction based on the residual signal generated by the lifting transformation prediction unit. The lifting transformation update weight (updateWeight) may be derived according to the vltp_log2_lifting_update_weight syntax.
[0377] According to the embodiments, the lifting transformation update may be performed by sharing the same weights for each LoD level of the mesh, or by using different weights for each LoD level.
[0378] According to the embodiments, when the adaptiveUpdateWeight value indicating whether to perform an adaptive update is 0, the same update weight can be used for each LoD level, and when the adaptiveUpdateWeight value is 1, different adaptive updates can be performed for each LoD level depending on the characteristics of the LoD level.
[0379] According to the embodiments, the displacement vector transformation coefficient quantization unit can perform quantization on the transformation coefficients transformed through the displacement vector transformation unit. The transformation coefficients can be quantized with different quantization parameters for each axis, and the quantization parameters or scaling parameters are derived by the encoder and decoder agreements to determine the quantization rate for each LoD level.
[0380] FIG. 30 shows a dynamic mesh decoder according to embodiments.
[0381] Referring to FIG. 30, a dynamic mesh decoder according to embodiments may include a motion vector decoder, a static mesh decoder, a base mesh restoration unit, a mesh subdivision unit, a displacement vector coordinate system inverse transformation unit, a mesh restoration unit, a texture map decoder, and a displacement vector decoder.
[0382] The motion vector decoder according to the embodiments can perform motion vector decoding when inter-frame prediction is performed on the current mesh. The motion vector decoder can decode residual motion vectors at the vertex or sub-block level through a motion vector bitstream, perform connection information-based prediction using the previously decoded motion vector as a predictor, and then restore the motion vector by combining it with the residual motion vector.
[0383] The static mesh decoding unit according to the embodiments can restore at least one of the connection information, vertex geometry information, vertex texture coordinates, and normal information of the base mesh.
[0384] According to the embodiments, the base mesh restoration unit can restore the current base mesh by applying the restored motion vector to the reference base mesh and then performing inverse quantization when the current base mesh is encoded based on the reference mesh. Additionally, when the current base mesh is decoded through the static mesh decoding unit, the restored base mesh can be generated by performing inverse quantization. According to the embodiments, the inverse quantization unit may be omitted.
[0385] According to the embodiments, the mesh subdivision unit can generate additional vertices by performing subdivision on a base mesh. The mesh subdivision unit can be generated by implicitly deriving geometric connectivity information, texture coordinate connectivity information, and texture coordinates according to the subdivision method. The mesh subdivision unit can perform subdivision through at least one of the mid-edge, Loop, and Catmul & Clark methods.
[0386] Mesh segmentation can be performed n times by user parameters or encoder and decoder promises, and according to the embodiment, the vertices of the base mesh , the newly generated vertices by performing subdivision 1 , vertices generated by performing subdivision n times When defined as It can be defined as follows.
[0387]
[0388] The mesh segmentation unit can perform segmentation as many times as the number of segmentations corresponding to the target LoD entered by the user. For example, if the target LoD is entered as 1, the number of segmentations can be set to 1 to perform 1 segmentation, and if the target LoD is entered as 2, the number of segmentations can be set to 2 to perform 2 segmentations.
[0389] FIGS. 31 and 32 illustrate the inverse transformation process of the displacement vector coordinate system according to the embodiments.
[0390] According to the embodiments, the displacement vector coordinate system inverse transformation unit can inversely transform the inversely quantized restored displacement vector into spatial coordinate axes (x, y, z) when the coordinate system transformation flag (applyLocalCoord), which is parsed in units of sequence, GOF (group of frame), frame, or submesh, is 1.
[0391] Referring to Fig. 31, a normal vector per vertex is calculated based on the reconstructed vertex position information of the reconstructed base mesh, and the normal value of a newly generated vertex can be assigned by interpolating with the calculated vertex normal vector of the reconstructed base mesh for the additional vertex generated through the subdivision process.
[0392] According to the embodiments, interpolation can be performed by averaging or distance-based weighted summing the normal information of the base mesh used for segmentation.
[0393] Referring to Fig. 32, after performing subdivision on the restored base mesh, normal vectors can be calculated for the vertices generated through the subdivision unit and the vertices of the base mesh.
[0394] According to the embodiments, tangential and bi-tangential vectors perpendicular to the normal vectors can be calculated using the calculated normal vectors per vertex, and the inverse transformation of the displacement vector coordinate system can be performed using the following formula.
[0395]
[0396] According to the embodiments, the inverse coordinate system transformation may always be performed without transmitting a flag.
[0397] Referring again to FIG. 30, the mesh restoration unit can calculate the vertex position information of the restored mesh by adding the restoration displacement vector to the vertices generated through the subdivision process in the mesh subdivision unit.
[0398] The texture map decoder can receive a texture map bitstream as input and decode the texture map. According to embodiments, the texture map decoder can decode the texture map bitstream using at least one of a video decoder, a zero-run length decoder, or an arithmetic decoder. According to embodiments, the texture map decoder can perform color space conversion of the texture map.
[0399] FIG. 33 shows the displacement vector decoding process of the displacement vector decoding unit according to the embodiments.
[0400] Referring to FIG. 33, the displacement vector decoder may be composed of a displacement vector transformation coefficient decoder, a displacement vector inverse quantization unit, and a displacement vector inverse transformation unit.
[0401] According to the embodiments, the displacement vector transformation coefficient decoder can perform displacement vector decoding through a 2D video decoder such as H.264, HEVC, VVC, etc., or perform decoding through a zero-run length decoder or an arithmetic decoder, etc.
[0402] According to the embodiments, the displacement vector transformation coefficient decoding method can be determined as a specific decoding method by an agreement between the encoder and the decoder, or the decoding method can be determined by the decoder by receiving the encoding method determined by the encoder as a flag or index (dispEncType).
[0403] According to the embodiments, the displacement vector transformation coefficient decoding method may be one of a method combining a video codec-based decoding method and a zero-run length decoding method, a method combining a video codec-based decoding method and an arithmetic decoding method, a video codec-based decoding method, a zero-run length decoding method, and an arithmetic decoding method.
[0404] According to the embodiments, a displacement vector transformation coefficient decoding method can be determined according to a profile defined in an encoder and a decoder, and an index (profileToolsetIdx) representing profile information is transmitted so that the decoder can determine the displacement vector decoding method according to profileToolsetIdx.
[0405] If the displacement vector decoding method is a video codec-based decoding method, the displacement vector transformation coefficient inverse packing unit can perform the process of inverse packing the packed displacement vector transformation coefficient frame.
[0406] According to the embodiments, the reverse packing method may be derived in the same way in the encoder and decoder, or the reverse packing method may be determined in the decoder by receiving the packing method.
[0407] According to the embodiments, the reverse packing method may be a raster-scan sequence or a rectangular shape.
[0408] The displacement vector inverse quantization unit can perform inverse quantization on the displacement vector. According to the embodiments, the transformation coefficients can be quantized through different quantization parameters for each axis, and the quantization parameter or scaling parameter can be derived by an agreement between the encoder and the decoder to determine the quantization rate per LoD level.
[0409] According to embodiments, the displacement vector inverse transform unit can perform an inverse transform corresponding to the transform performed in the encoder. Depending on the embodiments, the inverse transform may include a lifting inverse transform or a wavelet inverse transform.
[0410] FIG. 34 illustrates the lifting inverse conversion process according to the embodiments.
[0411] The lifting inverse transformation process of Fig. 34 can be performed in the displacement vector inverse transformation unit of Fig. 33.
[0412] Referring to Fig. 34, the number of lifting inverse transformations can be determined using the number of mesh subdivision levels (lodCount).
[0413] According to the embodiments, when a target LoD is received from a user, the number of lifting inverse transformations can be induced equal to the number of subdivisions corresponding to the target LoD.
[0414] Referring to Fig. 34, inverse lifting transformation prediction and inverse lifting transformation update can be performed during the inverse lifting transformation process.
[0415] The lifting inverse transform prediction unit can be performed on the displacement vector on which the lifting inverse transform update unit has been performed. When performing prediction of vertex displacement vectors at a refinement level, the lifting inverse transform prediction unit can perform displacement vector prediction by using vertex displacement vectors of the same or lower refinement level as predictors. Specifically, the vertex of the k-th refinement level When performing displacement vector prediction, t( or The vertex of the )th level of subdivision The displacement vector of can be used as a predictor to predict the displacement vector of the k-th level of refinement.
[0416] According to the embodiments, when performing displacement vector prediction, n vertices that are close based on connectivity information among vertices with a lower level of refinement than the current vertex can be selected, and prediction can be performed using an average or distance-based weighted average method.
[0417] According to the embodiments, a prediction can be performed based on the displacement vectors of n vertices used to generate the current vertex in the mesh subdivision step.
[0418] According to the embodiments, the vertex displacement vector of the refinement level can be restored through the sum of the predicted displacement vector and the parsed residual signal.
[0419] The lifting inverse transform update unit can perform the process of updating the displacement vector of the vertex used for prediction using the parsed residual signal.
[0420] According to the embodiments, the lifting inverse transformation update weight (updateWeight) can be derived by the vltp_log2_lifting_update_weight syntax.
[0421] According to the embodiments, the lifting inverse transformation update process may be performed by sharing the same weights for each LoD level of the mesh, or by using different weights for each LoD level of the mesh.
[0422] According to the embodiments, when the adaptiveUpdateWeight value is 0 depending on whether adaptive updates are performed, the same update weight can be used for each LoD level, and when the adaptiveUpdateWeight value is 1, different adaptive updates can be performed depending on the characteristics of the LoD level.
[0423] SEI message combination method for specific LoD-based mesh partial decoding or attribute partial decoding of a submesh
[0424] The embodiments may include a syntax combined into a single SEI message by taking into account redundancy information between the conventional F.2.6 LoD extraction information SEI payload syntax and the F.2.8 Attribute extraction information SEI payload syntax.
[0425] According to the embodiments, a combined SEI message may be referred to as extraction information and may be changed to another name. Some syntax of the extraction information SEI message may be omitted.
[0426] According to the embodiments, if an extraction information SEI message exists in an atlas bitstream, a receiver or decoder can parse the said SEI message.
[0427] According to the embodiments, the application period of the Extraction information SEI message may be determined according to ei_cancel_flag. If ei_cancel_flag is 1, the syntax related to extraction information may not be parsed, and if ei_cancel_flag is 0, the syntax related to extraction information may be parsed.
[0428] ei_extraction_type may indicate a target for partial decoding by partially extracting a bitstream of a portion of a video bitstream. If ei_extraction_type is 0, it may indicate that the purpose is to extract the LoD region from a displacement video bitstream. If ei_extraction_type is 1, it may indicate that the purpose is to extract an attribute region corresponding to a specific submesh from an attribute video bitstream.
[0429] ei_extractable_unit_type_idx can represent the type for extraction from the video bitstream, and if ei_extractable_unit_type_idx is 0, it may mean MCTS, and if ei_extractable_unit_type_idx is 1, it may mean sub-picture.
[0430] According to the embodiments, ei_extractable_unit_type_idx may differ between the attribute video and the geometry video, and in such cases, information about the bitstream extractable unit may be additionally signaled for each of the attribute video and the geometry video.
[0431] The value obtained by adding 1 to ei_number_of_submesh_minus1 can represent the number of submeshes. If ei_extraction_type is 1 or 2, attribute extraction information is being signaled, so the number of attributes (ei_attribute_count) can be parsed. If ei_extraction_type is 0 or 2, LoD extraction information of the displacement video is being signaled, so the number of subdivision iterations corresponding to the number of LoDs (ei_subdivision_iteration_count[i]) can be parsed.
[0432] Depending on ei_extractable_unit_type_idx, if the value is 0, the MCTS index (ei_mcts_idx) can be parsed, and if the value is 1, the subpicture index (ei_subpic_idx) can be parsed.
[0433] If Extraction information is an attribute, ei_attr_mcts_idx[i][j] or ei_attr_subpic_idx[i][j] may be the j-th attribute corresponding to the i-th submesh.
[0434] If the extraction information is LoD-based displacement, ei_LoD_mcts_idx[i][j] or ei_LoD_subpic_idx[i][j] may be the LoD refinement level j corresponding to the i-th submesh.
[0435] FIG. 35 shows the extraction information SEI payload syntax according to the embodiments.
[0436] The extraction information SEI payload syntax of FIG. 35 relates to embodiments for cases where MCTS or subpictures, etc., are distinguished and signaled as units for extracting bitstreams. According to the embodiments, extraction units such as MCTS or subpictures are distinguished through ei_extractable_unit_type_idx, which is type information of the bitstream extraction unit, and each MCTS index or subpicture index can be parsed.
[0437] According to the embodiments, the displacement sub-bitstream may include a displacement sequence parameter set low byte sequence payload (displ_sequence_parameter_set_rbsp), and displ_sequence_parameter_set_rbsp may include information about the number of displacement values within the displacement frame for the displacement data.
[0438] The semantics for the syntax element of Fig. 35 can be described as follows.
[0439] The cancellation flag (ei_cancel_flag) indicates whether to maintain persistence of the SEI message.
[0440] The extraction type (ei_extraction_type) may represent a target for partial extraction of a bitstream from a video bitstream. If the value of the extraction type (ei_extraction_type) is 0, it may indicate that the purpose is to extract the LoD region from the displacement video bitstream; if the value is 1, it may indicate that the purpose is to extract the attribute region corresponding to a specific submesh from the attribute video bitstream; and if the value is 2, it may indicate that the purpose is to extract the attribute region corresponding to the LoD and submesh from the displacement video bitstream and the attribute video bitstream, respectively.
[0441] The extractable unit type index (ei_extractable_unit_type_idx) may represent an extractable unit at the video codec level, and if the value is 0, it means that it is encoded as MCTS (Motion-Constrained Tile Sets), and if the value is 1, it means that it is encoded as a subpicture.
[0442] The value obtained by adding 1 to the number of submesh (ei_number_of_submesh_minus1) represents the number of submesh.
[0443] ei_attribute_count refers to the number of attributes.
[0444] The submesh ID (ei_submesh_id[i]) represents the ID of the i-th submesh.
[0445] The subdivision iteration count (ei_subdivision_iteration_count[i]) represents the number of subdivisions of the i-th submesh.
[0446] The MCTS index of the LoD region (ei_LoD_mcts_idx[i][j]) represents the MCTS index of the region packed with the displacement vector of the i-th submesh in the displacement video bitstream, where ei_extraction_type is 0 or 2 and ei_extractable_unit_type_idx is 0.
[0447] The subpicture index of the LoD region (ei_LoD_subpic_idx[i][j]) represents the subpicture index of the region packed with the displacement vector of the i-th submesh in the displacement video bitstream, where ei_extraction_type is 0 or 2 and ei_extractable_unit_type_idx is 1.
[0448] The MCTS index of the attribute region (ei_attr_mcts_idx[i][j]) represents the MCTS index of the region corresponding to the j-th attribute of the i-th submesh in the attribute video bitstream where ei_extraction_type is 1 or 2 and ei_extractable_unit_type_idx is 0.
[0449] The subpicture index of the attribute region (ei_attr_subpic_idx[i][j]) represents the subpicture index of the region corresponding to the j-th attribute of the i-th submesh in the attribute video bitstream where ei_extraction_type is 1 or 2 and ei_extractable_unit_type_idx is 1.
[0450] According to the embodiments, the extraction information SEI message may include extraction type information (ei_extraction_type) indicating a target for performing partial region extraction on at least one of the displacement sub-bitstream and the attribute sub-bitstream.
[0451] According to the embodiments, the extraction information SEI message may include information about the number of attributes associated with the basemesh data (ei_attribute_count).
[0452] According to the embodiments, the extraction information SEI message may include information about the ID of a submesh within the basemesh data (ei_submesh_id[i]).
[0453] According to the embodiments, based on a first or third value of the extraction type information (ei_extraction_type), the extraction information SEI message may further include information about the number of subdivisions of the submesh (ei_subdivision_iteration_count[i]) and an index of the extraction area for the LoD of the submesh within the displacement sub-bitstream (ei_LoD_extractable_unit_idx[i][j]).
[0454] According to the embodiments, based on a second or third value of the extraction type information (ei_extraction_type), the extraction information SEI message may further include an index of the extraction area (ei_attr_extractable_unit_idx[i][j]) for an attribute of a submesh within the attribute sub-bitstream.
[0455] According to the embodiments, based on a first value of information about some region extraction units (ei_extractable_unit_type_idx or a value derived from ei_extractable_unit_type_idx), at least one of ei_LoD_extractable_unit_idx[i][j] and ei_attr_extractable_unit_idx[i][j] may represent an index of a region encoded in MCTS.
[0456] According to the embodiments, based on a second value of information for some region extraction units, at least one of ei_LoD_extractable_unit_idx[i][j] and ei_attr_extractable_unit_idx[i][j] may represent the index of a region encoded as a subpicture.
[0457] According to the embodiments, when the displacement vector encoding result is transmitted as a displacement sub-bitstream, the displacement sub-bitstream may include a displacement sequence parameter set low byte sequence payload (displ_sequence_parameter_set_rbsp), and displ_sequence_parameter_set_rbsp may include information about the number of displacement values within the displacement frame for the displacement data.
[0458] According to the embodiments, a transmitter or encoding device may generate an extraction information SEI message according to the syntax of FIG. 35 and include it in an atlas sub-bitstream, and a receiver or decoding device may obtain the extraction information SEI message from the atlas sub-bitstream and then parse ei_extraction_type, ei_attribute_count, and ei_submesh_id[i].
[0459] According to the embodiments, the receiver or decoder may determine whether to parse ei_subdivision_iteration_count[i] and ei_LoD_extractable_unit_idx[i][j] based on the ei_extraction_type value, or whether to parse ei_attr_extractable_unit_idx[i][j].
[0460] According to the embodiments, the receiver or decoder can determine whether ei_LoD_extractable_unit_idx[i][j] and ei_attr_extractable_unit_idx[i][j] are MCTS area indices or subpicture area indices based on the ei_extractable_unit_type_idx value.
[0461] According to the embodiments, if an extraction information SEI message exists in the atlas bitstream, the receiver or decoder can parse the said SEI message, and the application period of the extraction information SEI message can be determined according to ei_cancel_flag.
[0462] If ei_cancel_flag is 1, syntax related to extraction information may not be parsed, and if ei_cancel_flag is 0, syntax related to extraction information may be parsed.
[0463] When ei_cancel_flag is 0, the receiver or decoder can parse ei_extraction_type to identify whether the extraction target is the LoD region of the displacement video, the submesh region of the attribute video, or both.
[0464] A receiver or decoder can parse ei_extractable_unit_type_idx to identify whether the video partition unit is MCTS or a subpicture, and then parse ei_number_of_submesh_minus1 to determine the range of iteration phrases for the number of submeshes.
[0465] In addition, if ei_extraction_type is 1 or 2, ei_attribute_count can be parsed as it contains attribute extraction information.
[0466] In addition, if ei_extraction_type is 0 or 2, the LoD extraction information of the displacement video is included, so the number of subdivision iterations per submesh (ei_subdivision_iteration_count[i]) can be parsed.
[0467] A receiver or decoder can parse ei_submesh_id[i] while iterating over the number of submeshes, and parse ei_subdivision_iteration_count[i] and an index by LoD level or an index by attribute according to the condition per submesh.
[0468] At this time, when ei_extractable_unit_type_idx is 0, ei_LoD_mcts_idx[i][j] for the LoD area and ei_attr_mcts_idx[i][j] for the attribute area can be parsed, and when ei_extractable_unit_type_idx is 1, ei_LoD_subpic_idx[i][j] for the LoD area and ei_attr_subpic_idx[i][j] for the attribute area can be parsed.
[0469] A transmitter or encoding device according to the embodiments can generate the syntax and syntax elements of FIG. 35 so that the transmitter or decoder can parse them as described above.
[0470] FIG. 36 shows the extraction information SEI payload syntax according to the embodiments.
[0471] The extraction information SEI payload syntax of FIG. 36 relates to embodiments for cases where the bitstream extraction unit (e.g., MCTS, subpicture, etc.) is not distinguished and signaled. According to the embodiments, compared to the case where the bitstream extraction unit is distinguished and signaled as in FIG. 35, FIG. 36 may omit ei_extractable_unit_type_idx, which is information about the bitstream extraction unit type.
[0472] According to the embodiments, information regarding video partition unit indices (ei_LoD_extractable_unit_idx and ei_attr_extractable_unit_idx) can be interpreted according to the type of video decoder, and if HEVC is used, it can be derived as an MCTS index, and if VVC is used, it can be derived as a subpicture index.
[0473] According to the embodiments, the type of video decoder can be identified based on profile information that can be parsed from a VPS header containing a V3C / V-DMC parameter set, and, for example, whether the video codec group is HEVC-based or VVC-based can be determined using ptl_profile_codec_group_idc, a profile-related field of the VPS header.
[0474] According to the embodiments, a receiver or decoder can parse ptl_profile_codec_group_idc from a VPS header to determine the type of codec applied to the displacement video or attribute video, and determine whether ei_LoD_extractable_unit_idx and ei_attr_extractable_unit_idx are interpreted as MCTS indexes or subpicture indexes according to the determined type of codec.
[0475] The semantics for the syntax element of Fig. 36 can be described as follows.
[0476] The cancellation flag (ei_cancel_flag) indicates whether to maintain persistence of the SEI message.
[0477] The extraction type (ei_extraction_type) refers to the target intended for partial extraction of a bitstream from a video bitstream. If the value of the extraction type (ei_extraction_type) is 0, it means the purpose is to extract the LoD region from the displacement video bitstream. If the value of the extraction type (ei_extraction_type) is 1, it means the purpose is to extract the attribute region corresponding to a specific submesh from the attribute video bitstream. If the value of the extraction type (ei_extraction_type) is 2, it means the purpose is to extract the attribute region corresponding to the LoD and submesh from the displacement video bitstream and the attribute video bitstream, respectively.
[0478] The value obtained by adding 1 to the number of submesh (ei_number_of_submesh_minus1) represents the number of submesh.
[0479] The submesh ID (ei_submesh_id[i]) represents the ID of the i-th submesh.
[0480] ei_attribute_count refers to the number of attributes.
[0481] The subdivision iteration count (ei_subdivision_iteration_count[i]) represents the number of subdivisions of the i-th submesh.
[0482] The LoD extraction area index (ei_LoD_extractable_unit_idx[i][j]) represents the video partition unit index of the region packed with a displacement vector of LoD level j in the i-th submesh within the displacement video bitstream, where the value of the extraction type (ei_extraction_type) is 0 or 2. The type of video partition unit for the LoD extraction area index (ei_LoD_extractable_unit_idx[i][j]) can be derived according to the displacement video codec type; in the case of the HEVC codec, the video partition unit refers to MCTS, and in the case of the VVC codec, the video partition unit refers to a subpicture.
[0483] The attribute extraction area index (ei_attr_extractable_unit_idx[i][j]) represents the video partition unit index of the area corresponding to the j-th attribute of the i-th submesh within the attribute video bitstream when the value of the extraction type (ei_extraction_type) is 1 or 2. The type of video partition unit for the attribute extraction area index (ei_attr_extractable_unit_idx[i][j]) can be derived according to the attribute video codec type, where the video partition unit is MCTS in the case of HEVC codec and the video partition unit is subpicture in the case of VVC codec.
[0484] According to the embodiments, the extraction information SEI message may include extraction type information (ei_extraction_type) indicating a target for partial region extraction for at least one of the displacement sub-bitstream and the attribute sub-bitstream.
[0485] According to the embodiments, the extraction information SEI message may include information about the number of attributes associated with the base mesh data (ei_attribute_count) and information about the ID of the submesh within the base mesh data (ei_submesh_id[i]).
[0486] According to the embodiments, based on a first or third value of the extraction type information (ei_extraction_type), the extraction information SEI message may further include information about the number of subdivisions of the submesh (ei_subdivision_iteration_count[i]) and an index of the extraction area for the LoD level of the submesh within the displacement sub-bitstream (ei_LoD_extractable_unit_idx[i][j]).
[0487] According to the embodiments, based on a second or third value of the extraction type information (ei_extraction_type), the extraction information SEI message may further include an index of the extraction area (ei_attr_extractable_unit_idx[i][j]) for an attribute of a submesh within the attribute sub-bitstream.
[0488] A transmitter or encoding device according to the embodiments may generate an extraction information SEI message according to the syntax of FIG. 36 and include it in an atlas sub-bitstream, and a receiver or decoding device according to the embodiments may obtain the extraction information SEI message from the atlas sub-bitstream and then parse ei_extraction_type, ei_attribute_count, and ei_submesh_id[i].
[0489] A receiver or decoder according to the embodiments can determine which of the first, second, and third values of the extraction type information (ei_extraction_type) is parsed, and which of the syntax elements to be parsed is ei_subdivision_iteration_count[i], ei_LoD_extractable_unit_idx[i][j], or ei_attr_extractable_unit_idx[i][j].
[0490] According to the embodiments, if an extraction information SEI message exists in the atlas bitstream, the receiver or decoder may parse the said SEI message, and whether to parse the extraction information SEI message may be determined according to the cancellation flag (ei_cancel_flag).
[0491] If the cancellation flag (ei_cancel_flag) is 1, subsequent syntax related to the extraction information may not be parsed, and if the cancellation flag (ei_cancel_flag) is 0, subsequent syntax related to the extraction information may be parsed.
[0492] When the cancellation flag (ei_cancel_flag) is 0, the receiver or decoder can parse the extraction type information (ei_extraction_type) to identify whether the extraction target is the LoD region of the displacement video, the submesh and attribute region of the attribute video, or both.
[0493] A receiver or decoder can determine the range of submesh repetition phrases by parsing the number of submeshes (ei_number_of_submesh_minus1).
[0494] In addition, if the value of the extraction type information (ei_extraction_type) is 1 or 2, the receiver or decoder can determine the range of attribute repetition phrases by parsing the attribute count (ei_attribute_count).
[0495] A receiver or decoder can parse the submesh ID (ei_submesh_id[i]) while iterating over the number of submeshes.
[0496] Additionally, if the value of the extraction type information (ei_extraction_type) is 0 or 2, the receiver or decoder can parse the subdivision iteration count per submesh (ei_subdivision_iteration_count[i]) and parse the LoD extraction area index (ei_LoD_extractable_unit_idx[i][j]) for the LoD level iteration phrase according to the subdivision iteration count.
[0497] In addition, if the value of the extraction type information (ei_extraction_type) is 1 or 2, the receiver or decoder can parse the attribute extraction area index (ei_attr_extractable_unit_idx[i][j]) for the repetition phrase according to the number of attributes.
[0498] A transmitter or encoding device according to the embodiments can generate the syntax and syntax elements of FIG. 35 so that the transmitter or decoder can parse them as described above.
[0499] SEI message mapping information between video codec partition information and mesh data
[0500] The embodiments may include a syntax combined into a single SEI message by taking into account redundant information between the conventional F.2.6 LoD extraction information SEI payload syntax, F.2.7 Tile submesh mapping SEI payload syntax, and F.2.8 Attribute extraction information SEI payload syntax.
[0501] According to the embodiments, a combined SEI message may be referred to as mapping information and may be changed to another name. Some syntax of the mapping information SEI message may be omitted.
[0502] The syntax may be a structure in which LoD extraction and attribute extraction SEI messages are added based on the tile submesh mapping information SEI message.
[0503] FIG. 37 (Fig. 37a, FIG. 37b, and FIG. 37c) shows the extraction information SEI payload syntax according to the embodiments.
[0504] According to the embodiments, the mapping information SEI message can be signaled in the atlas sub-bitstream and can provide mapping information associated with the partitioning information of the displacement sub-bitstream and the attribute sub-bitstream.
[0505] According to the embodiments, the mapping information SEI message may be configured to signal the correspondence relationship with the submesh on an atlas tile basis and to signal region index information required for LoD extraction or attribute extraction together.
[0506] The semantics for the syntax element of Fig. 37 can be described as follows.
[0507] The mapping persistence flag (mi_persistance_mapping_flag) indicates whether SEI messages are maintained.
[0508] The value obtained by adding 1 to the number of tiles (mi_num_tiles_minus1) represents the total number of atlas tiles and attribute tiles.
[0509] The value obtained by adding 1 to the tile ID length (mi_tile_id_length_minus1) represents the number of bits used to represent the tile ID.
[0510] The codec tile signal flag (mi_codec_tile_signal_flag) indicates whether tile information of the video codec corresponding to an atlas tile or attribute tile is signaled. A value of 1 for the codec tile signal flag (mi_codec_tile_signal_flag) means that mapping information between the atlas tile or attribute tile and the video tile is signaled, and a value of 0 means that it is not signaled. If the codec tile signal flag (mi_codec_tile_signal_flag) does not exist, the value is set to 0.
[0511] The LoD mapping signal flag (mi_LoD_mapping_signal_flag) indicates whether to signal LoD information mapped to displacement video partitioning information in order to perform partial extraction of the displacement video bitstream by LoD before performing displacement video decoding. If the value of the LoD mapping signal flag (mi_LoD_mapping_signal_flag) is 1, it means that mapping information between the displacement video partitioning information and the LoD is signaled. If the value of the LoD mapping signal flag (mi_LoD_mapping_signal_flag) is 0, it means that no signaling is performed. If the LoD mapping signal flag (mi_LoD_mapping_signal_flag) does not exist, the value is set to 0.
[0512] The attribute mapping signal flag (mi_attribute_mapping_signal_flag) indicates whether to signal submesh information mapped to attribute video partitioning information in order to perform partial extraction of the attribute video bitstream by submesh before performing attribute video decoding. If the value of the attribute mapping signal flag (mi_attribute_mapping_signal_flag) is 1, it means that the mapping information between the attribute video partitioning information and the submesh is signaled. If the value of the attribute mapping signal flag (mi_attribute_mapping_signal_flag) is 0, it means that it is not signaled. If the attribute mapping signal flag (mi_attribute_mapping_signal_flag) does not exist, the value is set to 0.
[0513] The geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the geometry video codec is the same as the atlas tile partitioning structure in V-DMC. If the value of the geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) is 1, it means that the partitioning structures of the codec tile and the atlas tile are the same. If the value of the geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) is 0, it means that they are not the same. If the geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) does not exist, the value is set to 0.
[0514] The attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the attribute video codec is identical to the attribute tile partitioning structure in V-DMC. If the value of the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) is 1, it means that the partitioning structures of the codec tile and the attribute tile are identical. If the value of the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) is 0, it means that they are not identical. If the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) does not exist, the value is set to 0.
[0515] The LoD extractable unit type index (mi_LoD_extractable_unit_type_idx) represents the extraction unit at the video codec level for geometry video. If the LoD extractable unit type index (mi_LoD_extractable_unit_type_idx) value is 0, it indicates that the extraction unit is MCTS (Motion-constrained tile sets), and if the LoD extractable unit type index (mi_LoD_extractable_unit_type_idx) value is 1, it indicates that the extraction unit is a sub-picture.
[0516] The attribute extractable unit type index (mi_attr_extractable_unit_type_idx) represents the extraction unit at the video codec level for the attribute video. If the attribute extractable unit type index (mi_attr_extractable_unit_type_idx) value is 0, it indicates that the extraction unit is MCTS (Motion-constrained tile sets), and if the attribute extractable unit type index (mi_attr_extractable_unit_type_idx) value is 1, it indicates that the extraction unit is a sub-picture.
[0517] The attribute count (mi_attribute_count) indicates the number of attributes.
[0518] Tile ID (mi_tile_id[ i ]) refers to the tile ID of the i-th atlas tile.
[0519] The tile type flag (mi_tile_type_flag[ i ]) is a flag indicating the type of the i-th tile. If the value of the tile type flag (mi_tile_type_flag[ i ]) is 0, it indicates that it is an atlas tile associated with geometry, and if the value of the tile type flag (mi_tile_type_flag[ i ]) is 1, it indicates that it is an attribute tile associated with attributes. An atlas tile can be P_TILE or I_TILE, and an attribute tile can be P_TILE_ATTR or I_TILE_ATTR.
[0520] The value obtained by adding 1 to the number of submeshes (mi_num_submeshes_minus1[i]) represents the number of submeshes corresponding to the i-th atlas tile ID.
[0521] The value obtained by adding 1 to the submesh ID length (mi_submesh_id_length_minus1[i]) represents the number of bits to represent the submesh ID corresponding to the i-th atlas tile ID.
[0522] The submesh ID (mi_submesh_id[i][j]) refers to the j-th submesh ID corresponding to the i-th atlas tile ID.
[0523] The subdivision iteration count (mi_subdivision_iteration_count[i]) represents the subdivision count of the i-th submesh.
[0524] The MCTS index of the LoD region (mi_LoD_mcts_idx[i][j]) represents the MCTS index of the region packed with a displacement vector of LoD level j of the i-th submesh in the displacement video bitstream, where mi_tile_type_flag[i] is 0 (i.e., the atlas tile is a geometry tile), mi_LoD_mapping_signal_flag is 1, and mi_LoD_extractable_unit_type_idx is 0 (i.e., encoded in MCTS).
[0525] The subpicture index of the LoD region (mi_LoD_subpic_idx[i][j]) represents the subpicture index of the region packed with a displacement vector of LoD level j of the i-th submesh in the displacement video bitstream, where mi_tile_type_flag[i] is 0, mi_LoD_mapping_signal_flag is 1, and mi_LoD_extractable_unit_type_idx is 1 (i.e., encoded as a subpicture).
[0526] The MCTS index of the attribute region (mi_attr_mcts_idx[i][j]) represents the MCTS index of the region corresponding to the j-th attribute of the i-th submesh in the attribute video bitstream when mi_tile_type_flag[i] is 1 (i.e., it is an atlas attribute tile) and mi_attribute_mapping_signal_flag is 1, and mi_attr_extractable_unit_type_idx is 0 (i.e., it is encoded in MCTS).
[0527] The subpicture index of the attribute region (mi_attr_subpic_idx[i][j]) represents the subpicture index of the region corresponding to the j-th attribute of the i-th submesh in the attribute video bitstream when mi_tile_type_flag[i] is 1, mi_attribute_mapping_signal_flag is 1, and mi_attr_extractable_unit_type_idx is 1 (i.e., encoded as a subpicture).
[0528] The attribute extraction info present flag (mi_extraction_info_present_flag[i]) indicates whether attribute extraction info corresponding to the submesh exists. If the attribute extraction info present flag (mi_extraction_info_present_flag[i]) does not exist, it can be set to 0.
[0529] The value obtained by adding 1 to the number of codec tiles in the tile (mi_num_codec_tiles_in_tile_minus1[i]) represents the number of tiles in the video codec corresponding to the i-th tile.
[0530] The codec tile index (mi_codec_tile_idx[i][j]) refers to the index of the j-th video codec tile among the video codec tiles corresponding to the i-th tile.
[0531] The codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]) indicates whether to initialize the variable currCount[1] to 0.
[0532] According to the embodiments, a mapping information SEI message may be included in an atlas sub-bitstream, and a receiver or decoder may obtain the mapping information SEI message from the atlas sub-bitstream.
[0533] According to the embodiments, the mapping information SEI message may include a first flag (mi_LoD_mapping_signal_flag) indicating whether to signal LoD information mapped to the partitioning information of the displacement sub-bitstream.
[0534] According to the embodiments, the mapping information SEI message may include a second flag (mi_attribute_mapping_signal_flag) indicating whether to signal submesh information mapped to the partitioning information of the attribute sub-bitstream.
[0535] According to the embodiments, the mapping information SEI message may include information (mi_tile_type_flag[i]) indicating the type of atlas tile for the atlas sub-bitstream. For example, the first value of the tile type information may indicate an atlas tile associated with geometry, and the second value of the tile type information may indicate an attribute tile associated with attribute.
[0536] According to the embodiments, the mapping information SEI message may include information about the ID of the submesh associated with the atlas tile (mi_submesh_id[i][j]).
[0537] According to the embodiments, based on the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include information (mi_attribute_count) about the number of attributes associated with the basemesh data.
[0538] According to the embodiments, based on the first value of tile type information (mi_tile_type_flag[i]) and the first value of the first flag (mi_LoD_mapping_signal_flag), the mapping information SEI message may further include information about the number of subdivisions of a submesh (mi_subdivision_iteration_count[submeshIdx]).
[0539] According to the embodiments, based on the first value of the tile type information and the first value of the first flag, the mapping information SEI message may further include an index of the extraction area (mi_LoD_extractable_unit_idx[submeshIdx][k]) for the LoD of the submesh within the displacement sub-bitstream.
[0540] According to the embodiments, based on the second value of the tile type information (mi_tile_type_flag[i]) and the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include an index of the extraction area (mi_attr_extractable_unit_idx[submeshIdx][k]) for the attribute of the submesh within the attribute sub-bitstream.
[0541] According to the embodiments, based on the first value of the first flag (mi_LoD_mapping_signal_flag), the mapping information SEI message may further include information (mi_LoD_extractable_unit_type_idx) indicating the area unit of the extraction area for the LoD of the submesh. For example, the first value of the information indicating the area unit may indicate MCTS (Motion-Constrained Tile Sets), and the second value of the information indicating the area unit may indicate the case where it is encoded as a subpicture.
[0542] According to the embodiments, based on the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include information (mi_attr_extractable_unit_type_idx) indicating the area unit of the extraction area for the attributes of the submesh. For example, the first value of the information indicating the area unit may indicate MCTS, and the second value of the information indicating the area unit may indicate the case where it is encoded as a subpicture.
[0543] According to the embodiments, mi_LoD_extractable_unit_idx[submeshIdx][k] included in the mapping information SEI message can be interpreted as an MCTS index or a subpicture index depending on the value of mi_LoD_extractable_unit_type_idx.
[0544] According to the embodiments, mi_attr_extractable_unit_idx[submeshIdx][k] included in the mapping information SEI message can be interpreted as an MCTS index or a subpicture index depending on the value of mi_attr_extractable_unit_type_idx.
[0545] According to embodiments, a transmitter or encoding device may generate a mapping information SEI message and include it in an atlas sub-bitstream, wherein the mapping information SEI message may be configured to include a first flag, a second flag, tile type information, and submesh ID information.
[0546] According to embodiments, based on the first value of the second flag, the transmitter or encoding device may be configured to further include attribute count information in the mapping information SEI message.
[0547] According to embodiments, based on a first value of tile type information and a first value of a first flag, a transmitter or encoding device may be configured to further include subdivision count information and an index of an extraction area for LoD in the mapping information SEI message.
[0548] According to embodiments, based on the second value of the tile type information and the first value of the second flag, the transmitter or encoding device may be configured to further include an index of the extraction area for the attribute in the mapping information SEI message.
[0549] According to the embodiments, based on the first value of the first flag, the transmitter or encoding device may be configured to further include information indicating the area unit of the extraction area for the LoD in the mapping information SEI message, and based on the first value of the second flag, the transmitter or encoding device may be configured to further include information indicating the area unit of the extraction area for the attribute in the mapping information SEI message.
[0550] According to the embodiments, if a mapping information SEI message exists in an atlas bitstream (or atlas sub-bitstream), a receiver or decoder can parse the said SEI message.
[0551] According to the embodiments, the receiver or decoder can parse the payload of the mapping information SEI message, first parse the persistence flag (mi_persistance_mapping_flag), and then parse the tile count information (mi_num_tiles_minus1), tile ID length information (mi_tile_id_length_minus1), and codec tile signal flag (mi_codec_tile_signal_flag) in order.
[0552] According to the embodiments, when mi_codec_tile_signal_flag is 1, the receiver or decoder may further parse the geocodec tile alignment flag (mi_geo_codec_tile_alignment_flag) and the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag), and then initialize variables (currCount[0], currCount[1]) for managing codec tile indices. Since mi_codec_tile_signal_flag may be set to 0 if it does not exist in the message, in this case, the receiver or decoder may not perform further parsing of the relevant syntax.
[0553] According to embodiments, a receiver or decoder may parse a LoD mapping signal flag (mi_LoD_mapping_signal_flag) and an attribute mapping signal flag (mi_attribute_mapping_signal_flag). If mi_LoD_mapping_signal_flag or mi_attribute_mapping_signal_flag is not present in the message, it may be set to 0, and if the flag is interpreted as 0, the receiver or decoder may omit parsing the corresponding mapping information.
[0554] According to the embodiments, when mi_LoD_mapping_signal_flag is 1, the receiver or decoder can parse the LoD extractable unit type index (mi_LoD_extractable_unit_type_idx). When mi_attribute_mapping_signal_flag is 1, the receiver or decoder can parse the attribute extractable unit type index (mi_attr_extractable_unit_type_idx) and the attribute count (mi_attribute_count).
[0555] According to the embodiments, a receiver or decoder can iterate over tile index i and sequentially parse for each tile a tile ID (mi_tile_id[i]), a tile type flag (mi_tile_type_flag[i]), the number of submeshes corresponding to the tile (mi_num_submeshes_minus1[i]), and submesh ID length information (mi_submesh_id_length_minus1[i]).
[0556] According to the embodiments, a receiver or decoder can parse a submesh ID (mi_submesh_id[i][j]) by iterating over a submesh index j included in each tile, and determine a subsequent parsing path based on a tile type flag (mi_tile_type_flag[i]) and the value of each mapping signal flag. Additionally, the receiver or decoder can manage submesh-specific parsing results by updating a global index (submeshIdx) at the submesh level.
[0557] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 0 and mi_LoD_mapping_signal_flag is 1, the receiver or decoder can parse the subdivision iteration count (mi_subdivision_iteration_count[submeshIdx]) for the corresponding submesh. Then, for an iteration index k, the receiver or decoder can parse the extraction region index for each LoD level as many times as the number of iterations determined by mi_subdivision_iteration_count[submeshIdx].
[0558] According to the embodiments, when mi_LoD_extractable_unit_type_idx is 0, the receiver or decoder can parse the MCTS index of the LoD region (mi_LoD_mcts_idx[submeshIdx][k]), and when mi_LoD_extractable_unit_type_idx is 1, the receiver or decoder can parse the subpicture index of the LoD region (mi_LoD_subpic_idx[submeshIdx][k]). In this case, the value of mi_LoD_extractable_unit_type_idx can determine whether the LoD extraction region index is interpreted as an index of a unit (MCTS or subpicture).
[0559] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 1 and mi_attribute_mapping_signal_flag is 1, the receiver or decoder may parse the attribute extraction information present flag (mi_attribute_extraction_info_present_flag[submeshIdx]) for attribute index k by repeating a number of times determined by mi_attribute_count. Since the flag may be set to 0 if it does not exist in the message, the receiver or decoder may omit parsing the attribute extraction area index.
[0560] According to the embodiments, when mi_attribute_extraction_info_present_flag[submeshIdx] is 1, the receiver or decoder may parse the MCTS index of the attribute region (mi_attr_mcts_idx[submeshIdx][k]) or the subpicture index of the attribute region (mi_attr_subpic_idx[submeshIdx][k]) according to the value of mi_attr_extractable_unit_type_idx. In this case, the value of mi_attr_extractable_unit_type_idx may determine which unit (MCTS or subpicture) the attribute extraction region index is interpreted as.
[0561] According to the embodiments, after submesh parsing for each tile is completed, if mi_codec_tile_signal_flag is 1, the receiver or decoder may additionally parse mapping information between the atlas tile (or attribute tile) and the video codec tile.
[0562] According to the embodiments, (when the tile type flag is 0 and mi_geo_codec_tile_alignment_flag is 0) or (when the tile type flag is 1 and mi_attr_codec_tile_alignment_flag is 0), the receiver or decoder may parse the number of codec tiles in the tile (mi_num_codec_tiles_in_tile_minus1[i]) and iteratively parse the codec tile index list (mi_codec_tile_idx[i][j]) based on the number.
[0563] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is interpreted as 1, the receiver or decoder may apply an automatic assignment or reset operation of the codec tile index based on the state of the tile type flag and the variable (currCount). For example, if the tile type flag is 1 and currCount[1] is not 0, the receiver or decoder may parse the codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]) and initialize currCount[1] to 0 if the reset flag is 1. Subsequently, the receiver or decoder may determine the tile-specific codec tile index by setting the value of mi_codec_tile_idx[i][0] to currCount[tile type] and increasing currCount[tile type].
[0564] A transmitter or encoding device according to the embodiments may include the syntax elements of the mapping information SEI message in an atlas sub-bitstream by configuring them according to the syntax of FIG. 37 so that a receiver or decoder can parse them in the order described above.
[0565] Video codec agnostic approach
[0566] FIG. 38 (Fig. 38a and FIG. 38b) shows mapping information SEI payload syntax according to embodiments.
[0567] The mapping information SEI payload syntax of Fig. 38 is an SEI message that represents information mapped to video partitioning information regardless of the type of video codec of the attribute video or / and displacement video in the configuration of the V-DMC.
[0568] The SEI message of Fig. 37 parses the extractable_unit_type syntax and, depending on the syntax value, parses information regarding whether it maps to MCTS information of HEVC or subpicture information of VVC. That is, it is a case of signaling information regarding whether MCTS was used or subpicture was used when partitioning to support partial extraction during attribute video or / and displacement video coding.
[0569] The SEI message of FIG. 38 does not separate information regarding MCTS and subpicture, thereby omitting syntax such as extraction unit type, which can save signaling bits. The type of partition type can be determined by checking the displacement video decoder or attribute video decoder type in the decoder. For example, if the video codec type is HEVC, it can be determined to MCTS, and if it is VVC, it can be determined to subpicture.
[0570] The semantics for the syntax element of Fig. 38 can be described as follows.
[0571] The persistence flag (mi_persistence_mapping_flag) indicates whether the SEI message is maintained.
[0572] The value obtained by adding 1 to the number of tiles (mi_num_tiles_minus1) represents the total number of atlas tiles and attribute tiles.
[0573] The value obtained by adding 1 to the tile ID length (mi_tile_id_length_minus1) represents the number of bits used to represent the tile ID.
[0574] The codec tile signal flag (mi_codec_tile_signal_flag) indicates whether tile information of the video codec corresponding to an atlas tile or attribute tile is being signaled. If mi_codec_tile_signal_flag is 1, it indicates that video tile information is being signaled for the atlas tile or atlas attribute tile, and if mi_codec_tile_signal_flag is 0, it indicates that it is not being signaled. If the flag does not exist, it may be set to 0.
[0575] The LoD mapping signal flag (mi_LoD_mapping_signal_flag) indicates whether to signal LoD information mapped to displacement video partitioning information in order to partially extract the displacement video bitstream by LoD before performing displacement video decoding. For example, if mi_LoD_mapping_signal_flag is 1, it means that mapping information between displacement video partitioning information and LoDs is signaled, and if mi_LoD_mapping_signal_flag is 0, it means that no signaling is given. If the flag does not exist, it can be set to 0.
[0576] The attribute mapping signal flag (mi_attribute_mapping_signal_flag) indicates whether to signal submesh information mapped to attribute video partitioning information in order to partially extract the attribute video bitstream by submesh before performing attribute video decoding. If mi_attribute_mapping_signal_flag is 1, it means that mapping information between attribute video partitioning information and submesh is signaled, and if mi_attribute_mapping_signal_flag is 0, it means that it is not signaled. If the flag does not exist, it can be set to 0.
[0577] The attribute count (mi_attribute_count) indicates the number of attributes. If mi_attribute_mapping_signal_flag is 1, the mapping information SEI message may include mi_attribute_count.
[0578] The geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the geometry video codec is the same as the atlas tile partitioning structure in V-DMC. If mi_geo_codec_tile_alignment_flag is 1, it means that the partitioning structures of the codec tiles and atlas tiles are the same, and if mi_geo_codec_tile_alignment_flag is 0, it means that they are not the same. If the flag does not exist, it can be set to 0.
[0579] The attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the attribute video codec is the same as the attribute tile partitioning structure in V-DMC. If mi_attr_codec_tile_alignment_flag is 1, it means that the partitioning structures of the codec tile and the attribute tile are the same, and if mi_attr_codec_tile_alignment_flag is 0, it means that they are not the same. If the flag does not exist, it can be set to 0.
[0580] Tile ID (mi_tile_id[i]) refers to the tile ID of the i-th atlas tile.
[0581] The tile type flag (mi_tile_type_flag[i]) is a flag indicating the type of the i-th tile, where a value of 0 indicates that it is an atlas tile associated with geometry, and a value of 1 indicates that it is an attribute tile associated with attributes. An atlas tile can be P_TILE or I_TILE, and an attribute tile can be P_TILE_ATTR or I_TILE_ATTR.
[0582] The value obtained by adding 1 to the number of submeshes (mi_num_submeshes_minus1[i]) represents the number of submeshes corresponding to the i-th atlas tile ID.
[0583] The value obtained by adding 1 to the submesh ID length (mi_submesh_id_length_minus1[i]) represents the number of bits to represent the submesh ID corresponding to the i-th atlas tile ID.
[0584] The submesh ID (mi_submesh_id[i][j]) refers to the j-th submesh ID corresponding to the i-th atlas tile ID.
[0585] The subdivision iteration count (mi_subdivision_iteration_count[i]) represents the number of subdivisions.
[0586] The LoD extraction unit index (mi_LoD_extractable_unit_idx[i][j]) represents the index of the video partition unit of the region packed with the displacement vector of the i-th submesh in the displacement video bitstream, where mi_LoD_mapping_signal_flag is 1. The type of video partition unit can be derived according to the displacement video codec type, and if the displacement video codec type is HEVC, the video partition unit may be MCTS, and if the displacement video codec type is VVC, the video partition unit may be a subpicture.
[0587] The attribute extraction unit index (mi_attr_extractable_unit_idx[i][j]) represents the index of the video partition unit corresponding to the j-th attribute of the i-th submesh within the attribute video bitstream when mi_attribute_mapping_signal_flag is 1. The type of video partition unit can be derived according to the attribute video codec type, and if the attribute video codec type is HEVC, the video partition unit may be MCTS, and if the attribute video codec type is VVC, the video partition unit may be a subpicture.
[0588] The attribute extraction info present flag (mi_extraction_info_present_flag[i]) indicates whether attribute extraction info corresponding to the submesh exists. If the flag does not exist, it can be set to 0.
[0589] The value obtained by adding 1 to the number of codec tiles in the tile (mi_num_codec_tiles_in_tile_minus1[i]) represents the number of tiles in the video codec corresponding to the i-th tile.
[0590] The codec tile index (mi_codec_tile_idx[i][j]) refers to the index of the j-th video codec tile among the video codec tiles corresponding to the i-th tile.
[0591] The codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]) indicates whether to initialize the variable currCount[1] to 0.
[0592] According to the embodiments, the atlas sub-bitstream may include a mapping information SEI message, and a receiver or decoder may obtain the mapping information SEI message from the atlas sub-bitstream.
[0593] According to the embodiments, the mapping information SEI message may include a first flag (mi_LoD_mapping_signal_flag) indicating whether to signal Level of Detail (LoD) information mapped to the partitioning information of the displacement sub-bitstream.
[0594] According to the embodiments, the mapping information SEI message may include a second flag (mi_attribute_mapping_signal_flag) indicating whether to signal submesh information mapped to the partitioning information of the attribute sub-bitstream.
[0595] According to the embodiments, the mapping information SEI message may include information (mi_tile_type_flag[i]) indicating the type of atlas tile for the atlas sub-bitstream. For example, the first value of the tile type information may indicate an atlas tile associated with geometry, and the second value of the tile type information may indicate an attribute tile associated with attribute.
[0596] According to the embodiments, the mapping information SEI message may include information about the ID of the submesh corresponding to the atlas tile (mi_submesh_id[i][j]).
[0597] According to the embodiments, based on the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include information (mi_attribute_count) about the number of attributes associated with the basemesh data.
[0598] According to the embodiments, based on the first value of tile type information (mi_tile_type_flag[i]) and the first value of the first flag (mi_LoD_mapping_signal_flag), the mapping information SEI message may further include information about the number of subdivisions of a submesh (mi_subdivision_iteration_count[submeshIdx]).
[0599] According to the embodiments, based on the first value of the tile type information and the first value of the first flag, the mapping information SEI message may further include an index of an extraction area (mi_LoD_extractable_unit_idx[submeshIdx][k]) corresponding to the area where the displacement vector by LoD level of the submesh in the displacement sub-bitstream is packed. In this case, mi_LoD_extractable_unit_idx[submeshIdx][k] may represent an index of a video partition unit of the area corresponding to LoD level k in the displacement video bitstream.
[0600] According to the embodiments, based on the second value of the tile type information (mi_tile_type_flag[i]) and the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include an index of the extraction area (mi_attr_extractable_unit_idx[submeshIdx][k]) of the area corresponding to the attribute of the submesh in the attribute sub-bitstream. In this case, mi_attr_extractable_unit_idx[submeshIdx][k] may represent the index of the video partition unit of the area corresponding to attribute k in the attribute video bitstream.
[0601] According to the embodiments, the type of video partition unit represented by mi_LoD_extractable_unit_idx and mi_attr_extractable_unit_idx can be derived according to the video codec type. For example, if the video codec type is HEVC, the video partition unit may be MCTS, and if the video codec type is VVC, the video partition unit may be a subpicture.
[0602] According to embodiments, a transmitter or encoding device may generate a mapping information SEI message to be included in an atlas sub-bitstream, and the generated mapping information SEI message may be configured to include a first flag (mi_LoD_mapping_signal_flag), a second flag (mi_attribute_mapping_signal_flag), tile type information (mi_tile_type_flag[i]), and submesh ID information (mi_submesh_id[i][j]).
[0603] According to the embodiments, based on the first value of the second flag, the transmitter or encoding device may be configured to further include attribute count information (mi_attribute_count) in the mapping information SEI message.
[0604] According to embodiments, based on a first value of tile type information and a first value of a first flag, a transmitter or encoding device may be configured to further include subdivision iteration count information (mi_subdivision_iteration_count[submeshIdx]) and an index of an extraction area for a LoD (mi_LoD_extractable_unit_idx[submeshIdx][k]) in the mapping information SEI message.
[0605] According to embodiments, based on the second value of the tile type information and the first value of the second flag, the transmitter or encoding device may be configured to further include the index of the extraction area for the attribute (mi_attr_extractable_unit_idx[submeshIdx][k]) in the mapping information SEI message.
[0606] According to the embodiments, if a mapping information (Mapping_information) SEI message exists in an atlas bitstream or an atlas sub-bitstream, a receiver or decoder can parse the said SEI message.
[0607] According to the embodiments, a receiver or decoder can parse the payload of a mapping information SEI message, first parse the persistence flag (mi_persistance_mapping_flag), and then parse the tile count information (mi_num_tiles_minus1), tile ID length information (mi_tile_id_length_minus1), codec tile signal flag (mi_codec_tile_signal_flag), LoD mapping signal flag (mi_LoD_mapping_signal_flag), and attribute mapping signal flag (mi_attribute_mapping_signal_flag) in order.
[0608] According to the embodiments, if at least one of mi_codec_tile_signal_flag, mi_LoD_mapping_signal_flag, mi_attribute_mapping_signal_flag is not present in the message, the flag may be set to 0, and the receiver or decoder may omit parsing additional syntax corresponding to the flag interpreted as 0.
[0609] According to the embodiments, if mi_attribute_mapping_signal_flag is 1, the receiver or decoder may further parse the attribute count (mi_attribute_count). If mi_attribute_mapping_signal_flag is interpreted as 0, the receiver or decoder may not perform parsing of mi_attribute_count.
[0610] According to the embodiments, when mi_codec_tile_signal_flag is 1, the receiver or decoder may further parse the geocodec tile alignment flag (mi_geo_codec_tile_alignment_flag) and the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag), and then initialize variables (currCount[0], currCount[1]) for managing codec tile indices.
[0611] According to the embodiments, the receiver or decoder may initialize the global index (submeshIdx) of the submesh unit to 0.
[0612] According to the embodiments, a receiver or decoder can iterate over tile index i and parse a tile ID (mi_tile_id[i]) for each tile, and set TileIdxToID[i] to mi_tile_id[i] for mapping between the tile index and the tile ID.
[0613] According to the embodiments, a receiver or decoder can sequentially parse a tile type flag (mi_tile_type_flag[i]), the number of submeshes corresponding to the tile (mi_num_submeshes_minus1[i]), and submesh ID length information (mi_submesh_id_length_minus1[i]) for each tile.
[0614] According to the embodiments, a receiver or decoder may parse a submesh ID (mi_submesh_id[i][j]) while iterating over a submesh index j, and may set SubmeshIdxToID[i][j] to mi_submesh_id[i][j] for mapping between the submesh index and the submesh ID.
[0615] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 0 and mi_LoD_mapping_signal_flag is 1, the receiver or decoder can parse the subdivision iteration count (mi_subdivision_iteration_count[submeshIdx]) for the corresponding submesh.
[0616] According to the embodiments, a receiver or decoder may iteratively parse a LoD extraction unit index (mi_LoD_extractable_unit_idx[submeshIdx][k]) for a LoD level index k for a number of iterations determined by mi_subdivision_iteration_count[submeshIdx]. In this case, mi_LoD_extractable_unit_idx[submeshIdx][k] may represent a video partition unit index of a region packed with a displacement vector corresponding to LoD level k in a displacement video bitstream.
[0617] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 1 and mi_attribute_mapping_signal_flag is 1, the receiver or decoder may iteratively parse the attribute extraction unit index (mi_attr_extractable_unit_idx[submeshIdx][k]) by iterating a number of times determined by mi_attribute_count for attribute index k. In this case, mi_attr_extractable_unit_idx[submeshIdx][k] may represent the video partition unit index of the region corresponding to attribute k in the attribute video bitstream.
[0618] According to the embodiments, whenever LoD or attribute-related parsing for one submesh is completed, the receiver or decoder may increment submeshIdx to update the parsing position of the next submesh.
[0619] According to the embodiments, after the submesh iteration for each tile is finished, if mi_codec_tile_signal_flag is 1, the receiver or decoder may further parse mapping information between the atlas tile or attribute tile and the video codec tile.
[0620] According to the embodiments, (when mi_tile_type_flag[i] is 0 and mi_geo_codec_tile_alignment_flag is interpreted as 0) or (when mi_tile_type_flag[i] is 1 and mi_attr_codec_tile_alignment_flag is interpreted as 0), the receiver or decoder may parse the number of codec tiles in the tile (mi_num_codec_tiles_in_tile_minus1[i]) and iteratively parse the codec tile index list (mi_codec_tile_idx[i][j]) based on the number.
[0621] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is interpreted as 1, the receiver or decoder may apply an automatic assignment or reset operation of the codec tile index based on the tile type and currCount state. For example, if the tile type flag is 1 and currCount[1] is not 0, the receiver or decoder may parse the codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]), and if the flag is 1, currCount[1] may be initialized to 0. Subsequently, the receiver or decoder may determine mi_codec_tile_idx[i][0] as currCount[mi_tile_type_flag[i]] and increase currCount[mi_tile_type_flag[i]].
[0622] A transmitter or encoding device according to the embodiments may configure the syntax elements of a mapping information SEI message according to the syntax of FIG. 38 and include them in an atlas sub-bitstream so that a receiver or decoder can parse them in the order described above.
[0623] Signaling method for partitioning information (video tile / MCTS index / subpicture index) corresponding to submeshes
[0624] FIG. 39 (Fig. 39a and FIG. 39b) shows mapping information SEI payload syntax according to embodiments.
[0625] The embodiments may include or perform a method of signaling a video tile index and / or MCTS index and / or subpicture index of geometry / attributes corresponding to a submesh.
[0626] In conventional Tile submesh mapping SEI messages, it is difficult to identify the video tile corresponding to the submesh when there are two or more submeshes within an atlas tile. Through the mapping information SEI message of FIG. 39, partial decoding based on the submesh can be performed by using video tile index information corresponding to the submesh.
[0627] The persistence flag (mi_persistence_mapping_flag) indicates whether the SEI message is maintained.
[0628] The value obtained by adding 1 to the number of tiles (mi_num_tiles_minus1) represents the total number of atlas tiles and attribute tiles.
[0629] The value obtained by adding 1 to the tile ID length (mi_tile_id_length_minus1) represents the number of bits used to represent the tile ID.
[0630] The codec tile signal flag (mi_codec_tile_signal_flag) indicates whether tile information of the video codec corresponding to an atlas tile or attribute tile is signaled. If mi_codec_tile_signal_flag is 1, it indicates that video tile mapping information between the atlas tile or atlas attribute tile is signaled, and if mi_codec_tile_signal_flag is 0, it indicates that it is not signaled. If the flag does not exist, it can be set to 0.
[0631] The LoD mapping signal flag (mi_LoD_mapping_signal_flag) indicates whether to signal LoD information mapped to displacement video partitioning information in order to extract the displacement video bitstream by LoD before performing displacement video decoding. If mi_LoD_mapping_signal_flag is 1, it means that mapping information between displacement video partitioning information and LoDs is signaled, and if mi_LoD_mapping_signal_flag is 0, it means that no signaling is given. If the flag does not exist, it can be set to 0.
[0632] The attribute mapping signal flag (mi_attribute_mapping_signal_flag) indicates whether to signal submesh information mapped to attribute video partitioning information in order to partially extract the attribute video bitstream by submesh before performing attribute video decoding. If mi_attribute_mapping_signal_flag is 1, it means that mapping information between attribute video partitioning information and submesh is signaled, and if mi_attribute_mapping_signal_flag is 0, it means that it is not signaled. If the flag does not exist, it can be set to 0.
[0633] The geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the geometry video codec is the same as the partitioning structure of the atlas tile in V-DMC. If mi_geo_codec_tile_alignment_flag is 1, it means that the partitioning structures of the codec tile and the atlas tile are the same, and if mi_geo_codec_tile_alignment_flag is 0, it means that they are not the same. If the flag does not exist, it can be set to 0.
[0634] The attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) indicates whether the tile partitioning structure in the attribute video codec is the same as the partitioning structure of the attribute tile in V-DMC. If mi_attr_codec_tile_alignment_flag is 1, it means that the partitioning structures of the codec tile and the attribute tile are the same, and if mi_attr_codec_tile_alignment_flag is 0, it means that they are not the same. If the flag does not exist, it can be set to 0.
[0635] Tile ID (mi_tile_id[i]) refers to the tile ID of the i-th atlas tile.
[0636] The tile type flag (mi_tile_type_flag[i]) is a flag indicating the type of the i-th tile, where a value of 0 indicates that it is an atlas tile associated with geometry, and a value of 1 indicates that it is an attribute tile associated with attributes. An atlas tile can be P_TILE or I_TILE, and an attribute tile can be P_TILE_ATTR or I_TILE_ATTR.
[0637] The value obtained by adding 1 to the number of submeshes (mi_num_submeshes_minus1[i]) represents the number of submeshes corresponding to the i-th tile ID.
[0638] The value obtained by adding 1 to the submesh ID length (mi_submesh_id_length_minus1[i]) represents the number of bits to represent the submesh ID corresponding to the i-th tile ID.
[0639] The submesh ID (mi_submesh_id[i][j]) refers to the j-th submesh ID corresponding to the i-th tile ID.
[0640] The subdivision iteration count (mi_subdivision_iteration_count[i]) represents the number of subdivisions.
[0641] The LoD extraction unit index (mi_LoD_extractable_unit_idx[i][j]) represents the index of the video partition unit of the region packed with the displacement vector of the i-th submesh in the displacement video bitstream, where mi_LoD_mapping_signal_flag is 1. The type of video partition unit can be derived based on the displacement video codec type, for example, if the displacement video codec type is HEVC, the video partition unit may be MCTS, and if the displacement video codec type is VVC, the video partition unit may be subpicture.
[0642] The attribute extraction unit index (mi_attr_extractable_unit_idx[i][j]) represents the index of the video partition unit corresponding to the j-th attribute of the i-th submesh in the attribute video bitstream when mi_attribute_mapping_signal_flag is 1. The type of video partition unit can be derived according to the attribute video codec type, for example, if the attribute video codec type is HEVC, the video partition unit may be MCTS, and if the attribute video codec type is VVC, the video partition unit may be a subpicture.
[0643] The attribute extraction info present flag (mi_extraction_info_present_flag[i]) indicates whether attribute extraction info corresponding to the submesh exists. If the flag does not exist, it can be set to 0.
[0644] The value obtained by adding 1 to the number of codec tiles in the tile (mi_num_codec_tiles_in_tile_minus1[i][j]) represents the number of tiles in the video codec corresponding to the j-th submesh of the i-th tile.
[0645] The codec tile index (mi_codec_tile_idx[i][j][k]) refers to the index of the k-th video codec tile among the video codec tiles corresponding to the j-th submesh of the i-th tile.
[0646] The codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]) indicates whether to initialize the variable currCount[1] to 0.
[0647] According to the embodiments, a mapping information SEI message may be included in an atlas sub-bitstream, a transmitter or encoding device may generate a mapping information SEI message to be included in an atlas sub-bitstream, and a receiver or decoder may obtain the mapping information SEI message from the atlas sub-bitstream.
[0648] According to embodiments, the mapping information SEI message may include a first flag (mi_LoD_mapping_signal_flag) indicating whether Level of Detail (LoD) information mapped to the partitioning information of the displacement sub-bitstream is signaled, and a second flag (mi_attribute_mapping_signal_flag) indicating whether sub-mesh information mapped to the partitioning information of the attribute sub-bitstream is signaled.
[0649] According to the embodiments, based on the first value of the first flag (mi_LoD_mapping_signal_flag), the mapping information SEI message may further include information (mi_LoD_extractable_unit_type_idx) indicating the area unit for the extraction area for the LoD of the submesh. mi_LoD_extractable_unit_type_idx may mean the extraction unit at the video codec level for the geometry video, the first value may indicate MCTS (Motion-Constrained Tile Sets), and the second value may indicate the case encoded as a subpicture.
[0650] According to the embodiments, based on the first value of the second flag (mi_attribute_mapping_signal_flag), the mapping information SEI message may further include information (mi_attr_extractable_unit_type_idx) indicating the area unit for the extraction area for the attributes of the submesh. mi_attr_extractable_unit_type_idx may represent the extraction unit at the video codec level for the attribute video, the first value may represent MCTS, and the second value may represent the case where it is encoded as a subpicture.
[0651] According to the embodiments, the mapping information SEI message may further include a third flag (mi_codec_tile_signal_flag) indicating whether the tile information of the video codec corresponding to the atlas tile is signaled.
[0652] According to embodiments, based on the first value of the third flag (mi_codec_tile_signal_flag), the mapping information SEI message may include a fourth flag (mi_geo_codec_tile_alignment_flag) indicating whether the tile alignment structure in the geometry video codec and the alignment structure of the atlas tile are the same, and a fifth flag (mi_attr_codec_tile_alignment_flag) indicating whether the tile alignment structure in the attribute video codec and the alignment structure of the atlas tile are the same.
[0653] According to the embodiments, the first value of the information (mi_tile_type_flag[i]) indicating the type of the atlas tile may indicate an atlas tile associated with geometry, and the second value of the type information may indicate an attribute tile associated with an attribute.
[0654] According to embodiments, based on the first value of type information and the first value of the fourth flag (mi_geo_codec_tile_alignment_flag), or based on the second value of type information and the first value of the fifth flag (mi_attr_codec_tile_alignment_flag), the mapping information SEI message may further include a video codec tile index (mi_codec_tile_idx[i][j][k]) for a submesh of an atlas tile.
[0655] According to the embodiments, based on the first value of the third flag, the mapping information SEI message may be configured to provide a video codec tile index in submesh units, and the video codec tile index may be included in a structure (mi_codec_tile_idx[i][j][k]) identified by tile index i, submesh index j, and codec tile index k.
[0656] According to embodiments, the transmitter or encoding device may be configured to include information indicating area units in a mapping information SEI message based on a first value of a first flag and a first value of a second flag, and may also generate a mapping information SEI message including a third flag, a fourth flag, a fifth flag and a video codec tile index and include it in an atlas sub-bitstream.
[0657] According to the embodiments, if a mapping information SEI message exists in an atlas bitstream or an atlas sub-bitstream, a receiver or decoder can parse the said SEI message.
[0658] According to the embodiments, a receiver or decoder can parse the payload of a mapping information SEI message, first parse the persistence flag (mi_persistance_mapping_flag), and then parse the tile count information (mi_num_tiles_minus1), tile ID length information (mi_tile_id_length_minus1), codec tile signal flag (mi_codec_tile_signal_flag), LoD mapping signal flag (mi_LoD_mapping_signal_flag), and attribute mapping signal flag (mi_attribute_mapping_signal_flag) in order.
[0659] According to the embodiments, if at least one of mi_codec_tile_signal_flag, mi_LoD_mapping_signal_flag, mi_attribute_mapping_signal_flag is not present in the message, the flag may be set to 0, and the receiver or decoder may omit parsing additional syntax corresponding to the flag interpreted as 0.
[0660] According to the embodiments, when mi_attribute_mapping_signal_flag is 1, the receiver or decoder may further parse the attribute count (mi_attribute_count), and when mi_attribute_mapping_signal_flag is interpreted as 0, the parsing of mi_attribute_count may not be performed.
[0661] According to the embodiments, when mi_codec_tile_signal_flag is 1, the receiver or decoder may further parse the geocodec tile alignment flag (mi_geo_codec_tile_alignment_flag) and the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag), and then initialize variables (currCount[0], currCount[1]) for managing codec tile indices.
[0662] According to the embodiments, the receiver or decoder may initialize the global index (submeshIdx) of the submesh unit to 0.
[0663] According to the embodiments, a receiver or decoder can iterate over tile index i and parse a tile ID (mi_tile_id[i]) for each tile, and set TileIdxToID[i] to mi_tile_id[i] for mapping between the tile index and the tile ID.
[0664] According to the embodiments, a receiver or decoder can sequentially parse a tile type flag (mi_tile_type_flag[i]), the number of submeshes corresponding to the tile (mi_num_submeshes_minus1[i]), and submesh ID length information (mi_submesh_id_length_minus1[i]) for each tile.
[0665] According to the embodiments, a receiver or decoder may parse a submesh ID (mi_submesh_id[i][j]) while iterating over a submesh index j, and may set SubmeshIdxToID[i][j] to mi_submesh_id[i][j] for mapping between the submesh index and the submesh ID.
[0666] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 0 and mi_LoD_mapping_signal_flag is 1, the receiver or decoder can parse the subdivision iteration count (mi_subdivision_iteration_count[submeshIdx]) for the corresponding submesh and iterate parse the LoD extraction unit index (mi_LoD_extractable_unit_idx[submeshIdx][k]) for the LoD level index k as many times as the number of iterations determined by mi_subdivision_iteration_count[submeshIdx].
[0667] According to the embodiments, when the tile type flag (mi_tile_type_flag[i]) is 1 and mi_attribute_mapping_signal_flag is 1, the receiver or decoder can iteratively parse the attribute extraction unit index (mi_attr_extractable_unit_idx[submeshIdx][k]) by iterating a number of times determined by mi_attribute_count for attribute index k.
[0668] According to the embodiments, whenever LoD or attribute-related parsing for one submesh is completed, the receiver or decoder may increment submeshIdx to update the parsing position of the next submesh.
[0669] According to the embodiments, when mi_codec_tile_signal_flag is 1, the receiver or decoder may parse the number of codec tiles corresponding to the submesh (mi_num_codec_tiles_in_tile_minus1[i][j]) within the submesh iteration phrase (where the tile type flag is 0 and mi_geo_codec_tile_alignment_flag is interpreted as 0) or (where the tile type flag is 1 and mi_attr_codec_tile_alignment_flag is interpreted as 0), and iterate through the codec tile index list (mi_codec_tile_idx[i][j][k]) based on the number.
[0670] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is interpreted as 1, the receiver or decoder may apply an automatic assignment or reset operation of the codec tile index based on the tile type and currCount state. For example, if the tile type flag is 1 and currCount[1] is not 0, the receiver or decoder may parse the codec tile index reset flag (mi_codec_tile_idx_reset_flag[i]), and if the flag is 1, currCount[1] may be initialized to 0. Subsequently, the receiver or decoder may determine mi_codec_tile_idx[i][j][0] as currCount[mi_tile_type_flag[i]] and increase currCount[mi_tile_type_flag[i]].
[0671] A transmitter or encoding device according to the embodiments may configure syntax elements of a mapping information SEI message and include them in an atlas sub-bitstream so that a receiver or decoder can parse them in the order described above.
[0672] FIG. 40 shows a configuration diagram of a dynamic mesh content receiver according to embodiments.
[0673] The receiver of FIG. 40 may include a delivery module, a file or segment decapsulation module, a mesh data decoding module, a mesh data processing / rendering module, an audio decoding module, an audio rendering module, and a sensing / tracking module, and may be coupled with a display and speakers or headphones. Additionally, metadata and orientation / viewport information may be transmitted between the modules to enable playback that adapts to the user's viewpoint.
[0674] The receiver operation according to the embodiments may correspond to the internal operation of the Mesh Data Decoding module of FIG. 40, and in particular, may be applied when information for partial extraction of displacement vector bitstreams by LoD level or partial extraction of attribute regions corresponding to submeshes is signaled as an SEI message.
[0675] According to the embodiments, the stored or received dynamic mesh content may be processed through a delivery module and a file or segment decapsulation module, and the result may include a bitstream in the form of a dynamic mesh bitstream structure such as that shown in FIG. 17.
[0676] According to the embodiments, a bitstream parser may be included within the mesh data decoding module, and the bitstream parser may perform the role of parsing a dynamic mesh content bitstream. The bitstream parser may parse the V3C unit header and V3C unit payload constituting the bitstream.
[0677] According to the embodiments, the bitstream parser can obtain data corresponding to the unit type V3C_VPS defined in FIG. 18, i.e., V3C parameter set (VPS) data, and data corresponding to the unit type V3C_GVD or V3C_AVD, i.e., video sub-bitstream.
[0678] According to embodiments, a receiver or decoder may obtain information for performing partial extraction by LoD level in a displacement video bitstream, or information for performing partial extraction by attribute region corresponding to a submesh in an attribute video bitstream, and / or information for partially decoding a video tile corresponding to an atlas tile, through an SEI message as described by the extraction information SEI payload syntax of FIGS. 35 and 36, or the mapping information SEI payload syntax of FIGS. 37, 38, and 39.
[0679] According to the embodiments, prior to direct decoding of displacement vector data as described above, the receiver may identify information related to displacement vector data encoded for partial extraction by LoD level within the dynamic mesh content bitstream, and / or information related to attribute data encoded for partial extraction by submesh, and / or information related to displacement or attribute data encoded for partial decoding of regions corresponding to atlas tiles within the geometry video bitstream or attribute video bitstream.
[0680] According to the embodiments, in the process of extracting a displacement vector video bitstream, an MCTS index or subpicture index corresponding to the target LoD selected by the user can be used as an input to the bitstream extraction process, and accordingly, a sub-video bitstream can be extracted. Additionally, in the process of extracting an attribute video bitstream, an MCTS index or subpicture index of the attribute region corresponding to the submesh selected by the user for decoding can be used as an input to the bitstream extraction process, and accordingly, a sub-video bitstream can be extracted.
[0681] According to the embodiments, the MCTS-based bitstream extraction process can extract a sub-bitstream corresponding to a specific MCTS according to the process specified in Annex D.3.43 (Motion-constrained tile sets extraction information sets SEI message semantics) of ISO / IEC 23008-2. Additionally, the subpicture-based bitstream extraction process can extract a sub-bitstream corresponding to a specific subpicture according to the process specified in Annex C.7 (Subpicture sub-bitstream extraction process) of ISO / IEC 23090-3. The extracted multiple sub-bitstreams can be merged to create a single sub-bitstream.
[0682] According to the embodiments, in the process of decoding displacement vector video and attribute video data through a video decoder, a video tile index corresponding to an atlas tile selected by the user for decoding may be used, and accordingly, a decoded image for a specific tile area may be obtained.
[0683] Below, the receiver operation process of the SEI message of FIGS. 41 to 43 will be described.
[0684] FIG. 41 illustrates the process of LoD extraction and attribute extraction being performed at a receiver when parsing an extraction information SEI message according to the embodiments.
[0685] FIG. 41 is a flowchart illustrating an example of a receiver operation process according to embodiments, showing the process in which LoD extraction and attribute extraction are performed in the receiver when the extraction information SEI message defined in FIG. 35 and FIG. 36 is parsed.
[0686] According to the embodiments, a receiver or decoder may determine whether an extraction information SEI message exists within an atlas bitstream (or atlas sub-bitstream). If an extraction information SEI message does not exist, the receiver or decoder may not support the extraction function or may be unable to perform the extraction function.
[0687] According to the embodiments, if an extraction information SEI message exists, the receiver or decoder may decide to use the extraction function by considering the user, the receiver environment or settings, and if it decides to use the extraction function, it may parse the extraction information SEI message.
[0688] According to the embodiments, a receiver or decoder can parse an extraction information SEI message and then determine whether the extraction information is information for supporting LoD extraction functions, information for supporting attribute extraction functions, or information for both, based on an extraction type (ei_extraction_type) indicating the type of extraction information. When ei_extraction_type is 0, the extraction information SEI message may include information for supporting LoD extraction functions; when ei_extraction_type is 1, the extraction information SEI message may include information for supporting attribute extraction functions; and when ei_extraction_type is 2, the extraction information SEI message may include information for supporting LoD extraction functions and information for supporting attribute extraction functions.
[0689] According to the embodiments, when information about a bitstream extraction unit is signaled as in the syntax of FIG. 35, a receiver or decoder can parse the extraction unit type index (ei_extractable_unit_type_idx) to identify the type of unit to be extracted from the video bitstream. If the value of ei_extractable_unit_type_idx is 0, the extraction unit may be MCTS (Motion-Constrained Tile Sets), and if the value is 1, the extraction unit may be a subpicture.
[0690] According to the embodiments, when information regarding the bitstream extraction unit is not signaled, such as the syntax of FIG. 36, the receiver or decoder may derive the extraction unit according to the type of video codec being decoded. By checking the type of displacement video or attribute video codec, if the video codec is HEVC, the extraction unit may be derived to MCTS, and if the video codec is VVC, the extraction unit may be derived to subpicture.
[0691] According to the embodiments, if the receiver or decoder supports an LoD extraction function and decides to perform LoD extraction, it may receive as input the target LoD level and target submesh ID information to be restored according to the user or receiver environment.
[0692] According to the embodiments, a receiver or decoder can obtain ID information of each submesh through a submesh ID (ei_submesh_id) from an extraction information SEI message, and can obtain information about the LoD level of each submesh through a subdivision iteration count (ei_subdivision_iteration_count).
[0693] According to the embodiments, a receiver or decoder may obtain an MCTS index or a subpicture index corresponding to each LoD level from an extraction information SEI message. When the syntax of FIG. 35 is applied, the receiver or decoder may obtain an index through the MCTS index of the LoD area (ei_LoD_mcts_idx) or the subpicture index of the LoD area (ei_LoD_subpic_idx), and when the syntax of FIG. 36 is applied, the receiver or decoder may obtain an index through the LoD extraction unit index (ei_LoD_extractable_unit_idx).
[0694] According to embodiments, a receiver or decoder can derive an MCTS index or a subpicture index for a target submesh ID received as input and a submesh LoD level corresponding to the target LoD using parsed extraction information.
[0695] According to the embodiments, a receiver or decoder inputs the derived target MCTS index or target subpicture index into an extraction process prior to decoding a displacement vector video bitstream, thereby extracting only a portion of the video bitstream corresponding to a specific MCTS or a specific subpicture, and decoding the extracted bitstream.
[0696] According to the embodiments, a receiver or decoder can obtain a restored displacement vector by restoring a decoded displacement vector video, and can restore a mesh by performing subdivision up to a target LoD on a decoded base mesh and then applying the restored displacement vector to the subdivided mesh.
[0697] According to the embodiments, if the receiver or decoding device supports an attribute extraction function and decides to perform attribute extraction, it may receive target submesh ID information to be restored as input depending on the user or receiver environment.
[0698] According to the embodiments, a receiver or decoder can obtain ID information of each submesh through a submesh ID (ei_submesh_id) from an extraction information SEI message, and can obtain information about the number of attributes for each submesh through an attribute count (ei_attribute_count).
[0699] According to the embodiments, a receiver or decoder may obtain an MCTS index or a subpicture index for an attribute region corresponding to a submesh from an extraction information SEI message. When the syntax of FIG. 35 is applied, the receiver or decoder may obtain an index through the MCTS index (ei_attr_mcts_idx) of the attribute region or the subpicture index (ei_attr_subpic_idx) of the attribute region, and when the syntax of FIG. 36 is applied, the receiver or decoder may obtain an index through the attribute extraction unit index (ei_attr_extractable_unit_idx).
[0700] According to embodiments, a receiver or decoder can derive an MCTS index or subpicture index of a submesh corresponding to a target submesh ID received as input using parsed extraction information.
[0701] According to the embodiments, a receiver or decoder inputs an induced target MCTS index or a target subpicture index into an extraction process prior to decoding an attribute video bitstream, thereby extracting only a portion of the video bitstream corresponding to a specific MCTS or a specific subpicture, and decoding the extracted bitstream.
[0702] According to embodiments, a receiver or a decoding device can recover a decoded attribute video to obtain a recovered attribute.
[0703] FIG. 42 (Fig. 42a and FIG. 42b) illustrates the process of LoD extraction, attribute extraction, and target atlas tile decoding being performed at a receiver when parsing a mapping information SEI message according to the embodiments.
[0704] Figure 42 is the receiver operation process for the mapping information SEI message of Figures 37 and 38.
[0705] FIG. 42 is a flowchart illustrating an example of a receiver operation process according to embodiments, showing the process in which LoD extraction, attribute extraction, and target atlas tile decoding are performed in the receiver when the mapping information SEI message defined in FIG. 37 to FIG. 39 is parsed.
[0706] According to the embodiments, a receiver or decoder may check whether a mapping information SEI message exists within an atlas bitstream (or atlas sub-bitstream). If a mapping information SEI message does not exist, the receiver or decoder may not support LoD extraction, attribute extraction, or target atlas tile decoding functions, or the performance of such functions may be impossible.
[0707] According to the embodiments, if a mapping information SEI message exists, the receiver or decoder may decide to use at least one of the LoD extraction, attribute extraction, or target atlas tile decoding functions in consideration of the user, receiver environment, or settings, and if it decides to use the function, it may parse the mapping information SEI message.
[0708] According to embodiments, a receiver or decoder can parse a mapping information SEI message and then determine whether the mapping information SEI message includes information for supporting a LoD extraction function, information for supporting an attribute extraction function, or information for supporting a target atlas tile-based decoding function based on flags (mi_LoD_mapping_signal_flag, mi_attribute_mapping_signal_flag, mi_codec_tile_signal_flag) indicating the type of mapping information.
[0709] When mi_LoD_mapping_signal_flag is 1, the mapping information SEI message may include information for supporting LoD extraction functions, when mi_attribute_mapping_signal_flag is 1, the mapping information SEI message may include information for supporting attribute extraction functions, and when mi_codec_tile_signal_flag is 1, the mapping information SEI message may include information for supporting target atlas tile-based decoding functions.
[0710] According to embodiments, when information about a bitstream extraction unit is signaled as in FIG. 37, a receiver or decoder can identify the type of unit to be extracted from the video bitstream by parsing the extraction unit type index (mi_LoD_extractable_unit_type_idx) for the displacement video and the extraction unit type index (mi_attr_extractable_unit_type_idx) for the attribute video. If the value of the extraction unit type index is 0, the extraction unit may be MCTS (Motion-Constrained Tile Sets), and if the value is 1, the extraction unit may be a subpicture.
[0711] According to the embodiments, when information regarding the bitstream extraction unit is not signaled as in FIG. 38, the receiver or decoder may derive the extraction unit according to the type of video codec being decoded. The receiver or decoder identifies the type of displacement video codec or attribute video codec, and if the video codec is HEVC, the extraction unit may be derived to MCTS, and if the video codec is VVC, the extraction unit may be derived to subpicture.
[0712] According to the embodiments, if the receiver or decoder supports an LoD extraction function and decides to perform LoD extraction, it may receive as input the target LoD level and target submesh ID information to be restored according to the user or receiver environment.
[0713] According to the embodiments, a receiver or decoder can obtain ID information of each submesh through the submesh ID (mi_submesh_id) from the mapping information SEI message, and can obtain information about the LoD level of each submesh through the subdivision iteration count (mi_subdivision_iteration_count).
[0714] According to the embodiments, a receiver or decoder may obtain an MCTS index or a subpicture index corresponding to each LoD level from a mapping information SEI message. As shown in FIG. 37, the receiver or decoder may obtain an index through the MCTS index of the LoD area (mi_LoD_mcts_idx) or the subpicture index of the LoD area (mi_LoD_subpic_idx), and as shown in FIG. 38, the receiver or decoder may obtain an index through the LoD extractable unit index (mi_LoD_extractable_unit_idx).
[0715] According to embodiments, a receiver or decoder can derive an MCTS index or a subpicture index for a target submesh ID received as input and a submesh LoD level corresponding to the target LoD using parsed mapping information.
[0716] According to the embodiments, a receiver or decoder inputs the derived target MCTS index or target subpicture index into an extraction process prior to decoding a displacement vector video bitstream, thereby extracting only a portion of the video bitstream corresponding to a specific MCTS or a specific subpicture, and decoding the extracted bitstream.
[0717] According to the embodiments, a receiver or decoder can obtain a restored displacement vector by restoring a decoded displacement vector video, and can restore a mesh by performing subdivision up to a specific LoD on a decoded base mesh and then applying the restored displacement vector to the subdivided mesh.
[0718] According to the embodiments, if the receiver or decoding device supports an attribute extraction function and decides to perform attribute extraction, it may receive target submesh ID information to be restored as input depending on the user or receiver environment.
[0719] According to the embodiments, a receiver or decoder can obtain ID information of each submesh through the submesh ID (mi_submesh_id) from the mapping information SEI message, and can obtain information about the number of attributes through the number of attributes (mi_attribute_count).
[0720] According to the embodiments, a receiver or decoder may obtain an MCTS index or a subpicture index for an attribute region corresponding to a submesh from a mapping information SEI message. As shown in FIG. 37, the receiver or decoder may obtain an index through the MCTS index (mi_attr_mcts_idx) of the attribute region or the subpicture index (mi_attr_subpic_idx) of the attribute region, and as shown in FIG. 38, the receiver or decoder may obtain an index through the attribute extraction unit index (mi_attr_extractable_unit_idx).
[0721] According to the embodiments, a receiver or decoder can derive an MCTS index or subpicture index of a submesh corresponding to a target submesh ID received as input using parsed mapping information.
[0722] According to the embodiments, a receiver or decoder inputs a derived target MCTS index or a derived target subpicture index into an extraction process prior to decoding an attribute video bitstream, thereby extracting only a portion of the video bitstream corresponding to a specific MCTS or a specific subpicture, and decoding the extracted bitstream.
[0723] According to embodiments, a receiver or a decoding device can recover a decoded attribute video to obtain a recovered attribute.
[0724] According to the embodiments, if the receiver or decoding device supports a target atlas tile decoding function and decides to perform the function, it may receive target atlas tile ID information to be restored as input depending on the user or receiver environment.
[0725] According to the embodiments, a receiver or decoder can obtain tile ID information of each atlas tile through a tile ID (mi_tile_id) from a mapping information SEI message, identify the tile type through tile type information (mi_tile_type_flag), and obtain ID information of a submesh within a tile through mi_submesh_id for each submesh.
[0726] According to the embodiments, when the tile type information (mi_tile_type_flag) is 0, the tile may represent an atlas tile associated with geometry, and when the tile type information (mi_tile_type_flag) is 1, the tile may represent an atlas attribute tile.
[0727] According to the embodiments, the geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) and the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) may indicate whether the division of atlas tiles and video tiles is performed identically for geometry video and attribute video, respectively.
[0728] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is 1, the same partitioning is performed, so the receiver or decoder can derive a video tile index corresponding to the atlas tile by increasing the index by the number of atlas tiles or attribute tiles using currCount[0] or currCount[1] depending on the tile type.
[0729] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is 0, the receiver or decoder can obtain video tile information corresponding to the atlas tile by parsing it from the mapping information SEI message.
[0730] A receiver or decoder can obtain the video tile index corresponding to an atlas tile through mi_codec_tile_idx, and if there are multiple video tiles corresponding to an atlas tile, it can obtain the number of video tiles through the number of video tiles information (mi_num_codec_tiles_in_tile_minus1).
[0731] According to the embodiments, a receiver or decoder can derive a video tile index corresponding to a target atlas tile ID based on a mapping information SEI message.
[0732] According to the embodiments, when the type of the target atlas tile is a geometry atlas tile, the receiver or decoder can perform partial decoding on the displacement video corresponding to the target atlas tile by using the target video tile index as the input for video decoding during the process of decoding the displacement video.
[0733] According to the embodiments, when the type of the target atlas tile is an attribute atlas tile, the receiver or decoder may perform partial decoding on the attribute video corresponding to the target atlas tile by using the target video tile index as the input for video decoding during the process of decoding the attribute video.
[0734] According to the embodiments, a receiver or a decoder can restore a partially decoded displacement video or a partially decoded attribute video to obtain the restored displacement or the restored attribute, respectively.
[0735] FIG. 43 (Fig. 43a and FIG. 43b) illustrates the process of decoding video tiles of target atlas tiles and sub-meshes at a receiver when parsing mapping information SEI messages according to the embodiments.
[0736] Figure 43 is the receiver operation process for the mapping information SEI message of Figure 39.
[0737] FIG. 43 is a flowchart illustrating an example of a receiver operation process according to embodiments, showing the process in which video tile decoding for a target atlas tile and a target submesh is performed at the receiver when the mapping information SEI message defined in FIG. 39 is parsed.
[0738] According to the embodiments, a receiver or decoder may check whether a mapping information SEI message exists within an atlas bitstream (or atlas sub-bitstream). If a mapping information SEI message does not exist, the receiver or decoder may not support an extraction function or a video tile decoding function for a target atlas tile and a target sub-mesh, or the performance of such functions may be impossible.
[0739] According to the embodiments, if a mapping information SEI message exists, the receiver or decoder may decide to use at least one of an extraction function or a video tile decoding function for a target atlas tile and a target submesh, taking into consideration the user, receiver environment or settings, and if it decides to use the function, it may parse the mapping information SEI message.
[0740] According to embodiments, a receiver or decoder can parse a mapping information SEI message and then determine whether the mapping information includes information for supporting LoD extraction, information for supporting attribute extraction, or information for supporting target atlas tile-based decoding, based on flags (mi_LoD_mapping_signal_flag, mi_attribute_mapping_signal_flag, mi_codec_tile_signal_flag) indicating the type of mapping information. When mi_LoD_mapping_signal_flag is 1, the mapping information SEI message may include information for supporting LoD extraction; when mi_attribute_mapping_signal_flag is 1, the mapping information SEI message may include information for supporting attribute extraction; and when mi_codec_tile_signal_flag is 1, the mapping information SEI message may include information for supporting target atlas tile-based decoding.
[0741] According to the embodiments, the operation process of LoD extraction and attribute extraction can be performed in the same or similar manner as the receiver operation process of FIG. 42 described above.
[0742] According to the embodiments, if the receiver or decoding device supports a video tile decoding function for a target atlas tile and a target submesh and decides to perform the function, it may receive target atlas tile ID and target submesh ID information to be restored as input depending on the user or receiver environment.
[0743] According to the embodiments, a receiver or decoder can obtain tile ID information of each atlas tile through a tile ID (mi_tile_id) from a mapping information SEI message, identify the tile type through tile type information (mi_tile_type_flag), and obtain ID information of a submesh within a tile through mi_submesh_id for each submesh.
[0744] According to the embodiments, when the tile type information (mi_tile_type_flag) is 0, it may represent an atlas tile associated with geometry, and when the tile type information (mi_tile_type_flag) is 1, the tile may represent an atlas attribute tile.
[0745] According to the embodiments, the geometry codec tile alignment flag (mi_geo_codec_tile_alignment_flag) and the attribute codec tile alignment flag (mi_attr_codec_tile_alignment_flag) may indicate whether the division of atlas tiles and video tiles is performed identically for geometry video and attribute video, respectively.
[0746] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is 1, the receiver or decoder can derive a video tile index corresponding to an atlas tile by increasing the index by the number of atlas tiles or attribute tiles using currCount[0] or currCount[1] depending on the tile type.
[0747] According to the embodiments, when mi_geo_codec_tile_alignment_flag or mi_attr_codec_tile_alignment_flag is 0, the receiver or decoder can obtain video tile information corresponding to the atlas tile by parsing it from the mapping information SEI message.
[0748] According to the embodiments, a receiver or decoder can parse video tile information corresponding to each submesh within each atlas tile. The receiver or decoder can obtain video tile index information corresponding to each submesh ID through mi_codec_tile_idx[i][j][k], and if there are multiple video tiles corresponding to each submesh ID, it can obtain the number of video tiles corresponding to the submesh through information on the number of video tiles (mi_num_codec_tiles_in_tile_minus1[i][j]).
[0749] According to the embodiments, the receiver or decoder can derive a video tile index corresponding to a target atlas tile ID and a target submesh ID based on parsed mapping information.
[0750] According to the embodiments, when the type of the target atlas tile is a geometry atlas tile, the receiver or decoder can perform partial decoding on the displacement video corresponding to the target atlas tile and the target submesh by using the target video tile index as an input for video decoding during the process of decoding the geometry video, i.e., the displacement video.
[0751] According to the embodiments, when the type of the target atlas tile is an attribute atlas tile, the receiver or decoder can perform partial decoding on the attribute video corresponding to the target atlas tile and the target submesh by using the target video tile index as an input for video decoding during the process of decoding the attribute video.
[0752] According to the embodiments, a receiver or a decoder can restore a partially decoded displacement video or a partially decoded attribute video to obtain the restored displacement or the restored attribute, respectively.
[0753] FIG. 44 illustrates an encoding method according to embodiments.
[0754] The encoding method according to the embodiments may include the step of encoding base mesh data of the mesh data (S4400); the step of encoding displacement data of the mesh data (S4410); and the step of encoding attribute data of the mesh data (S4420).
[0755] The encoding method according to the embodiments may include the method described in FIGS. 1 to 43 above. Specifically, the step of encoding base mesh data (S4400) may include the method described in FIGS. 1 to 7, FIG. 12, FIG. 17, FIG. 19, FIG. 35 to 39 above. The step of encoding displacement data of mesh data (S4410) may include the method described in FIGS. 1 to 4, FIGS. 6 to 9, FIG. 12, FIG. 17, FIG. 19, FIG. 20 to 29, FIG. 35 to 39 above. The step of encoding attribute data of mesh data (S4420) may include the method described in FIGS. 1 to 4, FIG. 10, FIG. 12, FIG. 17, FIG. 19, FIG. 20 to 29, FIG. 35 to 39 above.
[0756] Referring together to FIG. 1 and FIG. 17, the encoding method according to the embodiments further comprises the step of generating a bitstream containing encoded mesh data, and the bitstream may include an atlas sub-bitstream; a basemesh sub-bitstream containing basemesh data; a displacement sub-bitstream containing displacement data; and an attribute sub-bitstream containing attribute data.
[0757] Referring together to FIG. 35 and FIG. 36, the encoding method according to the embodiments includes a displacement sub-bitstream comprising a displacement sequence parameter set raw byte sequence payload (displ_sequence_parameter_set_rbsp), and the displacement sequence parameter set raw byte sequence payload may include first information regarding the number of displacement values within a displacement frame for displacement data.
[0758] Referring together with FIG. 35, the encoding method according to the embodiments further includes information about a partial area extraction unit in the extracted SEI message information, and based on a first value of the information about the partial area extraction unit, at least one of the index of the extraction area for the Level of Detail (LoD) of the sub-mesh in the displacement sub-bitstream or the index of the extraction area for an attribute of the attributes of the sub-mesh in the attribute sub-bitstream represents an index of the area encoded as MCTS (Motion-constrained Tile Sets), and based on a second value of the information about the partial area extraction unit, at least one of the index of the extraction area for the Level of Detail (LoD) of the sub-mesh in the displacement sub-bitstream or the index of the extraction area for an attribute of the attributes of the sub-mesh in the attribute sub-bitstream represents an index of the area encoded as a sub-picture.
[0759] Referring together to FIG. 37 and FIG. 38, the encoding method according to the embodiments further comprises the step of generating a mapping information SEI message within an atlas sub-bitstream, wherein the mapping information SEI message comprises: a first flag regarding whether to signal Level of Detail (LoD) information mapped to partitioning information of a displacement sub-bitstream; a second flag regarding whether to signal sub-mesh information mapped to partitioning information of an attribute sub-bitstream; and information for a type of an atlas tile for the atlas sub-bitstream. The mapping information SEI message includes information about the ID of the sub-mesh of the atlas tile and, based on the first value of the second flag, further includes information for a number of attributes related to the basemesh data; based on the first value of the information about the type of the atlas tile and the first value of the first flag, the mapping information SEI message further includes information about the number of subdivisions of the sub-mesh; and further includes an index of an extraction area for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream; and based on the second value of the information about the type of the atlas tile and the first value of the second flag, the mapping information SEI message may further include an index of an extraction area for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream.
[0760] Referring together with FIG. 37, the encoding method according to the embodiments may include, based on the first value of the first flag, the mapping information SEI message further includes information indicating the area unit for the extraction area for the Level of Detail (LoD) of the sub-mesh, and based on the first value of the second flag, the mapping information SEI message further includes information indicating the area unit for the extraction area for the attribute of the sub-mesh.
[0761] Referring together with FIG. 39, the encoding method according to the embodiments further includes a mapping information SEI message, which further includes a third flag indicating whether the tile information of the video codec corresponding to the atlas tile is signaled, and based on a first value of the third flag, the mapping information SEI message includes a fourth flag indicating whether the tile division structure in the geometry video codec and the division structure of the atlas tile are the same; and a fifth flag indicating whether the tile division structure in the attribute video codec and the division structure of the atlas tile are the same, and based on a first value of information about the type of the atlas tile and a first value of the fourth flag, or a second value of information about the type of the atlas tile and a first value of the fifth flag, the mapping information SEI message may further include a video codec tile index for the submesh of the atlas tile.
[0762] The encoding method is performed by an encoding device. The encoding device includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: encode base mesh data of the mesh data; encode displacement data of the mesh data; and encode attribute data of the mesh data.
[0763] The embodiments may further include a computer-readable storage medium that stores a bitstream generated by the method of FIG. 44.
[0764] The embodiments may further include a method comprising the steps of: acquiring a bitstream for mesh data; generating the bitstream based on the steps of: encoding base mesh data of the mesh data; encoding displacement data of the mesh data; and encoding attribute data of the mesh data; and transmitting data including the bitstream.
[0765] FIG. 45 illustrates a decoding method according to embodiments.
[0766] The decoding method according to the embodiments may include the step of decoding base mesh data in a bitstream (S4500); the step of decoding displacement data in a bitstream (S4510); and the step of decoding attribute data in a bitstream (S4520).
[0767] The decoding method according to the embodiments may include the method described in FIGS. 1 to 43 above. Specifically, the step of decoding basemesh data in the bitstream (S4500) may include the method described in FIGS. 1, 2, 11, 13, 17, 30, 40, etc. above. The step of decoding displacement data in the bitstream (S4510) may include the method described in FIGS. 1, 2, 11, 13, 17, 30 to 43, etc. above. The step of decoding attribute data in the bitstream (S4520) may include the method described in FIGS. 1, 2, 11, 13, 17, 30 to 43, etc. above.
[0768] Referring to FIG. 1 and FIG. 17 together, the decoding method according to the embodiments may include a bitstream comprising: an atlas sub-bitstream; a basemesh sub-bitstream comprising basemesh data; a displacement sub-bitstream comprising displacement data; and an attribute sub-bitstream comprising attribute data.
[0769] Referring together to FIG. 35 and FIG. 36, the decoding method according to the embodiments further comprises the step of obtaining an Extraction_information SEI message within an atlas sub-bitstream, wherein the Extraction SEI message information includes: extraction type information indicating a target for partial region extraction for at least one of a displacement sub-bitstream or an attribute sub-bitstream; information for a number of attributes related to the basemesh data; and information for an ID of a sub-mesh within the basemesh data, and based on a first value or a third value of the Extraction type information, the Extraction SEI information includes information for the number of sub-divisions of the sub-mesh; and further include an index of an extraction region for the Level of Detail (LoD) of a sub-mesh within a displacement sub-bitstream, and based on a second or third value of the extraction type information, the extraction SEI information may further include an index of an extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream.
[0770] Referring together with FIG. 35, the decoding method according to the embodiments further includes information about a partial area extraction unit in the extracted SEI message information, and based on a first value of the information about the partial area extraction unit, at least one of the index of the extraction area for the Level of Detail (LoD) of the sub-mesh in the displacement sub-bitstream or the index of the extraction area for an attribute of the attributes of the sub-mesh in the attribute sub-bitstream represents an index of the area encoded as MCTS (Motion-constrained Tile Sets), and based on a second value of the information about the partial area extraction unit, at least one of the index of the extraction area for the Level of Detail (LoD) of the sub-mesh in the displacement sub-bitstream or the index of the extraction area for an attribute of the attributes of the sub-mesh in the attribute sub-bitstream represents an index of the area encoded as a sub-picture.
[0771] Referring together to FIG. 37 and FIG. 38, the decoding method according to the embodiments further comprises the step of acquiring a mapping information SEI message within an atlas sub-bitstream, wherein the mapping information SEI message comprises: a first flag regarding whether to signal Level of Detail (LoD) information mapped to partitioning information of a displacement sub-bitstream; a second flag regarding whether to signal sub-mesh information mapped to partitioning information of an attribute sub-bitstream; and information for a type of an atlas tile for the atlas sub-bitstream. The mapping information SEI message includes information about the ID of the sub-mesh of the atlas tile and, based on the first value of the second flag, further includes information for a number of attributes related to the basemesh data; based on the first value of the information about the type of the atlas tile and the first value of the first flag, the mapping information SEI message further includes information about the number of subdivisions of the sub-mesh; and further includes an index of an extraction area for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream; and based on the second value of the information about the type of the atlas tile and the first value of the second flag, the mapping information SEI message may further include an index of an extraction area for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream.
[0772] Referring together with FIG. 37, the decoding method according to the embodiments may include, based on the first value of the first flag, the mapping information SEI message further includes information indicating the area unit for the extraction area for the Level of Detail (LoD) of the sub-mesh, and based on the first value of the second flag, the mapping information SEI message further includes information indicating the area unit for the extraction area for the attribute of the sub-mesh.
[0773] Referring together with FIG. 39, the decoding method according to the embodiments further includes a mapping information SEI message, a third flag indicating whether the tile information of the video codec corresponding to the atlas tile is signaled, and based on a first value of the third flag, the mapping information SEI message includes a fourth flag indicating whether the tile partitioning structure in the geometry video codec and the partitioning structure of the atlas tile are the same; and a fifth flag indicating whether the tile partitioning structure in the attribute video codec and the partitioning structure of the atlas tile are the same, and based on a first value of information about the type of the atlas tile and a first value of the fourth flag, or a second value of information about the type of the atlas tile and a first value of the fifth flag, the mapping information SEI message may further include a video codec tile index for the submesh of the atlas tile.
[0774] The decoding method is performed by a decoding device. The decoding device includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: decode basemesh data in a bitstream; decode displacement data in a bitstream; and decode attribute data in a bitstream.
[0775] The embodiments can reduce the number of bits for syntax (submesh count, submesh ID, etc.) that were redundantly signaled in the prior art by combining SEI messages for signaling mapping information between video codec partition information (e.g., tile, MCTS, subpicture, etc.) and V-DMC component information into a single message.
[0776] In addition, the embodiments can support a submesh-based video tile decoding function by solving the problem in which the video tile index corresponding to the submesh cannot be known when there are two or more submeshes in an atlas tile in a conventional tile submesh mapping SEI message.
[0777] The embodiments have been described in terms of methods and / or devices, and the description of the methods and the description of the devices may be applied complementarily.
[0778] Although the drawings have been described separately for the convenience of explanation, it is also possible to design a new embodiment by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by a person skilled in the art, falls within the scope of the claims of the embodiments. The apparatus and method according to the embodiments are not limited to the configuration and method of the embodiments described above; rather, the embodiments may be configured by selectively combining all or part of each embodiment to allow for various modifications. Although preferred embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above. It is not only possible for a person skilled in the art to make various modifications without departing from the essence of the embodiments claimed in the claims, but such modifications should not be understood individually from the technical concept or perspective of the embodiments.
[0779] Various components of the device of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented as a single chip, for example, a single hardware circuit. Depending on the embodiments, the components according to the embodiments may each be implemented as separate chips. Depending on the embodiments, at least one of the components of the device according to the embodiments may be composed of one or more processors capable of executing one or more programs, and one or more programs may include instructions for performing or executing any one or more of the operations / methods according to the embodiments. Executable instructions for performing the methods / operations of the device according to the embodiments may be stored in non-transient CRMs or other computer program products configured to be executed by one or more processors, or may be stored in transient CRMs or other computer program products configured to be executed by one or more processors. Additionally, memory according to the embodiments may be used as a concept that includes not only volatile memory (e.g., RAM, etc.) but also non-volatile memory, flash memory, PROM, etc. In addition, it may also include implementation in the form of carrier waves, such as transmission over the Internet. Furthermore, processor-readable recording media are distributed across networked computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[0780] In this document, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B and / or C." Also, "A, B, C" means "at least one of A, B and / or C." Additionally, in this document, "or" is interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." In other words, "or" in this document may mean "additionally or alternatively."
[0781] Terms such as "first," "second," etc., may be used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not mean the same user input signals unless clearly indicated in the context.
[0782] The terms used to describe the embodiments are intended for the purpose of describing specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless explicitly indicated in the context. Expressions of and / or are used to mean including all possible combinations between the terms. Expressions of include describe the presence of features, numbers, steps, elements, and / or components and do not imply the exclusion of additional features, numbers, steps, elements, and / or components. Conditional expressions such as "if" or "when" used to describe the embodiments are not limited to being optional. It is intended to be interpreted as "when a specific condition is satisfied," "when a related action is performed in response to a specific condition," or "when a related definition is interpreted."
[0783] Additionally, operations according to the embodiments described herein may be performed by a transmitting and receiving device including memory and / or a processor, depending on the embodiments. The memory may store programs for processing / controlling operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. Operations in the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in memory.
[0784] Meanwhile, the operation according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting and receiving device may include a transmitting and receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting and receiving devices.
[0785] The processor may be referred to as a controller, etc., and may correspond, for example, to hardware, software, and / or a combination thereof. The operation according to the embodiments described above may be performed by the processor. Additionally, the processor may be implemented as an encoder / decoder, etc., for the operation of the embodiments described above.
[0786]
[0787] As described above, the relevant details have been explained in the best mode for carrying out the embodiments.
[0788]
[0789] As described above, the embodiments may be applied wholly or partially to point cloud data transmission and reception devices and systems.
[0790] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.
[0791] The embodiments may include modifications / variations, and such modifications / variations do not exceed the scope of the claims and their equivalents.
Claims
1. A step of decoding basemesh data within a bitstream; A step of decoding displacement data within the bitstream; and A step of decoding attribute data within the bitstream; comprising Decryption method.
2. In Paragraph 1, The above bitstream is, Atlas sub-bitstream; A basemesh sub-bitstream including the above basemesh data; A displacement sub-bitstream including the above displacement data; and including an attribute sub-bitstream containing the above attribute data, Decryption method.
3. In Paragraph 2, The above method further includes the step of obtaining an Extraction_information SEI message within the atlas sub-bitstream, and The above extracted SEI message information is, Extraction type information indicating a target for partial region extraction for at least one of the displacement sub-bitstream or the attribute sub-bitstream; Information for a number of attributes related to the basemesh data; and Includes information on the ID of the sub-mesh within the above base mesh data, and Based on the first or third value of the above extraction type information, the above extraction SEI information is, Information regarding the number of subdivisions of the above sub-mesh; and It further includes an index of an extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream, and Based on the second or third value of the above extraction type information, the above extraction SEI information is, Further including an index of an extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream, Decryption method.
4. In Paragraph 3, The above extracted SEI message information further includes information regarding the above partial area extraction unit, and Based on the first value of the information regarding the above partial area extraction unit, At least one of the index of the extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream or the index of the extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream represents the index of the region encoded in Motion-constrained Tile Sets (MCTS), Based on the second value of the information regarding the partial region extraction unit, at least one of the index of the extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream or the index of the extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream represents the index of the region encoded as a sub-picture. Decryption method.
5. In Paragraph 2, The above method further includes the step of obtaining a mapping information SEI message within the atlas sub-bitstream, and The above mapping information SEI message is, A first flag regarding whether to signal Level of Detail (LoD) information mapped to the partitioning information of the above displacement sub-bitstream; A second flag regarding whether to signal sub-mesh information mapped to the partitioning information of the above attribute sub-bitstream; Information for a type of an atlas tile for the atlas sub-bitstream; and Includes information on the ID of the sub-mesh of the above atlas tile, and Based on the first value of the second flag above, the mapping information SEI message is, It further includes information for a number of attributes related to the basemesh data, and Based on the first value of the information regarding the type of the atlas tile and the first value of the first flag, the mapping information SEI message is, Information regarding the number of subdivisions of the above sub-mesh; and It further includes an index of an extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream, and Based on the second value of the information regarding the type of the atlas tile and the first value of the second flag, the mapping information SEI message is, Further including an index of an extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream, Decryption method.
6. In Paragraph 5, Based on the first value of the first flag, the mapping information SEI message is, It further includes information indicating area units for the extraction area for the Level of Detail (LoD) of the above sub-mesh, and Based on the first value of the second flag, the mapping information SEI message is, Further including information indicating the area unit of the extraction region for the attribute of the sub-mesh, Decryption method.
7. In Paragraph 5, The above mapping information SEI message further includes a third flag indicating whether tile information of a video codec corresponding to the atlas tile is signaled, and Based on the first value of the third flag above, the mapping information SEI message is, A fourth flag indicating whether the tile division structure in the geometry video codec and the division structure of the atlas tile are identical; and It includes a fifth flag indicating whether the tile partitioning structure in the attribute video codec and the partitioning structure of the atlas tile are identical, and Based on the first value of the information regarding the type of the atlas tile and the first value of the fourth flag, or the second value of the information regarding the type of the atlas tile and the first value of the fifth flag, the mapping information SEI message, A video codec tile index for the submesh of the above atlas tile, further comprising Decryption method.
8. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Decoding basemesh data within a bitstream; Decoding displacement data within the above bitstream; and Configured to decode attribute data within the above bitstream, Decoding device.
9. Step of encoding the base mesh data of the mesh data; A step of encoding displacement data of the above mesh data; and A step of encoding attribute data of the above mesh data; comprising Encoding method.
10. In Paragraph 9, The above method is, The method further includes the step of generating a bitstream containing the above-mentioned encoded mesh data, and The above bitstream is, Atlas sub-bitstream; A basemesh sub-bitstream including the above basemesh data; A displacement sub-bitstream including the above displacement data; and including an attribute sub-bitstream containing the above attribute data, Encoding method.
11. In Paragraph 10, The above displacement sub-bitstream includes a displacement sequence parameter set raw byte sequence payload (displ_sequence_parameter_set_rbsp), and The above displacement sequence parameter set low byte sequence payload is, including first information regarding the number of displacement values within the displacement frame for the above displacement data, Encoding method.
12. In Paragraph 11, The above extracted SEI message information further includes information regarding the above partial area extraction unit, and Based on the first value of the information regarding the above partial area extraction unit, At least one of the index of the extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream or the index of the extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream represents the index of the region encoded in Motion-constrained Tile Sets (MCTS), Based on the second value of the information regarding the partial region extraction unit, at least one of the index of the extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream or the index of the extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream represents the index of the region encoded as a sub-picture. Encoding method.
13. In Paragraph 10, The above method further includes the step of generating a mapping information SEI message within the atlas sub-bitstream, and The above mapping information SEI message is, A first flag regarding whether to signal Level of Detail (LoD) information mapped to the partitioning information of the above displacement sub-bitstream; A second flag regarding whether to signal sub-mesh information mapped to the partitioning information of the above attribute sub-bitstream; Information for a type of an atlas tile for the atlas sub-bitstream; and Includes information on the ID of the sub-mesh of the above atlas tile, and Based on the first value of the second flag above, the mapping information SEI message is, It further includes information for a number of attributes related to the basemesh data, and Based on the first value of the information regarding the type of the atlas tile and the first value of the first flag, the mapping information SEI message is, Information regarding the number of subdivisions of the above sub-mesh; and It further includes an index of an extraction region for the Level of Detail (LoD) of the sub-mesh within the displacement sub-bitstream, and Based on the second value of the information regarding the type of the atlas tile and the first value of the second flag, the mapping information SEI message is, Further including an index of an extraction region for an attribute of the attributes of the sub-mesh within the attribute sub-bitstream, Encoding method.
14. In Paragraph 13, Based on the first value of the first flag, the mapping information SEI message is, It further includes information indicating area units for the extraction area for the Level of Detail (LoD) of the above sub-mesh, and Based on the first value of the second flag, the mapping information SEI message is, Further including information indicating the area unit of the extraction region for the attribute of the sub-mesh, Encoding method.
15. In Paragraph 13, The above mapping information SEI message further includes a third flag indicating whether tile information of a video codec corresponding to the atlas tile is signaled, and Based on the first value of the third flag above, the mapping information SEI message is, A fourth flag indicating whether the tile division structure in the geometry video codec and the division structure of the atlas tile are identical; and It includes a fifth flag indicating whether the tile partitioning structure in the attribute video codec and the partitioning structure of the atlas tile are identical, and Based on the first value of the information regarding the type of the atlas tile and the first value of the fourth flag, or the second value of the information regarding the type of the atlas tile and the first value of the fifth flag, the mapping information SEI message, A video codec tile index for the submesh of the above atlas tile, further comprising Encoding method.
16. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Encoding the base mesh data of the mesh data; Encoding the displacement data of the above mesh data; and Configured to encode the attribute data of the above mesh data; Encoding device.
17. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 9.
18. Step of acquiring a bitstream for mesh data, The bitstream is generated based on the steps of: encoding base mesh data of the mesh data; encoding displacement data of the mesh data; and encoding attribute data of the mesh data; and A method comprising the step of transmitting data including the bitstream above.