Mesh data encoding apparatus, mesh data encoding method, mesh data decoding apparatus, and mesh data decoding method
Patent Information
- Application Number
- PCT/KR2025/099576
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-03-05
- Publication Date
- 2025-10-02
AI Technical Summary
The sheer number of points in 3D space makes it difficult to generate point cloud data, requiring significant processing power for transmission and reception, and existing technologies face challenges in latency and encoding/decoding complexity.
A method for encoding and decoding point cloud data using Video-based Dynamic Mesh (V-Mesh) compression, which includes preprocessing, encoding, transmission, and decoding processes to efficiently transmit and receive point clouds, utilizing techniques such as mesh decimation, UV parameterization, and displacement calculation.
Enables high-quality point cloud services with reduced latency and complexity, supporting applications like VR, AR, MR, and autonomous driving by optimizing point cloud data transmission and reception.
Smart Images

Figure KR2025099576_02102025_PF_FP_ABST
Abstract
Description
Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
[0001] The embodiments provide a method for providing Point Cloud content to provide users with various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and autonomous driving services.
[0002] A point cloud is a collection of points in 3D space. The sheer number of points in 3D space makes it difficult to generate point cloud data.
[0003] There is a problem that a lot of processing power is required to transmit and receive point cloud data.
[0004] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device and reception method for efficiently transmitting and receiving point clouds in order to solve the problems described above.
[0005] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device, and reception method for resolving latency and encoding / decoding complexity.
[0006] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0007] A decoding method according to embodiments may include the steps of: receiving a file including a bitstream including mesh data; decapsulating the file; and decoding the mesh data. An encoding method according to embodiments may include the steps of: encoding the mesh data; and encapsulating a file including a bitstream including mesh data; and transmitting the file.
[0008] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide a high-quality point cloud service.
[0009] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can achieve various video codec methods.
[0010] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide general-purpose point cloud content such as autonomous driving services.
[0011] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0012] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0013] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0014] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0015] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0016] Figure 5 shows a displacement generation process according to embodiments.
[0017] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0018] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0019] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0020] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0021] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0022] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0023] Fig. 12 shows a V-MES and Fig. 13 shows a point cloud data transmission device according to embodiments.
[0024] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0025] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0026] Figure 15 shows a V-DMC bitstream according to embodiments.
[0027] Figure 16 shows a V3C unit according to embodiments.
[0028] Figure 17 shows a mesh system according to embodiments.
[0029] Figure 18 shows the tracks of a file according to embodiments.
[0030] Figure 19 shows the tracks of a file according to embodiments.
[0031] Fig. 20 shows a base mesh track according to embodiments.
[0032] Figure 21 shows an atlas track according to embodiments.
[0033] Fig. 22 shows a receiving device according to embodiments.
[0034] Figure 23 shows the relationship between atlas tiles and sub-meshes according to embodiments.
[0035] Figure 24 shows an encoding method according to embodiments.
[0036] Figure 25 shows a decryption method according to embodiments.
[0037] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0038] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0039] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0040] The system of FIG. 1 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a reception unit (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or a dynamic mesh video encoder (hereinafter, referred to as an encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to a receiving device (110) or a dynamic mesh video decoder (hereinafter, decoder) (113).
[0041] The system of Fig. 1 can perform video-based dynamic mesh compression and decompression.
[0042] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.
[0043] First, the method of compressing dynamic mesh data starts with the V-PCC (Video-based point cloud compression) standard technology. Point cloud data is data that contains color information at the vertex coordinates (X, Y, Z). Mesh data refers to data in which connectivity information between vertices is added to this vertex information. When creating content, it can be created in the form of mesh data from the beginning. By adding connectivity information to point cloud data, it can be converted into mesh data and used.
[0044] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.
[0045] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.
[0046] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.
[0047] A mesh compression system may include a transmitting device and a receiving device. The transmitting device can encode mesh video to output a bitstream, which can be delivered to the receiving device via digital storage media or a network in the form of a file or streaming segment. The digital storage media may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, or SSD.
[0048] The transmitting device may roughly include a mesh video acquisition unit, a mesh video encoder, and a transmitting unit. The receiving device may roughly include a receiving unit, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / video / picture / frame encoding device, and the decoder may be referred to as a mesh video / video / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device and the receiving device may further include separate internal or external modules / units / components for a feedback process.
[0049] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by its vertices in 3D space and connection information that describes how those vertices are connected. It can also contain vertex properties such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included as a mesh property. The mapping is typically described as a set of parametric coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement.
[0050] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc. into a mesh data type with the properties described above through a series of processes and generating a video composed of such mesh data. The mesh video may have properties of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video with properties and connection information that change over time can be expressed as a dynamic mesh video.
[0051] A mesh video encoder can encode an input mesh video into one or more video streams. A single video can include multiple frames, and a single frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and the mesh video can be used interchangeably with the mesh image / frame / picture. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. A mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0052] The encapsulation processing unit (file / segment encapsulation module) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from the metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder, or may be configured as a separate component / module. The encapsulation processing unit may encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The encapsulation processing unit may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. According to an embodiment, the encapsulation processing unit may encapsulate mesh video-related metadata itself in a file.
[0053] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit, or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process the same for transmission.
[0054] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and can include an element for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.
[0055] The receiver can receive mesh video data transmitted by a mesh video transmission device. Depending on the transmission channel, the receiver can receive mesh video data via a broadcast network, via broadband, or via digital storage media.
[0056] The receiving processing unit can perform processing on the received mesh video data according to a transmission protocol. The receiving processing unit can be included in the receiving unit, or can be configured as a separate component / module. In response to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmitting processing unit. The receiving processing unit can transfer the acquired mesh video data to the decapsulation processing unit, and transfer the acquired mesh video-related metadata to a metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
[0057] A decapsulation processing unit (file / segment decapsulation module) can decapsulate mesh video data in file format received from a receiving processing unit. The decapsulation processing unit can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder, or may be configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation processing unit may be in the form of a box or track within a file format. If necessary, the decapsulation processing unit may receive metadata required for decapsulation from the metadata processing unit. Mesh video related metadata can be passed to a Mesh video decoder for use in the Mesh video decoding process, or passed to a renderer for use in the Mesh video rendering process.
[0058] A mesh video decoder can receive a bitstream and perform operations corresponding to those of a mesh video encoder to decode video / images. The decoded mesh video can be displayed via a display unit. Users can view all or part of the rendered result via a VR / AR display or a general display.
[0059] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
[0060] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
[0061] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0062] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.
[0063] This document relates to dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and properties that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-view video, and AR / VR.
[0064] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0065] In this document, picture / frame can generally mean a unit representing one video of a specific time period.
[0066] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0067] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0068] The encoding process of Fig. 1 is as follows.
[0069] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. The V-Mesh compression process receives the following data as input and performs compression.
[0070] Input mesh: Contains the 3D coordinates (geometry) of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.
[0071] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the properties (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video (or attribute for short). The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).
[0072] Material Library File: Contains information about the material properties used in a mesh, particularly information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
[0073] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0074] Base mesh: The input mesh is simplified (decimated) through a preprocessing process to express the objects of the input mesh using the minimum number of vertices determined by the user's standards.
[0075] Displacement: This is displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.
[0076] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.
[0077] Referring to FIGS. 2 to 7, a method for encoding mesh position information (vertex) is described, and referring to FIGS. 6-10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0078] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0079] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The encoder of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. The transmitting device of Fig. 1 may be broadly referred to as an encoder, and the dynamic mesh video encoder of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing (200) and an encoding (201) process as in Fig. 2. The pre-processor of Fig. 2 may be located in front of the encoder of Fig. 2. The pre-processor and the encoder of Fig. 2 may be referred to as a single encoder.
[0080] The preprocessor can receive a static dynamic mesh and / or an attribute map. The preprocessor can generate a base mesh and / or displacement through preprocessing. The preprocessor can receive feedback information from the encoder and generate the base mesh and / or displacement based on the feedback information.
[0081] The encoder can receive a base mesh, displacement mesh, static dynamic mesh, and / or attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.
[0082] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0083] Figure 3 shows the configuration and operation of the pre-processor of Figure 2.
[0084] Fig. 3 shows a process of performing preprocessing on an input mesh. The preprocessing process (200) can be broadly divided into four steps: 1) Group of Frame (GoF) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). The preprocessor (200) can receive an input mesh, generate a displacement and / or base mesh, and transmit the generated displacement and / or base mesh to the encoder (201). The preprocessor (200) can transmit GoF information related to GoF generation to the encoder (201).
[0085] Below, each step of Figure 3 is described.
[0086] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, number of texture coordinates, vertex connection information, and texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame performs intra-frame encoding.
[0087] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.
[0088] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0089] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.
[0090] Fitting subdivision surface: This is the process of performing subdivision on a simplified mesh. The subdivision method can be a user-defined method, such as the mid-edge method. The fitting process ensures that the input mesh and the subdivision mesh are similar to each other.
[0091] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0092] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex midway between the edges between the vertices.
[0093] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decrypted base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents a position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0094] Figure 5 shows a displacement generation process according to embodiments.
[0095] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.
[0096] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a reconstructed base mesh and generate a subdivided reconstructed base mesh. The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and transform a coordinate system of the mesh into a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.
[0097] The method and device for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data (which may be referred to as a point cloud for short) according to the embodiments can refer to data including vertex coordinates and color information. The term "point cloud" includes mesh data, and in this document, point cloud and mesh data can be used interchangeably.
[0098] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).
[0099] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.
[0100] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0101] The encoding process of Fig. 6 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an intra-frame method. The encoder of Fig. 6 may include a preprocessor (200) and / or an encoder (201).
[0102] The preprocessor can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivided mesh. The quantizer can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder can encode the static mesh. The static mesh encoder can generate a bitstream including the encoded base mesh. The static mesh decoder can decode the encoded static mesh. The inverse quantizer can inversely quantize the quantized static mesh. The displacement calculation unit can receive the reconstructed static mesh and generate displacement, which is a position difference, based on the fitted subdivided mesh. The forward linear lifting unit can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. The image packing unit can pack an image based on the quantized lifting coefficients. A video encoder can encode a packed image. A video decoder decodes the encoded video. An image unpacker can unpack a packed image. A dequantizer can inversely quantize an image. An inverse linear lifting unit applies inverse lifting to the image to generate a reconstructed displacement. A mesh restoration unit restores a warped mesh using the reconstructed displacement and the reconstructed base mesh. An attribute transfer unit receives an input mesh and / or an input attribute map, and generates an attribute map based on the reconstructed warped mesh. A push-pull padding unit can pad data in the attribute map based on a push-pull method. A color space transformation unit can transform the space of a color component, which is an attribute. A video encoder can encode an attribute. A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.
[0103] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0104] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a preprocessor (200) and / or an encoder (201).
[0105] Among the encoding operations of Fig. 7, the corresponding configuration of the encoding operation of Fig. 6 refers to the description of Fig. 7. For the inter-frame-based encoding of Fig. 7, the motion encoder can encode motion based on the restored quantized reference base mesh. The base mesh restoration unit can restore the base mesh based on the restored quantized reference base mesh.
[0106] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.
[0107] The encoding method according to the embodiments includes base mesh encoding (intra encoding). When performing intra frame encoding on the current input mesh frame, the base mesh generated in the preprocessing process can be encoded using a static mesh compression technique after going through a quantization process. In the V-Mesh compression method, for example, Draco technology is applied, and the vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh are compressed.
[0108] The encoding method according to the embodiments may include motion field encoding (inter encoding). Inter frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field, may be calculated and this information may be encoded. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may be encoded as a value. Alternatively, the predicted motion field may be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and this predicted motion field The residual motion field, which is the difference between the value and the motion field value of the current vertex, can be encoded. This value can be encoded using entropy coding. The process of encoding the displacement and attribute map, excluding the motion field encoding process of inter-frame encoding, is the same as the structure of the intra-frame encoding method except for the base mesh encoding.
[0109] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0110] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0111] Figures 8-9 show the process of converting the displacement of the encoding process of Figures 6-7 and the process of packing the conversion coefficients, respectively.
[0112] The encoding method according to the embodiments includes displacement encoding.
[0113] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information.
[0114] Figure 8 shows the process of transforming displacement information using lifting transform in V-Mesh. The transform coefficients generated through the transform process are quantized and then packed into a 2D image as shown in Figure 9. The transform coefficients are organized into one block for each 256 (=16×16) units, and each block can be packed in z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. The transform coefficients can be packed by sorting them with Morton code within a block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0115] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. The first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices (LoD0) of the base mesh. The first subdivision mesh may be generated by dividing the second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 including the vertices additionally generated from LoD0, and the vertices additionally divided from LoD1. LoD is a level indicating the degree of detail (Level of Detail), and as the level index increases, the distance between vertices becomes closer and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When a vertex is further divided through subdivision, considering the previous vertices v1, v2 and the subdivided vertex v, the mesh can be encoded based on a prediction and / or update method. Instead of still encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated, and the mesh can be encoded using the residual value to reduce the size of the bitstream. The prediction process means the operation of predicting the current vertex v using the previous vertices v1, v2. Since adjacent subdivision meshes have similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted as the residual for the previous vertex position information, and the previous vertex position information is updated using the residual.
[0116] Referring to Figure 9, the vertices have coefficients generated through the lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.
[0117] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0118] Figure 10 shows the detailed operation of attribute transfer of encoding such as Figures 6-7.
[0119] Encoding according to embodiments includes attribute map encoding.
[0120] Information about the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding processes, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has vertex position information, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is created through the attribute transfer process.
[0121] Attribute transfer first checks whether each point P(u, v) in the 2D texture domain belongs to a texture triangle of the reconstructed deformed mesh, and if it is in the texture triangle T, calculates the barycentric coordinate (α, β γ) of P(u, v) according to the triangle T. Then, using the 3D vertex position and (α, β γ) of triangle T, calculate the 3D coordinate M(x, y, z) of P(u, v). Find the vertex coordinate M'(x', y', z') that corresponds to the position most similar to the calculated M(x, y, z) in the input mesh domain and the triangle T' that contains this point. And in this triangle T', the center of mass coordinates (α', β', γ') of M'(x', y', z') are calculated. Using the texture coordinates corresponding to the three vertices of triangle T' and (α', β', γ'), the texture coordinates (u', v') are calculated, and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.
[0122] The new attribute map generated through attribute transfer is grouped into GoF units to form an attribute map video, which is then compressed using a video codec.
[0123] Referring to Figure 10, the reference relationship between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.
[0124] The decoding process of Fig. 1 can perform the reverse process of the corresponding encoding process of Fig. 1. The specific decoding process is as follows.
[0125] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0126] Fig. 11 shows the configuration and operation of a decoder of a receiving device such as Fig. 1.
[0127] Figure 11 illustrates the intra decoding process of V-Mesh technology. First, the input bitstream can be separated into a mesh sub-stream, a displacement sub-stream, an attribute map sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0128] The mesh sub-stream is decoded by the decoder of the static mesh codec used in encoding, such as Google Draco, and as a result, the connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be restored. The displacement sub-stream is decoded into a displacement video by the decoder of the video compression codec used in encoding, and goes through the image unpacking, inverse quantization, and inverse transform processes to restore displacement information for each vertex. Inverse quantization is applied to the restored base mesh, and this result is combined with the restored displacement information to generate the final decoded mesh.
[0129] The attribute map sub-stream is decoded through the decoder of the video compression codec used in encoding, and then restored to the final attribute map through processes such as color format conversion.
[0130] The restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0131] Referring to FIG. 11, the bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term "substream" is interpreted as referring to a portion of the bitstream included in the bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).
[0132] The decoder performs the following decoding operations within the frame. The static mesh decoder decodes the mesh to generate a reconstructed quantized base mesh, and the inverse quantizer applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder decodes the displacement, the unpacker unpacks the decoded video image, and the inverse quantizer inversely quantizes the quantized image. The linear lifting inverse transform applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit generates a warped mesh based on the base mesh and the displacement. The video decoder decodes the attribute map, and the color transformation unit transforms the color format and / or space to generate the decoded attribute map.
[0133] Figure 12 shows the inter-frame decoding process of the V-MESH compression method.
[0134] Fig. 12 shows the configuration and operation of the decoder of the receiving device of Fig. 1.
[0135] Figure 12 illustrates the inter-decoding process of V-Mesh technology. First, the input bitstream can be separated into a motion sub-stream, a displacement sub-stream, an attribute sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0136] The motion sub-stream is decoded through entropy decoding and inverse prediction processes, and the reconstructed motion information is combined with the already reconstructed reference base mesh to generate a reconstructed quantized base mesh for the current frame. The result of applying inverse quantization to this is combined with the displacement information reconstructed in the same way as the intra decoding described above to generate the final decoded mesh. The attribute map sub-stream is decoded in the same way as the intra decoding. The reconstructed decoded mesh and the decoded attribute map can be utilized by the receiver as the final mesh data that can be utilized by the user.
[0137] Referring to Fig. 12, the bitstream includes motion, displacement, and attribute maps. Since inter-frame decoding is performed, a process of decoding inter-frame motion information is further included. The motion is decoded, and a restored quantized base mesh for the motion is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operations in Fig. 12, which are identical to those in Fig. 11, refer to the description in Fig. 11.
[0138] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0139] Fig. 13 corresponds to the transmitting device (100), the dynamic mesh video encoder (102), the encoder (preprocessor and encoder) of Fig. 2, and / or the transmitting encoding device corresponding thereto of Fig. 13. Each component of Fig. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0140] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13.
[0141] The mesh preprocessor receives the original mesh as input and generates a simplified mesh (decimated mesh). Simplification can be performed based on the target number of vertices or the target number of polygons that constitute the mesh. Parameterization can be performed on the simplified mesh to generate texture coordinates and texture connection information per vertex. Additionally, quantization of floating-point mesh information into fixed-point information can be performed. This result can be encoded as a base mesh through a static mesh encoding unit. The mesh preprocessor can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The subdivided mesh can be fitted by adjusting the vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.
[0142] When performing intra-encoding on the corresponding mesh frame, the base mesh generated through the mesh preprocessing unit can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to the multiplexing unit.
[0143] When performing inter encoding for the corresponding mesh frame, a motion vector encoding unit is performed, which can calculate a motion vector between the base mesh and the reference reconstruction base mesh as input and encode the value. The motion vector encoding unit can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexing unit.
[0144] The encoded base mesh and motion vectors can be used to generate a restored base mesh through the base mesh restoration unit.
[0145] The displacement vector calculator can perform mesh refinement on the restored base mesh. The displacement vector can be calculated as the difference in vertex positions between the refined restored base mesh and the fitted subdivision mesh generated in the preprocessing unit. As a result, the number of displacement vectors can be calculated equal to the number of vertices in the refined mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0146] A displacement vector video generator can transform displacement vectors for effective encoding. The transformation can be performed by a lifting transformation, a wavelet transformation, etc., depending on the embodiment. In addition, quantization can be performed on the transformed displacement vector values, i.e., the transform coefficients. Different quantization parameters can be applied to each axis of the transform coefficients, and the quantization parameters can be derived according to the promise of the encoder / decoder. The transformed and quantized displacement vector information can be packed into a 2D image. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh.
[0147] A displacement vector video encoder can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer.
[0148] The displacement vector restored through the displacement vector restorer and the base mesh restored through the base mesh restorer and refined are restored through the mesh restorer, and the restored mesh has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0149] The texture map of the original mesh can be regenerated as a texture map for the restored mesh through the texture map video generation unit. The color information per vertex of the texture map of the original mesh can be assigned to the texture coordinates of the restored mesh. The regenerated texture maps for each frame can be bundled into GoF units to generate a texture map video.
[0150] The generated texture map video can be encoded using a video compression codec through a texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to a multiplexing unit.
[0151] The generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be multiplexed into a single bitstream and transmitted to a receiver via a transmitter. Alternatively, the generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via a transmitter.
[0152] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (displacement), and a texture map (attribute map). A transmitting device according to inter-encoding can generate a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map). A texture map obtained from a data input unit is generated and encoded based on a restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the segmented mesh. The base mesh is generated by preprocessing, simplifying, and encoding the original mesh. Motion is generated as a motion vector for the mesh of the current frame based on the reference base mesh of the previous frame.
[0153] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0154] Fig. 14 corresponds to the receiving device (110), the dynamic mesh video decoder (113), the decoder of Figs. 11-12, and / or the receiving decoding device corresponding thereto of Fig. 1. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.
[0155] The bitstream of the received Mesh is demultiplexed into a compressed motion vector bitstream or base mesh bitstream, displacement vector bitstream, and texture map bitstream after file / segment decapsulation.
[0156] If the current mesh has inter-frame encoding applied according to the frame header information, the motion vector decoding unit can perform decoding on the motion vector bitstream. The final motion vector can be reconstructed by adding the previously decoded motion vector to the residual motion vector decoded from the bitstream using the previously decoded motion vector as a predictor.
[0157] If the current mesh has been encoded within the screen according to the frame header information, the base mesh bitstream can restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh through the static mesh encoding unit.
[0158] In the base mesh restoration unit, if the current mesh has inter-frame encoding applied, the current base mesh can be restored by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. If the current mesh has intra-frame encoding applied, the static mesh decoding unit can perform inverse quantization on the decoded mesh to generate a restored base mesh.
[0159] The displacement vector bitstream can be decoded as a video bitstream using a video codec in a displacement vector video decoding unit.
[0160] In the displacement vector restoration unit, displacement vector transformation coefficients are extracted from the decoded displacement vector video, and the displacement vector is restored through the inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, a process of inverse transformation to the Cartesian coordinate system can be performed.
[0161] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Subdivision can generate vertex connection information, texture coordinates, and texture coordinate connection information, including the added vertices. The subdivided restored base mesh can be combined with the restored displacement vector to generate the final restored mesh.
[0162] The texture map bitstream can be decoded as a video bitstream using a video codec in a texture map video decoding unit. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of the corresponding vertex.
[0163] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.
[0164] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data. A receiving device according to inter-decoding can receive a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data.
[0165] A point cloud data transmission device and method according to embodiments can encode mesh data and transmit a bitstream including the encoded mesh data. A point cloud data reception device and method according to embodiments can receive a bitstream including mesh data and decode the mesh data. The point cloud data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The point cloud data transmission and reception method / device according to embodiments may also be referred to as mesh data transmission and reception method / device according to embodiments.
[0166] The encoding method / device according to the embodiments includes a transmitting device (100) in FIG. 1, an acquiring unit (101), an encoder (102), an encapsulator (103), a transmitter (104), FIG. 2, FIG. 3, FIG. 5, FIG. 7 pre-processing, an encoder, FIG. 13 an encoder, a multiplexing unit, a transmitting unit, FIG. 17 an acquiring unit, pre-processing, encoding, encapsulating, FIG. 24 an encoding method, etc.
[0167] The decryption method / device according to the embodiments includes a receiving device (110) in FIG. 1, a receiving unit (111), a decapsulator (112), a decoder (113), a renderer (114), FIG. 11, FIG. 12 decoding, FIG. 14 receiving unit, demultiplexing unit, renderer, FIG. 17 decapsulating, decoding, rendering, FIG. 20 file receiver, file parser, bitstream packager, decoder, FIG. 25 decryption method, etc.
[0168] Embodiments include a file encapsulation of submesh and atlas tile based dynamic mesh coding bitstream.
[0169] Embodiments include techniques for storing and signaling a Video-based Dynamic Mesh Coding (V-DMC) bitstream as multiple tracks within a file.
[0170] Embodiments include a method for encapsulating a file when a Video-based Dynamic Mesh Coding (V-DMC) bitstream is composed of sub-meshes.
[0171] Embodiments include a partial access signaling scheme based on sub-mesh and atlas tile units for a Video-based Dynamic Mesh Coding (V-DMC) bitstream.
[0172] Embodiments include a transmitter (encoder) or receiver (decoder) for providing a mesh content service that efficiently stores V-DMC or mesh bitstreams in multiple tracks within a file and provides signaling therefor.
[0173] Embodiments include a transmitter (encoder) or receiver (decoder) for providing mesh content services that handle file storage techniques to enable efficient access to V-DMC bitstreams stored within files.
[0174] Although the embodiments are described from a V-DMC perspective, the content described in this document can be applied to bitstreams using other mesh coding methods that are structured in the same or similar bitstream format as V-DMC.
[0175] Embodiments include a method for signaling sub-mesh related information to a sample entry and / or sample at the file level when the base mesh data of a V-DMC bitstream is composed of sub-meshes.
[0176] Embodiments include a method for signaling sub-mesh and atlas tile association information in a sample entry at the file level, when the base mesh data of a V-DMC bitstream is composed of sub-meshes and atlas tiles associated with each sub-mesh.
[0177] Embodiments may encapsulate mesh content encoded with V-DMC into multiple tracks based on ISOBMFF files to provide services such as streaming.
[0178] Due to the file-level signaling scheme according to the embodiments, when the base mesh data of a V-DMC bitstream is composed of sub-meshes, the receiver may be able to decode each sub-mesh unit sequentially or simultaneously in parallel. In addition, partial access and decoding may be possible for each sub-mesh unit. The file-level signaling scheme according to the embodiments may include a file-level signaling scheme for sub-mesh information, a track grouping scheme between one or more sub-mesh tracks, a signaling scheme for association information between sub-meshes and atlas tiles, etc.
[0179] Figure 15 shows a V-DMC bitstream according to embodiments.
[0180] The encoding method / device according to the embodiments (Fig. 1, transmitting device (100), obtaining unit (101), encoder (102), encapsulator (103), transmitter (104), Figs. 2, 3, 5, 7, pre-processing, encoder, Fig. 13 encoder, multiplexing unit, transmitting unit, Fig. 17 obtaining unit, pre-processing, encoding, encapsulating, Fig. 24 encoding method, etc.) can encode mesh data and generate a bitstream as shown in Fig. 15.
[0181] The decoding method / device according to the embodiments (Fig. 1 receiving device (110), receiving unit (111), decapsulator (112), decoder (113), renderer (114), Figs. 11 and 12 decoding, Fig. 14 receiving unit, demultiplexing unit, renderer, Fig. 17 decapsulating, decoding, rendering, Fig. 20 file receiver, file parser, bitstream packager, decoder, Fig. 25 decoding method, etc.) can receive a bitstream as in Fig. 15 and decode mesh data based on parameter information included in the bitstream.
[0182] Referring to Fig. 15(a), the bitstream may include a sequence header, an encoded base mesh, an encoded displacement (displacement), an encoded attribute (texture or attribute), etc. The sequence header may include decoder configuration information. The encoded base mesh includes an input mesh file. The encoded displacement information, which is difference information between vertices between a fitted and refined base mesh and a restored submesh, may be data packed as a 2D image and encoded using a video codec. The encoded attribute may be data encoded using a video codec such as HEVC, and may be an attribute (e.g., color) generated by matching the coordinates of the property (texture (2D)) of the restored base mesh using an input attribute map (e.g., PNG image).
[0183] In other words, FIG. 15(a) may be an example of a bitstream format resulting from video-based dynamic mesh encoding. The bitstream may include a sequence header, compressed basemesh data, compressed displacement data, and compressed texture data.
[0184] Basemesh data can be compressed using Edgebreaker or Draco tool.
[0185] Displacement data can be compressed using a video codec such as HEVC or VVC or arithmetic coding.
[0186] Texture / Attribute data can be compressed with video codecs such as HEVC, VVC, etc.
[0187] Referring to Fig. 15(b), as a result of video-based dynamic mesh encoding, a bitstream according to embodiments may be composed of sample stream units. For example, the bitstream may include a sample stream DMC header and at least one sample stream DMC unit. The sample stream DMC unit may include parameter information, mesh data, etc.
[0188] VPS (V-DMC Parameter Set): May include decoder configuration information and / or parameter set information related to mesh encoding / decoding. AD (Atlas Data): May include information related to 2D mapping or texture mapping for 3D objects. BMD (Base Mesh Data): Encoded base mesh information for mesh encoding / decoding. DD (Displacements Data): Displacement information encoded with arithmetic coding. GVD (Displacements Video Data): Displacement information encoded with a video codec. Displacement information may be referred to as geometry information, etc. AVD (Attribute Video Data): Attribute or texture information encoded with a video codec.
[0189] Figure 16 shows a V3C unit according to embodiments.
[0190] Specifically, FIG. 16 shows the header syntax of a V3C unit, which is a sample stream DMC unit (which may be referred to as a sample stream unit or a V3C unit) included in the bitstream of FIG. 15. The V3C unit (v3c_unit(numBytesInV3CUnit)), which is a data unit of the bitstream, includes a header (FIG. 16, v3c_unit_header( )) and a payload (v3c_unit_payload( numBytesInV3CUnit - 4)).
[0191] The unit type (vuh_unit_type) of the unit header can use values from 0 to 6 to indicate the unit type, such as VPS, AD, OVD, GVD, AVD, PVD, PVD, CAD, etc.
[0192] vuh_unit_typeIdentifierV3C unit typeDescription0V3C_VPSV3C parameter setV3C level parameters1V3C_ADAtlas dataAtlas information2V3C_OVDOccupancy video dataOccupancy information3V3C_GVDGeometry video dataGeometry information4V3C_AVDAttribute video dataAttribute information5V3C_PVDPacked video dataPacking information6V3C_CADCommon atlas dataInformation that is common for atlases in a CVS. Specified in ISO / IEC 23090-127...31V3C_RSVDReserved-
[0193] vuh_unit_type: Indicates the specified V3C unit type, as above. Values marked as reserved are reserved for future use in ISO / IEC.
[0194] vuh_v3c_parameter_set_id: Indicates the value of vps_v3c_parameter_set_id for the active V3C VPS. The value of vuh_v3c_parameter_set_id is in the range of 0 to 15.
[0195] vuh_atlas_id: Indicates the ID of the atlas corresponding to the current V3C unit. The value of vuh_atlas_id ranges from 0 to 63.
[0196] vuh_attribute_index: Indicates the index of the attribute data included in the attribute video data unit. The value of vuh_attribute_index is in the range of 0 to (ai_attribute_count[vuh_atlas_id] - 1).
[0197] vuh_attribute_partition_index: Indicates the index of the attribute dimension group included in the attribute video data unit. The value of vuh_attribute_partition_index is in the range of 0 to ai_attribute_dimension_partitions_minus1[vuh_atlas_id][vuh_attribute_index].
[0198] vuh_map_index: Indicates the map index of the current geometry or attribute stream. If this value is absent, the map index of the current geometry or attribute sub-bitstream is derived based on the type of the sub-bitstream and certain operations on the geometry and attribute video sub-bitstreams. The value of vuh_map_index, if present, is in the range 0 to vps_map_count_minus1[vuh_atlas_id].
[0199] vuh_auxiliary_video_flag: If 1, indicates that the associated geometry or attribute video data unit is a RAW and / or EOM-coded point video-only sub-bitstream. If vuh_auxiliary_video_flag is 0, indicates that the associated geometry or attribute video data unit may contain RAW and / or EOM-coded points. If vuh_auxiliary_video_flag is absent, its value is inferred to be 0.
[0200] vuh_reserved_zero_12bits: Equal to 0 in bitstreams that follow this version of this document.
[0201] vuh_reserved_zero_17bits: Equal to 0 in bitstreams that follow this version of this document.
[0202] vuh_reserved_zero_23bits: Equal to 0 in bitstreams that follow this version of this document.
[0203] vuh_reserved_zero_27bits: Equal to 0 in bitstreams that follow this version of this document.
[0204] The syntax of the sample stream DMC header and the sample stream DMC unit of the bitstream according to the embodiments can be defined as follows.
[0205] Syntax of the sample stream DMC header:
[0206]
[0207] sample_stream_dmc_header() {Descriptorsdmh_unit_size_precision_bytes_minus1u(3)sdmh_reserved_zero_5bitsu(5)}
[0208] sdmh_unit_size_precision_bytes_minus1: Adding 1 to this value indicates the precision (in bytes) of sdmu_dmc_unit_size elements in every sample stream DMC unit. sdmh_unit_size_precision_bytes_minus1 is in the range 0 to 7. The syntax of a sample stream DMC unit is as follows. Each sample stream DMC unit contains a DMC unit of one type: VPS, AD, BMD, DD, GVD, or AVD. The contents of each sample stream DMC unit are associated with the same access unit as the DMC unit contained in the sample stream DMC unit.
[0209] sample_stream_dmc_unit() {Descriptorsdmu_dmc_unit_sizeu(v)dmc_unit(sdmu_dmc_unit_size )}
[0210] sdmu_dmc_unit_size: Indicates the size (in bytes) of the subsequent dmc_unit. The number of bits used to express sdmu_dmc_unit_size is equal to (sdmh_unit_size_precision_bytes_minus1 + 1) * 8. A dmc_unit may be composed of a unit header and a unit payload. Type information about the unit payload may be inserted into the unit header, and data of the corresponding type may be inserted into the unit payload. The dmc_unit data syntax and semantics may be in the format of v3c_unit used in the V3C codec specification, i.e., ISO / IEC 23090-5 referenced in the V-DMC standard, but may not be limited to the format. The contents according to the embodiments may be implemented in a specific and independent manner to the V-DMC or V3C codec.
[0211] A bitstream according to embodiments may be composed of sample stream NAL (Network Abstraction Layer) units. A sample stream NAL unit may include a sample stream NAL header and a sample stream NAL unit.
[0212] Sample Stream NAL Header Syntax:
[0213] sample_stream_nal_header() {Descriptorssnh_unit_size_precision_bytes_minus1u(3)ssnh_reserved_zero_5bitsu(5)}
[0214] Sample Stream NAL Header Unit Syntax:
[0215] sample_stream_nal_unit() {Descriptorssnu_nal_unit_sizeu(v)nal_unit (ssnu_nal_unit_size)}
[0216] Semantics of the sample stream NAL header: The sample stream NAL header is always at the beginning of a NAL stream.
[0217] ssnh_unit_size_precision_bytes_minus1: Adding 1 to this value indicates the precision (in bytes) of ssnu_nal_unit_size elements in every sample stream NAL unit. ssnh_unit_size_precision_bytes_minus1 is in the range 0 to 7.
[0218] ssnh_reserved_zero_5bits: is equal to 0 in bitstreams that follow this version of this document.
[0219] Sample Stream NAL Unit Semantics:
[0220] The order of sample stream NAL units in a sample stream follows the decoding order of the NAL units contained in the sample stream NAL units.
[0221] A unit (nal_unit) included in a bitstream according to embodiments may be basemesh data that may be included in a sample of a basemesh track to be described later. That is, a unit of a bitstream may be a basemesh NAL unit (bmesh_nal_unit) and / or displacement data (displ_nal_unit) encoded using arithmetic coding.
[0222] Additionally, nal_unit may be submesh data that can be included in the samples of the submesh track described later. The submesh data may be defined in the form of bmesh_nal_unit.
[0223] ssnu_nal_unit_size represents the size (in bytes) of the subsequent NAL_unit. The number of bits used to represent ssnu_nal_unit_size is equal to (ssnh_unit_size_precision_bytes_minus1 + 1) * 8.
[0224] A bitstream may contain a basemesh sub-bitstream.
[0225] A NAL sample stream format can be constructed by arranging NAL units in decoding order and prefixing each NAL unit with a header indicating the exact size (in bytes) of the NAL unit. The sample stream header is included at the beginning of the sample stream bitstream, which indicates the precision (in bytes) of the signaled NAL unit size. The NAL unit stream format can be extracted from the sample stream format by traversing the sample stream format, reading the size information, and appropriately extracting each NAL unit.
[0226] General NAL unit syntax:
[0227] bmesh_nal_unit( NumBytesInNalUnit ) {Descriptorbmesh_nal_unit_header( )NumBytesInRbsp = 0for( i = 2; i < NumBytesInNalUnit; i++ )rbsp_byte[ NumBytesInRbsp++ ]b(8)}
[0228] NAL unit header syntax:
[0229] bmesh_nal_unit_header() {Descriptorbmesh_nal_forbidden_zero_bitf(1)bmesh_nal_unit_typeu(6)bmesh_nal_layer_idu(6)bmesh_nal_temporal_id_plus1u(3)}
[0230] Low-byte sequence payload, trailing bits, and byte alignment syntax BaseMesh Sequence Parameter Set RBSP syntax
[0231] General Basemesh Sequence Parameter Set RBSP Syntax
[0232] bmesh_sequence_parameter_set_rbsp( ) {Descriptorbmsps_sequence_parameter_set_idu(4)bmesh_profile_tier_level( )bmsps_intra_mesh_codec_idu(8)bmsps_inter_mesh_codec_idu(8)bmsps_inter_mesh_motion_group_size_minus1u(8)bmsps_inter_mesh_max_num_neighbours_minus1u(8)bmsps_geometry_3d_bit_depth_minus1u(5)bmsps_facegroup_segmentation_methodue(v)bmsps_mesh_attribute_countu(7)for( i = 0; i < bmsps_mesh_attribute_count; i++ ) {bmsps_mesh_attribute_type_id[ i ]u(4)bmsps_attribute_bit_depth_minus1[ i ]u(5)bmsps_attribute_msb_align_flag[ i ]u(1)}bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4ue(v)bmsps_max_dec_mesh_frame_buffering_minus1ue(v)bmsps_long_term_ref_mesh_frames_flagu(1)bmsps_num_ref_mesh_frame_lists_in_bmspsue(v)for( i = 0; i < bmsps_num_ref_mesh_frame_lists_in_bmsps; i++ )bmesh_ref_list_struct( i )bmsps_extension_present_flagu(1)if( bmsps_extension_present_flag ) {bmsps_extension_countu(8)}if( bmsps_extension_count ){bmsps_extensions_length_minus1ue(v)for( i = 0; i < bmsps_extension_count;i++ ) {bmsps_extension_type[ i ]u(8)bmsps_extension_length[ i ]u(16)bmsps_extension( bmsps_extension_type[ i ], bmsps_extension_length[ i ] )}}rbsp_trailing_bits( )};
[0233] 베이스메쉬 SPS 확장 신택스:
[0234] bmsps_extension( extension_type, extension_length ) {Descriptorfor( j = 0; j < extension_length; j++ )bmsps_extension_data_byteu(8)length_alignment( )}
[0235] 베이스메쉬 프로파일, 티어 및 레벨 신택스:
[0236] bmesh_profile_tier_level( ) {Descriptorbmptl_tier_flagu(1)bmptl_profile_codec_group_idcu(7)bmptl_reserved_zero_32bitsu(32)bmptl_level_idcu(8)bmptl_num_sub_profilesu(6)bmptl_extended_sub_profile_flagu(1)for( i = 0; I < bmptl_num_sub_profiles; i++ ) {bmptl_sub_profile_idc[ i ]u(v)}bmptl_toolset_constraints_present_flagu(1)if( bmptl_toolset_constraints_present_flag ) {bmesh_profile_toolset_constraints_information( )}}
[0237] 프로파일 툴셋 제약 정보(Profile toolset constraints information) 신택스:
[0238] bmesh_profile_toolset_constraints_information( ) {Descriptorbmptc_one_mesh_frame_only_flagu(1)bmptc_intra_frames_only_flagu(1)bmptc_reserved_zero_6bitsu(6)bmptc_num_reserved_constraint_bytesu(8)for( i = 0; i < bmptc_num_reserved_constraint_bytes; i++ )bmptc_reserved_constraint_byte[ i ]u(8)}
[0239] 베이스메쉬 프레임 파라미터 세트 RBSP 신택스:제너럴 베이스메쉬 프레임 파라미터 RBSP 신택스:
[0240] bmesh_frame_parameter_set_rbsp( ) {Descriptorbfps_mesh_sequence_parameter_set_idu(4)bfps_mesh_frame_parameter_set_idu(4)bmesh_sub_mesh_information( ) bfps_output_flag_present_flagu(1)bfps_num_ref_idx_default_active_minus1ue(v)bfps_additional_lt_mfoc_lsb_lenue(v)bfps_extension_present_flagu(1)if( bfps_extension_present_flag )bfps_extension_8bitsu(8)if( bfps_extension_8bits )while( more_rbsp_data( ) )bfps_extension_data_flagu(1)rbsp_trailing_bits( )}
[0241] 베이스메쉬 서브메쉬 정보:
[0242] bmesh_sub_mesh_information( ) {Descriptorbmsi_use_single_mesh_flagu(1)if(!bmsi_use_single_mesh_flag){bmsi_num_submeshes_minus1u(8)}elsebmsi_num_submeshes_minus1 = 0bmsi_signalled_submesh_id_flagu(1)if( bmsi_signalled_submesh_id_flag ) {bmsi_signalled_submesh_id_length_minus1ue(v)for( i = 0; < bmsi_num_submeshes_minus1 + 1; i++ )bmsi_submesh_id[ i ]u(v)SubMeshIDToIndex[ bmsi_submesh_id[ i ] ] = iSubMeshIndexToID[ i ] = bmsi_submesh_id[ i ]}}elsefor( i = 0; i < bmsi_num_submeshes_minus1 + 1; i++ ) { bmsi_submesh_id[ i ] = iSubMeshIDToIndex[ i ] = iSubMeshIndexToID[ i ] = i}}
[0243] Basemesh submesh layer RBSP 신태스
[0244] bmesh_submesh_layer_rbsp( ) {Descriptorsubmesh_header( )submesh_data_unit( mfh_submesh_type, SubMeshUnitSize )rbsp_trailing_bits( )}
[0245] Basemesh Submesh Header Syntax:
[0246] submesh_header( ) {Descriptorif( nal_unit_type >= NAL_BLA_W_LP && nal_unit_type <= NAL_RSV_IRAP_ACL_29 )smh_no_output_of_prior_mesh_frames_flagu(1)smh_basemesh_frame_parameter_set_idu(4)smh_idu(v)subMeshID = smh_idsmh_typeue(v)if( bfps_output_flag_present_flag )smh_mesh_output_flagu(1)smh_mesh_frm_order_cnt_lsbu(v)if( bmsps_num_ref_mesh_frame_lists_in_bmsps > 0 )smh_ref_mesh_frame_list_msps_flagu(1)if( smh_ref_basemesh_frame_list_msps_flag == 0 )basemesh_ref_list_struct( bmsps_num_ref_mesh_frame_lists_in_bmsps )else if( bmsps_num_ref_mesh_frame_lists_in_bmsps > 1 )smh_ref_mesh_frame_list_idxu(v)for( j = 0; j < NumLtrMeshFrmEntries[ RlsIdx ];j++ ) {smh_additional_mfoc_lsb_present_flag[ j ]u(1)if( smh_additional_mfoc_lsb_present_flag[ j ] )smh_additional_mfoc_lsb_val[ j ]u(v)}if( smh_type != SKIP_SUBMESH ) { if( smh_type == P_SUBMESH && num_ref_entries[ RlsIdx ] > 1 ) {smh_num_ref_idx_active_override_flagu(1)if( smh_num_ref_idx_active_override_flag )smh_num_ref_idx_active_minus1ue(v)}}byte_alignment( )};
[0247] 레퍼런스 리스트 구조 신택스:
[0248] bmesh_ref_list_struct( rlsIdx ) {Descriptornum_ref_entries[ rlsIdx ]ue(v)for( i = 0; i < num_ref_entries[ rlsIdx ]; i++ ) {if( bmsps_long_term_ref_mesh_frames_flag )st_ref_mesh_frame_flag[ rlsIdx ][ i ]u(1)if( st_ref_mesh_frame_flag[ rlsIdx ][ i ] ) {abs_delta_mfoc_st[ rlsIdx ][ i ]ue(v)if( abs_delta_mfoc_st[ rlsIdx ][ i ] > 0 )straf_entry_sign_flag[ rlsIdx ][ i ]u(1)} elsemfoc_lsb_lt[ rlsIdx ][ i ]u(v)}}
[0249] 베이스메쉬 서브메쉬 데이터 유닛 신택스:
[0250] submesh_data_unit( subMeshID, unitSize ) {Descriptorif( smh_type == I_SUBMESH ) {sdu_intra_sub_mesh_unit( subMeshID, unitSize )}else if( smh_type == P_SUBMESH ) {sdu_inter_sub_mesh_unit( subMeshID )}else if( smh_type == SKIP_SUBMESH ) {sdu_skip_sub_mesh_unit( )}}
[0251] Basemesh Intra Submesh Data Unit Syntax:
[0252] sdu_intra_sub_mesh_unit( subMeshID, vertexCount ) {Descriptorsismu_intra_unit( subMeshID )length_alignment( )}
[0253] sismu_intra_unit(subMeshID, sismu_intra_unit_size) contains a portion of mesh data of size sismu_intra_unit_size[subMeshID], which is an ordered string of bytes or bits that identifies the location of unit boundaries in a pattern of data. The format of this mesh data is identified by the bmptl_profile_codec_group_idc or component codec mapping SEI message. BaseMesh Inter Submesh Data Unit Syntax:
[0254] sdu_inter_sub_mesh_unit( subMeshID, vertexCount ) {Descriptorsismu_inter_vertex_count[ subMeshID ]ue(v)sismu_inter_unit( subMeshID , sismu_inter_vertex_count[ subMeshID ] )length_alignment( )}
[0255] sismu_inter_unit(sismu_inter_unit_size, sismu_inter_vertex_count) contains a portion of motion data of size sismu_inter_unit_size, which is an ordered stream of bytes or bits that can identify the locations of unit boundaries in a pattern of data. The format of this mesh data is identified by the bmptl_profile_codec_group_idc or component codec mapping SEI message. Basemesh Inter Submesh Data Unit Syntax:
[0256] sismu_inter_unit_default ( subMeshID, vertexCount ) {Descriptorif( vertexCount > 0 ) sismu_derived_mv_present_flag[ subMeshID ]ae(v)for( i = 0; i < vertexCount ; i++ ) {if(sismu_derived_mv_present_flag[ subMeshID ])sismu_mv_signalled_flag[ subMeshID ][ v ]ae(v)}groupSize = bmsps_inter_mesh_motion_group_size_minus1 + 1groupCount = ( vertexCount - 1) / groupSize + 1vStart = 0for( g = 0; g < groupCount: g++ ) {sismu_mv_pred_mode_group[ subMeshID ][ g ]ae(v)if ( g == (groupCount - 1) )groupSize = submeshMotionCount- groupSize *(groupCount - 1)for( v = vStart; v < (vStart+groupSize); v++ ) {if( sismu_mv_signalled_flag[ subMeshID ][ v ]) {for( k = 0; k < 3;k++ ) {sismu_mv_residual_abs_gt0[ subMeshID ][ v ][ k ]ae(v)if (sismu_mv_residual_abs_gt0[ subMeshID ][ v ][ k ]) {sismu_mv_residual_sign[ subMeshID ][ v ][ k ]ae(v)sismu_mv_residual_abs_gt1[ subMeshID ][ v ][ k ]ae(v)if (sismu_mv_residual_abs_gt1[ subMeshID ][ v ][ k ]) sismu_mv_residual_abs_rem[ subMeshID ][ v ][ k ]ae(v)}}}} / vvStart += groupSize}};
[0257] Basemesh Steep Submesh Data Unit Syntax:
[0258] sdu_skip_sub_mesh_unit( ) {Descriptor}
[0259] Below, we describe the semantics of the aforementioned syntax. Basemesh NAL unit semantics: General NAL unit semantics:
[0260] NumBytesInNalUnit represents the size of a NAL unit in bytes. This value is required for decoding a NAL unit. To enable inference of NumBytesInNalUnit, a form delimiting NAL unit boundaries is required.
[0261] rbsp_byte[ i ] is the ith byte of the RBSP. An RBSP is specified as an ordered sequence of bytes as follows:
[0262] An RBSP contains a string of data bits (SODBs) as follows:
[0263] - If the SODB is empty (i.e., has a length of 0 bits), the RBSP is also empty.
[0264] - Otherwise, the RBSP contains the SODB as follows:
[0265] 1) The first byte of the RBSP contains the first (most significant, leftmost) 8 bits of the SODB. The next byte of the RBSP contains the next 8 bits of the SODB, and so on until there are fewer than 8 bits of the SODB left.
[0266] 2) The rbsp_trailing_bits( ) syntax structure follows SODB as follows:
[0267] i) The first (most significant, leftmost) bit of the last RBSP byte contains the remaining bits of the SODB (if any).
[0268] ii) The next bit consists of a single bit equal to 1 (i.e. rbsp_stop_one_bit).
[0269] iii) If rbsp_stop_one_bit is not the last bit of a byte that is aligned, byte alignment is achieved by having at least one bit equal to 0 (i.e., an instance of rbsp_alignment_zero_bit).
[0270] Syntax structures with these RBSP properties are indicated in the syntax table using the "_rbsp" suffix. These structures are carried within the NAL unit as the contents of the rbsp_byte[ i ] data byte.
[0271] NAL unit header semantics:
[0272] As with atlases, similar NAL unit types are defined for base meshes, enabling similar functionality for random access and segmentation of the mesh. Unlike tile-divided atlases, the concept of sub-meshes is defined, with specific NAL units corresponding to coded mesh data. It also defines NAL units that can contain metadata, such as SEI messages.
[0273] The supported basic mesh NAL unit types are:
[0274] bmesh_nal_unit_typeName of bmesh_nal_unit_typeContent of base mesh NAL unit and RBSP syntax structureNAL unitype class01NAL_TRAIL_NNAL_TRAIL_RCoded sub-mesh of a non-TSA, non STSA trailing base mesh framesub_mesh_layer_rbsp( )BMCL23NAL_TSA_NNAL_TSA_RCoded sub-mesh of a TSA base mesh framesub_mesh_layer_rbsp( )BMCL45NAL_STSA_NNAL_STSA_RCoded sub-mesh of a STSA base mesh framesub_mesh_layer_rbsp( )BMCL67NAL_RADL_NNAL_RADL_RCoded sub-mesh of a RADL base mesh framesub_mesh_layer_rbsp( )BMCL89NAL_RASL_NNAL_RASL_RCoded sub-mesh of a RASL base mesh framesub_mesh_layer_rbsp( )BMCL1011NAL_SKIP_NNAL_SKIP_RCoded sub-mesh of a skipped base mesh framesub_mesh_layer_rbsp( )BMCL1214NAL_RSV_BMCL_N12NAL_RSV_BMCL_N14Reserved non-IRAP sub-layer non-reference BMCL mesh NAL unit typesBMCL1315NAL_RSV_BMCL_R13NAL_RSV_BMCL_R15Reserved non-IRAP sub-layer reference BMCL mesh NAL unit typesBMCL161718NAL_BLA_W_LPNAL_BLA_W_RADLNAL_BLA_N_LPCoded sub-mesh of a BLA base mesh framesub_mesh_layer_rbsp()BMCL1920NAL_IDR_W_RADLNAL_IDR_N_LPCoded sub-mesh of an IDR base mesh framesub_mesh_layer_rbsp( ) (BMCL)21NAL_CRACoded sub-mesh of a CRA base mesh framesub_mesh_layer_rbsp( )BMCL2223NAL_RSV_IRAP_BMCL_22NAL_RSV_IRAP_BMCL_23Reserved IRAP BMCL NAL unit typesBMCL24..29NAL_RSV_BMCL_24..NAL_RSV_BMCL_29Reserved non-IRAP BMCL NAL unit typesBMCL30NAL_BMSPSBase mesh sequence parameter setbmesh_sequence_parameter_set_rbsp( )non-BMCL31NAL_BMFPSBase mesh frame parameter setbmesh_frame_parameter_set_rbsp( )non-BMCL32NAL_AUDdelimiteraccess_unit_delimiter_rbsp( )non-BMCL33NAL_EOSEnd of sequenceend_of_sequence_rbsp( )non-BMCL34NAL_EOBEnd of bitstreamend_of_bmesh_sub_bitstream_rbsp( )non-BMCL35NAL_FDFillerfiller_data_rbsp( )non-BMCL3637NAL_PREFIX_NSEINAL_SUFFIX_NSEINon-essential supplemental enhancement informationsei_rbsp( )non-BMCL3839NAL_PREFIX_ESEINAL_SUFFIX_ESEIEssential supplemental enhancementinformationsei_rbsp( )non-BMCL40..4445..63NAL_RSV_NBMCL_40NAL_RSV_NBMCL_44NAL_UNSPEC_45NAL_UNSPEC_63Reserved non-BMCL NAL unit typesUnspecified non-BMCL NAL unit typesnon-BMCLnon-BMCL
[0275] Basemesh Raw byte sequence payloads, trailing bits, and byte alignment semantics: Basemesh Sequence Parameter Set RBSP Semantics: General Basemesh Sequence Parameter Set RBSP Semantics:
[0276] bmsps_sequence_parameter_set_id: An identifier for the basemesh sequence parameter set so that other syntax elements can reference it.
[0277] bmsps_intra_mesh_codec_id: Indicates the identifier of the codec used to compress the static mesh. bmsps_intra_mesh_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, a component codec mapping SEI message, or by means external to this document. A specific mesh or motion mesh codec may be associated with a profile specified in that specification, or may be explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0278] bmsps_inter_mesh_codec_id: Indicates the identifier of the codec used to compress the motion data. bmsps_intr_mesh_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, a component codec mapping SEI message, or by means external to this document. A specific mesh or motion mesh codec may be associated with a profile specified in that specification, or may be explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0279] bmsps_inter_mesh_motion_group_size_minus1: Adding 1 to this value indicates the size of the vertex grouping in motion vector coding. bmsps_inter_mesh_motion_group_size_minus1 is in the range of 0 to 255.
[0280] bmsps_inter_mesh_max_num_neighbours_minus1: Adding 1 to this value indicates the maximum number of vertex neighbors to use for motion vector estimation. bmsps_inter_mesh_max_num_neighbours_minus1 is in the range of 0 to 255.
[0281] bmsps_geometry_3d_bit_depth_minus1: Adding 1 to this value indicates the bit depth of the geometry coordinates of the reconstructed mesh. bmsps_geometry_3d_bit_depth_minus1 is in the range of 0 to 31.
[0282] bmsps_facegroup_segmentation_method: Indicates the identifier of the method for deriving facegroup IDs from a mesh.
[0283] bmsps_mesh_attribute_count: Indicates the number of attributes associated with the mesh. bmsps_mesh_attribute_count is in the range of 0 to 127.
[0284] bmsps_mesh_attribute_type_id[ i ]: Indicates the attribute type of the attribute with index i for the mesh. The list of supported attributes and their relationship to bmsps_mesh_attribute_type_id[ i ] are as follows:
[0285] bmsps_mesh_attribute_type_id[ i ]IdentifierAttribute type0ATTR_TEXTURE Texture1ATTR_MATERIAL_IDMaterial ID2ATTR_TRANSPARENCYTransparency3ATTR_REFLECTANCEReflectance4ATTR_NORMALNormals5ATTR_FACEGROUP_IDFacegroup ID6..14ATTR_RESERVEDReserved15ATTR_UNSPECIFIEDUnspecified
[0286] bmsps_attribute_bit_depth_minus1[ i ]: Adding 1 to this value indicates the bit depth of the attribute with index i for the mesh. bmsps_attribute_bit_depth_minus1[ i ] is in the range of 0 to 31. bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4: Adding 4 to this value indicates the values of the variables Log2MaxMeshFrmOrderCntLsb and MaxMeshFrmOrderCntLsb used in the decoding process for the atlas frame order count, as follows: Log2MaxMeshFrmOrderCntLsb =
[0287] bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4 + 4
[0288] MaxMeshFrmOrderCntLsb = 2Log2MaxMeshFrmOrderCntLsb
[0289] The value of bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4 is in the range of 0 to 12.
[0290] bmsps_max_dec_mesh_frame_buffering_minus1: Adding 1 to this value indicates the maximum size of the decoded atlas frame buffer required for CAS, in atlas frame storage buffer units. The value of bmps_max_dec_mesh_frame_buffering_minus1 is in the range of 0 to 15.
[0291] bmsps_long_term_ref_mesh_frames_flag: If this value is 0, it indicates that long-term reference atlas frames are not used for inter prediction of coded atlas frames in CAS. If bmsps_long_term_ref_mesh_frames_flag is 1, it indicates that long-term reference atlas frames can be used for inter prediction of one or more coded atlas frames in CAS.
[0292] bmsps_num_ref_mesh_frame_lists_in_bmsps: Indicates the number of bmesh_ref_list_struct(rlsIdx) syntax structures included in the atlas sequence parameter set. The value of bmsps_num_ref_mesh_frame_lists_in_bmsps ranges from 0 to 64.
[0293] NOTE: The decoder allocates memory for a total number of bmesh_ref_list_struct(rlsIdx) syntax structures equal to (bmsps_num_ref_mesh_frame_lists_in_bmsps + 1), since there can only be one bmesh_ref_list_struct(rlsIdx) syntax structure signaled directly from the atlas tile header of the current atlas tile.
[0294] bmsps_extension_present_flag: If this value is 1, it indicates that bmsps_extension_count_minus1 is in the basemesh sequence parameter set.
[0295] bmsps_extension_count: Indicates the number of BMSPS extensions in the v3c_parameter_set( ) syntax structure. If not present, bmsps_extension_count is inferred to be 0. If bmsps_extension_count is 0, BmspsExtensionsLength, which specifies the cumulative length in bytes of all extensions that follow this syntax element, is 0.
[0296] If bmsps_extensions_length_minus1 is present, it represents the cumulative length in bytes of all extensions following this syntax element. BmspsExtensionsLength is calculated as follows:
[0297] if( bmsps_extension_count == 0 )
[0298] BmspsExtensionsLength = 0
[0299] else
[0300] BmspsExtensionsLength = bmsps_extensions_length_minus1 + 1
[0301] If bmsps_extension_count is not 0, BmspsExtensionsLength is equal to 3 * bmsps_extension_count plus the sum of all bmsps_extension_length[ i ].
[0302] bmsps_extension_type[ i ]: Indicates the BMSPS extension type for the extension with index i as specified in ISO / IEC 23090-29.
[0303] bmsps_extension_length[ i ]: Indicates the number of bytes used to represent the payload size of the syntax structure of the associated extension with index i. If bmsps_extension_length[ i ] is 0, there is no extension payload for the extension with index i. Otherwise, the extension with index i has a payload size in bits in the range 8 * ( bmsps_extension_length[ i ] - 1 ) + 1 to 8 * bmsps_extension_length[ i ], inclusive.
[0304] BaseMesh SPS Extension Semantics:
[0305] bmsps_extension_data_byte can have any value.
[0306] Basemesh profile, hierarchy, and level semantics:
[0307] bmptl_tier_flag indicates the hierarchy context for interpreting bmptl_level_idc as specified in ISO / IEC 23090-29.
[0308] bmptl_profile_codec_group_idc: Indicates the codec group profile component that CVS complies with, as specified in ISO / IEC 23090-29.
[0309] bmptl_profile_toolset_idc: Represents a toolset combination profile component that CVS complies with as specified in ISO / IEC 23090-29.
[0310] bmptl_profile_reconstruction_idc: Indicates the reconstruction profile component that CVS is recommended to adhere to, as specified in ISO / IEC 23090-29.
[0311] bmptl_reserved_zero_16bits: equal to 0 in the bitstream.
[0312] bmptl_reserved_0xffff_16bits: Equivalent to 0xFFFF in the bitstream.
[0313] bmptl_level_idc: Indicates the level to which the CVS complies as specified in ISO / IEC 23090-29.
[0314] bmptl_num_sub_profiles represents the number of bmptl_sub_profile_idc[ i ] syntax elements.
[0315] bmptl_extended_sub_profile_flag: If this value is 1, it indicates that the bmptl_sub_profile_idc[ i ] syntax element should be represented using 64 bits, if present. If bmptl_extended_sub_profile_flag is 0, it indicates that the bmptl_sub_profile_idc[ i ] syntax element should be represented using 32 bits, if present.
[0316] bmptl_sub_profile_idc[ i ]: Indicates the ith registered interoperability metadata as specified in Rec. ITU-T T.35, the contents of which are not specified in this document. The number of bits used to represent bmptl_sub_profile_idc[ i ] is equal to (bmptl_extended_sub_profile_flag == 0 ? 32 : 64).
[0317] If bmptl_toolset_constraints_present_flag is 1, it indicates that the additional structure profile_toolset_constraints_information( ) is present in the bitstream. If bmptl_toolset_constraints_present_flag is 0, it indicates that the structure profile_toolset_constraints_information( ) is not present.
[0318] BaseMesh Profile Toolset Constraint Information Semantics:
[0319] If bmptc_one_mesh_frame_only_flag is present, it has the meaning specified in ISO / IEC 23090-29, and the profile indicated by bmptl_profile_toolset_idc is the profile specified in ISO / IEC 23090-29. If absent, ptc_one_mesh_frame_only_flag is inferred to be equal to 0.
[0320] If bmptc_intra_frames_only_flag is 1, it indicates that the bitstream contains only sdu_intra_sub_mesh_unit( ). If not present, bmptc_intra_frames_only_flag is inferred to be equal to 0.
[0321] bmptc_reserved_zero_6bits is equal to 0 in bitstreams that follow this version of this document.
[0322]
[0323] bmptc_num_reserved_constraint_bytes represents the number of reserved constraint bytes.
[0324] bmptc_reserved_constraint_byte[ i ] can have any value. Its presence and value do not affect the decoder's conformance to the profile specified in this version of this document. Decoders that conform to this version of this document ignore the value of any bmptc_reserved_constraint_byte[ i ] syntax element.
[0325] Basemesh Frame Parameter Set RBSP Semantics:
[0326] General Basemesh Frame Parameter Set RBSP Semantics:
[0327] bfps_mesh_sequence_parameter_set_id: Indicates the value of bmsps_sequence_parameter_set_id for the active Basemesh sequence parameter set.
[0328] bfps_mesh_parameter_set_id identifies the basemesh frame parameter set so that other syntax elements can reference it.
[0329] If bfps_output_flag_present_flag is 1, it indicates that the smh_output_flag syntax element is present in the associated submesh header. If bfps_output_flag_present_flag is 0, it indicates that the smh_output_flag syntax element is not present in the associated submesh header.
[0330] bfps_num_ref_idx_default_active_minus1 plus 1 represents the inferred value of the variable NumRefIdxActive for tiles where smh_num_ref_idx_active_override_flag is 0. The value of bfps_num_ref_idx_default_active_minus1 is in the range 0 to 14.
[0331] bfps_additional_lt_mfoc_lsb_len represents the value of the variable MaxLtMeshFrmOrderCntLsb used in the decoding process of the reference atlas frame list.
[0332] MaxLtMeshFrmOrderCntLsb =
[0333] 2 * (Log2MaxMeshFrmOrderCntLsb + bfps_additional_lt_mfoc_lsb_len)
[0334] The value of bfps_additional_lt_mfoc_lsb_len is in the range 0 to 32 - Log2MaxAtlasFrmOrderCntLsb.
[0335] If bmsps_long_term_ref_mesh_frames_flag is 0, the value of bfps_additional_lt_mfoc_lsb_len is equal to 0.
[0336] If bfps_extension_present_flag is 1, it indicates that the syntax element bfps_extension_8bits is present in the basemesh frame parameter set. If bfps_extension_present_flag is 0, it indicates that the syntax element bfps_extension_8bits is not present.
[0337] If bfps_extension_8bits is 0, it indicates that the AFPS RBSP syntax structure does not include the bfps_extension_data_flag syntax element. A non-zero value for bfps_extension_8bits is reserved for future use in ISO / IEC.
[0338] bfps_extension_data_flag can have any value.
[0339] Basemesh submesh information:
[0340] If bmsi_use_single_mesh_flag is 1, it indicates that each mesh frame has only one submesh referencing the BFPS. If bmsi_use_single_mesh_falg is 0, it indicates that each mesh frame can have more than one submesh referencing the BFPS.
[0341] bmsi_num_submeshes_minus1 plus 1 indicates the number of submeshes referencing BFPS in each mesh frame. The value of bmsi_num_submeshes_minus1 is in the range 0 to 63. If it is not present and bmsi_use_single_mesh_flag is 1, the value is inferred to be 1.
[0342] If bmsi_signalled_submesh_id_flag is 1, it indicates that the submesh ID of each mesh frame is signaled. If bmsi_signalled_tile_id_flag is 0, it indicates that the submesh ID is not signaled.
[0343] bmsi_signalled_submesh_id_length_minus1 plus 1 indicates the number of bits used to represent the syntax element bmsi_tile_id[ i ], if present, and the syntax element submesh_id in the submesh header. The value of bmsi_signalled_tile_id_length_minus1 is in the range 0 to 15. If absent, its value is inferred to be equal to Ceil(Log2(bmsi_num_submeshes_minus1 + 1 )) - 1.
[0344] bmsi_submesh_id[ i ] represents the tile ID of the ith submesh. The length of the bmsi_submesh_id[ i ] syntax element is bmsi_signalled_submesh_id_length_minus1 + 1 bit. If not present, the value of bmsi_submesh_id[ i ] is inferred to be equal to i for each i in the range 0 to bmsi_num_submeshes_minus1, inclusive. It is a bitstream conformance requirement that bmsi_submesh_id[ i ] must not be equal to bmsi_submesh_id[ j ] for all i != j. The length of the bmsi_submesh_id[ i ] syntax element is bmsi_signalled_submesh_id_length_minus1 + 1 bit.
[0345] The variable FirstSubmeshID is calculated as follows:
[0346] FirstSubmeshID=bmsi_submesh_id
[0000]
[0347] for ( i = 1; i < bmsi_num_submeshes_minus1+ 1; i++ )
[0348] FirstSubmeshID = Min(FirstSubmeshID, bmsi_submesh_id[ i ])
[0349] Basemesh submesh header semantics:
[0350] If present, the values of the atlas tile header syntax elements smh_basemesh_frame_parameter_set_id, smh_mesh_output_flag, smh_no_output_of_prior_mesh_frames_flag, and smh_mesh_frm_order_cnt_lsb are the same for all submesh headers of a coded mesh frame.
[0351] smh_no_output_of_prior_mesh_frames_flag affects the output of previously decoded mesh frames in DAB after decoding an atlas from a CAS AU that is not the first AU in the bitstream. If smh_no_output_of_prior_mesh_frames_flag is absent, its value is inferred to be equal to 0.
[0352] As a requirement for bitstream conformance, the value of smh_no_output_of_prior_mesh_frames_flag is the same for all mesh frames in an AU.
[0353] The smh_no_output_of_prior_mesh_frames_flag value in the submesh header is the output_of_prior_mesh_frames_flag value in the AU.
[0354] smh_basemesh_frame_parameter_set_id represents the value of bfps_basemesh_frame_parameter_set_id for the active basemesh frame parameter set for the current submesh.
[0355] smh_id indicates the submesh ID associated with the current submesh. If not present, the value of smh_id is assumed to be 0.
[0356] The following applies:
[0357] - The length of smh_id is bmsi_signalled_submesh_id_length_minus1 + 1 bit.
[0358] - The value of smh_id must be in the range of values specified in the SubMeshIndexToID [ i ] array, where i is in the range of 0 to bmsi_num_submeshes_minus1.
[0359] The following constraints apply to the requirements of bitstream conformance:
[0360] - The value of smh_id is not equal to the smh_id value of another coded atlas tile unit in the same coded atlas frame.
[0361] - Tiles in the atlas frame are sorted in increasing order of their smh_id values.
[0362] smh_type indicates the coding type of the current submesh, as follows. The value of smh_type is 0, 1, or 2 in bitstreams that follow this version of this document.
[0363] smh_typeName of smh_type0P_SUBMESH1I_SUBMESH2SKIP_SUBMESH3..RESERVED
[0364] smh_mesh_output_flag affects the decoded mesh output and removal process. If smh_mesh_output_flag is absent, it is inferred to be equal to 1. smh_mesh_frm_order_cnt_lsb represents the number of mesh frame orders modulo MaxMeshFrmOrderCntLsb for the current submesh. The length of the smh_mesh_frm_order_cnt_lsb syntax element is equal to Log2MaxMeshFrmOrderCntLsb bits. The value of smh_mesh_frm_order_cnt_lsb ranges from 0 to MaxMeshFrmOrderCntLsb - 1. If smh_ref_mesh_frame_list_bmsps_flag is 1, it indicates that the reference bmesh frame list of the current submesh is derived based on one of the bmesh_ref_list_struct(rlsIdx) syntax structures of the active BMSPS. If smh_ref_mesh_frame_list_bmsps_flag is 0, it indicates that the reference bmesh frame list of the current sub-mesh is derived based on the bmesh_ref_list_struct(rlsIdx) syntax structure directly included in the sub-mesh header of the current sub-mesh. When bmsps_num_ref_mesh_frame_lists_in_bmsps is 0, the value of smh_ref_mesh_frame_list_bmsps_flag is inferred to be 0.
[0365] smh_ref_mesh_frame_list_idx represents the index of the bmesh_ref_list_struct(rlsIdx) syntax structure used to derive the reference mesh frame list of the current submesh in the list of bmesh_ref_list_struct(rlsIdx) syntax structures contained in the active ASPS. The syntax element smh_ref_mesh_frame_list_idx is represented as Ceil(Log2(bmsps_num_ref_mesh_frame_lists_in_bmsps)) bits. If not present, the value of smh_ref_mesh_frame_list_idx is inferred to be equal to 0. The value of smh_ref_mesh_frame_list_idx ranges from 0 to bmsps_num_ref_mesh_frame_lists_in_bmsps - 1. If smh_ref_mesh_frame_list_bmsps_flag is 1 and bmsps_num_ref_mesh_frame_lists_in_bmsps is 1, the value of smh_ref_mesh_frame_list_idx is inferred to be equal to 0.
[0366] The variable RlsIdx of the current atlas tile is derived as follows:
[0367] RlsIdx = smh_ref_mesh_frame_list_bmsps_flag ?
[0368] smh_ref_mesh_frame_list_idx : bmsps_num_ref_mesh_frame_lists_in_bmsps
[0369] If smh_additional_mfoc_lsb_present_flag[ j ] is 1, it indicates that smh_additional_mfoc_lsb_val[ j ] exists in the current submesh. If smh_additional_mfoc_lsb_present_flag[ j ] is 0, it indicates that smh_additional_mfoc_lsb_val[ j ] does not exist.
[0370] smh_additional_mfoc_lsb_val[ j ] represents the value of FullMeshFrmOrderCntLsbLt[ RlsIdx ][ j ] for the current atlas tile as follows:
[0371] FullMeshFrmOrderCntLsbLt[RlsIdx][j] =
[0372] smh_additional_mfoc_lsb_val[ j ] * MaxMeshFrmOrderCntLsb+mfoc_lsb_lt[RlsIdx][j]
[0373] The syntax element smh_additional_mfoc_lsb_val[ j ] is represented by the bfps_additional_lt_mfoc_lsb_len bits. If it is not present, the value of smh_additional_mfoc_lsb_val[ j ] is inferred to be equal to 0.
[0374] If smh_num_ref_idx_active_override_flag is 1, it indicates that the syntax element smh_num_ref_idx_active_minus1 exists for the current submesh. If smh_num_ref_idx_active_override_flag is 0, it indicates that the syntax element smh_num_ref_idx_active_minus1 does not exist. If smh_num_ref_idx_active_override_flag is absent, its value is inferred to be equal to 0.
[0375] smh_num_ref_idx_active_minus1 is used to derive the variable NumRefIdxActive for the current submesh. The value of smh_num_ref_idx_active_minus1 is in the range 0 to 14.
[0376] If the current submesh is a P_SUBMESH submesh and smh_num_ref_idx_active_override_flag is 1 and smh_num_ref_idx_active_minus1 is absent, smh_num_ref_idx_active_minus1 is inferred to be 0.
[0377] The variable NumRefIdxActive is derived as follows:
[0378] if( smh_type == P_SUBMESH || smh_type == SKIP_SUBMESH ) {
[0379] if( smh_num_ref_idx_active_override_flag == 1 )
[0380] NumRefIdxActive = smh_num_ref_idx_active_minus1 + 1
[0381] else {
[0382] if( num_ref_entries[ RlsIdx ] >= bfps_num_ref_idx_default_active_minus1 + 1 )
[0383] NumRefIdxActive = bfps_num_ref_idx_default_active_minus1 + 1
[0384] else
[0385] NumRefIdxActive = num_ref_entries[RlsIdx]
[0386] }
[0387] }
[0388] else
[0389] NumRefIdxActive = 0
[0390] The value of NumRefIdxActive minus 1 represents the maximum number of atlas reference frame indices that can be used to decode the current atlas tile.
[0391] Reference list structure semantics:
[0392] num_ref_entries[rlsIdx] represents the number of entries in the bmesh_ref_list_struct(rlsIdx) syntax structure, where rlsIdx is the index into the mesh frame reference list. For P_SUBMESH and SKIP_SUBMESH, the value of num_ref_entries[rlsIdx] is in the range of 1 to bmsps_max_dec_mesh_frame_buffering_minus1 + 1. Otherwise, the value of num_ref_entries[rlsIdx] is in the range of 0 to bmsps_max_dec_mesh_frame_buffering_minus1 + 1.
[0393] If st_ref_mesh_frame_flag[rlsIdx][i] is 1, it indicates that the ith entry in the bmesh_ref_list_struct(rlsIdx) syntax structure is a short-term reference mesh frame entry. If st_ref_mesh_frame_flag[rlsIdx][i] is 0, it indicates that the ith entry in the ref_list_struct(rlsIdx) syntax structure is a long-term reference mesh frame entry. If not present, the value of st_ref_mesh_frame_flag[rlsIdx][i] is inferred to be equal to 1.
[0394] The variable NumLtrMeshFrmEntries[rlsIdx] is derived as follows:
[0395] NumLtrMeshFrmEntries[rlsIdx] = 0
[0396] for( i = 0; i < num_ref_entries[rlsIdx]; i++)
[0397] if(!st_ref_mesh_frame_flag[rlsIdx][i])
[0398] NumLtrMeshFrmEntries[rlsIdx]++
[0399] abs_delta_mfoc_st[rlsIdx][i] represents the absolute difference between the meshes if the ith entry is the first short-term reference mesh frame entry in the bmesh_ref_list_struct(rlsIdx) syntax structure, the frame order count value of the mesh frame referenced by the current mesh tile and the ith entry, or, if the ith entry is a short-term reference mesh frame entry but is not the first short-term reference mesh frame entry in the bmesh_ref_list_struct(rlsIdx) syntax structure, the absolute difference between the mesh frame order count value of the mesh frame referenced by the ith entry in the bmesh_ref_list_struct(rlsIdx) syntax structure and the previous short-term reference mesh frame entry. The value of abs_delta_mfoc_st[rlsIdx][i] is in the range of 0 to 215 - 1.
[0400] If straf_entry_sign_flag[rlsIdx][i] is 1, it indicates that the ith entry in the syntax structure bmesh_ref_list_struct(rlsIdx) has a value greater than or equal to 0. If straf_entry_sign_flag[rlsIdx][i] is 0, it specifies that the ith entry in the syntax structure bmesh_ref_list_struct(rlsIdx) has a value less than 0. If absent, the value of straf_entry_sign_flag[rlsIdx][i] is inferred to be 1.
[0401] The list DeltaMfocSt[rlsIdx][i] is derived as follows:
[0402] for( i = 0; i < num_ref_entries[rlsIdx]; i++ )
[0403] if( st_ref_mesh_frame_flag[ rlsIdx ][ i ] )
[0404] DeltaMfocSt[ rlsIdx ][ i ] =
[0405] ( 2 * straf_entry_sign_flag[ rlsIdx ][ i ] - 1 ) * abs_delta_mfoc_st[ rlsIdx ][ i ]
[0406] else
[0407] DeltaMfocSt[rlsIdx][i] = 0
[0408] mfoc_lsb_lt[ rlsIdx ][ i ] represents the mesh frame order count modulo MaxMeshFrmOrderCntLsb of the mesh frame referenced by the ith entry of the bmesh_ref_list_struct( rlsIdx ) syntax structure. The length of the mfoc_lsb_lt[ rlsIdx ][ i ] syntax element is Log2MaxMeshFrmOrderCntLsb bits.
[0409] Basemesh Submesh Data Unit Semantics:
[0410] Basemesh inter-submesh data unit semantics:
[0411] sismu_derived_mv_present_flag[ subMeshID ] indicates that sismu_mv_signalled_flag is present in the bitstream. If sismu_derived_mv_present_flag[ subMeshID ] is 0, sismu_mv_signalled_flag[ subMeshID ][ v ] is always inferred to be 1.
[0412] sismu_mv_signalled_flag[ subMeshID ][ v ] indicates that the motion vector for the vertex with index v is present in the bitstream. If sismu_mv_signalled_flag[ subMeshID ][ v ] is not present in the bitstream, sismu_mv_signalled_flag[ subMeshID ][ v ] is inferred to be 1.
[0413] sismu_mv_pred_mode_group[ subMeshID ][ g ] indicates the method used to predict the motion vectors associated with vertices in the group with index g of the current submesh. The submesh ID is the same as subMeshID.
[0414] sismu_mv_residual_abs_gt0[ subMeshID ][ v ][ k ] indicates whether the kth component of the motion vector prediction residual associated with the vertex with index v in the current submesh has an absolute value greater than 0 (if 1) or not (if 0).
[0415] sismu_mv_residual_sign[ subMeshID ][ v ][ k ] indicates whether the kth component of the motion vector prediction residual associated with the vertex with index v in the current submesh has a positive sign (if 1) or not (if 0). If sismu_mv_residual_sign[ v ][ k ] is absent, it is inferred to be equal to 1.
[0416] sismu_mv_residual_abs_gt1[ subMeshID ][ v ][ k ] indicates whether the kth component of the motion vector prediction residual associated with the vertex with index v in the current submesh has an absolute value greater than 1 (if 1) or not (if 0). If sismu_mv_residual_abs_gt1[ v ][ k ] is absent, it is inferred to be equal to 0.
[0417] sismu_mv_residual_abs_rem[ subMeshID ][ v ][ k ] represents the absolute value of the kth component of the motion vector prediction residual associated with the vertex with index v in the current submesh, where the submesh ID is subMeshID minus 2. If sismu_mv_residual_abs_rem[ v ][ k ] is absent, it is inferred to be equal to 0.
[0418] The kth component of the motion vector prediction residual VertexMotionVectorResiduals[ v ][ k ] associated with the vertex with index v in the current submesh is computed as follows:
[0419] VertexMotionVectorResiduals[ v ][ k ] = sismu_mv_residual_sign[ v ][ k ] ? 1:-1)*
[0420] (sismu_mv_ residual_sign_gt0[ v ][ k ] + sismu_mv_ residual_sign wk_gt1[ v ][ k ] +
[0421] sismu_mv_ residual_sign _rem[ v ][ k ])
[0422] Arthmetic Coded Displacement sub-bitstream:
[0423] A NAL sample stream format can be constructed from a NAL unit stream format by arranging NAL units in decoding order and prefixing each NAL unit with a header that specifies the exact size (in bytes) of the NAL unit. A sample stream header is included at the beginning of the sample stream bitstream, which specifies the precision (in bytes) of the signaled NAL unit size. A NAL unit stream format can be extracted from a sample stream format by traversing the sample stream format, reading the size information, and appropriately extracting each NAL unit.
[0424] Below, the syntax of the Arthmetic Coded Displacement sub-bitstream is described.
[0425] NAL unit syntax:
[0426] General NAL unit syntax:
[0427] displ_nal_unit( NumBytesInNalUnit ) {Descriptordispl_nal_unit_header( )NumBytesInRbsp = 0for( i = 2; i < NumBytesInNalUnit; i++ )rbsp_byte[ NumBytesInRbsp++ ]b(8)}
[0428] NAL unit header syntax:
[0429] displ_nal_unit_header() {Descriptordispl_nal_forbidden_zero_bitf(1)displ_nal_unit_typeu(6)displ_nal_layer_idu(6)displ_nal_temporal_id_plus1u(3)}
[0430] Raw byte sequence payload, trailing bits, and byte alignment syntax: Displacement sequence parameter set RBSP syntax: General displacement sequence parameter set RBSP syntax:
[0431] displ_sequence_parameter_set_rbsp( ) {Descriptordsps_sequence_parameter_set_idu(4)dsps_codec_idu(8)dsps_profile_tier_level( )dsps_range_log2_minus2u(3)dsps_single_dimension_flagu(1)dsps_msb_align_flagu(1)dsps_log2_max_displ_frame_order_cnt_lsb_minus4ue(v)dsps_max_dec_displ_frame_buffering_minus1ue(v)dsps_long_term_ref_displ_frames_flagu(1)dsps_num_ref_displ_frame_lists_in_dspsue(v)for( i = 0; i < dsps_num_ref_displ_frame_lists_in_dsps; i++ ) displ_ref_list_struct( i )dsps_extension_present_flagu(1)if( dsps_extension_present_flag ) {dsps_extension_count_minus1u(7)dsps_extension_length_minus1ue(v)while( more_rbsp_data( ) )dsps_extension_data_byte u(1)}rbsp_trailing_bits( )}
[0432] 디스플레이스먼트 프로파일, 티어, 및 레벨 신택스:
[0433] Dsps_profile_tier_level( ) {Descriptordptl_tier_flagu(1)dptl_profile_codec_group_idcu(7)dptl_profile_toolset_idcu(8)dptl_reserved_zero_32bitsu(32)dptl_level_idcu(8)dptl_num_sub_profilesu(6)dptl_extended_sub_profile_flagu(1)for( i = 0;
[0434] Profile toolset constraints information syntax:
[0435]
[0436] dptl_profile_toolset_constraints_information( ) {Descriptordptc_one_displacement_frame_only_flagu(1)dptc_reserved_zero_7bitsu(6)dptc_num_reserved_constraint_bytesu(8)for( i = 0; i < dptc_num_reserved_constraint_bytes; i++ )dptc_reserved_constraint_byte[ i ]u(8)}
[0437] Displacement Frame Parameter Set RBSP Syntax: General Displacement Frame Parameter Set RBSP Syntax:
[0438] displ_frame_parameter_set_rbsp( ) {Descriptordfps_displ_sequence_parameter_set_idu(4)dfps_displ_frame_parameter_set_idu(4)displ_information( ) dfps_output_flag_present_flagu(1)dfps_num_ref_idx_default_active_minus1ue(v)dfps_additional_lt_dfoc_lsb_lenue(v)dfps_extension_present_flagu(1)if( dfps_extension_present_flag )dfps_extension_8bitsu(8)if( dfps_extension_8bits )while( more_rbsp_data( ) )dfps_extension_data_flagu(1)rbsp_trailing_bits( )}
[0439] 디스플레이스먼트 레퍼런스 리스트 구조 신택스:
[0440] displ_ref_list_struct( rlsIdx ) {Descriptordrl_num_ref_entries[ rlsIdx ]ue(v)for( i = 0; i < drl_num_ref_entries[ rlsIdx ]; i++ ) {if( dsps_long_term_ref_displ_frames_flag )drl_st_ref_displ_frame_flag[ rlsIdx ][ i ]u(1)if( drl_st_ref_displ_frame_flag[ rlsIdx ][ i ] ) {drl_abs_delta_dfoc_st[ rlsIdx ][ i ]ue(v)if( drl_abs_delta_dfoc_st[ rlsIdx ][ i ] > 0 )drl_straf_entry_sign_flag[ rlsIdx ][ i ]u(1)} elsedrl_dfoc_lsb_lt[ rlsIdx ][ i ]u(v)}}
[0441] Displacement Layer RBSP Syntax:
[0442] displ_layer_rbsp() {Descriptordispl_header()displ_data_unit(displID)rbsp_trailing_bits()}
[0443] Displacement header syntax:
[0444] displ_header( ) {Descriptorif( nal_unit_type >= NAL_BLA_W_LP && nal_unit_type <= NAL_RSV_IRAP_DCL_29 )dh_no_output_of_prior_displ_frames_flagu(1)dh_frame_parameter_set_idu(4)dh_idu(v)displID = dh_iddh_typeue(v)if( dfps_output_flag_present_flag )dh_output_flagu(1)dh_frm_order_cnt_lsbu(v)if( dsps_num_ref_displ_frame_lists_in_dsps > 0 )dh_ref_displ_frame_list_dsps_flagu(1)if( dh_ref_displ_frame_list_dsps_flag == 0 )displ_ref_list_struct( dsps_num_ref_displ_frame_lists_in_dsps )else if( dsps_num_ref_displ_frame_lists_in_dsps > 1 )ref_displ_frame_list_idxu(v)for( j = 0; j < NumLtrDisplFrmEntries[ RlsIdx ]; j++ ) {dh_additional_dfoc_lsb_present_flag[ j ]u(1)if( additional_dfoc_lsb_present_flag[ j ] )dh_additional_dfoc_lsb_val[ j ]u(v)}if( dh_type == P_DISPLACEMENT && num_ref_entries[ RlsIdx ] > 1 ) {dh_num_ref_idx_active_override_flagu(1)if( num_ref_idx_active_override_flag )dh_num_ref_idx_active_minus1ue(v)}dh_log2_subblock_size_minus6u(4)byte_alignment( )}
[0445] Displacement Data Unit Syntax:
[0446] displ_data_unit( displID ) {Descriptorif( dh_type == I_DISPLACEMENT ) {displ_intra_unit( displID )}else if( dh_type == P_DISPLACEMENT ) {displ_inter_unit( displID )}}
[0447] Displacement Intra Data Unit Syntax:
[0448] displ_intra_unit( displID ) {Descriptordiu_lod_count[ displID ]for( i = 0; i < displ_lod_count; i++ ) {diu_vertex_count_lod[ displID ][ i ]subblockSize = 1 << dh_log2_subblock_size_minus6for( k = 0; k < 3; k++ ) {diu_last_sig_coeff[ k ]ae(v)for( b = 0; b < lodCount; b++ ) {diu_coded_block_flag[ k ][ b ]u(v)if( diu_coded_block_flag[ k ][ b] ) {subblockCountPerLevel =Ceil( diu_vertex_count_lod[ displID ][ b ] / subblockSize )for( s = 0; s < subblockCountPerLevel; s++ ) {diu_coded_subblock_flag[ k ][ b ][ s ]u(v)if( diu_coded_subblock_flag[ k ][ b ][ s ] ) {for( v = vStart; v < subblock_size; v++ ) {diu_coeff_abs_level_gt0[ k ][ b ][ s ][ v ]u(v)if( diu_coeff_abs_level_gt0[ k ][ b ][ s ][ v ] ) {diu_coeff_abs_level_gt1[ k ][ b ][ s ][ v ]u(v)diu_coeff_sign[ k ][ b ][ s ][ v ]u(1)if( diu_coeff_abs_level_gt1[ k ][ b ][ s ][ v ] ) {diu_coeff_abs_level_rem[ k ][ b ][ s ][ v ]ue(v)}}}}}}}if ( dsps_single_dimension_flag ) {break;}}}
[0449] Displacement Inter-Data Unit Syntax: The arithmetic decoding engine is a context-separated binary arithmetic decoder that performs binary renormalization and produces binary output. The displacement residual is derived from the arithmetic decoding.
[0450] displ_inter_unit( dispID ) {Descriptor / * Can be the same as specified in 4.3.1.3.7 * / }
[0451] Below, we describe the semantics of the Arthmetic Coded Displacement sub-bitstream. NAL Unit Semantics: General NAL Unit Semantics:
[0452] NumBytesInNalUnit represents the size of a NAL unit in bytes. This value is required for decoding the NAL unit. Some form of delimiting NAL unit boundaries is required to enable inference of NumBytesInNalUnit.
[0453] Note: The displacement coding layer (DCL) is specified to efficiently represent the contents of displacement data. The NAL is specified to format that data and provide header information in a manner suitable for transmission over various communication channels or storage media. All data is contained in NAL units, each of which contains an integer number of bytes. A NAL unit represents a common format that can be used in both packet-oriented systems and bitstream systems. The format of NAL units for both packet-oriented transport and sample streams is the same, but in sample stream formats, each NAL unit may be preceded by an additional element specifying its size.
[0454] rbsp_byte[ i ] is the ith byte of the RBSP. An RBSP is specified as an ordered sequence of bytes as follows:
[0455] An RBSP contains a string of data bits (SODBs) as follows:
[0456] - If the SODB is empty (i.e., has a length of 0 bits), the RBSP is also empty.
[0457] - Otherwise, the RBSP contains the SODB as follows:
[0458] 3) The first byte of the RBSP contains the first (most significant, leftmost) 8 bits of the SODB. The next byte of the RBSP contains the next 8 bits of the SODB, and so on until there are fewer than 8 bits of the SODB left.
[0459] 4) The rbsp_trailing_bits( ) syntax structure exists after SODB as follows.
[0460] iv) The first (most significant, leftmost) bit of the last RBSP byte contains the remaining bits of the SODB (if any).
[0461] v) The next bit consists of a single bit equal to 1 (i.e. rbsp_stop_one_bit).
[0462] vi) If rbsp_stop_one_bit is not the last bit of a byte that is aligned, byte alignment is achieved by the presence of one or more bits that are equal to 0 (i.e., instances of rbsp_alignment_zero_bit).
[0463] Syntax structures with these RBSP properties are indicated in the syntax table using the suffix "_rbsp". These structures are conveyed within a NAL unit as the contents of the rbsp_byte[ i ] data byte. The association between RBSP syntax structures and NAL units is shown in the table below.
[0464] Note: If the boundaries of an RBSP are known, a decoder can extract an SODB from an RBSP by concatenating the byte bits of the RBSP, discarding the rbsp_stop_one_bit whose last (least significant, rightmost) bit is 1, and discarding the following (less significant, rightmost) bits that are 0. The data required for the decoding process is contained in the SODB portion of the RBSP.
[0465] NAL unit header semantics:
[0466] displ_nal_forbidden_zero_bit
[0467] displ_nal_unit_type
[0468] As with the atlas case, a similar NAL unit type is defined for displacement, and similar functionality for random access defines a specific NAL unit corresponding to the coded displacement data. Additionally, NAL units that can contain metadata, such as SEI messages, are also defined.
[0469] In particular, the supported displacement NAL unit types are specified as follows:
[0470] NAL Unit Type Codecs and NAL Unit Type Classes:
[0471] displ_nal_unit_typeName of displ_nal_unit_typeContent of displacement NAL unit and RBSP syntax structureNAL unitype class01NAL_TRAIL_NNAL_TRAIL_RCoded displacement of a non-TSA, non STSA trailing displacement framedispl_layer_rbsp( )DCL23NAL_TSA_NNAL_TSA_RCoded displacement of a TSA displacement framedispl_layer_rbsp( )DCL45NAL_STSA_NNAL_STSA_RCoded displacement of a STSA displacement framedispl_layer_rbsp( )DCL67NAL_RADL_NNAL_RADL_RCoded displacement of a RADL displacement framedispl_layer_rbsp( )DCL89NAL_RASL_NNAL_RASL_RCoded displacement of a RASL displacement framedispl_layer_rbsp( )DCL1011NAL_SKIP_NNAL_SKIP_RCoded displacement of a skipped displacement framedispl_layer_rbsp( )DCL1214NAL_RSV_DCL_N12NAL_RSV_DCL_N14Reserved non-IRAP sub-layer non-reference DCL displacement NAL unit typesDCL1315NAL_RSV_DCL_R13NAL_RSV_DCL_R15Reserved non-IRAP sub-layer reference DCL displacement NAL unit typesDCL161718NAL_BLA_W_LPNAL_BLA_W_RADLNAL_BLA_N_LPCoded displacement of a BLA displacementframedispl_layer_rbsp( )DCL1920NAL_IDR_W_RADLNAL_IDR_N_LPCoded displacement of an IDR displacement framedispl_layer_rbsp( )DCL21NAL_CRACoded displacement of a CRA displacement framedispl_layer_rbsp( )DCL2223NAL_RSV_IRAP_DCL_22NAL_RSV_IRAP_DCL_23Reserved IRAP DCL NAL unit typesDCL24..29NAL_RSV_DCL_24..NAL_RSV_DCL_29Reserved non-IRAP DCL NAL unit typesDCL30NAL_DSPSDisplacement sequence parameter setdispl_sequence_parameter_set_rbsp( )non-DCL31NAL_DFPSDisplacement frame parameter setdispl_frame_parameter_set_rbsp( )non-DCL32NAL_DAUDAccess unit delimiteraccess_unit_delimiter_rbsp( )non-DCL33NAL_DEOSEnd of sequenceend_of_sequence_rbsp( )non-DCL34NAL_DEOBEnd of bitstreamend_of_displ_sub_bitstream_rbsp( )non-DCL35NAL_FDFillerfiller_data_rbsp( )non-DCL3637NAL_PREFIX_NSEINAL_SUFFIX_NSEINon-essential supplemental enhancement informationsei_rbsp( )non-DCL383940..4445..63NAL_PREFIX_ESEINAL_SUFFIX_ESEINAL_RSV_NDCL_40NAL_RSV_NDCL_44NAL_UNSPEC_45NAL_UNSPEC_63Essential supplemental enhancementinformationsei_rbsp( )Reserved non-DCL NAL unit typessnspecified non-DCL NAL unit typesnon-DCLnon-DCLnon-DCL
[0472] displ_nal_layer_iddispl_nal_temporal_id_plus1 The order of NAL units and displacement frames, and their relationship to coded displacement frames, access units, and coded displacement sequences: Low-byte sequence payload, trailing bits, and byte alignment semantics:
[0473] Displacement Sequence Parameter Set RBSP Semantics:
[0474] General displacement sequence parameter set RBSP semantics:
[0475] dsps_sequence_parameter_set_id: An identifier for the displacement sequence parameter set so that other syntax elements can reference it.
[0476] dsps_codec_id: The identifier of the codec used to compress the displacement. dsps_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, an SEI message mapping component codec, or by means external to this document. It may be associated with a specific displacement codec through a profile specified in that specification, or explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0477] dsps_range_log2_minus2: Adding 2 to this value gives the geometric displacement coordinate range of the displacement. dsps_range_log2_minus2 is in the range of 0 to 3.
[0478] dsps_single_dimension_flag: Indicates the number of dimensions for the displacement associated with the displacement. If dsps_single_dimension_flag is 0, it indicates that three components are used for the displacement. If dsps_single_dimension_flag is 1, it indicates that only the normal component is used for the displacement.
[0479] dsps_msb_align_flag: Indicates how decoded displacement samples are converted to samples of displacement range bit depth.
[0480] dsps_log2_max_displ_frame_order_cnt_lsb_minus4: Adding 4 to this value gives the values of the variables Log2MaxDisplFrmOrderCntLsb and MaxDisplFrmOrderCntLsb used in the decoding process for the displacement frame order count, as follows:
[0481] Log2MaxDisplFrmOrderCntLsb =
[0482] dsps_log2_max_displ_frame_order_cnt_lsb_minus4 + 4
[0483] MaxDisplFrmOrderCntLsb = 2Log2MaxDisplFrmOrderCntLsb
[0484] The value of dsps_log2_max_displ_frame_order_cnt_lsb_minus4 is in the range of 0 to 12.
[0485] dsps_max_dec_displ_frame_buffering_minus1 plus 1 indicates the maximum required size of the decoded disparity frame buffer for CDS, in disparity frame storage buffer units. The value of dsps_max_dec_displ_frame_buffering_minus1 is in the range 0 to 15.
[0486] If dsps_long_term_ref_displ_frames_flag is 0, it indicates that long-term reference disparity is not used for inter prediction of any coded disparity frames in the CDS. If dsps_long_term_ref_displ_frames_flag is 1, it indicates that long-term reference disparity frames can be used for inter prediction of one or more coded disparity frames in the CDS.
[0487] dsps_num_ref_displ_frame_lists_in_dsps represents the number of displ_ref_list_struct(rlsIdx) syntax structures included in the displacement sequence parameter set. The value of dsps_num_ref_displ_frame_lists_in_dsps is in the range of 0 to 64.
[0488] NOTE: The decoder allocates memory for a total number of displ_ref_list_struct(rlsIdx) syntax structures equal to (dsps_num_ref_displ_frame_lists_in_dsps + 1), since there can be only one displ_ref_list_struct(rlsIdx) syntax structure signaled directly in the displacement header of the current displacement frame.
[0489] If dsps_extension_present_flag is 1, it indicates that dsps_extension_count_minus1 and dsps_extension_length_minus1 are in the displacement sequence parameter set.
[0490] Adding 1 to dsps_extension_count_minus1 indicates the number of extensions in the current displacement sequence parameter set. If none exist, dsps_extension_count_minus1 is inferred to be equal to -1.
[0491] dsps_extension_length_minus1 plus 1 represents the length of the dsps_extension_data_byte element following this syntax element. If not present, dsps_extension_length_minus1 is inferred to be equal to -1.
[0492] dsps_extension_data_byte can have any value.
[0493] Displacement profile, tier, and level semantics:
[0494] dptl_tier_flag indicates the tier context for interpreting dptl_level_idc as specified in ISO / IEC 23090-29.
[0495] dptl_profile_codec_group_idc indicates the codec group profile component that the CDS complies with, as specified in ISO / IEC 23090-29.
[0496] dptl_profile_toolset_idc represents the toolset combination profile component that CDS complies with as specified in ISO / IEC 23090-29.
[0497] If dptl_reserved_zero_32bits is present, it is equal to 0 in a bitstream conforming to this version of this document.
[0498] dptl_level_idc indicates the level to which the CDS complies as specified in ISO / IEC 23090-29.
[0499] dptl_num_sub_profiles represents the number of dptl_sub_profile_idc[ i ] syntax elements.
[0500] If dptl_extended_sub_profile_flag is 1, it indicates that the dptl_sub_profile_idc[ i ] syntax element, if present, should be represented using 64 bits. If dptl_extended_sub_profile_flag is 0, it indicates that the dptl_sub_profile_idc[ i ] syntax element, if present, should be represented using 32 bits.
[0501] dptl_sub_profile_idc[ i ] represents the ith registered interoperability metadata as specified in Rec. ITU-T T.35. The number of bits used to represent dptl_sub_profile_idc[ i ] is equal to (dptl_extended_sub_profile_flag == 0 ? 32 : 64).
[0502] If dptl_toolset_constraints_present_flag is 1, it indicates that the additional structure dptl_profile_toolset_constraints_information( ) is present in the bitstream. If dptl_toolset_constraints_present_flag is 0, it indicates that the structure dptl_profile_toolset_constraints_information( ) is not present.
[0503] Displacement Profile Toolset Constraint Information Semantics:
[0504]
[0505] If dptc_one_displacement_frame_only_flag is present, it has the meaning specified in ISO / IEC 23090-29, where the profile indicated by dptl_profile_toolset_idc is the profile specified in ISO / IEC 23090-29. If absent, dptc_one_displacement_frame_only_flag is inferred to be equal to 0.
[0506] dptc_reserved_zero_7bits is equal to 0 in bitstreams that follow this version of this document.
[0507] dptc_num_reserved_constraint_bytes indicates the number of reserved constraint bytes.
[0508] dptc_reserved_constraint_byte[ i ] can have any value.
[0509] Displacement Frame Parameter Set RBSP Semantics
[0510] General displacement frame parameter set RBSP semantics
[0511] dfps_displ_sequence_parameter_set_id specifies the value of dsps_sequence_parameter_set_id for the active displacement sequence parameter set.
[0512] dfps_displ_parameter_set_id identifies the displacement frame parameter set so that other syntax elements can reference it.
[0513] If dfps_output_flag_present_flag is 1, it indicates that the displ_output_flag syntax element is present in the associated displacement header. If dfps_output_flag_present_flag is 0, it indicates that the displ_output_flag syntax element is not present in the associated displacement header.
[0514] dfps_num_ref_idx_default_active_minus1: Adding 1 to this value gives the inferred value of the variable NumRefIdxActive for tiles with displ_num_ref_idx_active_override_flag equal to 0. The value of dfps_num_ref_idx_default_active_minus1 is in the range 0 to 14.
[0515] dfps_additional_lt_dfoc_lsb_len represents the value of the variable MaxLtDisplFrmOrderCntLsb used in the decoding process of the reference atlas frame list.
[0516] MaxLtDisplFrmOrderCntLsb =
[0517] 2 * (Log2MaxDisplFrmOrderCntLsb + dfps_additional_lt_dfoc_lsb_len)
[0518] The value of dfps_additional_lt_dfoc_lsb_len is in the range 0 to 32 - Log2MaxDisplFrmOrderCntLsb.
[0519] If dsps_long_term_ref_displ_frames_flag is 0, the value of dfps_additional_lt_dfoc_lsb_len is equal to 0.
[0520] If dfps_extension_present_flag is 1, it indicates that the syntax element dfps_extension_8bits is present in the displacement frame parameter set. If dfps_extension_present_flag is 0, it indicates that the syntax element dfps_extension_8bits is not present. In this version of this document, the value of dfps_extension_present_flag is 0.
[0521] If dfps_extension_8bits is 0, it indicates that the DFPS RBSP syntax structure does not have a dfps_extension_data_flag syntax element. If dfps_extension_8bits is present, it is equal to 0 in bitstreams that follow this version of this document.
[0522] dfps_extension_data_flag can have any value.
[0523] The displ_no_output_of_prior_displ_frames_flag affects the output of previously decoded displaced frames in the DDB after decoding a displaced frame in a CDS AU that is not the first AU in the bitstream, as specified in ISO / IEC 23090-29. If no_output_of_prior_displ_frames_flag is absent, its value is inferred to be equal to 0.
[0524] As a requirement for bitstream conformance, the value of no_output_of_prior_displ_frames_flag is the same for all displacement frames in an AU.
[0525] The no_output_of_prior_displ_frames_flag value in the displacement header is the output_of_prior_displ_frames_flag value in the AU.
[0526] displ_frame_parameter_set_id indicates the value of dfps_displ_frame_parameter_set_id for the active displacement frame parameter set for the current displacement frame.
[0527] dislp_type indicates the coding type of the current displacement frame, as shown in the table below. The value of smh_type is 0, 1, or 2 in bitstreams that follow this version of this document. Other values of smh_type are reserved for future use in ISO / IEC.
[0528] Relationship to dislp_type
[0529] smh_typeName of smh_type0P_DISPLACEMENT1I_DISPLACEMENT2...RESERVED
[0530] utput_flag: Affects the decoded displacement output and removal process as specified in ISO / IEC 23090-29. If displ_output_flag is absent, it is inferred to be equal to 1. displ_frm_order_cnt_lsb: Indicates the number of displacement frame orders modulo MaxDisplFrmOrderCntLsb for the current displacement frame. The length of the displ_frm_order_cnt_lsb syntax element is equal to Log2MaxDisplFrmOrderCntLsb bits. The value of displ_frm_order_cnt_lsb is in the range 0 to MaxDisplFrmOrderCntLsb - 1. If ref_displ_frame_list_dsps_flag is 1, it indicates that the reference displacement frame list of the current displacement frame is derived based on one of the displ_ref_list_struct(rlsIdx) syntax structures of the active DSPS. If ref_displ_frame_list_dsps_flag is 0, it indicates that the reference displacement frame list of the current displacement frame is derived based on the displ_ref_list_struct(rlsIdx) syntax structure contained directly in the displacement frame header of the current displacement frame. When dsps_num_ref_displ_frame_lists_in_dsps is 0, the value of ref_displ_frame_list_dsps_flag is inferred to be 0.
[0531] ref_displ_frame_list_idx indicates the index of the displ_ref_list_struct(rlsIdx) syntax structure used to derive the reference displacement frame list for the current displacement frame in the list of displ_ref_list_struct(rlsIdx) syntax structures contained in the active DSPS. The syntax element ref_displ_frame_list_idx is expressed as Ceil(Log2(dsps_num_ref_displ_frame_lists_in_dsps)) bits. If not present, the value of ref_displ_frame_list_idx is inferred to be equal to 0. The value of ref_displ_frame_list_idx ranges from 0 to dsps_num_ref_displ_frame_lists_in_dsps - 1, inclusive. If ref_displ_frame_list_dsps_flag is 1 and dsps_num_ref_displ_frame_lists_in_dsps is 1, the value of ref_displ_frame_list_idx is inferred to be equal to 0.
[0532] The variable RlsIdx of the current atlas tile is derived as follows:
[0533] RlsIdx = ref_displ_frame_list_dsps_flag ?
[0534] ref_displ_frame_list_idx : dsps_num_ref_displ_frame_lists_in_dsps
[0535] If additional_dfoc_lsb_present_flag[ j ] is 1, it indicates that additional_dfoc_lsb_val[ j ] is present for the current displacement frame. If additional_dfoc_lsb_present_flag[ j ] is 0, it indicates that additional_dfoc_lsb_val[ j ] is not present.
[0536]
[0537] additional_dfoc_lsb_val[ j ] represents the value of FullFrmOrderCntLsbLt[ RlsIdx ][ j ] for the current atlas tile as follows:
[0538] FullDisplFrmOrderCntLsbLt[RlsIdx][j] =
[0539] additional_dfoc_lsb_val[ j ] * MaxDisplFrmOrderCntLsb +dfoc_lsb_lt[ RlsIdx ][ j ]
[0540] The syntax element additional_dfoc_lsb_val[ j ] is represented by the dfps_additional_lt_dfoc_lsb_len bits. If it is not present, the value of additional_dfoc_lsb_val[ j ] is inferred to be equal to 0.
[0541] If num_ref_idx_active_override_flag is 1, it indicates that the syntax element num_ref_idx_active_minus1 exists for the current displacement frame. If num_ref_idx_active_override_flag is 0, it indicates that the syntax element num_ref_idx_active_minus1 does not exist. If num_ref_idx_active_override_flag is absent, its value is inferred to be 0.
[0542] num_ref_idx_active_minus1 is used to derive the variable NumRefIdxActive as specified in Equation 5 for the current displacement frame. The value of num_ref_idx_active_minus1 is in the range 0 to 14.
[0543] If the current displacement frame is a P_DISPLACEMENT displacement frame, num_ref_idx_active_override_flag is 1, and num_ref_idx_active_minus1 is absent, then num_ref_idx_active_minus1 is inferred to be equal to 0.
[0544] The variable NumRefIdxActive is derived as follows:
[0545] if( displ_type == P_DISPLACEMENT ) {
[0546] if( num_ref_idx_active_override_flag == 1 )
[0547] NumRefIdxActive = num_ref_idx_active_minus1 + 1
[0548] else {
[0549] if( num_ref_entries[ RlsIdx ] >= dfps_num_ref_idx_default_active_minus1 + 1 )
[0550] NumRefIdxActive = dfps_num_ref_idx_default_active_minus1 + 1
[0551] else
[0552] NumRefIdxActive = num_ref_entries[RlsIdx]
[0553] }
[0554] }
[0555] else
[0556] NumRefIdxActive = 0
[0557] The value of NumRefIdxActive minus 1 represents the maximum number of displacement reference frame indices that can be used to decode the current displacement frame.
[0558] Displacement Reference List Structure Semantics:
[0559] drl_num_ref_entries[rlsIdx] represents the number of entries in the displ_ref_list_struct(rlsIdx) syntax structure, where rlsIdx is the index of the displacement frame reference list. For P_DISPLACEMENT, the value of num_ref_entries[rlsIdx] is in the range of 1 to dsps_max_dec_displ_frame_buffering_minus1 + 1. Otherwise, the value of num_ref_entries[rlsIdx] is in the range of 0 to dsps_max_dec_displ_frame_buffering_minus1 + 1.
[0560] If drl_st_ref_displ_frame_flag[rlsIdx][i] is 1, it indicates that the ith entry in the displ_ref_list_struct(rlsIdx) syntax structure is a short-term reference displacement frame entry. If st_ref_displ_frame_flag[rlsIdx][i] is 0, it indicates that the ith entry in the displ_ref_list_struct(rlsIdx) syntax structure is a long-term reference displacement frame entry. If not present, the value of drl_st_ref_displ_frame_flag[rlsIdx][i] is inferred to be equal to 1.
[0561] The variable NumLtrDisplFrmEntries[rlsIdx] is derived as follows:
[0562] NumLtrDisplFrmEntries[rlsIdx] = 0
[0563] for( i = 0; i < drl_num_ref_entries[rlsIdx]; i++)
[0564] if(!drl_st_ref_displ_frame_flag[rlsIdx][i])
[0565] NumLtrDisplFrmEntries[rlsIdx]++
[0566] drl_abs_delta_dfoc_st[rlsIdx][i], if the i-th item is the first short-term reference displacement frame entry in the displ_ref_list_struct(rlsIdx) syntax structure, specifies the absolute difference between the displacement frame order count value of the current displacement frame referenced by the i-th item, or if the i-th item is a short-term reference displacement frame entry but is not the first short-term reference displacement frame entry in the displ_ref_list_struct(rlsIdx) syntax structure, specifies the absolute difference between the displacement frame order count value of the displacement frame referenced by the i-th item in the displ_ref_list_struct(rlsIdx) syntax structure and the previous short-term reference displacement frame entry.
[0567] The value of drl_abs_delta_dfoc_st[rlsIdx][i] is in the range of 0 to 215-1.
[0568] If drl_straf_entry_sign_flag[rlsIdx][i] is 1, it indicates that the ith entry in the syntax structure displ_ref_list_struct(rlsIdx) has a value greater than or equal to 0. If drl_straf_entry_sign_flag[rlsIdx][i] is 0, it indicates that the ith entry in the syntax structure displ_ref_list_struct(rlsIdx) has a value less than 0. If absent, the value of drl_straf_entry_sign_flag[rlsIdx][i] is inferred to be 1.
[0569] The DeltaDfocSt[rlsIdx][i] list is derived as follows:
[0570] for( i = 0; i < drl_num_ref_entries[rlsIdx]; i++ )
[0571] if( drl_st_ref_displ_frame_flag[rlsIdx][i])
[0572] DeltaDfocSt[rlsIdx][i] =
[0573] ( 2 * drl_straf_entry_sign_flag[rlsIdx][i] - 1) * drl_abs_delta_dfoc_st[rlsIdx][i]
[0574] else
[0575] DeltaDfocSt[rlsIdx][i] = 0
[0576] drl_dfoc_lsb_lt[rlsIdx][i] represents the displacement frame order count value for MaxDisplFrmOrderCntLsb of the displacement frame referenced by the ith item of the displ_ref_list_struct(rlsIdx) syntax structure. The length of the drl_dfoc_lsb_lt[ rlsIdx ][ i ] syntax element is Log2MaxDisplFrmOrderCntLsb bits.
[0577] Displacement Layer RBSP Semantics:
[0578] Displacement header semantics:
[0579] The dh_no_output_of_prior_displ_frames_flag affects the output of previously decoded displacement frames in the DDB after decoding displacement frames in a CDS AU that is not the first AU in the bitstream, as specified in ISO / IEC 23090-29. If no_output_of_prior_displ_frames_flag is absent, its value is inferred to be 0.
[0580] As a requirement for bitstream conformance, the value of no_output_of_prior_displ_frames_flag is the same for all displacement frames in an AU.
[0581] The no_output_of_prior_displ_frames_flag value in the displacement header is the output_of_prior_displ_frames_flag value in the AU.
[0582] dh_frame_parameter_set_id indicates the dfps_displ_frame_parameter_set_id value for the active displacement frame parameter set for the current displacement frame.
[0583] dh_id: Indicates the displacement header ID.
[0584] dh_type indicates the coding type of the current displacement frame, as shown in the table below. The value of dh_type is 0 or 1 in bitstreams that follow this version of this document. Other values of dh_type are reserved for future use in ISO / IEC.
[0585] dh_type relationship
[0586] smh_typeName of smh_type0P_DISPLACEMENT1I_DISPLACEMENT2...RESERVED
[0587] dh_frm_order_cnt_lsb: Indicates the number of displacement frame orders modulo MaxDisplFrmOrderCntLsb for the current displacement frame. The length of the dh_frm_order_cnt_lsb syntax element is equal to Log2MaxDisplFrmOrderCntLsb bits. The value of dh_frm_order_cnt_lsb ranges from 0 to MaxDisplFrmOrderCntLsb - 1. If dh_ref_displ_frame_list_dsps_flag is 1, it indicates that the reference displacement frame list of the current displacement frame is derived based on one of the displ_ref_list_struct(rlsIdx) syntax structures of the active DSPS. If dh_ref_displ_frame_list_dsps_flag is 0, it indicates that the reference displacement frame list of the current displacement frame is derived based on the displ_ref_list_struct(rlsIdx) syntax structure directly included in the displacement frame header of the current displacement frame. If dsps_num_ref_displ_frame_lists_in_dsps is 0, the value of dh_ref_displ_frame_list_dsps_flag is inferred to be 0. dh_ref_displ_frame_list_idx specifies the index of the list of displ_ref_list_struct(rlsIdx) syntax structures included in the active DSPS, and is the displ_ref_list_struct(rlsIdx) syntax structure used to derive the reference displacement frame list of the current displacement frame. The syntax element dh_ref_displ_frame_list_idx is expressed in bits as Ceil(Log2(dsps_num_ref_displ_frame_lists_in_dsps)). If not present, the value of dh_ref_displ_frame_list_idx is inferred to be equal to 0. The value of dh_ref_displ_frame_list_idx is in the range 0 to dsps_num_ref_displ_frame_lists_in_dsps - 1.If dh_ref_displ_frame_list_dsps_flag is 1 and dsps_num_ref_displ_frame_lists_in_dsps is 1, the value of ref_displ_frame_list_idx is inferred to be 0. The variable RlsIdx of the current atlas tile is derived as follows.
[0588] RlsIdx = dh_ref_displ_frame_list_dsps_flag ?
[0589] ref_displ_frame_list_idx : dsps_num_ref_displ_frame_lists_in_dsps
[0590] If dh_additional_dfoc_lsb_present_flag[ j ] is 1, it indicates that dh_additional_dfoc_lsb_val[ j ] is present in the current displacement frame. If dh_additional_dfoc_lsb_present_flag[ j ] is 0, it indicates that dh_additional_dfoc_lsb_val[ j ] is not present.
[0591] dh_additional_dfoc_lsb_val[ j ] represents the value of FullFrmOrderCntLsbLt[ RlsIdx ][ j ] for the current atlas tile, as follows:
[0592] FullDisplFrmOrderCntLsbLt[ RlsIdx ][ j ] = dh_additional_dfoc_lsb_val[ j ] *
[0593] MaxDisplFrmOrderCntLsb +dfoc_lsb_lt[ RlsIdx ][ j ]
[0594] The syntax element dh_additional_dfoc_lsb_val[ j ] is represented by the dfps_additional_lt_dfoc_lsb_len bits. If it is not present, the value of dh_additional_dfoc_lsb_val[ j ] is inferred to be equal to 0.
[0595] If dh_num_ref_idx_active_override_flag is 1, it indicates that the syntax element num_ref_idx_active_minus1 exists for the current displacement frame. If dh_num_ref_idx_active_override_flag is 0, it indicates that the syntax element num_ref_idx_active_minus1 does not exist. If dh_num_ref_idx_active_override_flag is absent, its value is inferred to be equal to 0.
[0596] dh_num_ref_idx_active_minus1 is used to derive the variable NumRefIdxActive for the current displacement frame. The value of dh_num_ref_idx_active_minus1 is in the range 0 to 14.
[0597] If the current displacement frame is a P_DISPLACEMENT displacement frame, dh_num_ref_idx_active_override_flag is 1, and dh_num_ref_idx_active_minus1 is absent, then dh_num_ref_idx_active_minus1 is inferred to be equal to 0.
[0598] The variable NumRefIdxActive is derived as follows:
[0599] if( dh_type == P_DISPLACEMENT ) {
[0600] if( dh_num_ref_idx_active_override_flag == 1 )
[0601] NumRefIdxActive = dh_num_ref_idx_active_minus1 + 1
[0602] else {
[0603] if( num_ref_entries[ RlsIdx ] >= dfps_num_ref_idx_default_active_minus1 + 1 )
[0604] NumRefIdxActive = dfps_num_ref_idx_default_active_minus1 + 1
[0605] else
[0606] NumRefIdxActive = num_ref_entries[RlsIdx]
[0607] }
[0608] }
[0609] else
[0610] NumRefIdxActive = 0
[0611] Subtracting 1 from NumRefIdxActive indicates the maximum number of displacement reference frame indices that can be used to decode the current displacement frame.
[0612] Adding 6 to dh_log2_subblock_size_minus6 gives the value of the variable subblockSize as follows:
[0613] subblockSize = 1 << ( log2_subblock_size_minus6 + 6 )
[0614] Displacement Data Unit Semantics:
[0615] displ_intra_unit( displID ) contains a displacement unit stream, which is an aligned stream of bytes or bits, within which the location of unit boundaries can be identified from a pattern in the data. The format of this displacement unit stream is identified by the 4CC code defined by dptl_profile_codec_group_idc or by the component codec mapping SEI message.
[0616] displ_inter_unit( displID ) contains a displacement unit stream, which is an aligned stream of bytes or bits, within which the location of unit boundaries can be identified from a pattern in the data. The format of this displacement unit stream is identified by the 4CC code defined by dptl_profile_codec_group_idc or by the component codec mapping SEI message.
[0617] Displacement Intra Data Unit Semantics:
[0618] The arithmetic decoding engine is a context-separated binary arithmetic decoder that performs binary renormalization and produces binary output.
[0619] The displacement values are derived from arithmetic decoding.
[0620] diu_lod_count[ displID ] indicates the number of granularity levels used for signaled displacements in the data unit associated with displId displID.
[0621] diu_vertex_count_lod[ displID ] [ i ] represents the displacement count for the i-th level of the wavelet transform for the data unit associated with displId displID.
[0622] diu_last_sig_coeff[ k ] represents the index of the last non-zero displacement coefficient level in the kth component.
[0623] diu_coded_block_flag[ k ][ b ] indicates whether the block with index b has a non-zero displacement coefficient level in the kth component (if 1) or not (if 0).
[0624] diu_coded_subblock_flag[ k ][ b ][ s ] indicates whether the subblock with index s of the block with index b has a non-zero displacement coefficient level in the kth component (if 1) or not (if 0).
[0625] diu_coeff_abs_level_gt0[ k ][ b ][ s ][ v ] indicates whether the kth component of the displacement coefficient level associated with the vertex with index v in the subblock with index s of the block with index b has an absolute value greater than 0 (if 1) or not (if 0).
[0626] diu_coeff_abs_level_gt1[ k ][ b ][ s ][ v ] indicates whether the kth component of the displacement coefficient level associated with the vertex with index v in the subblock with index s of the block with index b has an absolute value greater than 1 (if 1) or not (if 0). If diu_coeff_abs_level_gt1[ k ][ b ][ s ][ v ] is absent, it is inferred to be equal to 0.
[0627] diu_coeff_sign[ k ][ b ][ s ][ v ] indicates whether the kth component of the displacement coefficient level associated with the vertex with index v in the subblock with index s of the block with index b has a positive sign (if 1) or not (if 0). If diu_coeff_sign[ k ][ b ][ s ][ v ] is absent, it is inferred to be equal to 1.
[0628] diu_coeff_abs_level_rem[ k ][ b ][ s ][ v ] represents the absolute value of the kth component of the displacement coefficient level associated with the vertex with index v in the block whose index is b minus 2. If diu_coeff_abs_level_rem[ k ][ b ][ s ][ v ] is absent, it is inferred to be equal to 0.
[0629] Displacement Inter-Data Unit Semantics:
[0630] The arithmetic decoding engine is a context-separated binary arithmetic decoder that performs binary renormalization and produces binary output.
[0631] The displacement residual is derived from arithmetic decoding. It can be the same as described above.
[0632] Figure 17 shows a mesh system according to embodiments.
[0633] Referring to the overall structure of the mesh system in Fig. 17, scenes and / or objects acquired in the real world using multiple cameras, sensors, and / or virtual cameras are output in the form of a V-DMC bitstream after going through mesh data pre-processing and mesh data encoding processes. This mesh bitstream can be converted into a file format suitable for storage and / or transmission through a file encapsulation process.
[0634] The dynamic mesh player can restore the file acquired and / or received through the above process into a mesh bitstream form through a file decapsulation process, and the restored mesh bitstream can be displayed on a display device, etc. through a mesh data decoding process and a mesh data processing / rendering process.
[0635] The mesh data encapsulation / decryption method according to the embodiments includes a method related to file / segment encapsulation or encapsulator and file / segment decapsulation or encapsulator.
[0636] A track of a file according to embodiments may include a common data structure.
[0637] Common Data Structure:
[0638] DMC Decoder Configuration Record:
[0639] This information represents decoder configuration information for mesh-based point cloud content. This record contains a version field. The specification for this version defines version 1 of this record. Any incompatible changes to the record are indicated by a change in the version number.
[0640] Syntax:
[0641] aligned(8) class DMCDecoderConfigurationRecord {
[0642] unsigned int(8) configurationVersion = 1;
[0643] unsigned int(8) num_of_setup_units;
[0644] for (i=0; I < num_of_setup_units; i++) {
[0645] unsigned int(8) setup_unit_type; / VPS, SPS, FPS, GPS, APS…
[0646] SetupUnit setup_unit;
[0647] }
[0648] / There may be additional fields.
[0649] }
[0650] Semantics:
[0651] configurationVersion: This is the version field. Incompatible changes to a record are indicated by a change in the version number.
[0652] num_of_setup_units: Indicates the number of DMC setup units in the decoder configuration record.
[0653] setup_unit_type represents a set of parameters of DMC-related types.
[0654] A SetupUnit is an instance of an encapsulating structure that carries a VPS (V-DMC Parameter Set), SPS (Sequence Parameter Set), FPS (Frame Parameter Set), DPS (Displacement Parameter Set), GPS (Geometry Parameter Set), and / or APS (Attribute Parameter Set). These parameter sets may be based on parameter sets defined in the V-DMC specification.
[0655] DMC decoder configuration box
[0656] The DMC Decoder Configuration box contains a DMCDecoderConfigurationRecord. The version is 0.
[0657] Syntax:
[0658] class DMCConfigurationBox extends FullBox('dmcC', version = 0, 0) {
[0659] DMCDecoderConfigurationRecord();
[0660] }
[0661] Semantics:
[0662] DMCDecoderConfigurationRecord follows the description above.
[0663] DMC component information record:
[0664] A DMC component information record represents DMC component information, including the type of DMC component (e.g., geometry or properties).
[0665] Syntax:
[0666] aligned(8) class DMCComponentInfoRecord(){
[0667] unsigned int(8) component_type;
[0668] if(component_type == 4){ / property component
[0669] unsigned int(8) attr_index;
[0670] utf8string attr_name;
[0671] }
[0672] / Additional fields may exist.
[0673] }
[0674] Semantics:
[0675] component_type: Identifies the type of DMC component specified in the table below. In this version of this document, the value of this field is 1, 2, or 4.
[0676] DMC component types
[0677] component_type valueDescription1Displacement component2Geometry component3Reserved4Attribute component5..31Reserved
[0678] attr_index is the type of the attribute, i.e. it can indicate its kind and can be mapped to the bmsps_mesh_attribute_type_id value of the basemesh sequence parameter set. attr_name specifies a human-readable name for the type of the DMC attribute component. DMC Component Information Box
[0679] If this box is present in a sample entry on a track, it indicates the type of DMC component that is being carried by that track. If this box is present in a sample entry on a DMC Attribute track, it also provides the attribute name and optional attribute type information.
[0680] Syntax:
[0681] aligned(8) class DMCComponentInfoBox extends FullBox('dcin', 0, flags){
[0682] DMCComponentInfoRecord();
[0683] }
[0684] Semantics:
[0685] DMCComponentInfoRecord follows the description above.
[0686] Multiple Track Encapsulation:
[0687] When a mesh data encoding method according to embodiments encodes mesh data to generate a V-DMC bitstream and encapsulates the bitstream into one or more tracks, a track type can be configured according to a data type included in the bitstream. Each of these can process a basemesh track, an atlas track, a geometry track or a displacement track, and an attribute track. The geometry track may correspond to a case where displacement data is recorded using a video codec, and the displacement track may correspond to a case where arithmetic coding is used. In addition, when the displacement data is processed using arithmetic coding, the basemesh data, the atlas data, and the displacement data can be configured into one track and encapsulated.
[0688] Basemesh track sample entry
[0689] Sample Entry Type: 'bmc1', 'bmcg'
[0690] Container: SampleDescriptionBox
[0691] Mandatory: A 'bmc1' or 'bmcg' sample entry is mandatory
[0692] Quantity: One or more
[0693] An encoding method according to embodiments may encapsulate basemesh data of a V-DMC bitstream into a basemesh track. A decoding method according to embodiments may decapsulate a file to parse a basemesh track including basemesh data of a V-DMC bitstream. A sample entry of a basemesh track may include configuration box (DMCConfigurationBox) information. If the sample entry type is 'bmc1', all parameter sets related to basemesh data may be included in a setup unit (setup_unit) of the DMCConfigurationBox. If the sample entry type is 'bmcg', all parameter sets related to basemesh data may be included in the setup unit of the DMCConfigurationBox and / or may be included in a sample of the basemesh track. A receiver (decoder) may recognize a track whose sample entry type is 'bmc1' or 'bmcg' as an entry point and perform a decoding operation.
[0694] The values of the sample entry type according to the embodiments, such as 'bmc1', 'bmcg', and 'bmcb', can be expressed according to, for example, the 4CC (Four Character code) of MP4RA. Depending on the type of file format, they can be used to identify the type of video or audio stream, etc. The notations 'bmc1', 'bmcg', and 'bmcb' according to the embodiments can be replaced with other name formats.
[0695] The sample entry type 'bmcb' indicates that the basemesh track references one or more submesh tracks. The references between basemesh tracks and submesh tracks are described in detail in the submesh data encapsulation description below.
[0696] Syntax:
[0697] aligned(8) class DMCBaseMeshSampleEntry() extends VolumetricVisualSampleEntry (type) {
[0698] / type is 'bmc1' or 'bmcg'
[0699] DMCConfigurationBox config; / as defined in section 4.5.2
[0700] }
[0701] Semantics:
[0702] config is the decoder configuration box mentioned above.
[0703] If the basemesh track is an entry point, the config information may include a V-DMC or V3C parameter set (VPS) and / or parameter sets such as SPS, FPS, etc. associated with the basemesh bitstream.
[0704] Basemesh track sample format
[0705] As in ISO / IEC 23090-29, each sample within a basemesh track corresponds to a single coded basemesh access unit.
[0706] Syntax
[0707] aligned(8) class DMCBaseMeshSample {
[0708] / sample_size size of sample from SampleSizeBox
[0709] for (int i = 0; i < sample_size; ) {
[0710] sample_stream_nal_unit ss_nal_unit; / See the sample stream NAL unit description above.
[0711] i += ss_nal_unit.ssnu_nal_unit_size; / You can create a nal unit of the base mesh sample by increasing the count i by the sample stream nal unit size.
[0712] }
[0713] Semantics:
[0714] A sample stream NAL unit (ss_nal_unit) contains a single NAL unit (bmesh_nal_unit) within the NAL unit sample stream format as defined in ISO / IEC 23090-29.
[0715] NAL unit size (ssnu_nal_unit_size) indicates the size in bytes of the sample stream NAL unit.
[0716] Displacement track sample entry
[0717] Sample Entry Type: 'dpc1', 'dpcg'
[0718] Container: SampleDescriptionBox
[0719] Mandatory: A 'dpc1' or 'dpcg' sample entry is mandatory
[0720] Quantity: One or more
[0721] If the displacement data of a DMC bitstream is encoded with arithmetic coding, it can be encapsulated in a displacement track. A sample entry of the displacement track can contain DMCConfigurationBox information. If the sample entry type is 'dpc1', all related parameter sets can be included in the setup unit of the DMCConfigurationBox. If the sample entry type is 'dpcg', all related parameter sets can be included in the setup unit (setup_unit) of the DMCConfigurationBox and / or can be included in a sample of the displacement track.
[0722] Syntax:
[0723] aligned(8) class DMCDisplSampleEntry() extends VolumetricVisualSampleEntry (type) {
[0724] / type is 'dpc1' or 'dpcg'
[0725] DMCConfigurationBox config; / as defined in section 4.5.2
[0726] }
[0727] Semantics:
[0728] config is the decoder configuration box mentioned above.
[0729] Displacement Track Sample Format:
[0730] Syntax:
[0731] aligned(8) class DMCDisplSample {
[0732] / sample_size size of sample from SampleSizeBox
[0733] for (int i = 0; i < sample_size; ) {
[0734] sample_stream_nal_unit ss_nal_unit; / See the sample stream NAL unit description above.
[0735] i += ss_nal_unit.ssnu_nal_unit_size; / You can create a nal unit of the base mesh sample by increasing the count i by the sample stream nal unit size.
[0736] }
[0737] }
[0738] Semantics:
[0739] A sample stream NAL unit (ss_nal_unit) contains a single NAL unit (displ_nal_unit) within the NAL unit sample stream format, as defined in ISO / IEC 23090-29.
[0740] The sample stream NAL unit size (ssnu_nal_unit_size) indicates the size in bytes of the sample stream NAL unit.
[0741] Atlas Track
[0742] The V-DMC specification, ISO / IEC 23090-29, is currently being developed and standardized by MPEG regarding the utilization of atlas data that constitutes a dynamic mesh bitstream. If atlas data is used for the same or similar purposes as in the V3C specification, ISO / IEC 23090-5, the file encapsulation method for the atlas data can follow the syntax and semantics of the atlas sample entry and sample format defined in the Carriage of V3C specification, ISO / IEC 23090-10. However, the sample entry type can be newly defined in V-DMC, such as 'dmc1' if the parameter set does not change within the stream and is the same, or 'dmcg' if the parameter set changes within the stream. The receiver can recognize and operate a track with a sample entry type of 'dmc1' or 'dmcg' as an entry point. If an atlas track is an entry point, the config information that can be included in the sample entry can include a V-DMC or V3C parameter set (VPS) and / or a parameter set such as SPS, FPS, etc. related to the atlas bitstream.
[0743] DMC video component track
[0744] Displacement data that constitutes a dynamic mesh bitstream can follow the Carriage of V3C specification, i.e. the file encapsulation method for 2D video of ISOBMFF referenced in ISO / IEC 23090-10, for geometry data and attribute data encoded by a video codec. The receiver can determine information about each data type contained in the corresponding track through the V3CUnitHeaderBox information included in the SchemeInformationBox.
[0745] The syntax and semantics of V3CunitHeaderBox can follow the V3C specification, i.e. ISO / IEC 23090-5, as described above.
[0746] The V3C video component track carries 2D video encoded data of a V3C video component. Storage of the V3C video component track leverages existing features of the ISO Base Media File Format and its derived specifications. For example, ISO / IEC 14496-15 defines a mechanism for carrying V3C video components encoded in ISO / IEC 14496-10 and ISO / IEC 23008-2.
[0747] V3C video component tracks must be represented as constrained video in the file, using the common constrained sample item 'resv' with additional requirements.
[0748] SchemeTypeBox is in RestrictedSchemeInfoBox and scheme_type is set to 'vvvc'
[0749] SchemeInformationBox is in RestrictedSchemeInfoBox and contains V3CUnitHeaderBox.
[0750] In the track header, the track_in_movie flag is set to 0 to indicate that this track should not be displayed alone.
[0751] Through the aforementioned DMCComponentInfoBox signaling, the receiver can determine the data component type included in the corresponding track and information about each data.
[0752] Submesh track sample entry
[0753] Sample Entry Type: 'smc1'
[0754] Container: SampleDescriptionBox
[0755] Mandatory: Yes
[0756] Quantity: One or more
[0757] If the base mesh data consists of one or more submesh data, the submesh data can be encapsulated into submesh tracks, i.e., tracks with a sample entry type of 'smc1'. The sample entry of each submesh track can contain information about the submesh data it contains.
[0758] Syntax
[0759] aligned(8) class DMCSubMeshSampleEntry() extends VolumetricVisualSampleEntry (type) {
[0760] / type is 'smc1'
[0761] DMCSubMeshConfigurationBox submesh_info
[0762] }
[0763] Semantics:
[0764] Submesh_info is information about the sub-mesh included in the track described later (see DMCSubMeshConfigurationBox).
[0765] Submesh track sample format
[0766] Each sample within a submesh track corresponds to a single coded submesh access unit, as defined in ISO / IEC 23090-29.
[0767] Syntax:
[0768] aligned(8) class DMCSubMeshSample {
[0769] / sample_size size of sample from SampleSizeBox
[0770] for (int i = 0; i < sample_size; ) {
[0771] sample_stream_nal_unit ss_nal_unit; The aforementioned sample stream NAL unit
[0772] i += ss_nal_unit.ssnu_nal_unit_size; / You can create a nal unit of the base mesh sample by increasing the count i by the sample stream nal unit size.
[0773] }
[0774] }
[0775] Semantics:
[0776] A NAL unit (ss_nal_unit) contains a single NAL unit (bmesh_nal_unit) within the NAL unit sample stream format as defined in ISO / IEC 23090-29.
[0777] The content for the base mesh sample and the content for the sub mesh sample may be the same.
[0778] Figure 18 shows the tracks of a file according to embodiments.
[0779] Track References
[0780] If the basemesh track is an entry track:
[0781] Among tracks composed of each track type, a basemesh track can be designated as the entry point where the file parser can begin parsing. References between tracks can be made using the TrackReferenceBox of the TrackBox defined in the ISOBMFF specification (ISO / IEC 14496-12) within each track.
[0782] The possible reference relationships between the tracks are as follows:
[0783] Track Reference Method 1
[0784] Figure 18(a) is an example in which a base mesh track references an atlas track, and the atlas track references a geometry track and an attribute track.
[0785] To connect different types of tracks, use the track reference tool from ISO / IEC 14496-12.
[0786] Add a TrackReferenceTypeBox to the TrackReferenceBox within the TrackBox of the basemesh track. The TrackReferenceTypeBox contains an array of track_IDs that specify the tracks referenced by the DMC track. To associate a basemesh track with an atlas track, the reference type (reference_type) of the basemesh track's TrackReferenceTypeBox identifies the associated atlas track, and the atlas track identifies the associated DMC track. The 4CCs for these track reference types are as follows:
[0787] 'bmct': the referenced atlas track(s).
[0788] 'atcg': the referenced geometry track(s).
[0789] 'atca': the referenced attribute track(s).
[0790] Track Reference Method 2
[0791] Figure 18(b) is an example in which a base mesh track references an atlas track, a geometry track, and an attribute track.
[0792] To associate a basemesh track with each track, the reference_type of the basemesh track's TrackReferenceTypeBox identifies the associated DMC track. The 4CCs for these track reference types are as follows:
[0793] 'bmct': referenced atlas track(s)
[0794] 'bmcg': Referenced geometry track(s)
[0795] 'bmca': Referenced attribute track(s)
[0796] Figure 19 shows the tracks of a file according to embodiments.
[0797] If the Atlas track is an entry track:
[0798] Among tracks composed of each track type, an atlas track can be designated as the entry point where the file parser can begin parsing. References between tracks can be made using the TrackReferenceBox of the TrackBox defined in the ISOBMFF specification (ISO / IEC 14496-12) within each track.
[0799] The reference relationship between possible tracks according to this is as shown in Fig. 19.
[0800] To connect different types of tracks, use the track reference tool from ISO / IEC 14496-12.
[0801] Add a TrackReferenceTypeBox to the TrackReferenceBox within the TrackBox of the basemesh track. The TrackReferenceTypeBox contains an array of track_IDs that specify the tracks referenced by the V-DMC track. To associate a basemesh track with an atlas track, the reference_type of the basemesh track's TrackReferenceTypeBox identifies the associated atlas track, and the atlas track identifies the associated V-DMC track. The 4CCs for these track reference types are as follows:
[0802] 'atcb': Referenced basemesh track(s)
[0803] 'atcg': referenced geometry track(s)
[0804] 'atca': Referenced attribute track(s)
[0805] An encoding method according to embodiments may generate a file based on grouping of tracks within the file. A decoding method according to embodiments may receive a file and decapsulate tracks based on the track grouping.
[0806] Track grouping:
[0807] Submesh track group:
[0808] A submesh track group is a way to indicate that each identical submesh track represents the same basemesh when the file is encapsulated into one or more submesh tracks.
[0809] definition:
[0810] Box Types: 'sutg'
[0811] Container: TrackGroupBox
[0812] Mandatory: No
[0813] Quantity: Zero or more
[0814] Submesh track groups are defined using a SubmeshTrackGroupBox, a track group type that extends the TrackGroupTypeBox defined in ISO / IEC 14496-12. A SubmeshTrackGroupBox indicates that a track belongs to a set of tracks that constitute a submesh group. For each submesh group a track belongs to, there is a corresponding instance of a SubmeshTrackGroupBox with a unique track group ID (track_group_id) for that playout group in the TrackGroupBox for that track.
[0815] Syntax:
[0816] aligned(8) class SubmeshTrackGroupBox extends TrackGroupTypeBox('sutg') {
[0817] / Track group ID (track_group_id) can be inherited from Track Group Type Box (TrackGroupTypeBox)}
[0818] Submesh data encapsulation scheme:
[0819] The encoding method according to the embodiments can encapsulate submesh data of mesh data within a file. The decoding method according to the embodiments can receive a file and decode submesh data.
[0820] How to signal by defining a submesh configuration box (DMCSubMeshConfigurationBox):
[0821] Submesh related information included in a V-DMC bitstream can be signaled in the sample entries of the submesh tracks when the V-DMC tracks are encapsulated into multiple tracks.
[0822] Syntax:
[0823] aligned(8) class DMCSubMeshConfigurationBox {
[0824] unsigned int(8) num_of_submeshes;
[0825] for (i=0; i < num_of_submeshes; i++) {
[0826] unsigned int(8) submesh_id;
[0827] }
[0828] / Additional fields
[0829] }
[0830] In addition, when the atlas data corresponding to the submesh, i.e., the mapping information with information such as patches and / or atlas tiles, is signaled together, the receiver (decoder) can acquire the atlas tile data unit associated with one or more submeshes. In addition, since the atlas data information corresponding to one or more submeshes can include position and size information in the geometry and attribute frames encoded by the video codec, the receiver can perform partial access and decoding based on the submesh and atlas tile. Therefore, the information for mapping the submesh and atlas tile data can be signaled by defining it in the submesh configuration box (DMCSubMeshConfigurationBox) as follows.
[0831] aligned(8) class DMCSubMeshConfigurationBox {
[0832] unsigned int(8) num_of_submeshes;
[0833] for (i=0; i < num_of_submeshes; i++) {
[0834] unsigned int(8) submesh_id;
[0835] unsigned int(8) num_of_tiles;
[0836] for (j=0; j < num_of_tiles; j++) {
[0837] unsigned int(16) tile_id;
[0838] }
[0839] }
[0840] / Additional fields
[0841] }
[0842] In some embodiments, sub-meshes may be included within a tile.
[0843] Embodiments may signal using the atlas tile track number containing the tiles, rather than directly specifying the tile ID.
[0844] aligned(8) class DMCSubMeshConfigurationBox {
[0845] unsigned int(8) num_of_submeshes;
[0846] for (i=0; i < num_of_submeshes; i++) {
[0847] unsigned int(8) submesh_id;
[0848] unsigned int(8) num_of_tile_tracks;
[0849] for (j=0; j < num_of_tile_tracks; j++) {
[0850] unsigned int(16) tile_track_id;
[0851] }}
[0852] / Additional fields
[0853] }
[0854] Semantics:
[0855] Number of submeshes (num_of_submeshes): Indicates the number of submeshes included in the bitstream. The num_of_submeshes value can be mapped to the number of submeshes (bmsi_num_submeshes_minus1) value in the base mesh submesh information (bmesh_sub_mesh_information()).
[0856] Submesh ID (submesh_id): Identifies each submesh included in the bitstream. The submesh_id value can be mapped to the submesh ID (bmsi_submesh_id) value in the bmesh_sub_mesh_information() information.
[0857] num_of_tiles: Indicates the number of atlas tiles associated with this track.
[0858] Tile ID (tile_id): Indicates the atlas tile ID of the tile associated with this track.
[0859] num_of_tile_tracks: Indicates the number of atlas tile tracks associated with this track.
[0860] Tile Track ID (tile_track_id): Indicates the atlas tile track associated with this track.
[0861] Fig. 20 shows a base mesh track according to embodiments.
[0862] The encoding method according to the embodiments may define the basemesh track of the file as an entry point. The decoding method according to the embodiments may receive the file and first decapsulate the basemesh track of the file as the entry point of the file.
[0863] As described above, a basemesh track can serve as an entry point for a content file. As described above, the sample entry type of a basemesh track including submesh tracks can be defined as 'bmcb'. In this case, the role of the basemesh track can express its relationship with the submesh tracks using the track reference type 'bmcs'. The values of the sample entry types according to embodiments, such as 'bmc1', 'bmcg', and 'bmcb', can be expressed according to, for example, the 4CC (Four Character code) of MP4RA. Depending on the type of the file format, it can be used to identify the type of a video or audio stream, etc. The notations 'bmc1', 'bmcg', and 'bmcb' according to embodiments can be replaced with other name formats.
[0864] For example, a file may include a basemesh track, a submesh track, an atlas track, a geometry track, and an attribute track, as shown in FIG. 20. The basemesh track may include basemesh data, and the submesh track may include submesh data of the basemesh. The atlas track may include atlas data, the geometry track may include geometry data encoded based on a video codec and / or an arithmetic encoding method, and the attribute track may include attribute data encoded based on a video codec. When the basemesh track is defined as the entry point of the file, the reference relationship between the basemesh track and other tracks may be defined, as shown in FIG. 20. The decoder may parse the basemesh track and decode submeshes, atlases, geometry, attributes, etc. related to the basemesh using each reference relationship.
[0865] Figure 21 shows an atlas track according to embodiments.
[0866] The encoding method according to the embodiments may define an atlas track of a file as an entry point. The decoding method according to the embodiments may receive a file and first decapsulate the atlas track of the file as an entry point of the file.
[0867] As described above, an atlas track can serve as an entry point of a content file. The sample entry type of an atlas track including atlas tile tracks is defined as 'v3cb', and this atlas track can reference a basemesh track with a sample entry type of 'bmcb' using a track reference type of 'atcb'. At this time, the role of the basemesh track can express the relationship with the submesh tracks using the track reference type of 'bmcs'. Even when the atlas track is an entry point, it can indicate the association information between the submesh and the atlas tile signaled by the DMCSubMeshConfigurationBox of the submesh track. The values of the sample entry type according to embodiments, such as 'bmc1', 'bmcg', and 'bmcb', can be expressed according to, for example, the 4CC (Four Character code) of MP4RA. Depending on the type of the file format, it can be used to identify the type of video or audio stream, etc. The notations 'bmc1', 'bmcg', and 'bmcb' according to the embodiments may be replaced with other name formats.
[0868] For example, a file may include an atlas track, a basemesh track, an atlas tile track for an atlas tile included in the atlas, a submesh track for a submesh of the basemesh, a geometry track, and an attribute track, as shown in FIG. 20. Since an atlas includes one or more atlas tiles, an atlas track and one or more atlas tile tracks may be referenced, as shown in FIG. 20. For example, if an atlas has four tiles, it may include an atlas tile track 2 carrying data for tile 0, an atlas tile track 3 carrying data for tile 1, and an atlas tile track 4 carrying data for tiles 2 and 3.
[0869] A basemesh track can reference submesh tracks that carry data about the basemesh's submeshes. For example, there might be a submesh track 6 that carries data about submesh 0, a submesh track 7 that carries data about submesh 1 and submesh 2, a submesh track 8 that carries data about submesh 3, and so on.
[0870] The geometry track may include geometry data encoded based on a video codec and / or an arithmetic encoding method, and the attribute track may include attribute data encoded based on a video codec. When the basemesh track is defined as an entry point of a file, the reference relationship between the atlas track and other tracks may be defined as in Fig. 21. The decoder may parse the atlas track and decode the atlas tile track, basemesh track, submesh track, geometry track, attribute track, etc. using each reference relationship.
[0871] The encoding method according to the embodiments may signal subsamples within a file.
[0872] In order to use the SubSampleInformationBox in a V-DMC bitstream, subsamples are defined based on the value of the flag field of the SubSampleInformationBox. The flag field can indicate the type of subsample information provided in this box as follows.
[0873] If the flag is 0: Indicates a submesh-based subsample. A subsample contains one or more contiguous submesh data units corresponding to one V-DMC submesh.
[0874] Other values for the flag may be reserved.
[0875] The subsample_priority field can be set to a value according to the specification for this field in ISO / IEC 14496-12 [ISOBMFF].
[0876] The codec-specific parameters of the SubsampleInformationBox can be defined as follows:
[0877] if (flags == 0) {
[0878] unsigned int(1) submesh_data_present;
[0879] bit(7) reserved = 0;
[0880] if (submesh_data_present)
[0881] unsigned int(24) submesh_id;
[0882] else
[0883] bit(24) reserved = 0;
[0884] }
[0885] submesh_data_present: If this value is 1, it indicates that the subsample contains a submesh data unit.
[0886] Submesh ID (submesh_id): Indicates the identifier of each submesh included in the bitstream. The submesh_id value can be mapped to the bmsi_submesh_id value in the bmesh_sub_mesh_information() information.
[0887] Fig. 22 shows a receiving device according to embodiments.
[0888] As shown in Fig. 22, the receiving device (decoder) can decode mesh data by parsing information about the file, the track within the file, and the decoder configuration within the track.
[0889] File Receiver: The receiver (decoder) receives dynamic mesh content in file form and can parse one or more tracks contained within the file.
[0890] For example, when dynamic mesh content is composed of submesh tracks and atlas tile tracks, as shown in Figures 20 and 21 in the file, each submesh track and atlas tile track can be related as follows.
[0891] Submesh track 2 (submesh id 0) and Atlas tile track 6 (tile 0): Submesh track 2, which carries data about the submesh identified by submesh id 0, and Atlas tile track 6, which carries data about the atlas tile for tile 0, can be associated with each other.
[0892] Submesh track 3 (submesh id 1 and 2) and Atlas tile track 7 (tile 1): Submesh track 3, which carries data about submeshes identified by submesh id 1 and submesh id 2, and Atlas tile track 7, which carries data about the atlas tile for tile 1, can be associated with each other.
[0893] Submesh track 4 (submesh id 3) and Atlas tile track 8 (tile 2 and 3): Submesh track 4, which carries data about the submesh identified by submesh id 3, and Atlas tile track 8, which carries data about the atlas tiles for tile 2 and tile 3, can be associated with each other.
[0894] In the above example, the submesh configuration box (DMCSubMeshConfigurationBox) can have the following values:
[0895] There are four submeshes, and their submesh IDs are 0, (1, 2), and 3. The atlas tile IDs associated with each submesh are 0, 1, and (2, 3), respectively. Parentheses are used to distinguish that they are composed of the same track. The receiver (decoder) can only parse, decode, and render data corresponding to submesh IDs 1, 2, and 3.
[0896] File Parser: The receiver's file parser can parse the sample entries of one or more tracks within a dynamic mesh content file to find submesh tracks with a sample entry type of 'smc1'.
[0897] The file parser can find out information such as the number and ID of submeshes contained in the track by parsing the submesh configuration box (DMCSubMeshConfigurationBox) contained in the sample entry of each track.
[0898] The receiver (decoder) can know information such as the number of atlas tiles associated with each submesh, the atlas tile ID, and / or the number of atlas tiles and the atlas tile track ID. As mentioned above, it can be known that submesh IDs 1, 2, and 3 and their associated atlas tile IDs are 1, 2, and 3, and / or the atlas tile ID is 7, 8. Therefore, the receiver can only parse the corresponding atlas tile tracks.
[0899] Bitstream Packager: The receiver's bitstream packager can collect data for each submesh parsed through file parsing and form a decodable dynamic mesh basemesh bitstream.
[0900] As a result, the receiver can partially configure the bitstream form, excluding the submesh data and atlas tile data contained in Tracks 2 and 6.
[0901] Decoder (V-DMC Decoder): Can decode a dynamic mesh bitstream configured through the above bitstream packager process.
[0902] Embodiments include a multiple submesh track encapsulation of a dynamic mesh coding bitstream.
[0903] Embodiments include a partial access signaling scheme for a Video-based Dynamic Mesh Coding (VDMC) bitstream.
[0904] Embodiments include a method for signaling scene object and submesh connection information of a Video-based Dynamic Mesh Coding (VDMC) bitstream.
[0905] Embodiments include a multiple submesh signaling scheme for a Video-based Dynamic Mesh Coding (VDMC) bitstream.
[0906] The embodiments include a method of signaling relationship information between a scene object and a submesh in spatial region information when basemesh data of a V-DMC bitstream is composed of submeshes and scene object information within the content is signaled.
[0907] The embodiments include a signaling scheme for dynamically changing spatial region information using a sample grouping method when base mesh data of a V-DMC bitstream is composed of sub-meshes and scene object information is signaled together with spatial region information.
[0908] Embodiments include a signaling scheme for a connection relationship between a basemesh track and a submesh track according to a sample entry type when a V-DMC bitstream includes one or more submeshes and a basemesh bitstream and a submesh bitstream are stored as separate tracks.
[0909] V-DMC can be partially decoded at the submesh level at the bitstream level. However, submeshes at the bitstream level can only be decoded independently and cannot become meaningful 3D objects. Therefore, spatial region signaling is configured to enable partial access at the file level for meaningful spatial region units, and mapping signaling with one or more submeshes associated with each spatial region is included.
[0910] For example, it includes a signaling scheme for spatial regions for partial access, a signaling scheme for mapping information between spatial regions and sub-meshes, a signaling scheme for scene object and sub-mesh configuration information, and a sample grouping signaling scheme for dynamically changing spatial region configuration information.
[0911] When a base mesh bitstream is composed of one or more sub-mesh bitstreams, the encoding method according to the embodiments can define and signal the connection relationship and constraints between the base mesh track and the sub-mesh track according to the sample entry type of each track when encapsulating mesh data in a multiple track file format.
[0912] The encoding method according to the embodiments can generate submesh SOI relationship indication information and transmit it as SEI information (Submesh SOI relationship indication SEI payload) in a bitstream.
[0913] Syntax of the Submesh SOI relationship indication SEI payload:
[0914] submesh_soi_relationship_indication( payloadSize ) {Descriptorssr_persistence_association_flagu(1)ssr_number_of_active_scene_objectsue(v)ssr_submesh_id_length_minus1ue(v)for( i = 0; i < ssr_number_of_active_scene_object; i++ ) {ssr_soi_object_idx[ i ]u(v)ssr_number_of_submesh_included[ i ]ue(v)for( j = 0; j < ssr_number_of_submesh_included[ i ]; j++ ) { ssr_submesh_id[ i ][ j ]u(v)ssr_completely_included[ i ][ j ]u(v)}}}TK1926
[0915] Semantics:
[0916] The submesh SOI relationship indication SEI message according to embodiments indicates the relationship between a scene object and a submesh. One or more submeshes can be connected to each scene object, and a submesh can be connected to two or more scene objects.
[0917] Relationship Persistence Flag (ssr_persistence_association_flag): Indicates that the relationship between a scene object and a submesh is persistent. If the value of this flag is '0', the relationship is valid only for the current frame.
[0918] Number of active scene objects (ssr_number_of_active_scene_object): Indicates the number of active scene objects defined for the mesh when the SEI message is signaled.
[0919] Submesh ID length (ssr_submesh_id_length_minus1): Adding 1 to this value indicates the number of bits used to represent the syntax element ssr_submesh_id[i][j].
[0920] Scene object index (ssr_soi_object_idx[i]): Indicates the index of the ith scene object (object) to be connected to the submesh.
[0921] Number of sub-meshes in the scene object (ssr_number_of_submesh_included[i]: This indicates the number of sub-meshes completely included in the i-th scene object.
[0922] Submesh ID (ssr_submesh_id[i][j]): Indicates the identifier of the jth submesh contained in the ith scene object. The number of bits used to represent ssr_submesh_id[i][j] is ssr_submesh_id_length_minus1 + 1.
[0923] Whether the submesh is included in the scene object (ssr_completely_included[i][j]: Indicates that the j-th submesh is completely included in the i-th scene object.
[0924] The encoding method according to the embodiments may encode mesh data and encapsulate a bitstream including the encoded mesh data and a parameter set into a file. The file may consist of a single track or multiple tracks.
[0925] At least one track within a file according to embodiments may include the following common data structure:
[0926] Basemesh decoder configuration record:
[0927] The BaseMeshDecoderConfiguration record contains decoder configuration information for mesh-based point cloud content. This record contains a version field. The specification for this version defines version 1 of this record. Incompatible changes to the record are indicated by a change in the version number.
[0928] 신택스:
[0929] aligned(8) class BaseMeshDecoderConfigurationRecord {
[0930] unsigned int(3) unit_size_precision_bytes_minus1;
[0931] bit (5) reserved = 0;
[0932] unsigned int(8) num_of_setup_unit_arrays;
[0933] for (int i=0; i < num_of_setup_unit_arrays; i++) {
[0934] unsigned int(1) array_completeness;
[0935] bit(1) reserved = 0;
[0936] unsigned int(6) nal_unit_type;
[0937] unsigned int(8) num_nal_units;
[0938] for (int j=0; j < num_nal_units; j++) {
[0939] unsigned int(16) setup_unit_length;
[0940] bmesh_nal_unit setup_unit(setup_unit_length);
[0941] }
[0942] }
[0943] / additional fields
[0944] }
[0945] 시맨틱스:
[0946] Unit size precision (unit_size_precision_bytes_minus1): Adding 1 to this value indicates the precision (in bytes) of the NAL units of the sample stream to which this configuration record applies.
[0947] Number of setup unit arrays (num_of_setup_unit_arrays): Indicates the number of next basemesh NAL unit arrays of the specified type, greater than or equal to 1.
[0948] Array completeness: A value of 1 indicates that all basemesh NAL units of the specified type are present in the following array and none are in the stream. A value of 0 indicates that additional basemesh NAL units of the specified type may be present in the stream. The default and allowed values are constrained by the sample entry name.
[0949] nal_unit_type: Indicates the basemesh NAL unit type of the following array. It can take the values defined in ISO / IEC 23090-29. It can take one of the values representing the BNAL_BMSPS, BNAL_BMFPS, BNAL_PREFIX_ESEI, BNAL_PREFIX_NSEI, BNAL_SUFFIX_ESEI, or BNAL_SUFFIX_NSEI basemesh NAL unit.
[0950] Number of NAL Units (num_nal_units): Indicates the number of basemesh NAL units of the type indicated by nal_unit_type in the following array.
[0951] setup_unit_length: Indicates the size (in bytes) of the setup_unit field. The length field includes the sizes of both the basemesh NAL unit header and the basemesh NAL unit payload, but does not include the length field itself.
[0952] setup_unit contains basemesh NAL units according to their associated nal_unit_type.
[0953] Displacement decoder configuration record
[0954] Syntax:
[0955] aligned(8) class DisplacementDecoderConfigurationRecord {
[0956] unsigned int(3) unit_size_precision_bytes_minus1;
[0957] bit (5) reserved = 0;
[0958] unsigned int(8) num_of_setup_unit_arrays;
[0959] for (int i=0; i < num_of_setup_unit_arrays; i++) {
[0960] unsigned int(1) array_completeness;
[0961] bit(1) reserved = 0;
[0962] unsigned int(6) nal_unit_type;
[0963] unsigned int(8) num_nal_units;
[0964] for (int j=0; j < num_nal_units; j++) {
[0965] unsigned int(16) setup_unit_length;
[0966] bmesh_nal_unit setup_unit(setup_unit_length);
[0967] }
[0968] }
[0969] / additional fields
[0970] }
[0971] Semantics:
[0972] Unit Size Precision (unit_size_precision_bytes_minus1): Adding 1 to this value indicates the precision (in bytes) of the NAL units of the sample stream to which this configuration record applies.
[0973] Number of setup unit arrays (num_of_setup_unit_arrays): Indicates the number of next displacement NAL unit arrays of the specified type, which is greater than or equal to 1.
[0974] Array completeness: A value of 1 indicates that all displacement NAL units of the specified type are in the following array and not in the stream. A value of 0 indicates that additional displacement NAL units of the specified type may be in the stream. The default and allowed values are limited to sample entry names.
[0975] nal_unit_type: Indicates the displacement NAL unit type of the following array. Uses the values defined for the nal units described above. One of the values representing displacement NAL units can be DNAL_DSPS, DNAL_DFPS, DNAL_PREFIX_ESEI, DNAL_PREFIX_NSEI, DNAL_SUFFIX_ESEI, or DNAL_SUFFIX_NSEI.
[0976] Number of NAL units (num_nal_units): Indicates the number of displacement NAL units of the type indicated by nal_unit_type in the following array.
[0977] setup_unit_length: Indicates the size (in bytes) of the setup_unit field. The length field includes the sizes of both the displacement NAL unit header and the displacement NAL unit payload, but does not include the length field itself.
[0978] Setup unit (setup_unit): Contains displacement NAL units according to their associated nal_unit_type.
[0979] DMC decoder configuration box:
[0980] The DMC Decoder Configuration box contains either a BaseMeshDecoderConfigurationRecord or a DisplacementDecoderConfigurationRecord, depending on the sample entry in the track.
[0981] In this document, the version may be 0.
[0982] Syntax:
[0983] class BaseMeshConfigurationBox extends FullBox('bmcC', version = 0, 0) {
[0984] BaseMeshDecoderConfigurationRecord();
[0985] }
[0986] class DisplacementConfigurationBox extends FullBox('dmcC', version = 0, 0) {
[0987] DisplacementDecoderConfigurationRecord();
[0988] }
[0989] Semantics:
[0990] The syntax of the BaseMeshDecoderConfigurationRecord and DisplacementDecoderConfigurationRecord follows the description described above.
[0991] A sample entry of a basemesh track of a file according to embodiments may be defined as follows.
[0992] Basemesh track sample entry:
[0993] Sample Entry Type: 'bmcb', 'bmc1' or 'bmcg'
[0994] Container: SampleDescriptionBox
[0995] Mandatory: A 'bmcb','bmc1' or 'bmcg' sample entry is mandatory
[0996] Quantity: One or more
[0997] A BaseMesh track uses a BaseMeshSampleEntry, which extends a VolumetricVisualSampleEntry with a sample entry type of 'bmcb', 'bmc1', or 'bmcg'. The following restrictions can be set for a BaseMesh track:
[0998] If the basemesh bitstream contains a single submesh, a basemesh track with sample entry type 'bmc1' or 'bmcg' can be used.
[0999] When a basemesh bitstream contains multiple submeshes, a basemesh track with a sample entry type of 'bmcb' may be used, and each submesh bitstream may be stored as a separate submesh track with a sample entry type of 'smc1'. The values of the sample entry types according to embodiments, such as 'bmc1', 'bmcg', and 'bmcb', may be expressed according to, for example, the 4CC (Four Character code) of MP4RA. Depending on the type of the file format, it may be used to identify the type of a video or audio stream, etc. The notations 'bmc1', 'bmcg', and 'bmcb' according to embodiments may be replaced with other name formats.
[1000] In the 'bmcb' sample entry, the basemesh track does not contain a BMCL NAL unit.
[1001] In the 'bmc1' sample entry, all basemesh parameter sets are stored in the setup_unit array. SEI messages that apply to the entire stream can also be stored in the setup_unit array.
[1002] In the 'bmcg' sample entry, parameter sets and SEI messages can exist in the setup_unit array or in the samples of the basemesh track.
[1003] Syntax:
[1004] aligned(8) class DMCBaseMeshSampleEntry() extends VolumetricVisualSampleEntry (type) {
[1005] / type is 'bmcb', 'bmc1' or 'bmcg'
[1006] BaseMeshConfigurationBox config;
[1007] }
[1008] Semantics:
[1009] config is the decoder configuration box mentioned above.
[1010] If the basemesh track is an entry point, the config information may include a V-DMC or V3C parameter set (VPS) and / or parameter sets such as SPS, FPS, etc. associated with the basemesh bitstream.
[1011] The aforementioned displacement track sample entry can additionally convey displacement configuration box information as follows.
[1012] Displacement track sample entry
[1013] Sample Entry Type: 'dpc1', 'dpcg'
[1014] Container: SampleDescriptionBox
[1015] Mandatory: A 'dpc1' or 'dpcg' sample entry is mandatory
[1016] Quantity: One or more
[1017] Syntax:
[1018] aligned(8) class DMCDisplSampleEntry() extends VolumetricVisualSampleEntry (type) {
[1019] / type is 'dpc1' or 'dpcg'
[1020] DisplacementConfigurationBox config; / as defined in section 4.6.3
[1021] }
[1022] Semantics:
[1023] config is the decoder configuration box.
[1024] The encoding method according to the embodiments can generate a submesh track within a file. A sample entry of the submesh track can be defined as follows.
[1025] Submesh track sample entry:
[1026] Sample Entry Type: 'smc1'
[1027] Container: SampleDescriptionBox
[1028] Mandatory: Yes
[1029] Quantity: One or more
[1030] A basemesh may contain one or more submeshes, each contained in a separate track, a submesh track, with sample entry type 'smc1'. A submesh track sample contains only BMCL NAL units belonging to the same basemesh. The track reference type 'bmcs' is used to indicate the relationship from a basemesh track to its associated submesh track.
[1031] Syntax:
[1032] aligned(8) class DMCSubMeshSampleEntry() extends VolumetricVisualSampleEntry (type) {
[1033] / type is 'smc1'
[1034] DMCSubMeshConfigurationBox submesh_info
[1035] }
[1036] Semantics:
[1037] Submesh_info is information about the submesh included in the track.
[1038] The encoding method according to the embodiments can generate signaling information for partial access to mesh data and transmit it by including it in a file structure. The decoding method according to the embodiments can partially decode mesh data based on partial access-related information within the file.
[1039] Partial Access Signaling Scheme:
[1040] Submesh information structure:
[1041] Syntax:
[1042] aligned(8) class SubmeshInfoStruct() {
[1043] unsigned int(16) num_submeshes;
[1044] unsigned int(1) tile_mapping_present_flag;
[1045] unsigned int(7) reserved;
[1046] for (i=0; i < num_submeshes; i++) {
[1047] unsigned int(16) submesh_id;
[1048] if(tile_mapping_present_flag) {
[1049] TileMappingInfoStruct();
[1050] }
[1051] }
[1052] }
[1053] Semantics:
[1054] Number of submeshes (num_submeshes): Indicates the number of submeshes signaled in this information structure.
[1055] Tile Mapping Present Flag (tile_mapping_present_flag): Indicates the presence of dimension information for each submesh.
[1056] Submesh ID (submesh_id): This is the identifier of the submesh.
[1057] The tile mapping information structure (TileMappingInfoStruct()) provides information about the V-DMC atlas tiles associated with a submesh.
[1058] Tile mapping information structure:
[1059] Syntax:
[1060] aligned(8) class TileMappingInfoStruct() {
[1061] unsigned int(16) num_tiles;
[1062] for (j=0; j < num_tiles; j++) {
[1063] unsigned int(16) tile_id;
[1064] }
[1065] }
[1066] Semantics:
[1067] Number of tiles (num_tiles): Indicates the number of V-DMC atlas tiles signaled in this information structure.
[1068] Tile ID (tile_id): Identifier of the V-DMC atlas tile being signaled.
[1069] VDMC Object Information Box:
[1070] Syntax:
[1071] aligned(8) class VDMCObjectInfo() {
[1072] unsigned int(16) object_idx;
[1073] unsigned inf(16) num_of_submeshes;
[1074] for(i=0; i < num_of_submeshes; i++) {
[1075] unsigned int(15) submesh_id;
[1076] unsigned int(1) sumesh_completely_included_flag;
[1077] }
[1078] / additional fields
[1079] }
[1080] aligned(8) class ObjectInformationBox() {
[1081] unsigned int(16) num_of_objects;
[1082] for (j=0; j < num_of_objects; j++) {
[1083] VDMCObjectInfo object_info;
[1084] }
[1085] }
[1086] Semantics:
[1087] Object Index (object_idx): Indicates the value of the object index as defined in the scene object information SEI message of the V3C or VDMC specification.
[1088] Number of submeshes (num_of_submeshes): Indicates the number of submeshes contained in the scene object.
[1089] Submesh ID (submesh_id): Indicates the identifier of the submesh included in the scene object.
[1090] Submesh_completely_included_flag: Indicates that the submesh is completely included in the scene object.
[1091] num_of_objects: Indicates the number of scene objects in the object information box.
[1092] Object information (object_info): Contains information related to scene objects.
[1093] Spatial region information structure:
[1094] Syntax:
[1095] aligned(8) class VDMCSpatialRegionStruct() {
[1096] unsigned int(32) size;
[1097] unsigned int(16) region_id;
[1098] unsigned int(1) bounding_box_present_flag;
[1099] unsigned int(1) dimensions_included_flag;
[1100] unsigned int(1) submesh_info_present_flag;
[1101] unsigned int(5) reserved;
[1102] if(bounding_box_present_flag) {
[1103] VDMCBoundingBox(dimensions_included_flag);
[1104] }
[1105] if(submesh_info_present_flag) {
[1106] SubmeshInfoStruct();
[1107] }
[1108] }
[1109] Semantics:
[1110] size: An integer value specifying the number of bytes in this element, including all fields and contained elements.
[1111] Region ID (region_id): An identifier for a 3D spatial region.
[1112] Bounding box presence flag (bounding_box_present_flag): Indicates whether there is 3D bounding box information for the signaled area.
[1113] dimensions_included_flag: Indicates whether the signaled spatial domain has dimension information.
[1114] Submesh information presence flag (submesh_info_present_flag): Indicates whether submesh information exists in the signaled spatial region.
[1115] The VDMC bounding box (VDMCBoundingBox) information can basically follow the information defined in the V3C carriage (ISO / IEC 23090-10) or G-PCC carriage (ISO / IEC 23090-18).
[1116] The submesh information structure (SubmeshInfoStruct()) information may be information related to the submesh as described above.
[1117] As follows, the mapping information between scene objects and submeshes defined in the VDMC codec can be signaled by adding it to spatial region information. Scene object information refers to additional information such as 3D bounding boxes, labels, priorities, etc. related to one or more 3D objects included in a frame of VDMC content, and can be signaled by being distinguished as the SEI (Supplemental Enhancement Information) type of the SetupUnit in the aforementioned decoder configuration record (DMCDecoderConfigurationRecord).
[1118] There is a submesh information structure (SubmeshInfoStruct()) that can be included in a 3D spatial region defined at the file system level, but since this can be used to signal the relationship between one or more submeshes and one or more V-DMC atlas tiles, as described above, object information (VDMCObjectInfo()) can be added to signal the relationship between one or more scene objects and one or more submeshes included in the 3D spatial region.
[1119] Syntax:
[1120] aligned(8) class VDMCSpatialRegionStruct() {
[1121] unsigned int(32) size;
[1122] unsigned int(16) region_id;
[1123] unsigned int(1) bounding_box_present_flag;
[1124] unsigned int(1) dimensions_included_flag;
[1125] unsigned int(1) submesh_info_present_flag;
[1126] unsigned int(1) object_info_present_flag;
[1127] unsigned int(4) reserved;
[1128] if(bounding_box_present_flag) {
[1129] VDMCBoundingBox(dimensions_included_flag);
[1130] }
[1131] if(submesh_info_present_flag) {
[1132] SubmeshInfoStruct();
[1133] }
[1134] if(object_info_present_flag) {
[1135] VDMCObjectInfo();
[1136] }
[1137] }
[1138] Semantics:
[1139] Object information presence flag (object_info_present_flag): Indicates whether VDMC object information exists in the signaled spatial area.
[1140] VDMCObjectInfo() contains information related to the object as described above.
[1141] Static spatial region signaling
[1142] Box Types: 'vdsr'
[1143] Container: DMCBaseMeshSampleEntry ('bmc1', 'bmcg')
[1144] Mandatory: No
[1145] Quantity: Zero or one
[1146] The Spatial Region Information Box (DMCSpatialRegionInfoBox) contains information about one or more 3D spatial regions of the V-DMC bitstream carried by each track.
[1147] Syntax:
[1148] aligned(8) class DMCSpatialRegionInfoBox extends FullBox('vdsr',0,0){
[1149] unsigned int(16) num_regions;
[1150] for (int i=0; i < num_regions; i++) {
[1151] VDMCSpatialRegionStruct();
[1152] }
[1153] }
[1154] Semantics:
[1155] Number of regions (num_regions): Indicates the number of signaled 3D spatial regions.
[1156] Spatial region structure (VDMCSpatialRegionStruct()): Contains the 3D spatial region information described above.
[1157] Dynamic spatial region signaling:
[1158] This metadata track with sample entry type 'vddr' represents dynamically changing 3D spatial region information corresponding to part or all of the V-DMC bitstream, or the associations between 3D spatial regions and sub-meshes and atlas tiles over time. When a V-DMC track or a basemesh track is associated with a dynamic spatial region temporal metadata track, the 3D spatial region information of the V-DMC bitstream carried by the track, or the associations between 3D spatial regions and sub-meshes and atlas tiles, is considered dynamic.
[1159] Syntax:
[1160] aligned(8) class DynamicDMCSpatialRegionSampleEntry
[1161] extends MetaDataSampleEtnry('vddr',0,0){
[1162] DMCSpatialRegionInfoBox region_info;
[1163] }
[1164] Semantics:
[1165] Region information (region_info): As described above, it represents the initial 3D spatial region information.
[1166] Additionally, signaling for dynamic spatial region information can use the submesh method or sample grouping method described above.
[1167] Sample group:
[1168] definition:
[1169] Group Types: 'sgsr'
[1170] Container: Sample Group Description Box ('sgpd')
[1171] Mandatory: No
[1172] Quantity: Zero or more
[1173] Using grouping_type 'sgsr' in sample grouping indicates the spatial domain information of the samples in the V-DMC basemesh track or V-DMC atlas track.
[1174] Syntax:
[1175] aligned(8) class VDMCSpatialRegionSampleGroupDescriptionEntry() extends VolumetricVisualSampleGroupEntry('sgsr') {
[1176] unsigned int(16) num_regions;
[1177] for (int i=0; i < num_regions; i++) {
[1178] VDMCSpatialRegionStruct();
[1179] }
[1180] }
[1181] Semantics:
[1182] Number of regions (num_regions): Indicates the number of signaled 3D spatial regions.
[1183] Spatial region structure (VDMCSpatialRegionStruct()): Provides 3D spatial region information related to sub-mesh.
[1184] Referring to Figure 22, the decoder can decode submesh tracks and submesh information as follows.
[1185] File Receiver: The receiver (decoder) can receive dynamic mesh content consisting of submesh data in file form, and can parse one or more tracks contained within the file.
[1186] The received dynamic mesh content may include signaling information including information about one or more scene objects.
[1187] File Parser: The receiver's file parser can parse the sample entries of one or more tracks within a dynamic mesh content file to find base mesh tracks with a sample entry type of 'bmc1' or 'bmcg'.
[1188] The receiver can obtain decoder configuration information by parsing the configuration box (DMCConfigurationBox) included in the sample entry of the basemesh track. In addition, if the content supports partial access, the sample entry can include a spatial region information box (DMCSpatialRegionInfoBox). A receiver that supports partial access can obtain information about each spatial region, one or more submeshes included in the spatial region, and / or one or more atlas tiles by parsing this DMCSpatialRegionInfoBox.
[1189] Additionally, by parsing the scene object information included in each spatial region, connection information between each scene object and submeshes can be obtained.
[1190] Figure 23 shows the relationship between atlas tiles and sub-meshes according to embodiments.
[1191] The receiver's bitstream packager may be responsible for collecting submesh-specific and / or atlas tile-specific data for each spatial region through file parsing and configuring it into a decodable dynamic mesh basemesh bitstream.
[1192] Fig. 23 is an embodiment showing the relationship between texture and displacement data related to a base mesh composed of three sub-meshes. For example, if tile 0 includes sub-mesh 0 and sub-mesh 1, and tile 1 includes sub-mesh 2, the decoder can decode displacement (geometry) data for the three sub-meshes associated with the two tiles. In addition, the decoder can decode texture data for the three sub-meshes associated with the two tiles. If one spatial region includes a base mesh including a sub-mesh identified by sub-mesh index 0 (subIndex0) and a sub-mesh identified by sub-mesh index 2 (subIndex2), the receiver can select and decode texture and displacement data corresponding to the sub-meshes.
[1193] A bitstream packager can decode a dynamic mesh bitstream composed of each spatial region. It can provide the submesh information associated with the scene object information contained in each spatial region to the receiver (decoder).
[1194] Figure 24 shows an encoding method according to embodiments.
[1195] The encoding method according to the embodiments may include a step of encoding mesh data (S2400), a step of encapsulating a file including a bitstream including mesh data (S2410), and / or a step of transmitting the file (S2420).
[1196] The step of encoding mesh data (S2400) may include a step of encoding base mesh data of the mesh data; a step of encoding displacement data of the mesh data; and a step of encoding attribute data of the mesh data.
[1197] The step of encapsulating a file (S2410) can generate information within the file and generate a file including a bitstream containing encoded mesh data.
[1198] The file includes a geometry track conveying geometry data about mesh data, an attribute track conveying attribute data about the mesh data, a basemesh track conveying basemesh data about the mesh data, and an atlas track conveying atlas data about the mesh data, wherein the basemesh data includes at least one submesh data, the file further includes an atlas tile track conveying atlas tile information about at least one tile, the file further includes a submesh track conveying data about at least one submesh data of the basemesh data, the atlas tile track is referenced by the atlas track, the submesh track is referenced by the basemesh track, the bitstream includes subsample information, and the subsample information may include identifier information about submesh data present in the bitstream.
[1199] A submesh track that carries submesh data identified by identifier information for the submesh data is associated with an atlas tile track that carries atlas tile information about a tile associated with the submesh data identified by identifier information for the submesh data, and the tile can be identified by the atlas tile identifier.
[1200] The basemesh data includes at least one submesh data, and the at least one submesh data is associated with an atlas tile, and a particular submesh track contained in the file is partially decoded based on a submesh ID, and an atlas tile track identified by an atlas tile ID associated with the submesh data identified by the submesh ID can be decoded.
[1201] The file includes a geometry track that conveys geometry data about mesh data, an attribute track that conveys attribute data about mesh data, a basemesh track that conveys basemesh data about mesh data, and an atlas track that conveys atlas data about mesh data, wherein a sample entry of the basemesh track includes basemesh configuration information, and the basemesh configuration information may include a setup unit that includes information indicating a number related to the setup unit and parameter information about the basemesh data.
[1202] The base mesh configuration information includes at least one of information indicating the type of a unit containing base mesh data, or information indicating the number of units, and a unit related to the information indicating the type of unit may be included in the setup unit.
[1203] The file may include at least one of information indicating the number of units containing displacement data of mesh data, information indicating the type of units containing displacement data, or units containing displacement data related to the type of units.
[1204] An encoding device for performing an encoding method includes a memory; and at least one processor connected to the memory; wherein the at least one processor can be configured to: encode mesh data; and encapsulate a file including a bitstream including the mesh data; and transmit the file.
[1205] Embodiments further include a computer-readable storage medium storing a bitstream generated by the encoding method.
[1206] Embodiments further include a method comprising: obtaining a bitstream for mesh data, the bitstream being generated based on: encoding basemesh data of the mesh data; encoding displacement data of the mesh data; and encoding attribute data of the mesh data; encapsulating the bitstream into a file; and transmitting data including the file.
[1207] Figure 25 shows a decryption method according to embodiments.
[1208] A decryption method according to embodiments may include a step of receiving a file including a bitstream including mesh data (S2500), a step of decapsulating the file (S2510), and / or a step of decoding the mesh data (S2520).
[1209] The file includes a geometry track conveying geometry data about mesh data, an attribute track conveying attribute data about the mesh data, a basemesh track conveying basemesh data about the mesh data, and an atlas track conveying atlas data about the mesh data, wherein the basemesh data includes at least one submesh data, the file further includes an atlas tile track conveying atlas tile information about at least one tile, the file further includes a submesh track conveying data about at least one submesh data of the basemesh data, the atlas tile track is referenced by the atlas track, the submesh track is referenced by the basemesh track, the bitstream includes subsample information, and the subsample information may include identifier information about submesh data present in the bitstream.
[1210] A submesh track that carries submesh data identified by identifier information for the submesh data is associated with an atlas tile track that carries atlas tile information about a tile associated with the submesh data identified by identifier information for the submesh data, and the tile can be identified by the atlas tile identifier.
[1211] The basemesh data includes at least one submesh data, and the at least one submesh data is associated with an atlas tile, and a particular submesh track contained in the file is partially decoded based on a submesh ID, and an atlas tile track identified by an atlas tile ID associated with the submesh data identified by the submesh ID can be decoded.
[1212] The file includes a geometry track that conveys geometry data about mesh data, an attribute track that conveys attribute data about mesh data, a basemesh track that conveys basemesh data about mesh data, and an atlas track that conveys atlas data about mesh data, wherein a sample entry of the basemesh track includes basemesh configuration information, and the basemesh configuration information may include a setup unit that includes information indicating a number related to the setup unit and parameter information about the basemesh data.
[1213] The base mesh configuration information includes at least one of information indicating the type of a unit containing base mesh data, or information indicating the number of units, and a unit related to the information indicating the type of unit may be included in the setup unit.
[1214] The file may include at least one of information indicating the number of units containing displacement data of mesh data, information indicating the type of units containing displacement data, or units containing displacement data related to the type of units.
[1215] The decryption method may be performed by a decryption device (decoder). The decryption device includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: receive a file including a bitstream including mesh data; decapsulate the file; and decode the mesh data.
[1216] The decryption method can efficiently decode mesh data contained in a file based on the aforementioned information contained in the file.
[1217] The method and device according to the embodiments provide the following technical effects.
[1218] The embodiments include a transmitter or receiver for providing mesh content services, and can construct and store a V-DMC bitstream in a file as described above. This provides the effect of effectively multiplexing a V-DMC bitstream.
[1219] Metadata for data processing and rendering within a V-DMC bitstream can be transmitted within a file. Embodiments include a VDMC (Video-based Dynamic Mesh Compression) processing device, a transmitter, a receiver, a mesh player encoder or decoder, and the like, and provide the above-described effects.
[1220] The data representation method within a file according to the embodiments provides the effect of efficiently accessing a V-DMC bitstream. A transmitter or receiver according to the embodiments can efficiently store and transmit a file of a V-DMC bitstream through a storage technique and signaling of the V-DMC bitstream as multiple tracks within a file.
[1221] As described above, if information about basemesh data composed of submeshes is provided at the file level, the receiver (decoder) can efficiently manage available resources based on this information. Furthermore, by encapsulating each submesh data into a submesh track, partial access and decoding can be enabled for each submesh. By signaling the association information with the atlas tile corresponding to each submesh, partial access and decoding of one or more submesh units and / or one or more atlas tile units can be enabled.
[1222] In addition, the signaling method according to the embodiments transmits one or more spatial regions and one or more scene objects and submesh-related linkage information that can be included in the spatial regions, so that the receiver (decoder) can provide a service such as partial access based on spatial regions including scene object information.
[1223] Due to the configuration of files and tracks according to the embodiments, a dynamic mesh bitstream consisting of one or more submeshes can be encapsulated into multiple tracks. Depending on the newly defined sample entry type of the base mesh track of the file, the submesh tracks associated with the base mesh track can be encapsulated and decapsulated, and receivers (decoders) can perform the same parsing and decoding operations according to the constraints of each sample entry type.
[1224] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.
[1225] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[1226] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[1227] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.
[1228] Terms such as "first," "second," etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.
[1229] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[1230] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.
[1231] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.
[1232] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.
[1233] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[1234] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
Claims
1. A step of receiving a file including a bitstream containing mesh data; a step of decapsulating the above file; and A step of decoding the above mesh data; comprising: How to decrypt.
2. In paragraph 1, The file includes a geometry track that transmits geometry data regarding the mesh data, an attribute track that transmits attribute data regarding the mesh data, a basemesh track that transmits basemesh data regarding the mesh data, and an atlas track that transmits atlas data regarding the mesh data. The above base mesh data includes at least one submesh data, The above file further comprises an atlas tile track conveying atlas tile information about at least one tile, The file further includes a submesh track that transmits data regarding at least one submesh data of the basemesh data, The above atlas tile track is referenced by the above atlas track, The above submesh track is referenced by the above basemesh track, The bitstream includes sub-sample information, and the sub-sample information includes identifier information for sub-mesh data existing in the bitstream. How to decrypt.
3. In paragraph 2, A submesh track that carries submesh data identified by identifier information for the submesh data is associated with an atlas tile track that carries atlas tile information about a tile associated with the submesh data identified by the identifier information for the submesh data, wherein the tile is identified by an atlas tile identifier. How to decrypt.
4. In paragraph 2, The base mesh data includes at least one submesh data, and the at least one submesh data is associated with an atlas tile, Based on the submesh ID, a specific submesh track contained in the above file is partially decoded, An atlas tile track identified by an atlas tile ID associated with submesh data identified by the above submesh ID is decoded, How to decrypt.
5. In paragraph 1, The file includes a geometry track that transmits geometry data regarding the mesh data, an attribute track that transmits attribute data regarding the mesh data, a basemesh track that transmits basemesh data regarding the mesh data, and an atlas track that transmits atlas data regarding the mesh data. The sample entry of the above basemesh track contains basemesh configuration information, The above base mesh configuration information includes information indicating the number of setup units and a setup unit including parameter information regarding the base mesh data. How to decrypt.
6. In paragraph 5, The above base mesh configuration information includes at least one of information indicating the type of the unit including the base mesh data or information indicating the number of the units, The unit related to the information indicating the type of the unit is included in the setup unit, How to decrypt.
7. In paragraph 5, The file includes at least one of information indicating the number of units containing displacement data of the mesh data, information indicating the type of the unit containing the displacement data, or a unit containing displacement data related to the type of the unit. How to decrypt.
8. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Receive a file containing a bitstream containing mesh data; Decapsulating the above file; and configured to decode the above mesh data; Decryption device.
9. In paragraph 8, The file includes a geometry track that transmits geometry data regarding the mesh data, an attribute track that transmits attribute data regarding the mesh data, a basemesh track that transmits basemesh data regarding the mesh data, and an atlas track that transmits atlas data regarding the mesh data. The sample entry of the above basemesh track contains basemesh configuration information, The above base mesh configuration information includes information indicating the number of setup units and a setup unit including parameter information regarding the base mesh data. Decryption device.
10. In paragraph 9, The above base mesh configuration information includes at least one of information indicating the type of the unit including the base mesh data or information indicating the number of the units, The unit related to the information indicating the type of the unit is included in the setup unit, Decryption device.
11. Step of encoding mesh data; and A step of encapsulating a file including a bitstream including the above mesh data; and a step of transmitting the above file; comprising; Encoding method.
12. In paragraph 11, The file includes a geometry track that transmits geometry data regarding the mesh data, an attribute track that transmits attribute data regarding the mesh data, a basemesh track that transmits basemesh data regarding the mesh data, and an atlas track that transmits atlas data regarding the mesh data. The sample entry of the above basemesh track contains basemesh configuration information, The above base mesh configuration information includes information indicating the number of setup units and a setup unit including parameter information regarding the base mesh data. Encoding method.
13. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Encode mesh data; and Encapsulating a file containing a bitstream including the above mesh data; and configured to transmit the above file; Encoding device.
14. A computer-readable storage medium storing a bitstream generated by the method according to Article 11.
15. Step of obtaining bitstream for mesh data, The bitstream is generated based on the steps of: encoding base mesh data of the mesh data; encoding displacement data of the mesh data; and encoding attribute data of the mesh data; a step of encapsulating the above bitstream into a file; and A method comprising the step of transmitting data including the above file.