Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
V-Mesh technology addresses the challenges of generating and processing point cloud data by employing efficient encoding and decoding methods, enabling high-quality point cloud services for VR, AR, MR, and autonomous driving.
Patent Information
- Application Number
- PCT/KR2025/004880
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
The sheer number of points in 3D space makes it difficult to generate point cloud data, requiring significant processing power for transmission and reception, and existing methods face challenges in latency and encoding/decoding complexity.
A method for encoding and decoding mesh data using Video-based Dynamic Mesh Compression (V-Mesh) technology, which includes preprocessing, encoding, transmission, and decoding processes to efficiently transmit and receive point clouds, utilizing techniques like intra-frame and inter-frame encoding, displacement encoding, and attribute map transfer.
Enables high-quality point cloud services with reduced latency and improved encoding/decoding efficiency, supporting applications such as VR, AR, MR, and autonomous driving.
Smart Images

Figure KR2025004880_16102025_PF_FP_ABST
Abstract
Description
Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
[0001] The embodiments provide a method for providing Point Cloud content to provide users with various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and autonomous driving services.
[0002] A point cloud is a collection of points in 3D space. The sheer number of points in 3D space makes it difficult to generate point cloud data.
[0003] There is a problem that a lot of processing power is required to transmit and receive point cloud data.
[0004] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device and reception method for efficiently transmitting and receiving point clouds in order to solve the problems described above.
[0005] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device, and reception method for resolving latency and encoding / decoding complexity.
[0006] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0007] In order to achieve the above-described purpose and other advantages, a decoding method according to embodiments may include a step of decapsulating a file including a bitstream including mesh data by a decoder; and a step of decoding the mesh data. An encoding method according to embodiments may include a step of encoding the mesh data by an encoder; and a step of encapsulating a file including a bitstream including mesh data.
[0008] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide a high-quality point cloud service.
[0009] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can achieve various video codec methods.
[0010] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide general-purpose point cloud content such as autonomous driving services.
[0011] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0012] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0013] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0014] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0015] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0016] Figure 5 shows a displacement generation process according to embodiments.
[0017] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0018] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0019] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0020] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0021] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0022] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0023] Fig. 12 shows a V-MES and Fig. 13 shows a point cloud data transmission device according to embodiments.
[0024] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0025] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0026] Figure 15 shows a V-DMC bitstream according to embodiments.
[0027] Figure 16 shows a V3C unit according to embodiments.
[0028] Figure 17 shows the structure of a scene description system according to embodiments.
[0029] Figure 18 shows an example of a file format for a scene description according to embodiments.
[0030] Figure 19 shows a pipeline for V-DMC (Video-based Dynamic Mesh Coding) according to embodiments.
[0031] Figure 20 shows a pipeline for V-DMC according to embodiments.
[0032] Figure 21 shows a pipeline for V-DMC according to embodiments.
[0033] Figure 22 shows a buffer format for V-DMC according to embodiments.
[0034] Figure 23 shows an encoding method according to embodiments.
[0035] Figure 24 shows a decryption method according to embodiments.
[0036] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0037] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0038] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0039] The system of FIG. 1 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a reception unit (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or a dynamic mesh video encoder (hereinafter, referred to as an encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to a receiving device (110) or a dynamic mesh video decoder (hereinafter, decoder) (113).
[0040] The system of Fig. 1 can perform video-based dynamic mesh compression and decompression.
[0041] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.
[0042] First, the method of compressing dynamic mesh data starts with the V-PCC (Video-based point cloud compression) standard technology. Point cloud data is data that contains color information at the vertex coordinates (X, Y, Z). Mesh data refers to data in which connectivity information between vertices is added to this vertex information. When creating content, it can be created in the form of mesh data from the beginning. By adding connectivity information to point cloud data, it can be converted into mesh data and used.
[0043] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.
[0044] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.
[0045] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.
[0046] A mesh compression system may include a transmitting device and a receiving device. The transmitting device can encode mesh video to output a bitstream, which can be delivered to the receiving device via digital storage media or a network in the form of a file or streaming segment. The digital storage media may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, or SSD.
[0047] The transmitting device may roughly include a mesh video acquisition unit, a mesh video encoder, and a transmitting unit. The receiving device may roughly include a receiving unit, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / video / picture / frame encoding device, and the decoder may be referred to as a mesh video / video / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device and the receiving device may further include separate internal or external modules / units / components for a feedback process.
[0048] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by its vertices in 3D space and connection information that describes how those vertices are connected. It can also contain vertex properties such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included as a mesh property. The mapping is typically described as a set of parameter coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement.
[0049] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc. into a mesh data type with the properties described above through a series of processes and generating a video composed of such mesh data. The mesh video may have properties of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video with properties and connection information that change over time can be expressed as a dynamic mesh video.
[0050] A mesh video encoder can encode an input mesh video into one or more video streams. A single video can include multiple frames, and a single frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and the mesh video can be used interchangeably with the mesh image / frame / picture. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. A mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0051] The encapsulation processing unit (file / segment encapsulation module) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from the metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder, or may be configured as a separate component / module. The encapsulation processing unit may encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The encapsulation processing unit may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. According to an embodiment, the encapsulation processing unit may encapsulate mesh video-related metadata itself in a file.
[0052] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit, or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process the same for transmission.
[0053] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and can include an element for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.
[0054] The receiver can receive mesh video data transmitted by a mesh video transmission device. Depending on the transmission channel, the receiver can receive mesh video data via a broadcast network, via broadband, or via digital storage media.
[0055] The receiving processing unit can perform processing on the received mesh video data according to a transmission protocol. The receiving processing unit can be included in the receiving unit, or can be configured as a separate component / module. In response to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmitting processing unit. The receiving processing unit can transfer the acquired mesh video data to the decapsulation processing unit, and transfer the acquired mesh video-related metadata to a metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
[0056] A decapsulation processing unit (file / segment decapsulation module) can decapsulate mesh video data in file format received from a receiving processing unit. The decapsulation processing unit can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder, or may be configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation processing unit may be in the form of a box or track within a file format. If necessary, the decapsulation processing unit may receive metadata required for decapsulation from the metadata processing unit. Mesh video related metadata can be passed to a Mesh video decoder for use in the Mesh video decoding process, or passed to a renderer for use in the Mesh video rendering process.
[0057] A mesh video decoder can receive a bitstream and perform operations corresponding to those of a mesh video encoder to decode video / images. The decoded mesh video can be displayed via a display unit. Users can view all or part of the rendered result via a VR / AR display or a general display.
[0058] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
[0059] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
[0060] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0061] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.
[0062] This document relates to dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and properties that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-view video, and AR / VR.
[0063] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0064] In this document, picture / frame can generally mean a unit representing one video of a specific time period.
[0065] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0066] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0067] The encoding process of Fig. 1 is as follows.
[0068] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. The V-Mesh compression process receives the following data as input and performs compression.
[0069] Input mesh: Contains the 3D coordinates (geometry) of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.
[0070] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the properties (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video (or attribute for short). The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).
[0071] Material Library File: Contains information about the material properties used in a mesh, particularly information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
[0072] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0073] Base mesh: The input mesh is simplified (decimated) through a preprocessing process, and the objects of the input mesh are expressed using the minimum number of vertices determined by the user's standards.
[0074] Displacement: This is displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.
[0075] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.
[0076] Referring to FIGS. 2 to 7, a method for encoding mesh position information (vertex) is described, and referring to FIGS. 6-10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0077] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0078] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The encoder of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. The transmitting device of Fig. 1 may be broadly referred to as an encoder, and the dynamic mesh video encoder of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing (200) and an encoding (201) process as in Fig. 2. The pre-processor of Fig. 2 may be located in front of the encoder of Fig. 2. The pre-processor and the encoder of Fig. 2 may be referred to as a single encoder.
[0079] The preprocessor can receive a static dynamic mesh and / or an attribute map. The preprocessor can generate a base mesh and / or displacement through preprocessing. The preprocessor can receive feedback information from the encoder and generate the base mesh and / or displacement based on the feedback information.
[0080] The encoder can receive a base mesh, displacement mesh, static dynamic mesh, and / or attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.
[0081] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0082] Figure 3 shows the configuration and operation of the pre-processor of Figure 2.
[0083] Fig. 3 shows a process of performing preprocessing on an input mesh. The preprocessing process (200) can be broadly divided into four steps: 1) Group of Frame (GoF) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). The preprocessor (200) can receive an input mesh, generate a displacement and / or base mesh, and transmit the generated displacement and / or base mesh to the encoder (201). The preprocessor (200) can transmit GoF information related to GoF generation to the encoder (201).
[0084] Below, each step of Figure 3 is described.
[0085] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, number of texture coordinates, vertex connection information, and texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame performs intra-frame encoding.
[0086] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.
[0087] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0088] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.
[0089] Fitting subdivision surface: This is the process of performing subdivision on a simplified mesh. The subdivision method can be a user-defined method, such as the mid-edge method. The fitting process ensures that the input mesh and the subdivision mesh are similar to each other.
[0090] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0091] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex midway between the edges between the vertices.
[0092] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decrypted base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents the position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0093] Figure 5 shows a displacement generation process according to embodiments.
[0094] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.
[0095] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a reconstructed base mesh and generate a subdivided reconstructed base mesh. The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and transform a coordinate system of the mesh into a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.
[0096] The method and device for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data (which may be referred to as a point cloud for short) according to the embodiments can refer to data including vertex coordinates and color information. The term "point cloud" includes mesh data, and in this document, point cloud and mesh data can be used interchangeably.
[0097] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).
[0098] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.
[0099] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0100] The encoding process of Fig. 6 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an intra-frame method. The encoder of Fig. 6 may include a preprocessor (200) and / or an encoder (201).
[0101] The preprocessor can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivided mesh. The quantizer can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder can encode the static mesh. The static mesh encoder can generate a bitstream including the encoded base mesh. The static mesh decoder can decode the encoded static mesh. The inverse quantizer can inversely quantize the quantized static mesh. The displacement calculation unit can receive the reconstructed static mesh and generate displacement, which is a position difference, based on the fitted subdivided mesh. The forward linear lifting unit can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. The image packing unit can pack an image based on the quantized lifting coefficients. A video encoder can encode a packed image. A video decoder decodes the encoded video. An image unpacker can unpack a packed image. A dequantizer can inversely quantize an image. An inverse linear lifting unit applies inverse lifting to the image to generate a reconstructed displacement. A mesh restoration unit restores a warped mesh using the reconstructed displacement and the reconstructed base mesh. An attribute transfer unit receives an input mesh and / or an input attribute map, and generates an attribute map based on the reconstructed warped mesh. A push-pull padding unit can pad data in the attribute map based on a push-pull method. A color space transformation unit can transform the space of a color component, which is an attribute. A video encoder can encode an attribute. A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.
[0102] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0103] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a preprocessor (200) and / or an encoder (201).
[0104] Among the encoding operations of Fig. 7, the corresponding configuration of the encoding operation of Fig. 6 refers to the description of Fig. 7. For the inter-frame-based encoding of Fig. 7, the motion encoder can encode motion based on the restored quantized reference base mesh. The base mesh restoration unit can restore the base mesh based on the restored quantized reference base mesh.
[0105] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.
[0106] The encoding method according to the embodiments includes base mesh encoding (intra encoding). When performing intra frame encoding on the current input mesh frame, the base mesh generated in the preprocessing process can be encoded using a static mesh compression technique after going through a quantization process. In the V-Mesh compression method, for example, Draco technology is applied, and the vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh are compressed.
[0107] The encoding method according to the embodiments may include motion field encoding (inter encoding). Inter frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field, may be calculated and this information may be encoded. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may be encoded as a value. Alternatively, the predicted motion field may be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and this predicted motion field The residual motion field, which is the difference between the value and the motion field value of the current vertex, can be encoded. This value can be encoded using entropy coding. The process of encoding the displacement and attribute map, excluding the motion field encoding process of inter-frame encoding, is the same as the structure of the intra-frame encoding method except for the base mesh encoding.
[0108] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0109] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0110] Figures 8-9 show the process of converting the displacement of the encoding process of Figures 6-7 and the process of packing the conversion coefficients, respectively.
[0111] The encoding method according to the embodiments includes displacement encoding.
[0112] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information.
[0113] Fig. 8 shows the process of transforming displacement information using lifting transform in V-Mesh. The transform coefficients generated through the transform process are quantized and then packed into a 2D image as shown in Fig. 9. The transform coefficients are organized into one block for every 256 (=16X16) units, and each block can be packed in z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. The transform coefficients can be packed by sorting them with Morton code within a block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0114] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. The first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices (LoD0) of the base mesh. The first subdivision mesh may be generated by dividing the second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 including the vertices additionally generated from LoD0, and the vertices additionally divided from LoD1. LoD is a level indicating the degree of detail (Level of Detail), and as the level index increases, the distance between vertices becomes closer and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When a vertex is further divided through subdivision, considering the previous vertices v1, v2 and the subdivided vertex v, the mesh can be encoded based on a prediction and / or update method. Instead of still encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated, and the mesh can be encoded using the residual value to reduce the size of the bitstream. The prediction process means the operation of predicting the current vertex v using the previous vertices v1, v2. Since adjacent subdivision meshes have similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted as the residual for the previous vertex position information, and the previous vertex position information is updated using the residual.
[0115] Referring to Figure 9, the vertices have coefficients generated through the lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.
[0116] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0117] Figure 10 shows the detailed operation of attribute transfer of encoding such as Figures 6-7.
[0118] Encoding according to embodiments includes attribute map encoding.
[0119] Information about the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has vertex position information, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is created through the attribute transfer process.
[0120] Attribute transfer first checks whether each point P(u, v) in the 2D texture domain belongs to a texture triangle of the reconstructed deformed mesh, and if it is in a texture triangle T, then the barycentric coordinate (α, α) of P(u, v) according to the triangle T is calculated. , γ) is calculated. And the 3D vertex positions of triangle T and (α, , γ) to compute the 3D coordinates M(x, y, z) of P(u, v). Find the vertex coordinates M'(x', y', z') and the triangle T' containing this point that corresponds to the most similar position to the computed M(x, y, z) in the input mesh domain. Then, in this triangle T', the center of mass coordinates (α', , γ') are calculated. The texture coordinates corresponding to the three vertices of Triangle T' and (α', ', γ') is used to calculate the texture coordinates (u', v'), and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.
[0121] The new attribute map generated through attribute transfer is grouped into GoF units to form an attribute map video, which is then compressed using a video codec.
[0122] Referring to Figure 10, the reference relationship between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.
[0123] The decoding process of Fig. 1 can perform the reverse process of the corresponding encoding process of Fig. 1. The specific decoding process is as follows.
[0124] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0125] Fig. 11 shows the configuration and operation of a decoder of a receiving device such as Fig. 1.
[0126] Figure 11 illustrates the intra decoding process of V-Mesh technology. First, the input bitstream can be separated into a mesh sub-stream, a displacement sub-stream, an attribute map sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0127] The mesh sub-stream is decoded by the decoder of the static mesh codec used in encoding, such as Google Draco, and as a result, the connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be restored. The displacement sub-stream is decoded into a displacement video by the decoder of the video compression codec used in encoding, and goes through the image unpacking, inverse quantization, and inverse transform processes to restore displacement information for each vertex. Inverse quantization is applied to the restored base mesh, and this result is combined with the restored displacement information to generate the final decoded mesh.
[0128] The attribute map sub-stream is decoded through the decoder of the video compression codec used in encoding, and then restored to the final attribute map through processes such as color format conversion.
[0129] The restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0130] Referring to FIG. 11, the bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term "substream" is interpreted as referring to a portion of the bitstream included in the bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).
[0131] The decoder performs the following decoding operations within the frame. The static mesh decoder decodes the mesh to generate a reconstructed quantized base mesh, and the inverse quantizer applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder decodes the displacement, the unpacker unpacks the decoded video image, and the inverse quantizer inversely quantizes the quantized image. The linear lifting inverse transform applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit generates a warped mesh based on the base mesh and the displacement. The video decoder decodes the attribute map, and the color transformation unit transforms the color format and / or space to generate the decoded attribute map.
[0132] Figure 12 shows the inter-frame decoding process of the V-MESH compression method.
[0133] Fig. 12 shows the configuration and operation of the decoder of the receiving device of Fig. 1.
[0134] Figure 12 illustrates the inter-decoding process of V-Mesh technology. First, the input bitstream can be separated into a motion sub-stream, a displacement sub-stream, an attribute sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0135] The motion sub-stream is decoded through entropy decoding and inverse prediction processes, and the reconstructed motion information is combined with the already reconstructed reference base mesh to generate a reconstructed quantized base mesh for the current frame. The result of applying inverse quantization to this is combined with the displacement information reconstructed in the same way as the intra decoding described above to generate the final decoded mesh. The attribute map sub-stream is decoded in the same way as the intra decoding. The reconstructed decoded mesh and the decoded attribute map can be utilized by the receiver as the final mesh data that can be utilized by the user.
[0136] Referring to Fig. 12, the bitstream includes motion, displacement, and attribute maps. Since inter-frame decoding is performed, a process of decoding inter-frame motion information is further included. The motion is decoded, and a restored quantized base mesh for the motion is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operations in Fig. 12, which are identical to those in Fig. 11, refer to the description in Fig. 11.
[0137] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0138] Fig. 13 corresponds to the transmitting device (100), the dynamic mesh video encoder (102), the Fig. 2 encoder (preprocessor and encoder), and / or the transmitting encoding device corresponding thereto of Fig. 1. Each component of Fig. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0139] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13.
[0140] The mesh preprocessor receives the original mesh as input and generates a simplified mesh (decimated mesh). Simplification can be performed based on the target number of vertices or the target number of polygons that constitute the mesh. Parameterization can be performed on the simplified mesh to generate texture coordinates and texture connection information per vertex. Additionally, quantization of floating-point mesh information into fixed-point information can be performed. This result can be encoded as a base mesh through a static mesh encoding unit. The mesh preprocessor can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The subdivided mesh can be fitted by adjusting the vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.
[0141] When performing intra-encoding on the corresponding mesh frame, the base mesh generated through the mesh preprocessing unit can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to the multiplexing unit.
[0142] When performing inter encoding for the corresponding mesh frame, a motion vector encoding unit is performed, which can calculate a motion vector between the base mesh and the reference reconstruction base mesh as input and encode the value. The motion vector encoding unit can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexing unit.
[0143] The encoded base mesh and motion vectors can be used to generate a restored base mesh through the base mesh restoration unit.
[0144] The displacement vector calculator can perform mesh refinement on the restored base mesh. The displacement vector can be calculated as the difference in vertex positions between the refined restored base mesh and the fitted subdivision mesh generated in the preprocessing unit. As a result, the number of displacement vectors can be calculated equal to the number of vertices in the refined mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0145] A displacement vector video generator can transform displacement vectors for effective encoding. The transformation can be performed by a lifting transformation, a wavelet transformation, etc., depending on the embodiment. In addition, quantization can be performed on the transformed displacement vector values, i.e., the transform coefficients. Different quantization parameters can be applied to each axis of the transform coefficients, and the quantization parameters can be derived according to the promise of the encoder / decoder. The transformed and quantized displacement vector information can be packed into a 2D image. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh.
[0146] A displacement vector video encoder can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer.
[0147] The displacement vector restored through the displacement vector restorer and the base mesh restored through the base mesh restorer and refined are restored through the mesh restorer, and the restored mesh has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0148] The texture map of the original mesh can be regenerated as a texture map for the restored mesh through the texture map video generation unit. The color information per vertex of the texture map of the original mesh can be assigned to the texture coordinates of the restored mesh. The regenerated texture maps for each frame can be bundled into GoF units to generate a texture map video.
[0149] The generated texture map video can be encoded using a video compression codec through a texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to a multiplexing unit.
[0150] The generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be multiplexed into a single bitstream and transmitted to a receiver via a transmitter. Alternatively, the generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via a transmitter.
[0151] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (displacement), and a texture map (attribute map). A transmitting device according to inter-encoding can generate a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map). A texture map obtained from a data input unit is generated and encoded based on a restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the segmented mesh. The base mesh is generated by preprocessing, simplifying, and encoding the original mesh. Motion is generated as a motion vector for the mesh of the current frame based on the reference base mesh of the previous frame.
[0152] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0153] Fig. 14 corresponds to the receiving device (110), the dynamic mesh video decoder (113), the decoder of Figs. 11-12, and / or the receiving decoding device corresponding thereto of Fig. 1. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.
[0154] The bitstream of the received Mesh is demultiplexed into a compressed motion vector bitstream or base mesh bitstream, displacement vector bitstream, and texture map bitstream after file / segment decapsulation.
[0155] If the current mesh has inter-frame encoding applied according to the frame header information, the motion vector decoding unit can perform decoding on the motion vector bitstream. The final motion vector can be reconstructed by adding the previously decoded motion vector to the residual motion vector decoded from the bitstream using the previously decoded motion vector as a predictor.
[0156] If the current mesh has been encoded within the screen according to the frame header information, the base mesh bitstream can restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh through the static mesh encoding unit.
[0157] In the base mesh restoration unit, if the current mesh has inter-frame encoding applied, the current base mesh can be restored by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. If the current mesh has intra-frame encoding applied, the static mesh decoding unit can perform inverse quantization on the decoded mesh to generate a restored base mesh.
[0158] The displacement vector bitstream can be decoded as a video bitstream using a video codec in a displacement vector video decoding unit.
[0159] In the displacement vector restoration unit, displacement vector transformation coefficients are extracted from the decoded displacement vector video, and the displacement vector is restored through the inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, a process of inverse transformation to the Cartesian coordinate system can be performed.
[0160] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Subdivision can generate vertex connection information, texture coordinates, and texture coordinate connection information, including the added vertices. The subdivided restored base mesh can be combined with the restored displacement vector to generate the final restored mesh.
[0161] The texture map bitstream can be decoded as a video bitstream using a video codec in a texture map video decoding unit. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of the corresponding vertex.
[0162] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.
[0163] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data. A receiving device according to inter-decoding can receive a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data.
[0164] A point cloud data transmission device and method according to embodiments can encode mesh data and transmit a bitstream including the encoded mesh data. A point cloud data reception device and method according to embodiments can receive a bitstream including mesh data and decode the mesh data. The point cloud data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The point cloud data transmission and reception method / device according to embodiments may also be referred to as mesh data transmission and reception method / device according to embodiments.
[0165] The encoding method / device according to the embodiments includes and performs the following: Fig. 1 transmitting device (100), acquisition unit (101), encoder (102), encapsulator (103), transmitter (104), Figs. 2, 3, 5, 7 pre-processing, encoder, Fig. 13 encoder, multiplexing unit, transmitter, Figs. 15 to 16 bitstream and syntax generation, Figs. 17 to 18 scene description information generation, Fig. 23 encoding method, etc.
[0166] The decoding method / device according to the embodiments includes and performs the following: a receiving device (110), a receiving unit (111), a decapsulator (112), a decoder (113), a renderer (114), decoding in FIGS. 11 and 12, a receiving unit, a demultiplexer, a renderer in FIGS. 14, bitstream and syntax parsing in FIGS. 15 to 16, scene description information parsing in FIGS. 17 to 18, mesh content parsing and scene description information parsing in FIGS. 19 to 22, and a decoding method in FIG. 24.
[0167] The method and device according to the embodiments may include and perform a Scene Description based Dynamic Mesh Coding Bitstream Support method.
[0168] The method and device according to the embodiments can provide a pipeline configuration and buffer format for each VDMC component to support decoding of a Video-based Dynamic Mesh Coding bitstream in a Scene Description System.
[0169] Embodiments include a transmitter (encoder) and / or receiver (decoder) for providing mesh content services that decode and render a V-DMC or mesh bitstream in a scene description system and provide signaling therefor.
[0170] Embodiments define a buffer format to support efficient access to stored V-DMC bitstreams and include a transmitter (encoder) and / or receiver (decoder) for providing mesh content services.
[0171] Although this document describes an embodiment from the perspective of V-DMC, the contents of this document can be applied not only to V-DMC but also to bitstreams using other mesh codings that are structured in the same manner or in a similar bitstream format.
[0172] In order to provide mesh content encoded with V-DMC as a service such as streaming, it can be encapsulated based on a single track and / or multiple tracks based on an ISOBMFF file. In order to provide V-DMC content encapsulated in a file as a service based on a scene description, the embodiments propose the following configuration. For example, 1) Pipeline setup for V-DMC content (data), 2) Buffer format for -DMC content (data).
[0173] Figure 15 shows a V-DMC bitstream according to embodiments.
[0174] Referring to Fig. 15(a), the bitstream may include a sequence header, an encoded base mesh, an encoded displacement (displacement), an encoded attribute (texture or attribute), etc. The sequence header may include decoder configuration information. The encoded base mesh includes an input mesh file. The encoded displacement information, which is difference information between vertices between a fitted and refined base mesh and a restored submesh, may be data packed as a 2D image and encoded using a video codec. The encoded attribute may be data encoded using a video codec such as HEVC, and may be an attribute (e.g., color) generated by matching the coordinates of the property (texture (2D)) of the restored base mesh using an input attribute map (e.g., PNG image).
[0175] In other words, FIG. 15(a) may be an example of a bitstream format resulting from video-based dynamic mesh encoding. The bitstream may include a sequence header, compressed basemesh data, compressed displacement data, and compressed texture data.
[0176] Basemesh data can be compressed using Edgebreaker or Draco tool.
[0177] Displacement data can be compressed using a video codec such as HEVC or VVC or arithmetic coding.
[0178] Texture / Attribute data can be compressed with video codecs such as HEVC, VVC, etc.
[0179] Referring to Fig. 15(b), as a result of video-based dynamic mesh encoding, a bitstream according to embodiments may be composed of sample stream units. For example, the bitstream may include a sample stream DMC header and at least one sample stream DMC unit. The sample stream DMC unit may include parameter information, mesh data, etc.
[0180] VPS (V-DMC Parameter Set): May include decoder configuration information and / or parameter set information related to mesh encoding / decoding. AD (Atlas Data): May include information related to 2D mapping or texture mapping for 3D objects. BMD (Base Mesh Data): Encoded base mesh information for mesh encoding / decoding. DD (Displacements Data): Displacement information encoded with arithmetic coding. GVD (Displacements Video Data): Displacement information encoded with a video codec. Displacement information may be referred to as geometry information, etc. AVD (Attribute Video Data): Attribute or texture information encoded with a video codec.
[0181] Figure 16 shows a V3C unit according to embodiments.
[0182] Specifically, FIG. 16 shows the header syntax of a V3C unit, which is a sample stream DMC unit (which may be referred to as a sample stream unit or a V3C unit) included in the bitstream of FIG. 15. The V3C unit (v3c_unit(numBytesInV3CUnit)), which is a data unit of the bitstream, includes a header (FIG. 16, v3c_unit_header( )) and a payload (v3c_unit_payload( numBytesInV3CUnit - 4)).
[0183] The unit type (vuh_unit_type) of the unit header can use values from 0 to 6 to indicate the unit type, such as VPS, AD, OVD, GVD, AVD, PVD, PVD, CAD, etc.
[0184] vuh_unit_typeIdentifierV3C unit typeDescription0V3C_VPSV3C parameter setV3C level parameters1V3C_ADAtlas dataAtlas information2V3C_OVDOccupancy video dataOccupancy information3V3C_GVDGeometry video dataGeometry information4V3C_AVDAttribute video dataAttribute information5V3C_PVDPacked video dataPacking information6V3C_CADCommon atlas dataInformation that is common for atlases in a CVS. Specified in ISO / IEC 23090-127...31V3C_RSVDReserved-
[0185] vuh_unit_type: Indicates the specified V3C unit type, as above. Values marked as reserved are reserved for future use in ISO / IEC.
[0186] vuh_v3c_parameter_set_id: Indicates the value of vps_v3c_parameter_set_id for the active V3C VPS. The value of vuh_v3c_parameter_set_id is in the range of 0 to 15.
[0187] vuh_atlas_id: Indicates the ID of the atlas corresponding to the current V3C unit. The value of vuh_atlas_id ranges from 0 to 63.
[0188] vuh_attribute_index: Indicates the index of the attribute data included in the attribute video data unit. The value of vuh_attribute_index is in the range of 0 to (ai_attribute_count[vuh_atlas_id] - 1).
[0189] vuh_attribute_partition_index: Indicates the index of the attribute dimension group included in the attribute video data unit. The value of vuh_attribute_partition_index is in the range of 0 to ai_attribute_dimension_partitions_minus1[vuh_atlas_id][vuh_attribute_index].
[0190] vuh_map_index: Indicates the map index of the current geometry or attribute stream. If this value is absent, the map index of the current geometry or attribute sub-bitstream is derived based on the type of the sub-bitstream and certain operations on the geometry and attribute video sub-bitstreams. The value of vuh_map_index, if present, is in the range 0 to vps_map_count_minus1[vuh_atlas_id].
[0191] vuh_auxiliary_video_flag: If 1, indicates that the associated geometry or attribute video data unit is a RAW and / or EOM-coded point video-only sub-bitstream. If vuh_auxiliary_video_flag is 0, indicates that the associated geometry or attribute video data unit may contain RAW and / or EOM-coded points. If vuh_auxiliary_video_flag is absent, its value is inferred to be 0.
[0192] vuh_reserved_zero_12bits: Equal to 0 in bitstreams that follow this version of this document.
[0193] vuh_reserved_zero_17bits: Equal to 0 in bitstreams that follow this version of this document.
[0194] vuh_reserved_zero_23bits: Equal to 0 in bitstreams that follow this version of this document.
[0195] vuh_reserved_zero_27bits: Equal to 0 in bitstreams that follow this version of this document.
[0196] The syntax of the sample stream DMC header and the sample stream DMC unit of the bitstream according to the embodiments can be defined as follows.
[0197] Syntax of the sample stream DMC header:
[0198] sample_stream_dmc_header() {Descriptorsdmh_unit_size_precision_bytes_minus1u(3)sdmh_reserved_zero_5bitsu(5)}
[0199] sdmh_unit_size_precision_bytes_minus1: Adding 1 to this value indicates the precision (in bytes) of sdmu_dmc_unit_size elements in every sample stream DMC unit. sdmh_unit_size_precision_bytes_minus1 is in the range 0 to 7.
[0200] The syntax of a sample stream DMC unit is as follows. Each sample stream DMC unit contains one type of DMC unit among VPS, AD, BMD, DD, GVD, and AVD. The contents of each sample stream DMC unit are associated with the same access unit as the DMC unit contained in the sample stream DMC unit.
[0201] sample_stream_dmc_unit() {Descriptorsdmu_dmc_unit_sizeu(v)dmc_unit(sdmu_dmc_unit_size )}
[0202] sdmu_dmc_unit_size: Indicates the size (in bytes) of the subsequent dmc_unit. The number of bits used to express sdmu_dmc_unit_size is equal to (sdmh_unit_size_precision_bytes_minus1 + 1) * 8. A dmc_unit can consist of a unit header and a unit payload. The unit header can contain type information about the unit payload, and the unit payload can contain data of the corresponding type.
[0203] The dmc_unit data syntax and semantics may be in the format of v3c_unit used in the V3C codec specification, i.e., ISO / IEC 23090-5, which is referenced in the V-DMC standard, but may not be limited to that format. The contents according to the embodiments are specific to the V-DMC or V3C codec and can be implemented independently.
[0204] A bitstream according to embodiments may be composed of sample stream NAL (Network Abstraction Layer) units. A sample stream NAL unit may include a sample stream NAL header and a sample stream NAL unit.
[0205] Sample Stream NAL Header Syntax:
[0206] sample_stream_nal_header() {Descriptorssnh_unit_size_precision_bytes_minus1u(3)ssnh_reserved_zero_5bitsu(5)}
[0207] Sample Stream NAL Header Unit Syntax:
[0208] sample_stream_nal_unit() {Descriptorssnu_nal_unit_sizeu(v)nal_unit (ssnu_nal_unit_size)}
[0209] Semantics of the sample stream NAL header:
[0210] The sample stream NAL header is always at the beginning of the NAL stream.
[0211] ssnh_unit_size_precision_bytes_minus1: Adding 1 to this value indicates the precision (in bytes) of ssnu_nal_unit_size elements in every sample stream NAL unit. ssnh_unit_size_precision_bytes_minus1 is in the range 0 to 7.
[0212] ssnh_reserved_zero_5bits: is equal to 0 in bitstreams that follow this version of this document.
[0213] Sample Stream NAL Unit Semantics:
[0214] The order of sample stream NAL units in a sample stream follows the decoding order of the NAL units contained in the sample stream NAL units.
[0215] A unit (nal_unit) included in a bitstream according to embodiments may be basemesh data that may be included in a sample of a basemesh track to be described later. That is, a unit of a bitstream may be a basemesh NAL unit (bmesh_nal_unit) and / or displacement data (displ_nal_unit) encoded using arithmetic coding.
[0216] Additionally, nal_unit may be submesh data that can be included in the samples of the submesh track described later. The submesh data may be defined in the form of bmesh_nal_unit.
[0217] ssnu_nal_unit_size represents the size (in bytes) of the subsequent NAL_unit. The number of bits used to represent ssnu_nal_unit_size is equal to (ssnh_unit_size_precision_bytes_minus1 + 1) * 8.
[0218] A bitstream may contain a basemesh sub-bitstream.
[0219] A NAL sample stream format can be constructed by arranging NAL units in decoding order and prefixing each NAL unit with a header indicating the exact size (in bytes) of the NAL unit. The sample stream header is included at the beginning of the sample stream bitstream, which indicates the precision (in bytes) of the signaled NAL unit size. The NAL unit stream format can be extracted from the sample stream format by traversing the sample stream format, reading the size information, and appropriately extracting each NAL unit.
[0220] General NAL unit syntax:
[0221] bmesh_nal_unit( NumBytesInNalUnit ) {Descriptorbmesh_nal_unit_header( )NumBytesInRbsp = 0for( i = 2; i < NumBytesInNalUnit; i++ )rbsp_byte[ NumBytesInRbsp++ ]b(8)}
[0222] NAL unit header syntax:
[0223] bmesh_nal_unit_header() {Descriptorbmesh_nal_forbidden_zero_bitf(1)bmesh_nal_unit_typeu(6)bmesh_nal_layer_idu(6)bmesh_nal_temporal_id_plus1u(3)}
[0224] Low-byte sequence payload, trailing bits, and byte alignment syntax
[0225] Basemesh Sequence Parameter Set RBSP Syntax
[0226] General Basemesh Sequence Parameter Set RBSP Syntax
[0227] bmesh_sequence_parameter_set_rbsp( ) {Descriptorbmsps_sequence_parameter_set_idu(4)bmesh_profile_tier_level( )bmsps_intra_mesh_codec_idu(8)bmsps_inter_mesh_codec_idu(8)bmsps_inter_mesh_motion_group_size_minus1u(8)bmsps_inter_mesh_max_num_neighbours_minus1u(8)bmsps_geometry_3d_bit_depth_minus1u(5)bmsps_facegroup_segmentation_methodue(v)bmsps_mesh_attribute_countu(7)for( i = 0; i < bmsps_mesh_attribute_count; i++ ) {bmsps_mesh_attribute_type_id[ i ]u(4)bmsps_attribute_bit_depth_minus1[ i ]u(5)bmsps_attribute_msb_align_flag[ i ]u(1)}bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4ue(v)bmsps_max_dec_mesh_frame_buffering_minus1ue(v)bmsps_long_term_ref_mesh_frames_flagu(1)bmsps_num_ref_mesh_frame_lists_in_bmspsue(v)for( i = 0; i < bmsps_num_ref_mesh_frame_lists_in_bmsps; i++ )bmesh_ref_list_struct( i )bmsps_extension_present_flagu(1)if( bmsps_extension_present_flag ) {bmsps_extension_countu(8)}if( bmsps_extension_count ){bmsps_extensions_length_minus1ue(v)for( i = 0; i < bmsps_extension_count;i++ ) {bmsps_extension_type[ i ]u(8)bmsps_extension_length[ i ]u(16)bmsps_extension( bmsps_extension_type[ i ], bmsps_extension_length[ i ] )}}rbsp_trailing_bits( )};
[0228] 베이스메쉬 SPS 확장 신택스:
[0229] bmsps_extension( extension_type, extension_length ) {Descriptorfor( j = 0; j < extension_length; j++ )bmsps_extension_data_byteu(8)length_alignment( )}
[0230] 베이스메쉬 프레임 파라미터 세트 RBSP 신택스:
[0231] 제너럴 베이스메쉬 프레임 파라미터 RBSP 신택스:
[0232] bmesh_frame_parameter_set_rbsp( ) {Descriptorbfps_mesh_sequence_parameter_set_idu(4)bfps_mesh_frame_parameter_set_idu(4)bmesh_sub_mesh_information( ) bfps_output_flag_present_flagu(1)bfps_num_ref_idx_default_active_minus1ue(v)bfps_additional_lt_mfoc_lsb_lenue(v)bfps_extension_present_flagu(1)if( bfps_extension_present_flag )bfps_extension_8bitsu(8)if( bfps_extension_8bits )while( more_rbsp_data( ) )bfps_extension_data_flagu(1)rbsp_trailing_bits( )}
[0233] 베이스메쉬 서브메쉬 정보:
[0234] bmesh_sub_mesh_information( ) {Descriptorbmsi_use_single_mesh_flagu(1)if(!bmsi_use_single_mesh_flag){bmsi_num_submeshes_minus1u(8)}elsebmsi_num_submeshes_minus1 = 0bmsi_signalled_submesh_id_flagu(1)if( bmsi_signalled_submesh_id_flag ) {bmsi_signalled_submesh_id_length_minus1ue(v)for( i = 0; < bmsi_num_submeshes_minus1 + 1; i++ )bmsi_submesh_id[ i ]u(v)SubMeshIDToIndex[ bmsi_submesh_id[ i ] ] = iSubMeshIndexToID[ i ] = bmsi_submesh_id[ i ]}}elsefor( i = 0; i < bmsi_num_submeshes_minus1 + 1; i++ ) {bmsi_submesh_id[ i ] = iSubMeshIDToIndex[ i ] = iSubMeshIndexToID[ i ] = i}}
[0235] 베이스메쉬 서브메쉬 데이터 유닛 신택스:
[0236] submesh_data_unit( subMeshID, unitSize ) {Descriptorif( smh_type == I_SUBMESH ) {sdu_intra_sub_mesh_unit( subMeshID, unitSize )}else if( smh_type == P_SUBMESH ) {sdu_inter_sub_mesh_unit( subMeshID )}else if( smh_type == SKIP_SUBMESH ) {sdu_skip_sub_mesh_unit( )}}
[0237] Basemesh Intra Submesh Data Unit Syntax:
[0238] sdu_intra_sub_mesh_unit( subMeshID, vertexCount ) {Descriptorsismu_intra_unit( subMeshID )length_alignment( )}
[0239] sismu_intra_unit(subMeshID, sismu_intra_unit_size) contains a portion of mesh data of size sismu_intra_unit_size[subMeshID], which is an ordered string of bytes or bits that identifies the location of unit boundaries in a pattern of data. The format of this mesh data is identified by the bmptl_profile_codec_group_idc or component codec mapping SEI message.
[0240] VDMC Atlas Tile Data Unit Syntax:
[0241] General VDMC Atlas Tile Data Unit Syntax:
[0242] vmdc_atlas_tile_data_unit( tileID ) { Descriptor( ath_type == SKIP_TILE ) { for( p = 0; p < RefAtduTotalNumMeshpatches[ tileID ]; p++ )skip_meshpatch_data_unit( )} else { p = 0do { atdu_meshpatch_mode[ tileID ][ p ]ue(v)isEnd = ( ath_type == P_TILE && atdu_meshpatch_mode[ tileID ][ p ] == P_END) ||( ath_type == I_TILE && atdu_meshpatch_mode[ tileID ][ p ] == I_END )if( !isEnd ) {meshpatch_information_data( tileID , p , atdu_meshpatch_mode[ tileID ][ p ] )p++}} while( !isEnd )}AtduNumTotalMeshpatches[ tileID ] = p}
[0243] Mesh Patch Information Data Syntax:
[0244] meshpatch_information_data( tileID, patchIdx, meshpatchMode ) {Descriptorif( ath_type == P_TILE ) {if( meshpatchMode == P_SKIP )skip_meshpatch_data_unit( )else if( meshpatchMode == P_MERGE )merge_meshpatch_data_unit( tileID, patchIdx )else if( meshpatchMode == P_INTRA )meshpatch_data_unit( tileID, patchIdx )else if( meshpatchMode == P_INTER )inter_meshpatch_data_unit( tileID, patchIdx )}else if( ath_type == I_TILE ) {if( meshpatchMode == I_INTRA )meshpatch_data_unit( tileID, patchIdx )}}
[0245] 메쉬패치 데이터 유닛 신택스:
[0246]
[0247] Skip Mesh Patch Data Unit Syntax:
[0248] skip_meshpatch_data_unit() {Descriptor|
[0249] Merge mesh patch data unit syntax:
[0250] merge_meshpatch_data_unit(tileID, patchIdx) {Descriptorif(NumRefIdxActive)mmdu_ref_index[tileID][patchIdx]ue(v)mmdu_patch_index[tileID][patchIdx]se(v)}
[0251] Inter-MeshPatch Data Unit Syntax:
[0252] inter_meshpatch_data_unit( tileID, patchIdx ) {Descriptorif( NumRefIdxActive )imdu_ref_index[ tileID ][ patchIdx ]ue(v)imdu_patch_index[ tileID ][ patchIdx ]se(v)imdu_delta_vertex_count_minus1[ tileID ][ patchIdx ]se(v)imdu_delta_face_count_minus1[ tileID ][ patchIdx ]se(v)imdu_2d_delta_pos_x[ tileID ][ patchIdx ]se(v)imdu_2d_delta_pos_y[ tileID ][ patchIdx ]se(v)imdu_2d_delta_size_x[ tileID ][ patchIdx ]se(v)imdu_2d_delta_size_y[ tileID ][ patchIdx ]se(v)for(i=0; i< asve_num_attribute_video; i++ ) {if( asve_attribute_subtexture_enabled_flag[ i ] ) {imdu_attributes_2d_delta_pos_x[ tileID ][ patchIdx ][ i ]se(v)imdu_attributes_2d_delta_pos_y[ tileID ][ patchIdx ][ i ]se(v)imdu_attributes_2d_delta_size_x[ tileID ][ patchIdx ][ i ]se(v)imdu_attributes_2d_delta_size_y[ tileID ][ patchIdx ][ i ]se(v)}}
[0253] 텍스쳐 프로젝션 정보 신택스:
[0254] texture_projection_information( tileID, patchIdx ) {Descriptortpi_face_id_present_flag[ tileID ][ patchIdx ]u(1)tpi_frame_scale[ tileID ][ patchIdx ]fl(64)tpi_subpatch_count_minus1[ tileID ][ patchIdx ]ue(v)numSubpatches = tpi_subpatch_count_minus1[ tileID ][ patchIdx ] + 1for( subIdx = 0; subIdx < numSubpatches; subIdx++)subpatch_information( tileID, patchIdx , subIdx )}
[0255] 서브-패치 정보 신택스:
[0256] subpatch_information( tileID, patchIdx , subIdx ) {Descriptorif( !tpi_face_id_present_flag[ tileID ][ patchIdx ] )si_face_id[ tileID ][ patchIdx ][ subIdx ]ue(v)si_projection_id[ tileID ][ patchIdx ][ subIdx ]u(v)si_orientation_id[ tileID ][ patchIdx ][ subIdx ]ue(v)si_2d_pos_x[ tileID ][ patchIdx ][ subIdx ]ue(v)si_2d_pos_y[ tileID ][ patchIdx ][ subIdx ]ue(v)si_2d_size_x_minus1_diff[ tileID ][ patchIdx ][ subIdx ]se(v)si_2d_size_y_minus1_diff[ tileID ][ patchIdx ][ subIdx ]ve(v)si_scale_present_flag[ tileID ][ patchIdx ][ subdx ]u(1)if( si_scale_present_flag[ tileID ][ patchIdx ][ subIdx ] )si_scale_power_factor[ tileID ][ patchIdx ][ subIdx ]ue(v)}
[0257] Below, the semantics of the aforementioned syntax are explained.
[0258] BaseMesh NAL Unit Semantics:
[0259] General NAL unit semantics:
[0260] NumBytesInNalUnit represents the size of a NAL unit in bytes. This value is required for decoding a NAL unit. To enable inference of NumBytesInNalUnit, a form delimiting NAL unit boundaries is required.
[0261] rbsp_byte[ i ] is the ith byte of the RBSP. An RBSP is specified as an ordered sequence of bytes as follows:
[0262] An RBSP contains a string of data bits (SODBs) as follows:
[0263] - If the SODB is empty (i.e., has a length of 0 bits), the RBSP is also empty.
[0264] - Otherwise, the RBSP contains the SODB as follows:
[0265] 1) The first byte of the RBSP contains the first (most significant, leftmost) 8 bits of the SODB. The next byte of the RBSP contains the next 8 bits of the SODB, and so on until there are fewer than 8 bits of the SODB left.
[0266] 2) The rbsp_trailing_bits( ) syntax structure follows SODB as follows:
[0267] i) The first (most significant, leftmost) bit of the last RBSP byte contains the remaining bits of the SODB (if any).
[0268] ii) The next bit consists of a single bit equal to 1 (i.e. rbsp_stop_one_bit).
[0269] iii) If rbsp_stop_one_bit is not the last bit of a byte that is aligned, byte alignment is achieved by having at least one bit equal to 0 (i.e., an instance of rbsp_alignment_zero_bit).
[0270] Syntax structures with these RBSP properties are indicated in the syntax table using the "_rbsp" suffix. These structures are carried within the NAL unit as the contents of the rbsp_byte[ i ] data byte.
[0271] NAL unit header semantics:
[0272] As with atlases, similar NAL unit types are defined for base meshes, enabling similar functionality for random access and segmentation of the mesh. Unlike tile-divided atlases, the concept of sub-meshes is defined, with specific NAL units corresponding to coded mesh data. It also defines NAL units that can contain metadata, such as SEI messages.
[0273] Basemesh Raw byte sequence payloads, trailing bits, and byte alignment semantics:
[0274] Basemesh Sequence Parameter Set RBSP Semantics:
[0275] General Basemesh Sequence Parameter Set RBSP Semantics:
[0276] bmsps_sequence_parameter_set_id: An identifier for the basemesh sequence parameter set so that other syntax elements can reference it.
[0277] bmsps_intra_mesh_codec_id: Indicates the identifier of the codec used to compress the static mesh. bmsps_intra_mesh_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, a component codec mapping SEI message, or by means external to this document. A specific mesh or motion mesh codec may be associated with a profile specified in that specification, or may be explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0278] bmsps_inter_mesh_codec_id: Indicates the identifier of the codec used to compress the motion data. bmsps_intr_mesh_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, a component codec mapping SEI message, or by means external to this document. A specific mesh or motion mesh codec may be associated with a profile specified in that specification, or may be explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0279] bmsps_inter_mesh_motion_group_size_minus1: Adding 1 to this value indicates the size of the vertex grouping in motion vector coding. bmsps_inter_mesh_motion_group_size_minus1 is in the range of 0 to 255.
[0280] bmsps_inter_mesh_max_num_neighbours_minus1: Adding 1 to this value indicates the maximum number of vertex neighbors to use for motion vector estimation. bmsps_inter_mesh_max_num_neighbours_minus1 is in the range of 0 to 255.
[0281] bmsps_geometry_3d_bit_depth_minus1: Adding 1 to this value indicates the bit depth of the geometry coordinates of the reconstructed mesh. bmsps_geometry_3d_bit_depth_minus1 is in the range of 0 to 31.
[0282] bmsps_facegroup_segmentation_method: Indicates the identifier of the method for deriving facegroup IDs from a mesh.
[0283] bmsps_mesh_attribute_count: Indicates the number of attributes associated with the mesh. bmsps_mesh_attribute_count is in the range of 0 to 127.
[0284] bmsps_mesh_attribute_type_id[i]: Indicates the attribute type of the attribute with index i for the mesh. The list of supported attributes and their relationship to bmsps_mesh_attribute_type_id[i] are as follows:
[0285] bmsps_mesh_attribute_type_id[ i ]IdentifierAttribute type0ATTR_TEXTURE Texture1ATTR_MATERIAL_IDMaterial ID2ATTR_TRANSPARENCYTransparency3ATTR_REFLECTANCEReflectance4ATTR_NORMALNormals5ATTR_FACEGROUP_IDFacegroup ID6..14ATTR_RESERVEDReserved15ATTR_UNSPECIFIEDUnspecified
[0286] bmsps_attribute_bit_depth_minus1[ i ]: Adding 1 to this value indicates the bit depth of the attribute with index i for the mesh. bmsps_attribute_bit_depth_minus1[ i ] is in the range 0 to 31.
[0287] bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4: Adding 4 to this value gives the values of the variables Log2MaxMeshFrmOrderCntLsb and MaxMeshFrmOrderCntLsb used in the decoding process for the atlas frame order count, as follows:
[0288] Log2MaxMeshFrmOrderCntLsb =
[0289] bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4 + 4
[0290] MaxMeshFrmOrderCntLsb = 2Log2MaxMeshFrmOrderCntLsb
[0291] The value of bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4 is in the range of 0 to 12.
[0292] bmsps_max_dec_mesh_frame_buffering_minus1: Adding 1 to this value indicates the maximum size of the decoded atlas frame buffer required for CAS, in atlas frame storage buffer units. The value of bmps_max_dec_mesh_frame_buffering_minus1 is in the range of 0 to 15.
[0293] bmsps_long_term_ref_mesh_frames_flag: If this value is 0, it indicates that long-term reference atlas frames are not used for inter prediction of coded atlas frames in CAS. If bmsps_long_term_ref_mesh_frames_flag is 1, it indicates that long-term reference atlas frames can be used for inter prediction of one or more coded atlas frames in CAS.
[0294] bmsps_num_ref_mesh_frame_lists_in_bmsps: Indicates the number of bmesh_ref_list_struct(rlsIdx) syntax structures included in the atlas sequence parameter set. The value of bmsps_num_ref_mesh_frame_lists_in_bmsps ranges from 0 to 64.
[0295] NOTE: The decoder allocates memory for a total number of bmesh_ref_list_struct(rlsIdx) syntax structures equal to (bmsps_num_ref_mesh_frame_lists_in_bmsps + 1), since there can only be one bmesh_ref_list_struct(rlsIdx) syntax structure signaled directly from the atlas tile header of the current atlas tile.
[0296] bmsps_extension_present_flag: If this value is 1, it indicates that bmsps_extension_count_minus1 is in the basemesh sequence parameter set.
[0297] bmsps_extension_count: Indicates the number of BMSPS extensions in the v3c_parameter_set( ) syntax structure. If not present, bmsps_extension_count is inferred to be 0. If bmsps_extension_count is 0, BmspsExtensionsLength, which specifies the cumulative length in bytes of all extensions that follow this syntax element, is 0.
[0298] If bmsps_extensions_length_minus1 is present, it represents the cumulative length in bytes of all extensions following this syntax element. BmspsExtensionsLength is calculated as follows:
[0299] if( bmsps_extension_count == 0 )
[0300] BmspsExtensionsLength = 0
[0301] else
[0302] BmspsExtensionsLength = bmsps_extensions_length_minus1 + 1
[0303] If bmsps_extension_count is not 0, BmspsExtensionsLength is equal to 3 * bmsps_extension_count plus the sum of all bmsps_extension_length[ i ].
[0304] bmsps_extension_type[ i ]: Indicates the BMSPS extension type for the extension with index i as specified in ISO / IEC 23090-29.
[0305] bmsps_extension_length[ i ]: Indicates the number of bytes used to represent the payload size of the syntax structure of the associated extension with index i. If bmsps_extension_length[ i ] is 0, there is no extension payload for the extension with index i. Otherwise, the extension with index i has a payload size in bits in the range 8 * ( bmsps_extension_length[ i ] - 1 ) + 1 to 8 * bmsps_extension_length[ i ], inclusive.
[0306] BaseMesh SPS Extension Semantics:
[0307] bmsps_extension_data_byte can have any value.
[0308] Basemesh Frame Parameter Set RBSP Semantics:
[0309] General Basemesh Frame Parameter Set RBSP Semantics:
[0310] bfps_mesh_sequence_parameter_set_id: Indicates the value of bmsps_sequence_parameter_set_id for the active Basemesh sequence parameter set.
[0311] bfps_mesh_parameter_set_id identifies the basemesh frame parameter set so that other syntax elements can reference it.
[0312] If bfps_output_flag_present_flag is 1, it indicates that the smh_output_flag syntax element is present in the associated submesh header. If bfps_output_flag_present_flag is 0, it indicates that the smh_output_flag syntax element is not present in the associated submesh header.
[0313] bfps_num_ref_idx_default_active_minus1 plus 1 represents the inferred value of the variable NumRefIdxActive for tiles where smh_num_ref_idx_active_override_flag is 0. The value of bfps_num_ref_idx_default_active_minus1 is in the range 0 to 14.
[0314] bfps_additional_lt_mfoc_lsb_len represents the value of the variable MaxLtMeshFrmOrderCntLsb used in the decoding process of the reference atlas frame list.
[0315] MaxLtMeshFrmOrderCntLsb =
[0316] 2 * (Log2MaxMeshFrmOrderCntLsb + bfps_additional_lt_mfoc_lsb_len)
[0317] The value of bfps_additional_lt_mfoc_lsb_len is in the range 0 to 32 - Log2MaxAtlasFrmOrderCntLsb.
[0318] If bmsps_long_term_ref_mesh_frames_flag is 0, the value of bfps_additional_lt_mfoc_lsb_len is equal to 0.
[0319] If bfps_extension_present_flag is 1, it indicates that the syntax element bfps_extension_8bits is present in the basemesh frame parameter set. If bfps_extension_present_flag is 0, it indicates that the syntax element bfps_extension_8bits is not present.
[0320] If bfps_extension_8bits is 0, it indicates that the AFPS RBSP syntax structure does not include the bfps_extension_data_flag syntax element. A non-zero value for bfps_extension_8bits is reserved for future use in ISO / IEC.
[0321] bfps_extension_data_flag can have any value.
[0322] Basemesh submesh information:
[0323] If bmsi_use_single_mesh_flag is 1, it indicates that each mesh frame has only one submesh referencing the BFPS. If bmsi_use_single_mesh_falg is 0, it indicates that each mesh frame can have more than one submesh referencing the BFPS.
[0324] bmsi_num_submeshes_minus1 plus 1 indicates the number of submeshes referencing BFPS in each mesh frame. The value of bmsi_num_submeshes_minus1 is in the range 0 to 63. If it is not present and bmsi_use_single_mesh_flag is 1, the value is inferred to be 1.
[0325] If bmsi_signalled_submesh_id_flag is 1, it indicates that the submesh ID of each mesh frame is signaled. If bmsi_signalled_tile_id_flag is 0, it indicates that the submesh ID is not signaled.
[0326] bmsi_signalled_submesh_id_length_minus1 plus 1 indicates the number of bits used to represent the syntax element bmsi_tile_id[ i ], if present, and the syntax element submesh_id in the submesh header. The value of bmsi_signalled_tile_id_length_minus1 is in the range 0 to 15. If absent, its value is inferred to be equal to Ceil(Log2(bmsi_num_submeshes_minus1 + 1 )) - 1.
[0327] bmsi_submesh_id[ i ] represents the tile ID of the ith submesh. The length of the bmsi_submesh_id[ i ] syntax element is bmsi_signalled_submesh_id_length_minus1 + 1 bit. If not present, the value of bmsi_submesh_id[ i ] is inferred to be equal to i for each i in the range 0 to bmsi_num_submeshes_minus1, inclusive. It is a bitstream conformance requirement that bmsi_submesh_id[ i ] must not be equal to bmsi_submesh_id[ j ] for all i != j. The length of the bmsi_submesh_id[ i ] syntax element is bmsi_signalled_submesh_id_length_minus1 + 1 bit.
[0328] The variable FirstSubmeshID is calculated as follows:
[0329] FirstSubmeshID=bmsi_submesh_id
[0000]
[0330] for ( i = 1; i < bmsi_num_submeshes_minus1+ 1; i++ )
[0331] FirstSubmeshID = Min(FirstSubmeshID, bmsi_submesh_id[ i ])
[0332] VDMC Atlas Tile Data Unit Semantics:
[0333] General VDMC Atlas Tile Data Unit Semantics:
[0334] atdu_meshpatch_mode[ tileID ][ p ]: Indicates the mesh patch mode for the mesh patch with patch index p in the current atlas tile with tile ID equal to tileID. Allowed values for atdu_meshpatch_mode[ tileID ][ p ] are shown in Table 33 for atlas tiles with ath_type of I_TILE, Table 34 for atlas tiles with ath_type of P_TILE, and Table 35 for atlas tiles with ath_type of SKIP_TILE. If not present, the value of atdu_meshpatch_mode[ tileID ][ p ] is inferred to be equal to P_SKIP.
[0335] Meshpatch mode for I_TILE type atlas tiles:
[0336] atdu_meshpatch_mode[ tileID ][ p ]IdentifierDescription0I_INTRANon-predicted meshpatch mode1-13I_RESERVEDReserved modes for future use by ISO / IEC14I_ENDMeshpatch termination mode
[0337] Meshpatch mode for P_TILE type atlas tiles:
[0338] atdu_meshpatch_mode[ tileID ][ p ]IdentifierDescription0P_SKIPMeshpatch Skip mode1P_MERGEMeshpatch Merge mode2P_INTERInter predicted Meshpatch mode3P_INTRANOn-predicted Meshpatch mode4-13P_RESERVEDReserved modes for future use by ISO / IEC14P_ENDPatch termination mode
[0339] Meshpatch mode for SKIP_TILE type atlas tiles:
[0340] atdu_meshpatch_mode[ tileID ][ p ]IdentifierDescription0P_SKIPMeshpatch Skip mode
[0341] Mesh patch information data may include configuration information about the mesh patch.
[0342] MeshPatch Data Unit Semantics:
[0343] mdu_submesh_id[ tileID ][ patchIdx ] represents the associated submesh ID assigned to the current mesh patch with index patchIdx in the current atlas tile, where the tile ID is equal to tileID. The value of mdu_submesh_id[ tileID ][ patchIdx ] must be one of afmi_submesh_id[ i ], where i is in the range 0 to ath_submesh_count - 1.
[0344] mdu_vertex_count_minus1[ tileID ][ patchIdx ] specifies the number of vertices associated with the current mesh patch with index patchIdx in the current atlas tile, where tileID is equal to tileID.
[0345] mdu_face_count_minus1[ tileID ][ patchIdx ] specifies the number of faces associated with the current mesh patch with index patchIdx in the current atlas tile, where tile ID is equal to tileID.
[0346] mdu_2d_pos_x[ tileID ][ patchIdx ] represents the x-coordinate of the upper-left corner of the mesh patch bounding box of the current mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID and is expressed as a multiple of PatchPackingBlockSize.
[0347] mdu_2d_pos_y[ tileID ][ patchIdx ] specifies the y-coordinate of the upper-left corner of the mesh patch bounding box of the current mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID and is expressed as a multiple of PatchPackingBlockSize.
[0348] mdu_2d_size_x_minus1[ tileID ][ patchIdx ] plus 1 specifies the width value of the mesh patch bounding box of the mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID.
[0349] mdu_2d_size_y_minus1[ tileID ][ patchIdx ] plus 1 represents the bounding box height value of the current mesh patch for the mesh patch with index patchIdx whose tile ID is equal to tileID in the current atlas tile.
[0350] If mdu_parameters_override_flag[ tileID ][ patchIdx ] is 1, it indicates that parameters mdu_subdivision_override_flag, mdu_quantization_override_flag, mdu_transform_method_override_flag, and mdu_transform_parameters_override_flag are present in the meshpatch with index patchIdx of the current atlas tile and the tile ID is equal to tileID.
[0351] If mdu_subdivision_override_flag[ tileID ][ patchIdx ] is 1, it indicates that mdu_subdivision_method and mdu_subdivision_iteration_count are in the meshpatch with the index patchIdx of the current atlas tile and that the tile ID is equal to tileID. If mdu_subdivision_override_flag[ tileID ][ patchIdx ] is absent, its value is inferred to be equal to 0.
[0352] mdu_quantization_override_flag[ tileID ][ patchIdx ] equal to 1 indicates that the vdmc_quantization_parameters(qpIndex, subdivisionCount) syntax construct exists in the mesh patch with the index patchIdx of the current atlas tile, and the tile ID is equal to tileID. If mdu_quantization_override_flag[ tileID ][ patchIdx ] is absent, its value is inferred to be equal to 0.
[0353] The variable QpIndex in the current patch is derived as follows:
[0354] QpIndex = mdu_quantization_override_flag[ tileID ][ patchIdx ] ? 2:
[0355] afve_quantization_parameters_enable_flag ? 1: 0
[0356] If mdu_transform_method_override_flag[ tileID ][ patchIdx ] is 1, it indicates that mdu_transform_method is present in the mesh patch with index patchIdx of the current atlas tile, and the tile ID is equal to tileID. If mdu_transform_method_override_flag[ tileID ][ patchIdx ] is absent, the value is inferred to be 0.
[0357] If mdu_transform_parameters_override_flag[ tileID ][ patchIdx ] is 1, it indicates that the vdmc_lifting_transform_parameters(lptIndex, subdivisionCount) syntax structure exists in the mesh patch with index patchIdx of the current atlas tile, and the tile ID is equal to tileID. If mdu_transform_parameters_override_flag[ tileID ][ patchIdx ] is absent, its value is inferred to be equal to 0.
[0358] The variable LtpIndex in the current patch is derived as follows:
[0359] LtpIndex = mdu_transform_parameters_override_flag[ tileID ][ patchIdx ] ? 2:
[0360] afve_transform_parameters_enable_flag ? 1: 0
[0361] mdu_subdivision_method[ tileID ][ patchIdx ] represents the identifier of the method for subdividing a mesh associated with the mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID. If mdu_subdivision_method[ tileID ][ patchIdx ] is absent, its value is inferred to be equal to afve_subdivision_method.
[0362] mdu_subdivision_iteration_count[ tileID ][ patchIdx ] represents the number of iterations used for subdivision in the mesdhpatch with index patchIdx in the current atlas tile, where tile ID is equal to tileID. If mdu_subdivision_iteration_count[ tileID ][ patchIdx ] is absent, its value is inferred to be 0 if mdu_subdivision_method[ tileID ][ patchIdx ] is 0, otherwise its value is inferred to be equal to afve_subdivision_iteration_count.
[0363] mdu_displacement_coordinate_system[ tileID ][ patchIdx ] represents the coordinate system identifier of the mesh subpart associated with the mesh patch whose index patchIdx has the same tile ID as tileID in the current atlas tile. Table 368 lists the supported displacement coordinate systems and their relationship to mdu_displacement_coordinate_system[ tileID ][ patchIdx ].
[0364] mdu_displacement_coordinate_system[ tileID ][ patchIdx ]Name of displacement coordinate system0CANNONICAL1LOCAL
[0365] mdu_transform_method[ tileID ][ patchIdx ] represents the identifier of the transformation applied to the displacement associated with the mesh patch with index patchIdx in the current atlas tile, where tileID is equal to tileID. If mdu_transform_method[ tileID ][ patchIdx ] is absent, its value is inferred to be equal to afve_transform_method.
[0366] mdu_attributes_2d_pos_x[ tileID ][ patchIdx ][ i ] specifies the x-coordinate of the upper-left corner of the attribute bounding box of the mesh patch for the current mesh patch with index patchIdx. This coordinate is for the attribute signaled in the attribute video data unit with tile ID equal to tileID and index i, expressed as a multiple of PatchPackingBlockSize. If mdu_attributes_2d_pos_x[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to 0.
[0367] mdu_attributes_2d_pos_y[ tileID ][ patchIdx ][ i ] specifies the y-coordinate of the upper-left corner of the mesh patch attributes bounding box of the current mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID and is expressed as a multiple of the PatchPackingBlockSize for the attribute signaled in the attribute video data unit with index i. If mdu_attributes_2d_pos_y[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be 0.
[0368] mdu_attributes_2d_size_x_minus1[ tileID ][ patchIdx ][ i ] plus 1 specifies the width value of the attribute bounding box of the meshpatch for the meshpatch of the current atlas tile with index patchIdx. This value is the value for the attribute signaled in the attribute video data unit with tile ID equal to tileID and index i. If mdu_attributes_2d_size_x_minus1[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to asve_attribute_frame_width[ i ] - 1.
[0369] mdu_attributes_2d_size_y_minus1[ tileID ][ patchIdx ][ i ] plus 1 specifies the height value of the bounding box of the mesh patch attribute of the mesh patch with index patchIdx. In the current atlas tile, the tile ID is equal to tileID and specifies the attribute signaled in the attribute video data unit with index i. If mdu_attributes_2d_size_y_minus1[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to asve_attribute_frame_height[ i ] - 1.
[0370]
[0371] Skip MeshPatch Data Unit Semantics:
[0372] Merge MeshPatch Data Unit Semantics:
[0373] mmdu_ref_index[ tileID ][ p ] specifies the atlas reference frame index, refIdx, for the current mesh patch with index p in the current atlas tile. The tile ID is equal to tileID. The value of mmdu_ref_index[ tileID ][ p ] ranges from 0 to NumRefIdxActive - 1. If mmdu_ref_index[ tileID ][ p ] is absent, it is inferred to be equal to 0.
[0374] mmdu_patch_index[ tileID ][ p ] specifies the index, RefPatchIdx, of the mesh patch with index p in the atlas tile with an ID equal to the current tile address of the atlas frame corresponding to index refIdx in the current reference atlas frame list. The tile ID is equal to tileID.
[0375] Inter-MeshPatchData Unit Semantics:
[0376] imdu_ref_index[ tileID ][ p ] specifies the atlas reference frame index, refIdx, for the current meshpatch with index p in the current atlas tile. The tile ID is equal to tileID. The value of imdu_ref_index[ tileID ][ p ] ranges from 0 to NumRefIdxActive - 1. If imdu_ref_index[ tileID ][ p ] is absent, it is inferred to be equal to 0.
[0377] imdu_patch_index[ tileID ][ p ] specifies the index, RefPatchIdx, of the meshpatch with index p in the atlas tile with an ID equal to the current tile address of the atlas frame corresponding to index refIdx in the current reference atlas frame list.
[0378] imdu_delta_vertex_count_minus1[ tileID ][ p ] specifies the difference between the number of vertices in the mesh patch with index p whose tile ID is equal to tileID in the current atlas tile and the number of vertices in the mesh patch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx.
[0379] imdu_delta_face_count_minus1[ tileID ][ p ] specifies the difference between the number of faces in the mesh patch with index p whose tile ID is equal to tileID in the current atlas tile and the number of faces in the mesh patch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx.
[0380] imdu_2d_delta_pos_x[ tileID ][ p ] specifies the difference between the x-coordinate of the upper-left corner of the meshpatch bounding box of the meshpatch with index p in the current atlas tile and the x-coordinate of the upper-left corner of the meshpatch bounding box of the meshpatch with index RefPatchIdx in the atlas tile with tile ID equal to tileID. This value is expressed as a multiple of the PatchPackingBlockSize in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx.
[0381] imdu_2d_delta_pos_y[ tileID ][ p ] is the difference between the y-coordinate of the upper left corner of the meshpatch bounding box of the meshpatch with index p in the current atlas tile whose tile ID is equal to tileID, and the y-coordinate of the upper left corner of the meshpatch bounding box of the meshpatch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx, expressed as a multiple of PatchPackingBlockSize.
[0382] imdu_2d_delta_size_x[ tileID ][ p ] specifies the difference between the width value of the mesh patch with index p in the current atlas tile whose tile ID is equal to tileID and the mesh patch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx.
[0383] imdu_2d_delta_size_y[ tileID ][ p ] specifies the difference between the height value of the mesh patch with index p in the current atlas tile whose tile ID is equal to tileID and the height value of the mesh patch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame corresponding to reference refIdx.
[0384] imdu_attributes_2d_delta_pos_x[ tileID ][ p ][ i ] is the difference between the x-coordinate of the upper left corner of the bounding box of the ith attribute meshpatch of the meshpatch with index p in the current atlas tile whose tile ID is equal to tileID, and the x-coordinate of the upper left corner of the bounding box of the ith attribute meshpatch of index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx, expressed as a multiple of PatchPackingBlockSize.
[0385] imdu_attributes_2d_delta_pos_y[ tileID ][ p ][ i ] is the difference between the y-coordinate of the upper left corner of the bounding box of the ith attribute meshpatch of the meshpatch with index p, whose tile ID is equal to tileID in the current atlas tile, and the y-coordinate of the upper left corner of the bounding box of the ith attribute meshpatch of the meshpatch with index RefPatchIdx, in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx, expressed as a multiple of PatchPackingBlockSize.
[0386] imdu_attributes_2d_delta_size_x[ tileID ][ p ][ i ] specifies the difference between the width value of the ith attribute meshpatch with index p in the current atlas tile whose tile ID is equal to tileID, and the width value of the meshpatch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame associated with reference refIdx.
[0387] imdu_attributes_2d_delta_size_y[ tileID ][ p ][ i ] specifies the difference between the height value of the ith attribute meshpatch with index p in the current atlas tile whose tile ID is equal to tileID, and the height value of the meshpatch with index RefPatchIdx in the atlas tile with the same ID as the current tile in the atlas frame corresponding to reference refIdx.
[0388] Texture projection information:
[0389] If tpi_face_id_present_flag[ tileID ][ patchIdx ] is 1, then the syntax element si_face_id[ tileId ][ patchIdx ][ i ] specifies that the subpatch with index i, the mesh patch with index patchIdx, and the tile ID in the current atlas tile are at the same location as tileID. If tpi_face_id_present_flag[ tileID ][ patchIdx ] is 0, then si_face_id[ tileId ][ patchIdx ][ i ] is absent. If tpi_face_id_present_flag[ tileID ][ patchIdx ] is absent, then it is assumed to be 0.
[0390] tpi_frame_scale[ tileID ][ patchIdx ] represents the frame scale value for the mesh patch with index patchIdx in the current atlas tile whose tile ID is equal to tileID. The variable TexCoordProjectionFrameScale[ tileID ][ patchIdx ] is used to derive texture coordinates from the geometry projection and is defined as follows.
[0391] TexCoordProjectionFrameScale[ tileID ][ patchIdx ] = tpi_frame_scale[ tileID ][ patchIdx ]
[0392] tpi_subpatch_count_minus1[ tileID ][ patchIdx ] plus 1 represents the number of subpatches in the mesh patch with tile ID equal to tileID for the mesh patch with index patchIdx in the current atlas tile.
[0393] Sub-patch information semantics:
[0394] si_face_id[ tileID ][ patchIdx ][ i ] specifies the face ID for the sub-patch with index i of the mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID and the sub-mesh index is equal to smIdx. If si_face_id[ tileID ][ patchIdx ][ i ] is not present, it is assumed to be equal to i.
[0395] si_projection_id[ tileID ][ patchIdx ][ i ] specifies the projection mode and normal index values for the projection plane for the sub-patch in the mesh patch with index patchIdx in the current atlas tile. The tile ID is equal to tileID and is the value for the sub-patch with index i. The value of si_projection_id[ tileID ][ patchIdx ][ i ] is in the range 0 to asps_max_number_projections_minus1. The number of bits used to represent si_projection_id[ tileID ][ patchIdx ][ i ] is Ceil( Log2( asps_max_number_projections_minus1 + 1) ). If si_projection_id[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to 0.
[0396] si_orientation_id[ tileID ][ patchIdx ][ i ] specifies the orientation index of a sub-patch in the mesh patch with index patchIdx of the current atlas tile. The tile ID for sub-patch index i is equal to tileID, which is used to determine the sub-patch rotation homography transformation used to transform the 3D space coordinates of the vertex to texture coordinates (u, v) as shown in Table 11. If si_orientation_id[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to 0.
[0397] si_2d_pos_x[ tileID ][ patchIdx ][ i ] specifies the x-coordinate of the upper-left corner of the sub-patch bounding box size of the sub-patch in the meshpatch with index patchIdx of the current atlas tile. The tile ID is equal to tileID for the sub-patch with index i. If si_2d_pos_x[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to 0.
[0398] si_2d_pos_y[ tileID ][ patchIdx ][ i ] specifies the y-coordinate of the upper-left corner of the sub-patch bounding box size for the sub-patch in the mesh patch with index patchIdx of the current atlas tile. The tile ID is equal to tileID and is for the sub-patch with index i. If si_2d_pos_y[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to 0.
[0399] si_2d_size_x_minus1_diff[ tileID ][ patchIdx ][ i ] specifies the difference between the width of the subpatch bounding box size of the meshpatch with index patchIdx and the width of the subpatch with tile ID equal to tileID and submesh index equal to smIdx in the current atlas tile. Applies to subpatches with index i and index i - 1. If i = 0, adding 1 to si_2d_size_x_minus1_diff[ tileID ][ patchIdx ][ i ] specifies the width of the subpatch bounding box size. If si_2d_size_x_minus1[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to TexCoordProjectionWidth[ smIdx ] - 1 for the submesh with index smIdx.
[0400] si_2d_size_y_minus1_diff[ tileID ][ patchIdx ][ i ] specifies the difference between the height of the sub-patch bounding box size of the sub-patch in the mesh patch with index patchIdx of the current atlas tile and the height of the sub-patch bounding box size of the sub-patch associated with index i whose tile ID is equal to tileID. If i = 0, adding 1 to si_2d_size_y_minus1_diff[ tileID ][ patchIdx ][ i ] specifies the height of the sub-patch bounding box size. If si_2d_size_y_minus1[ tileID ][ patchIdx ][ i ] is absent, its value is inferred to be equal to TexCoordProjectionHeight[ smIdx ] - 1 for the sub-mesh with index smIdx.
[0401] If si_scale_present_flag[ tileID ][ patchIdx ][ i ] is 1, it specifies that sub-patch scale is present for the sub-patch in the mesh patch with index patchIdx of the current atlas tile whose tile ID is equal to tileID for the sub-patch with index i. If si_scale_present_flag[ tileID ][ patchIdx ][ i ] is 0, there is no sub-patch scale information. If si_scale_present_flag[ tileID ][ patchIdx ][ i ] is not present, it is assumed to be 0.
[0402] If si_scale_power_factor[ tileID ][ patchIdx ][ i ] is present, it specifies the scaling power factor for a subpatch in the mesh patch with index patchIdx of the current atlas tile. For the subpatch SubpatchScalingFactor[ tileID ][ patchIdx ][ i ], where tileID is equal to tileID and index i is:
[0403] SubpatchScalingFactor[ tileID ][ patchIdx ][ i ] = FrameScale[ tileID ][ patchIdx ]
[0404] for( i = 0; i <= si_scale_power_factor[ tileID ][ patchIdx ][ i ]; i++)
[0405] SubpatchScalingFactor[ tileID ][ patchIdx ][ i ] *= TextCoordProjectionScaleFactor
[0406] If si_scale_power_factor[ tileID ][ patchIdx ][ i ] is absent, SubpatchScalingFactor[ tileID ][ patchIdx ][ i ] is equal to FrameScale[ tileID ][ patchIdx ].
[0407] Arthmetic Coded Displacement sub-bitstream:
[0408] A NAL sample stream format can be constructed from a NAL unit stream format by arranging NAL units in decoding order and prefixing each NAL unit with a header that specifies the exact size (in bytes) of the NAL unit. A sample stream header is included at the beginning of the sample stream bitstream, which specifies the precision (in bytes) of the signaled NAL unit size. A NAL unit stream format can be extracted from a sample stream format by traversing the sample stream format, reading the size information, and appropriately extracting each NAL unit.
[0409] Below, the syntax of the Arthmetic Coded Displacement sub-bitstream is described.
[0410] NAL unit syntax:
[0411] General NAL unit syntax:
[0412] displ_nal_unit( NumBytesInNalUnit ) {Descriptordispl_nal_unit_header( )NumBytesInRbsp = 0for( i = 2; i < NumBytesInNalUnit; i++ )rbsp_byte[ NumBytesInRbsp++ ]b(8)}
[0413] NAL unit header syntax:
[0414] displ_nal_unit_header() {Descriptordispl_nal_forbidden_zero_bitf(1)displ_nal_unit_typeu(6)displ_nal_layer_idu(6)displ_nal_temporal_id_plus1u(3)}
[0415] Low byte sequence payload, trailing bits, and byte alignment syntax:
[0416] Displacement Sequence Parameter Set RBSP Syntax:
[0417] General Displacement Sequence Parameter Set RBSP Syntax:
[0418] displ_sequence_parameter_set_rbsp( ) {Descriptordsps_sequence_parameter_set_idu(4)dsps_codec_idu(8)dsps_profile_tier_level( )dsps_range_log2_minus2u(3)dsps_single_dimension_flagu(1)dsps_msb_align_flagu(1)dsps_log2_max_displ_frame_order_cnt_lsb_minus4ue(v)dsps_max_dec_displ_frame_buffering_minus1ue(v)dsps_long_term_ref_displ_frames_flagu(1)dsps_num_ref_displ_frame_lists_in_dspsue(v)for( i = 0; i < dsps_num_ref_displ_frame_lists_in_dsps; i++ ) displ_ref_list_struct( i )dsps_extension_present_flagu(1)if( dsps_extension_present_flag ) {dsps_extension_count_minus1u(7)dsps_extension_length_minus1ue(v)while( more_rbsp_data( ) )dsps_extension_data_byte u(1)}rbsp_trailing_bits( )}
[0419] 디스플레이스먼트 레이어 RBSP 신택스:
[0420] displ_layer_rbsp( ) {Descriptordispl_header( )displ_data_unit( displID )rbsp_trailing_bits( )}
[0421] 디스플레이스먼트 헤더 신택스:
[0422] displ_header( ) {Descriptorif( nal_unit_type >= NAL_BLA_W_LP && nal_unit_type <= NAL_RSV_IRAP_DCL_29 )dh_no_output_of_prior_displ_frames_flagu(1)dh_frame_parameter_set_idu(4)dh_idu(v)displID = dh_iddh_typeue(v)if( dfps_output_flag_present_flag )dh_output_flagu(1)dh_frm_order_cnt_lsbu(v)if( dsps_num_ref_displ_frame_lists_in_dsps > 0 )dh_ref_displ_frame_list_dsps_flagu(1)if( dh_ref_displ_frame_list_dsps_flag == 0 )displ_ref_list_struct( dsps_num_ref_displ_frame_lists_in_dsps )else if( dsps_num_ref_displ_frame_lists_in_dsps > 1 )ref_displ_frame_list_idxu(v)for( j = 0; j < NumLtrDisplFrmEntries[ RlsIdx ]; j++ ) {dh_additional_dfoc_lsb_present_flag[ j ]u(1)if( additional_dfoc_lsb_present_flag[ j ] )dh_additional_dfoc_lsb_val[ j ]u(v)}if( dh_type == P_DISPLACEMENT && num_ref_entries[ RlsIdx ] > 1 ) {dh_num_ref_idx_active_override_flagu(1)if( num_ref_idx_active_override_flag )dh_num_ref_idx_active_minus1ue(v)}dh_log2_subblock_size_minus6u(4)byte_alignment( )}
[0423] Displacement Data Unit Syntax:
[0424] displ_data_unit( displID ) {Descriptorif( dh_type == I_DISPLACEMENT ) {displ_intra_unit( displID )}else if( dh_type == P_DISPLACEMENT ) {displ_inter_unit( displID )}}
[0425] Below, the semantics of the Arthmetic Coded Displacement sub-bitstream are described.
[0426] NAL unit semantics:
[0427] General NAL unit semantics:
[0428] NumBytesInNalUnit represents the size of a NAL unit in bytes. This value is required for decoding the NAL unit. Some form of delimiting NAL unit boundaries is required to enable inference of NumBytesInNalUnit.
[0429] Note: The displacement coding layer (DCL) is specified to efficiently represent the contents of displacement data. The NAL is specified to format that data and provide header information in a manner suitable for transmission over various communication channels or storage media. All data is contained in NAL units, each of which contains an integer number of bytes. A NAL unit represents a common format that can be used in both packet-oriented systems and bitstream systems. The format of NAL units for both packet-oriented transport and sample streams is the same, but in sample stream formats, each NAL unit may be preceded by an additional element specifying its size.
[0430] rbsp_byte[ i ] is the ith byte of the RBSP. An RBSP is specified as an ordered sequence of bytes as follows:
[0431] An RBSP contains a string of data bits (SODBs) as follows:
[0432] - If the SODB is empty (i.e., has a length of 0 bits), the RBSP is also empty.
[0433] - Otherwise, the RBSP contains the SODB as follows:
[0434] 3) The first byte of the RBSP contains the first (most significant, leftmost) 8 bits of the SODB. The next byte of the RBSP contains the next 8 bits of the SODB, and so on until there are fewer than 8 bits of the SODB left.
[0435] 4) The rbsp_trailing_bits( ) syntax structure exists after SODB as follows.
[0436] iv) The first (most significant, leftmost) bit of the last RBSP byte contains the remaining bits of the SODB (if any).
[0437] v) The next bit consists of a single bit equal to 1 (i.e. rbsp_stop_one_bit).
[0438] vi) If rbsp_stop_one_bit is not the last bit of a byte that is aligned, byte alignment is achieved by the presence of one or more bits that are equal to 0 (i.e., instances of rbsp_alignment_zero_bit).
[0439] Syntax structures with these RBSP properties are indicated in the syntax table using the suffix "_rbsp". These structures are conveyed within a NAL unit as the contents of the rbsp_byte[ i ] data byte. The association between RBSP syntax structures and NAL units is shown in the table below.
[0440] Note: If the boundaries of an RBSP are known, a decoder can extract an SODB from an RBSP by concatenating the byte bits of the RBSP, discarding the rbsp_stop_one_bit whose last (least significant, rightmost) bit is 1, and discarding the following (less significant, rightmost) bits that are 0. The data required for the decoding process is contained in the SODB portion of the RBSP.
[0441] NAL unit header semantics:
[0442] displ_nal_forbidden_zero_bit
[0443] displ_nal_unit_type
[0444] As with the atlas case, a similar NAL unit type is defined for displacement, and similar functionality for random access defines a specific NAL unit corresponding to the coded displacement data. Additionally, NAL units that can contain metadata, such as SEI messages, are also defined.
[0445] The order of NAL units and displacement frames, and their relationship to coded displacement frames, access units, and coded displacement sequences:
[0446] Low byte sequence payload, trailing bits, byte alignment semantics:
[0447] Displacement Sequence Parameter Set RBSP Semantics:
[0448] General displacement sequence parameter set RBSP semantics:
[0449] dsps_sequence_parameter_set_id: An identifier for the displacement sequence parameter set so that other syntax elements can reference it.
[0450] dsps_codec_id: The identifier of the codec used to compress the displacement. dsps_codec_id is in the range 0 to 255. This codec may be identified by a profile defined in ISO / IEC 23090-29, an SEI message mapping component codec, or by means external to this document. It may be associated with a specific displacement codec through a profile specified in that specification, or explicitly indicated in an SEI message, as is done in the V3C specification for video sub-bitstreams.
[0451] dsps_range_log2_minus2: Adding 2 to this value gives the geometric displacement coordinate range of the displacement. dsps_range_log2_minus2 is in the range of 0 to 3.
[0452] dsps_single_dimension_flag: Indicates the number of dimensions for the displacement associated with the displacement. If dsps_single_dimension_flag is 0, it indicates that three components are used for the displacement. If dsps_single_dimension_flag is 1, it indicates that only the normal component is used for the displacement.
[0453] dsps_msb_align_flag: Indicates how decoded displacement samples are converted to samples of displacement range bit depth.
[0454] dsps_log2_max_displ_frame_order_cnt_lsb_minus4: Adding 4 to this value gives the values of the variables Log2MaxDisplFrmOrderCntLsb and MaxDisplFrmOrderCntLsb used in the decoding process for the displacement frame order count, as follows:
[0455] Log2MaxDisplFrmOrderCntLsb =
[0456] dsps_log2_max_displ_frame_order_cnt_lsb_minus4 + 4
[0457] MaxDisplFrmOrderCntLsb = 2Log2MaxDisplFrmOrderCntLsb
[0458] The value of dsps_log2_max_displ_frame_order_cnt_lsb_minus4 is in the range of 0 to 12.
[0459] dsps_max_dec_displ_frame_buffering_minus1 plus 1 indicates the maximum required size of the decoded disparity frame buffer for CDS, in disparity frame storage buffer units. The value of dsps_max_dec_displ_frame_buffering_minus1 is in the range 0 to 15.
[0460] If dsps_long_term_ref_displ_frames_flag is 0, it indicates that long-term reference disparity is not used for inter prediction of any coded disparity frames in the CDS. If dsps_long_term_ref_displ_frames_flag is 1, it indicates that long-term reference disparity frames can be used for inter prediction of one or more coded disparity frames in the CDS.
[0461] dsps_num_ref_displ_frame_lists_in_dsps represents the number of displ_ref_list_struct(rlsIdx) syntax structures included in the displacement sequence parameter set. The value of dsps_num_ref_displ_frame_lists_in_dsps is in the range of 0 to 64.
[0462] NOTE: The decoder allocates memory for a total number of displ_ref_list_struct(rlsIdx) syntax structures equal to (dsps_num_ref_displ_frame_lists_in_dsps + 1), since there can be only one displ_ref_list_struct(rlsIdx) syntax structure signaled directly in the displacement header of the current displacement frame.
[0463] If dsps_extension_present_flag is 1, it indicates that dsps_extension_count_minus1 and dsps_extension_length_minus1 are in the displacement sequence parameter set.
[0464] Adding 1 to dsps_extension_count_minus1 indicates the number of extensions in the current displacement sequence parameter set. If none exist, dsps_extension_count_minus1 is inferred to be equal to -1.
[0465] dsps_extension_length_minus1 plus 1 represents the length of the dsps_extension_data_byte element following this syntax element. If not present, dsps_extension_length_minus1 is inferred to be equal to -1.
[0466] dsps_extension_data_byte can have any value.
[0467] Displacement Frame Parameter Set RBSP Semantics
[0468] Displacement Layer RBSP Semantics:
[0469] Displacement header semantics:
[0470] The dh_no_output_of_prior_displ_frames_flag affects the output of previously decoded displacement frames in the DDB after decoding displacement frames in a CDS AU that is not the first AU in the bitstream, as specified in ISO / IEC 23090-29. If no_output_of_prior_displ_frames_flag is absent, its value is inferred to be 0.
[0471] As a requirement for bitstream conformance, the value of no_output_of_prior_displ_frames_flag is the same for all displacement frames in an AU.
[0472] The no_output_of_prior_displ_frames_flag value in the displacement header is the output_of_prior_displ_frames_flag value in the AU.
[0473] dh_frame_parameter_set_id indicates the dfps_displ_frame_parameter_set_id value for the active displacement frame parameter set for the current displacement frame.
[0474] dh_id: Indicates the displacement header ID.
[0475] dh_type indicates the coding type of the current displacement frame, as shown in the table below. The value of dh_type is 0 or 1 in bitstreams that follow this version of this document. Other values of dh_type are reserved for future use in ISO / IEC.
[0476] dh_type relationship
[0477] smh_typeName of smh_type0P_DISPLACEMENT1I_DISPLACEMENT2...RESERVED
[0478] dh_frm_order_cnt_lsb: Indicates the number of displacement frame orders modulo MaxDisplFrmOrderCntLsb for the current displacement frame. The length of the dh_frm_order_cnt_lsb syntax element is equal to Log2MaxDisplFrmOrderCntLsb bits. The value of dh_frm_order_cnt_lsb ranges from 0 to MaxDisplFrmOrderCntLsb - 1.
[0479] If dh_ref_displ_frame_list_dsps_flag is 1, it indicates that the reference displacement frame list of the current displacement frame is derived based on one of the displ_ref_list_struct(rlsIdx) syntax structures of the active DSPS. If dh_ref_displ_frame_list_dsps_flag is 0, it indicates that the reference displacement frame list of the current displacement frame is derived based on the displ_ref_list_struct(rlsIdx) syntax structure directly included in the displacement frame header of the current displacement frame. If dsps_num_ref_displ_frame_lists_in_dsps is 0, the value of dh_ref_displ_frame_list_dsps_flag is inferred to be 0.
[0480] dh_ref_displ_frame_list_idx specifies the index into the list of displ_ref_list_struct(rlsIdx) syntax structures contained in the active DSPS, which are used to derive the reference displ_ref_list_struct(rlsIdx) syntax structures of the current displ_ref_list frame. The syntax element dh_ref_displ_frame_list_idx is expressed in bits as Ceil(Log2(dsps_num_ref_displ_frame_lists_in_dsps)). If not present, the value of dh_ref_displ_frame_list_idx is inferred to be equal to 0. The value of dh_ref_displ_frame_list_idx is in the range 0 to dsps_num_ref_displ_frame_lists_in_dsps - 1, inclusive. If dh_ref_displ_frame_list_dsps_flag is 1 and dsps_num_ref_displ_frame_lists_in_dsps is 1, the value of ref_displ_frame_list_idx is inferred to be 0.
[0481] The variable RlsIdx of the current atlas tile is derived as follows:
[0482] RlsIdx = dh_ref_displ_frame_list_dsps_flag ?
[0483] ref_displ_frame_list_idx : dsps_num_ref_displ_frame_lists_in_dsps
[0484] If dh_additional_dfoc_lsb_present_flag[ j ] is 1, it indicates that dh_additional_dfoc_lsb_val[ j ] is present in the current displacement frame. If dh_additional_dfoc_lsb_present_flag[ j ] is 0, it indicates that dh_additional_dfoc_lsb_val[ j ] is not present.
[0485] dh_additional_dfoc_lsb_val[ j ] represents the value of FullFrmOrderCntLsbLt[ RlsIdx ][ j ] for the current atlas tile, as follows:
[0486] FullDisplFrmOrderCntLsbLt[ RlsIdx ][ j ] = dh_additional_dfoc_lsb_val[ j ] *
[0487] MaxDisplFrmOrderCntLsb +dfoc_lsb_lt[ RlsIdx ][ j ]
[0488] The syntax element dh_additional_dfoc_lsb_val[ j ] is represented by the dfps_additional_lt_dfoc_lsb_len bits. If it is not present, the value of dh_additional_dfoc_lsb_val[ j ] is inferred to be equal to 0.
[0489] If dh_num_ref_idx_active_override_flag is 1, it indicates that the syntax element num_ref_idx_active_minus1 exists for the current displacement frame. If dh_num_ref_idx_active_override_flag is 0, it indicates that the syntax element num_ref_idx_active_minus1 does not exist. If dh_num_ref_idx_active_override_flag is absent, its value is inferred to be equal to 0.
[0490] dh_num_ref_idx_active_minus1 is used to derive the variable NumRefIdxActive for the current displacement frame. The value of dh_num_ref_idx_active_minus1 is in the range 0 to 14.
[0491] If the current displacement frame is a P_DISPLACEMENT displacement frame, dh_num_ref_idx_active_override_flag is 1, and dh_num_ref_idx_active_minus1 is absent, then dh_num_ref_idx_active_minus1 is inferred to be equal to 0.
[0492] The variable NumRefIdxActive is derived as follows:
[0493] if( dh_type == P_DISPLACEMENT ) {
[0494] if( dh_num_ref_idx_active_override_flag == 1 )
[0495] NumRefIdxActive = dh_num_ref_idx_active_minus1 + 1
[0496] else {
[0497] if( num_ref_entries[ RlsIdx ] >= dfps_num_ref_idx_default_active_minus1 + 1 )
[0498] NumRefIdxActive = dfps_num_ref_idx_default_active_minus1 + 1
[0499] else
[0500] NumRefIdxActive = num_ref_entries[RlsIdx]
[0501] }
[0502] }
[0503] else
[0504] NumRefIdxActive = 0
[0505] Subtracting 1 from NumRefIdxActive indicates the maximum number of displacement reference frame indices that can be used to decode the current displacement frame.
[0506] Adding 6 to dh_log2_subblock_size_minus6 gives the value of the variable subblockSize as follows:
[0507] subblockSize = 1 << ( log2_subblock_size_minus6 + 6 )
[0508] Displacement Data Unit Semantics:
[0509] displ_intra_unit( displID ) contains a displacement unit stream, which is an aligned stream of bytes or bits, within which the location of unit boundaries can be identified from a pattern in the data. The format of this displacement unit stream is identified by the 4CC code defined by dptl_profile_codec_group_idc or by the component codec mapping SEI message.
[0510] displ_inter_unit( displID ) contains a displacement unit stream, which is an aligned stream of bytes or bits, within which the location of unit boundaries can be identified from a pattern in the data. The format of this displacement unit stream is identified by the 4CC code defined by dptl_profile_codec_group_idc or by the component codec mapping SEI message.
[0511] Figure 17 shows the structure of a scene description system according to embodiments.
[0512] Scene Description (ISO / IEC 23090-14) specifies extensions to the existing scene description format to support MPEG media, particularly immersive media. MPEG media includes, but is not limited to, media encoded with MPEG codecs, media stored in MPEG containers, MPEG media and application formats, and media delivered via MPEG delivery mechanisms. The extensions include the scene description format syntax and semantics, and a processing model for presentation engines when using these extensions. It also defines the Media Access Function (MAF) API for communication between presentation engines and media access functions for these extensions. While the extensions defined in this document can be applied to other scene description formats, the functionality is defined as an extension to, for example, the glTF format defined in ISO / IEC 12113.
[0513] Scene descriptions are used by the Presentation Engine to render 3D scenes to viewers. The extensions defined in this document enable the use of timed media to create immersive experiences. The Scene Description extension is designed to decouple the Presentation Engine from the Media Access Function. The Presentation Engine and the Media Access Function communicate via the Media Access Function API, which allows the Presentation Engine to request timed media required for scene rendering. The Media Access Function retrieves the requested timed media and provides it in a format that the Presentation Engine can immediately process. For example, the requested timed media asset may be compressed and reside on the network. The Media Access Function retrieves the asset, decodes it, and passes the resulting decoded media data to the Presentation Engine for rendering. The decoded media data is then passed from the Media Access Function to the Presentation Engine in a buffer. Requests for timed media are passed from the Presentation Engine to the Media Access Function via the Media Access Function API. Table 52 lists several extensions related to MPEG media.
[0514] Extension NameExtension NameTypeSubclauseMPEG_mediaExtension for referencing external media sources.GenericMPEG_accessor_timedAn accessor extension to support timed media.GenericMPEG_buffer_circularA buffer extension to support circular buffers.GenericMPEG_scene_dynamicAn extension to support dynamic scenes.GenericMPEG_texture_videoA texture extension to support video textures.VisualMPEG_mesh_linkingAn extension to link two meshes and provide mapping informationVisualMPEG_mesh_linkingAdds support for spatial audioAudioMPEG_viewport_recommendedAn extension to describe a recommended viewport.MetadataMPEG_animation_timingAn extension to control animation timelines.Metadata
[0515] The definitions for each extension name are as follows:
[0516] MPEG_media: Extension for referencing external media sources.
[0517] MPEG_accessor_timed: Accessor extension supporting timed media.
[0518] MPEG_buffer_circular: Buffer extension supporting circular buffers.
[0519] MPEG_scene_dynamic: Extension supporting dynamic scenes.
[0520] MPEG_texture_video: Texture extension supporting video textures.
[0521] MPEG_mesh_linking: Extension that links two meshes and provides mapping information.
[0522] MPEG_audio_spatial: Adds spatial audio support.
[0523] MPEG_viewport_recommended: Extension describing the recommended viewport.
[0524] MPEG_animation_timing: Extension for controlling animation timelines.
[0525] Definition of top-level objects in MPEG media extensions:
[0526] NameTypeDefaultUsageUsagemediaarrayN / AMan array of items that describe the external media, referenced in this scene description document.
[0527] Media is an array of items that describe external media referenced in the scene description.
[0528] Definition of items in the media array of MPEG media extensions:
[0529] NameTypeUsageDefaultDescriptionnamestringON / AThe user-defined name of the media.ON / AThe startTime gives the time at which the rendering of the timed media will begin. The value is provided in seconds.In the case of timed textures, the static image should be rendered as a texture until the startTime is reached. A startTime of 0 means the presentation time of the current scene.startTimeOffsetnumber00The startTimeOffset indicates the time offset into the source, starting from which the timed media shall be generated. The value is provided in seconds, where 0 corresponds to the start of the source.endTimeOffsetnumber0N / AThe endTimeOffset indicates the end time offset into the source, up to which the timed media shall be generated. The value is provided in seconds. If not present, the endTimeOffset corresponds to the end of the source media.autoplaybooleanOFalseWhen present and set to True, it specifies that the media will start playing as soon as it is ready.When present and set to True, startTime shall not be present or shall be ignored by the client.Rendering of all media for which the autoplay flag is set to True should happen simultaneously.When set to False, it indicates that the media may be controlled by the functions in Scene Description such as interactivity.autoplayGroupintegerON / AAll media that have the same autoplayGroup identifier shall start playing synchronously as soon as all autoplayGroup media are ready.autoplayGroup is only allowed if autoplay is set to True.loopbooleanOFalseSpecifies that the media will start over again, every time it is finished. The timestamp in the buffer shall be continuously increasing when the media source loops, i.e. the playback duration prior to looping shall be added to the media time after looping.controlsbooleanOFalseSpecifies that media controls should be displayed (such as a play / pause button etc).alternativesarrayMAn array of items that indicate alternatives of the same media (e.g.different video codecs used)NOTE The client can select items (ie U and track) included in alternatives depending on the client's capability.
[0530] startTime: startTime provides the time at which timed media rendering begins. The value is provided in seconds.
[0531] For timed textures, static images are rendered to the texture until the startTime is reached. If startTime is 0, it refers to the presentation time of the current scene.
[0532] startTimeOffset: startTimeOffset indicates the time offset from the source where the timed media is generated. The value is given in seconds, with 0 corresponding to the beginning of the source.
[0533] endTimeOffset: endTimeOffset represents the end time offset from the source where the timed media is generated. The value is given in seconds. If not present, endTimeOffset corresponds to the end of the source media.
[0534] autoplay: If this value exists and is set to True, it specifies that the media will begin playing as soon as it is ready. If this value exists and is set to True, startTime is either not present or is ignored by the client. All media with the autoplay flag set to True will be rendered simultaneously. If this value is set to False, it indicates that the media can be controlled by scene description features, such as interactivity.
[0535] autoplayGroup: All media with the same autoplayGroup identifier will begin playing synchronously as soon as all autoplayGroup media are ready.
[0536] autoplayGroup is only allowed if autoplay is set to True.
[0537] loop: Specifies that the media will restart whenever it ends. When the media source loops, the buffer's timestamp will continuously increase. The playback time before the loop is added to the media time after the loop.
[0538] controls: Specifies that media controls should be displayed (e.g., play / pause buttons).
[0539] alternatives: An array of items representing alternatives to the same media (e.g., a different video codec used).
[0540] Note that clients can select items included in the alternatives (e.g. U and Track) depending on their capabilities.
[0541] Definitions of items in the alternatives array of MPEG_media extension:
[0542] NameTypeDefaultUsageDescriptionmimeTypestringN / AMThe media's MIME type.The profiles parameter, as defined in IETF RFC 6381, may be included as a part of the mimeType to specify the profile of the media container. (e.g. the profiles parameter indicates the DASH profile when the uri specifies a DASH manifest)uristringN / AMThe uri of the media. Relative paths are relative to the .gltf file. If the reference media is a real-time media stream, then the uri shall follow the referencing scheme as specified inAnnex Cin ISO / IEC 23090-14. If the tracks element is present, the last part of the URI (i.e. the stream identifier such as the mid) is provided by the tracks information.tracksarrayN / AOAn array of items that lists the components of the referenced media source that are to be used. These can e.g. be a track number of an ISOBMFF, a DASH / CMAF SwitchingSet identifier, or a media id of an RTP stream.extraParamsobjectN / AOAn object that may contain any additional media-specific parameters.
[0543] The definition of each item name is as follows:
[0544] mimeType: The MIME type of the media.
[0545] The profiles parameter, defined in IETF RFC 6381, can be included as part of the mimeType to specify the profile of the media container. (For example, the profiles parameter indicates the DASH profile when the uri specifies a DASH manifest.)
[0546] uri: The URI of the media. Relative paths are relative to the .gltf file. If the referenced media is a live media stream, the URI follows the referencing scheme specified in Annex C of ISO / IEC 23090-14. If the tracks element is present, the last part of the URI (i.e., the stream identifier, such as mid ) is provided by the tracks information.
[0547] tracks: An array of items listing the components of the reference media source to use. For example, this could be a track number from an ISOBMFF, a DASH / CMAF SwitchingSet identifier, or a media ID from an RTP stream.
[0548] extraParams: An object that can contain additional media-specific parameters.
[0549] Definition of items in the track array of the MPEG Media Alternative Extension:
[0550] NameTypeDefaultUsageDescriptiontrackstringN / AMURL fragment to access the track within the media alternative.The URL structure is defined for the following formats:DASH: Using MPD Anchors (URL fragments) as defined in ISO / IEC 23009-1:2022, Annex C (Table C.1).ISOBMFF: URL fragments as specified in ISO / IEC 14496-12:2022, Annex C.SDP: stream identifier of the media stream as defined inAnnex Cin ISO / IEC 23090-14.When V3C data is referenced in the scene description document as in item in MPEG_media.alternative.tracks and the referenced item corresponds to an ISBOBMFF track, the following applies:-For single-track encapsulated V3C data, the referenced track in MPEG_media shall be the V3C bitstream track.-For multi-track encapsulated V3C data, the referenced track in MPEG_media shall be the V3C atlas track.When VDMC data is referenced in the scene description document as in item in MPEG_media.alternative.tracks and the referenced item corresponds to an ISBOBMFF track, the following applies:-For single-track encapsulated VDMC data, the referenced track in MPEG_media shall be the VDMC bitstream track.-For multi-track encapsulated VDMC data, the referenced track in MPEG_media shall be the VDMC atlas track or VDMC basemesh track.codecsstringN / AMThe codecs parameter, as defined in IETF RFC 6381, of the media included in the track. When the track includes different types of codecs (eg the AdaptationSet includes Representations with different codecs), the codecs parameter may be signaled by comma-separated list of values of the codecs.
[0551] The definition of each item is as follows:
[0552] track: A URL fragment to access a track within a media alternative.
[0553] The URL structure is defined for the following format:
[0554] DASH: Uses MPD anchors (URL fragments) as defined in ISO / IEC 23009-1:2022, Annex C (Table C.1).
[0555] ISOBMFF: URL fragment specified in ISO / IEC 14496-12:2022, Annex C.
[0556] SDP: Stream identifier of a media stream as defined in Annex C of ISO / IEC 23090-14.
[0557] When V3C data is referenced in a scene description, such as in an entry in MPEG_media.alternative.tracks, and the referenced entry corresponds to an ISBOBMFF track, the following applies:
[0558] - For single-track encapsulated V3C data, the referenced track of MPEG_media is the V3C bitstream track.
[0559] - For multi-track encapsulated V3C data, the reference track of MPEG_media is the V3C atlas track.
[0560] When VDMC data is referenced in a scene description document, such as an entry in MPEG_media.alternative.tracks, and the referenced entry corresponds to an ISBOBMFF track, the following applies:
[0561] - For single-track encapsulated VDMC data, the reference track of MPEG_media is the VDMC bitstream track.
[0562] - For multi-track encapsulated VDMC data, the reference track of MPEG_media is the VDMC atlas track or VDMC basemesh track.
[0563] Codec: The codec parameter defined in IETF RFC 6381 for the media contained in the track.
[0564] If a track contains multiple types of codecs (e.g., its AdaptationSet contains Representations with different codecs), the codec parameter can be signaled as a comma-separated list of codec values.
[0565] MPEG_accessor_timed extension definition:
[0566] NameTypeDefaultUsageDescriptionimmutablebooleanTrueOThis flag equal to false indicates the accessor information componentType, type, and normalize may change over time. The changing values of componentType, type and normalize are provided through accessor information header.This flag equal to true indicates the accessor information componentType, type, and normalize do not change over time and are not present in the accessor information header.bufferViewintegerN / AOThis property provides the index in the bufferViews array to a bufferView element that points to the timed accessor information header as described in section 4.4.6. byteLength field of the bufferView element indicates the size of the timed accessor information header. The buffer properties in the bufferView element shall point to the same buffer as the bufferView in the containing accessor object.In the absence of the bufferView attribute, it shall be assumed that the buffer has no dynamic header.In that case, the immutable flag shall be present and shall be set to True.suggestedUpdateRatenumber25.0OThe suggestedUpdateRate provides the frequency at which the Presentation Engine is recommended to poll the underlying buffer for new data. The rate is provided in number of changes per second.
[0567] The definition of each item is as follows:
[0568] immutable: If this flag is false, it indicates that the accessor information componentType, type, and normalize can change over time. The changing values of componentType, type, and normalize are provided through the accessor information header.
[0569] If this flag is true, it indicates that the accessor information componentType, type, and normalize do not change over time and are not present in the accessor information header.
[0570] bufferView: This property provides the index into the bufferViews array for the bufferView element that points to the timed accessor information header, as described in Section 4.4.6. The byteLength field of the bufferView element indicates the size of the timed accessor information header. The buffer property of the bufferView element points to the same buffer as the bufferView of the containing accessor object.
[0571] If the bufferView attribute is missing, the buffer is assumed to have no dynamic header. In this case, the immutable flag is present and set to True.
[0572] suggestedUpdateRate: suggestedUpdateRate specifies how often the presentation engine polls the underlying buffer to retrieve new data. This rate is expressed in changes per second.
[0573] Definition of the timed accessor information header fields:
[0574] timed_accessor_information_header() {Descriptortimestamp_deltaf(32) if (!immutable) { componentTypeu(32) typeu(8) normalizedu(1) reserved_zero_bitu(7)} byteOffsetu(32) countu(32) maxsize(componentType)*components minsize(componentType)*components bufferViewByteOffsetu(32) bufferViewByteLengthu(32) bufferViewByteStrideu(32)}u(32)
[0575] timestamp_delta: Provides a delta in seconds added to the timestamp field of the corresponding buffer frame of the referenced buffer to determine the timestamp of the referenced time media. If the accessor information header is absent, the value of timestamp_delta is inferred to be equal to 0. The sum of timestamp_delta and the timestamp field of the corresponding buffer frame is less than the timestamp field of any subsequent buffer frame in the buffer.
[0576] componentType: Corresponds to the accessor property componentType defined in ISO / IEC 12113.
[0577] type: This field corresponds to the accessor attribute type defined in ISO / IEC 12113 and is modified as follows:
[0578] If type is 0, it indicates SCALAR as defined in ISO / IEC 12113.
[0579] If type is 1, it indicates VEC2 as defined in ISO / IEC 12113.
[0580] If type is 2, it indicates VEC3 as defined in ISO / IEC 12113.
[0581] If type is 3, it indicates VEC4 as defined in ISO / IEC 12113.
[0582] If type is 4, it indicates MAT2 as defined in ISO / IEC 12113.
[0583] If type is 5, it indicates MAT3 as defined in ISO / IEC 12113.
[0584] If type is 6, it indicates MAT4 as defined in ISO / IEC 12113.
[0585] normalized: Corresponds to the accessor attribute normalized defined in ISO / IEC 12113.
[0586] byteOffset: Corresponds to the accessor property byteOffset defined in ISO / IEC 12113.
[0587] count: Corresponds to the accessor attribute count defined in ISO / IEC 12113.
[0588] max: Corresponds to the accessor attribute max as defined in ISO / IEC 12113. The max array size depends on the number of components defined in the type as defined in ISO / IEC 12113.
[0589] min: Corresponds to the accessor attribute min as defined in ISO / IEC 12113. The min array size depends on the number of components defined in the type as defined in ISO / IEC 12113.
[0590] bufferViewByteOffset: Corresponds to the bufferView property byteOffset defined in ISO / IEC 12113.
[0591] bufferViewByteLength: Corresponds to the bufferView property byteLength defined in ISO / IEC 12113.
[0592] bufferViewByteStride: Corresponds to the byteStride field of the bufferView property defined in ISO / IEC 12113. The size() function returns the number of bits for the specified componentType as defined in the Accessor Data Types table of ISO / IEC 12113.
[0593] The fields bufferViewByteOffset, bufferViewByteLength, and bufferViewByteStride update information in the bufferView referenced by an accessor that includes the MPEG_accessor_timed extension and provide a description of how to access the corresponding media data in the buffer.
[0594] Definition of MPEG_buffer_circular extension:
[0595] NameTypeDefaultUsageDescriptioncountinteger2OThe count field provides the recommended number of sequential buffer frames to be offered by a circular buffer to the presentation engine.This information may be used by the MAF to setup the circular buffer towards the Presentation Engine.mediaintegerN / AMIndex of the media entry in the MPEG_media extension, which is used as the source for the input data to the buffer.tracksarrayN / AOIndex of a track of a media entry, indicated by media and listed by MPEG_media extension, used as the source for the input data to this buffer.When tracks element is not present, the media pipeline should perform the necessary processing of all tracks of the MPEG_media entry, referenced by the media property, to generate the requested data format of the buffer.When tracks array contains multiple tracks, the media pipeline should perform the necessary processing of all referenced tracks to generate the requested data format of the buffer.If the track attribute is present and there are multiple "alternatives" (ie indicating equivalent content) in the referenced media, then the selected track shall be present in all alternatives.NOTE: When more than one track is listed by tracks element, the corresponding buffer is in active state and the MAF is informed that the corresponding tracks are needed as source for the input buffer, then the MAF can optimize the delivery of multiple tracks.
[0596] count: The count field provides a recommended number of sequential buffer frames that the circular buffer should provide to the presentation engine.
[0597] This information can be used by MAF to set up a circular buffer towards the presentation engine.
[0598] media: The index of the media item in the MPEG_media extension, used as the source of input data for the buffer.
[0599] tracks: The track index of a media item marked with media and listed with the MPEG_media extension, which is used as the source of input data for this buffer.
[0600] If the tracks element is absent, the media pipeline processes all tracks in the MPEG_media item referenced in the media property as needed to generate the requested data format for the buffer.
[0601] If the tracks array contains multiple tracks, the media pipeline processes all referenced tracks as needed to produce the requested data format for the buffer.
[0602] If the track attribute is present and the referenced media has multiple "alternatives" (i.e., representing equivalent content), the selected track is in all of the alternatives.
[0603] Note: If more than one track is listed for a track element, and the corresponding buffer is active and MAF is informed that the track is required as a source for an input buffer, MAF can optimize the delivery of multiple tracks.
[0604] Definition of top-level objects in the MPEG_scene_dynamic extension:
[0605] NameTypeDefaultUsageDescriptionmediaintegerN / AMProvides the index of the media described in the MPEG_media extension and which will contain the scene update data.trackintegerN / AOProvides the index of a track of a media object, referenced by media attribute and listed by MPEG_media extension. The track samples contain scene description updates and provide timing to perform these updates.If track is not provided, it shall be assumed that all tracks provided by the referenced media object are used to provide the update samples.
[0606] media: Provides an index into the media described in the MPEG_media extension, which contains scene update data.
[0607] track: Provides the track index of the media object referenced by the media property and listed by the MPEG_media extension. Track samples contain scene description updates and provide the timing for performing these updates.
[0608] If no track is provided, it is assumed that all tracks provided by the referenced media object are used to provide update samples.
[0609] Definition of top-level objects in the MPEG_texture_video extension:
[0610] NameTypeDefaultUsageDescriptionaccessorintegerN / AMProvides a reference to the accessor, by specifying the accessor's index in accessors array, that describes the buffer where the decoded timed texture will be made available.The accessor shall have the MPEG_accessor_timed extension.The type, componentType, and count of the accessor depend on the width, height, and format.widthintegerN / AMProvides the maximum width of the texture.heightintegerN / AMProvides the maximum height of the textureformatstringRGBOIndicates the format of the pixel data for this video texture. The allowed values are: RED, GREEN, BLUE, RG, RGB, RGBA, BGR, BGRA, DEPTH_COMPONENT. The semantics of these values are defined in Table 8.3 of OpenGL® aSpecification.Additionally, YCbCr formats are supported. The semantics for the YCbCr formats are defined in Table 76 in Vulkan specification [Vulkan 1.3]. A sampler with the MPEG_sampler_YCbCr extension shall be linked to a YCbCr texture. Note that the number of components shall match the type indicated by the referenced accessor. Normalization of the pixel data shall be indicated by the normalized attribute of the accessor.
[0611] OpenGL® is a feature provided by Khronos and is described as an example in this document, but is not limited to Khronos.
[0612] accessor: Provides a reference to an accessor by specifying its index in the accessor array, describing the buffer for which the decoded timed texture is available.
[0613] The accessor has the MPEG_accessor_timed extension.
[0614] The type, componentType, and count of the accessor depend on the width, height, and format.
[0615] width: Provides the maximum width of the texture.
[0616] height: Provides the maximum height of the texture.
[0617] format: Indicates the pixel data format of this video texture. Allowed values are RED, GREEN, BLUE, RG, RGB, RGBA, BGR, BGRA, and DEPTH_COMPONENT. The meanings of these values are defined in Table 8.3 of the OpenGL® a specification.
[0618] Additionally, the YCbCr format is supported. The semantics of the YCbCr format are defined in Table 76 of the Vulkan specification [Vulkan 1.3]. A sampler with the MPEG_sampler_YCbCr extension is connected to a YCbCr texture.
[0619] The number of components matches the type specified in the referenced accessor. The normalization of the pixel data is indicated by the normalized property of the accessor.
[0620] Definition of top-level objects of the MPEG_mesh_linking extension:
[0621] NameTypeDefaultUsageDescriptioncorrespondenceintegerN / AMProvides a reference to the accessor, by specifying the accessor's index in accessors array, that describe the buffer where the correspondence values between the dependent mesh and its associated shadow mesh will be made available.meshintegerN / AMProvides a reference to the shadow mesh, by specifying the mesh index in meshes array, associated to the dependent mesh for which the correspondence values are established.poseintegerN / AMProvides a reference to the accessor, by specifying the accessor's index in accessors array, that describe the buffer where the transformation of the nodes associated to the dependent mesh will be made available.The componentType of the referenced accessor shall be FLOAT and the type shall be MAT4.weightsintegerN / AOProvides a reference to the accessor, by specifying the accessor's index in accessors array, that describe the buffer where the "weights" to be applied to the morph targets of the shadow mesh associated to the dependent mesh will be made available.The componentType of the referenced accessor shall be FLOAT and the type shall be SCALAR.
[0622] correspondence: Provides a reference to an accessor by specifying its index in the accessors array, and describes a buffer where correspondence values between the dependent mesh and the associated shadow mesh are available.
[0623] mesh: Provides a reference to a shadow mesh by specifying the mesh index in the meshes array associated with the dependent mesh that has the corresponding value set.
[0624] pose: Provides a reference to an accessor by specifying its index in the accessors array, and describes a buffer in which the transform of the node associated with the dependent mesh is available.
[0625] The referenced accessor's componentType is FLOAT and its type is MAT4.
[0626] weights: Provides a reference to an accessor by specifying its index in the accessors array, and describes a buffer where "weights" are available to be applied to the morph targets of the shadow mesh associated with the dependent mesh.
[0627] The referenced accessor's componentType is FLOAT and its type is SCALAR.
[0628] Definition of objects in the MPEG_audio_spatial extension:
[0629] NameTypeDefaultUsageDescriptionsourcesarrayN / AOan array of source objects that are attached to the current node. listenerobjectN / AOa listener object that places an audio listener node in the scene that should be attached to a parent camera node. The audio listener characteristics depend on the available audio output devices.reverbsarrayN / AOan array of reverb objects.
[0630] sources: An array of source objects attached to the current node.
[0631] listener: A listener object that places an audio listener node in the scene, which should be attached to the parent camera node. The audio listener properties depend on the available audio output devices.
[0632] reverbs: An array of reverb objects.
[0633] Definition of the source object in the MPEG_audio_spatial.source extension:
[0634] NameTypeUsageDefaultDescriptionidintegerMUnique identifier of the audio source in the scene.typestringMIndicates the type of the audio source.The value "Object" indicates mono objectThe value "HOA" indicates HOA objecttargetSampleRatenumberMProvides the target audio sampling rate that is expected to be supplied by the media pipeline of the corresponding audio source. If the sampling rate of one of the selected media source alternatives differs from this value, then the media pipeline shall perform resampling to match the target sample rate.pregainnumberO0.0Provides a level-adjustment in dB for the signal associated with the sourceplaybackSpeednumberODefines the playback speed of the audio signal. A value of 1.0 corresponds to playback at normal speed. The value shall be between 0.5 and 2.0.attenuationenumerationO"linearDistance"Indicates the function used to calculate the attenuation of the audio signal based on the distance to the source.The value "noAttenuation" indicates that no attenuation function should be used.The value "inverseDistance" indicates that the inverse distance function should be used.The value "linearDistance" indicates that the linear distance function should be used.The value "exponentialDistance" indicates that the exponential distance function should be used.The value "custom" indicates that a custom function should be used. The definition of custom functions is outside of the scope of this document.The attenuation functions and their parameters are specified in Annex D.attenuationParametersarrayON / AArray of parameters that are input to the attenuation function. The semantics of these parameters depend on the attenuation function itself.referenceDistancenumberO1.0Provides the distance in meters for which the distance gain is implicitly included in the source signal after application of pregain.When type equals 'HOA', the element shall not be present.accessorsarrayMProvides an array of accessor references, by specifying the accessors indices in accessors array, that describe the buffers where the decoded audio will be made available.reverbFeedarrayON / AProvides one or more pointers to reverb units, optionally extended by a floating point scaling factor.A reverb unit represents a reverberation audio processor that is configured by the metadata from a single reverb object. Typically, a reverb object represents reverberation properties of a single roomreverbFeedGainarrayON / AProvides an array of gain [dB] values to be applied to the source's signal(s) when feeding it to the corresponding reverbFeed.The array shall have the same number of elements as the reverbFeed array field.isClusterbooleanFalseOSpecifies if the audio source is a pre-mixed representation of a selection of audio sources.clusterPropertiesobjectN / AOA sourceClusterProperties object that contains cluster properties.This object must be defined, when the isCluster attribute is set to True.
[0635] id: A unique identifier for the audio source in the scene.
[0636] type: Indicates the type of the audio source. A type value of "Object" indicates a mono object. A type value of "HOA" indicates an HOA object.
[0637] targetSampleRate: Provides the target audio sampling rate expected to be provided by the media pipeline for the given audio source. If one of the selected media source alternatives has a sampling rate different from this value, the media pipeline will resample to match the target sample rate.
[0638] pregain: Provides level adjustment in dB for the signal relative to the source.
[0639] playbackSpeed: Defines the playback speed of the audio signal. A value of 1.0 corresponds to normal playback speed. Values range from 0.5 to 2.0.
[0640] attenuation: Indicates the function used to calculate the attenuation of the audio signal based on the distance to the source. If this value is "noAttenuation," it indicates that no attenuation function should be used. If this value is "inverseDistance," it indicates that the inverse distance function should be used. If this value is "linearDistance," it indicates that the linear distance function should be used. If this value is "exponentialDistance," it indicates that the exponential distance function should be used. If this value is "custom," it indicates that a user-defined function should be used.
[0641] The damping function and its parameters can be found in Appendix D.
[0642] attenuationParameters: An array of parameters input to the attenuation function. The meaning of these parameters depends on the attenuation function itself.
[0643] referenceDistance: Provides the distance in meters at which the distance gain is implicitly included in the source signal after applying pregain.
[0644] If type is equal to 'HOA', the element does not exist.
[0645] accessors: Provides an array of accessor references describing the buffers to which the decoded audio will be provided, by specifying the accessors indices into the accessors array.
[0646] reverbFeed: Provides one or more pointers to reverb units, optionally expanded with a floating-point scaling factor.
[0647] A reverb unit represents a reverb audio processor consisting of the metadata of a single reverb object. Typically, a reverb object represents the reverb properties of a single room.
[0648] reverbFeedGain: Provides an array of gain [dB] values to apply to the source signal when feeding it to the given reverbFeed.
[0649] The array has the same number of elements as the reverbFeed array field.
[0650] isCluster: Specifies whether the audio source is a pre-mixed representation of a selection of audio sources.
[0651] clusterProperties: A sourceClusterProperties object containing cluster properties. This object is defined when the isCluster property is set to True.
[0652] Definition of listener objects in the MPEG_audio_spatial.listener extension:
[0653] NameTypeDefaultUsageDescriptionidintegerN / AMunique identifier of the audio listener in the scene.
[0654] id: A unique identifier for the audio listener within the scene.
[0655]
[0656] Definition of the reverb object in the MPEG_audio_spatial.reverb extension:
[0657] NameTypeDefaultUsageDescriptionidintegerN / AMunique identifier of the audio reverb unit in the scene.bypassbooleanTrueOindicates if the reverb unit can be bypassed if the audio renderer does not support it.propertiesarrayMArray of items that contains reverbProperties objects describing reverb unit specific parameterspredelaynumber0.0ODelay [seconds] from onset of source to onset of late reverberation for which DSR is provided.
[0658] id: Unique identifier of the audio reverb unit in the scene.
[0659] bypass: Indicates whether the reverb unit can be bypassed if the audio renderer does not support it.
[0660] properties: An array of items containing reverbProperties objects describing reverb unit-specific parameters.
[0661] predelay: Indicates the delay time [in seconds] from the start of the source provided by DSR to the start of the late reverb.
[0662]
[0663] Definition of the reverb property object of the MPEG_audio_spatial extension
[0664] NameTypeDefaultUsageDescriptionfrequencynumberN / AMFrequency for the provided RT60 and DSR values.RT60numberN / AMRT60 values (in seconds) for the frequency provided in the 'frequency' field.
[0665] frequency: The frequency for the provided RT60 and DSR values.
[0666] RT60: The RT60 value (in seconds) for the frequency provided in the 'Frequency' field.
[0667] DSR: The spread-source ratio value [dB] for the frequency provided in the 'Frequency' field.
[0668]
[0669] Definition of the source cluster property object of the MPEG_audio_spatial extension:
[0670] NAmeTypeDefaultUsageDescriptionsourceidarrayMArray of integers that contains the unique identifiers of the audio sources contained in this cluster.radiusnumberO0.0Distance in meters used to encompass the audio sources contained in this cluster.
[0671] sourceId: An array of integers containing unique identifiers of the audio sources contained in this cluster.
[0672] radius: The distance in meters used to encompass audio sources included in this cluster.
[0673]
[0674] Definition of MPEG_viewport_recommended extension:
[0675] NameTypeDefaultUsageDescriptionnamestringN / AOLabel of the recommended viewporttranslationintegerN / AOProvides a reference to accessor where the timed data for the translation of camera object will be made available. The componentType of the referenced accessor shall be FLOAT and the type shall be VEC3, (x, y, z).rotationTypeparametersintegerstringintegerN / A'perspective'N / AOOOProvides a reference to accessor where the timed data for the rotation of camera object will be made available. The componentType of the referenced accessor shall be FLOAT and the type shall be VEC4, as a unit quaternion, (x, y, z, w).provides the type of camera.Provides a reference to a timed accessor where the timed data for the perspective or orthographic camera parameters will be made available. The componentType of the referenced accessor shall be FLOAT and the type shall be VEC4.When the type of the camera object which includes this extension is perspective, FLOAT_VEC4 means (aspectRatio, yfov, zfar, znear).When orthographic type, FLOAT_VEC4 means (xmag, ymag, zfar, znear).The requirements on the camera parameters from ISO / IEC 12113 shall apply.
[0676] name: The label of the recommended viewport.
[0677] translation: Provides a reference to an accessor that can access timing data for the camera object's translation. The referenced accessor has a componentType of FLOAT and a type of VEC3, (x, y, z).
[0678] rotation: Provides a reference to an accessor that can access timing data about the rotation of the camera object. The referenced accessor has a componentType of FLOAT and a type of VEC4, a unit quaternion, (x, y, z, w).
[0679] type: Provides the type of camera.
[0680] parameters: Provides a reference to a timing accessor that can access timing data for perspective or orthographic camera parameters. The referenced accessor has a componentType of FLOAT and a type of VEC4.
[0681] If the type of the camera object containing this extension is perspective, FLOAT_VEC4 means (aspectRatio, yfov, zfar, znear).
[0682] In the case of regular expression, FLOAT_VEC4 means (xmag, ymag, zfar, znear).
[0683] The requirements for camera parameters in ISO / IEC 12113 apply.
[0684]
[0685] Definition of the MPEG_animation_timing extension:
[0686] NameTypeDefaultUsageDescriptionaccessorintegerN / AMProvides a reference to the accessor, by specifying the accessor's index in accessors array, that describes the buffer where the animation timing data will be made available. The componentType of the referenced accessor shall be BYTE and the type shall be SCALAR.
[0687] accessor: Provides a reference to an accessor by specifying its index in the accessors array, which describes the buffer where animation timing data will be provided. The referenced accessor has a componentType of BYTE and a type of SCALAR.
[0688] Figure 18 shows an example of a file format for a scene description according to embodiments.
[0689] ISO / IEC 12113 defines an extension mechanism that allows glTF 2.0 to be extended with new features. A glTF node can have an optional extensions property that lists the extensions used by the node. All extensions used in a glTF document are listed in the top-level extensionsUsed array object, and extensions required to properly load / render a scene are also listed in the extensionsRequired array.
[0690] glTF 2.0 is an example of a media file format, and the media file format according to embodiments is not limited to glTF 2.0 and may refer to a format related to MPEG media.
[0691] Figure 18 shows the structure of a media file format including the MPEG media-related extensions described in Table 52.
[0692] The encoding method and / or decoding method according to the embodiments can generate a media file regarding a scene description as shown in FIG. 18 and store it in a buffer. Various types of information constituting the scene description can be generated based on a hierarchical structure as shown in FIG. 18 and stored in a buffer. The encoding method according to the embodiments can encode mesh data and generate additional MPEG media extension information regarding the mesh data in a media file format. The decoding method according to the embodiments can decode mesh data and provide an immersive content service based on the scene description information stored in a buffer. The information structure of the scene description is as follows: extension information regarding the scene (MPEG_scene_dynamic, MPEG_viewport_recommended, MPEG_animation_timing), extension information regarding MEPG media (MPEG_media), extension information for spatial audio (MPEG_audio_spatial_reverb), etc. can be stored in the buffer. Extended information about a scene can refer to the parent node or child node of the graph related to the extended information for spatial audio (MPEG_audio_spatial_listener, MPEG_audio_spatial_source) for the nodes of the graph that contain objects included in the scene. A node can refer to light information related to the node. A node can refer to a camera node related to the node. An object is expressed as a mesh, and a node that contains an object can refer to mesh information that includes extended information about the mesh (MPEG_mesh_linking). The mesh information can refer to material information about the object. The material information can refer to texture (attribute) information that includes extended information about the texture (MPEG_texture_video). Animation information and / or skin information can refer to node information.Mesh information, animation information, and / or skin information may reference time-dependent access information that dynamically changes over time. The time-dependent access information may include extended information (MPEG_accessor_timed). Scene description information may be stored in a circular buffer. It may include extended information (MPEG_buffer_circular) that represents the circular buffer.
[0693] Hereinafter, with reference to FIGS. 19 to 22, a data processing structure for decoding mesh data based on a scene description by a decoding method and device according to embodiments will be described.
[0694] Figure 19 shows a pipeline for V-DMC (Video-based Dynamic Mesh Coding) according to embodiments.
[0695] Pipeline #1 (first pipeline) of Fig. 19 is an exemplary pipeline for single-track V-DMC content. When the encoding method according to the embodiments encodes mesh data and encapsulates the mesh data into a single-track file, the single track can be parsed and, based on a buffer containing a scene description, immersive mesh data can be provided through the first pipeline of Fig. 19.
[0696] The Media Access Function (MAF) can receive V-DMC content, parse basemesh data, geometry data, attribute data, and atlas data contained in a single track of the file through a demultiplexer (or demuxer), and decode each data. The atlas decoder can decode the atlas data, the basemesh decoder can decode the basemesh, the attribute decoder can decode the attributes, the geometry decoder can decode the geometry data, or the displacement decoder can decode the displacement data. The basemesh data can be decoded through an Edge Breaker (EB) decoder, the geometry data and attribute data can be decoded through a video decoder such as an HEVC decoder, and the atlas data can be decoded through the atlas decoder. Geometry data can be encoded using arithmetic coding in addition to video codecs, and then decoded using an arithmetic decoder. Each piece of decoded data can then be reconstructed into 3D data through subsequent processing. For example, each piece of decoded data can undergo processing such as format conversion, which converts it into a buffer format appropriate for the type of data.
[0697] Data restored as 3D data is transferred to a buffer, and the V-DMC data stored in the buffer can be accessed and rendered in PE.
[0698] Figure 20 shows a pipeline for V-DMC according to embodiments.
[0699] Pipeline #2 is an embodiment of a pipeline for multiple track V-DMC content.
[0700] MA can receive V-DMC content and decode each track containing data using the corresponding decoder. Basemesh data can be decoded through an Edge Breaker (EB) decoder, and geometry data and attribute data can be decoded through a video decoder such as an HEVC decoder. Atlas data can be decoded through an atlas decoder. Geometry data can be encoded using arithmetic coding in addition to a video codec. Geometry data can be decoded through an arithmetic decoder. Each decoded data can be restored to 3D data through subsequent processing. For example, each decoded data can undergo processing such as format conversion, which converts it into a buffer format according to the type of each data.
[0701] Data restored as 3D data is input into a buffer, and the V-DMC data stored in this buffer can be accessed and rendered in PE.
[0702] While Fig. 19 shows a pipeline in which all V-DMC-related data is received in a single track, Fig. 20 shows a pipeline in which each V-DMC-related data is received individually in multiple tracks.
[0703] Figure 21 shows a pipeline for V-DMC according to embodiments.
[0704] Pipeline #3 is another example of a pipeline for multi-track V-DMC content.
[0705] The MAF can receive V-DMC content and decode the data within the track based on the corresponding decoder for each track containing each data. For example, basemesh data can be decoded through an Edge Breaker (EB) decoder, and geometry data and attribute data can be decoded through a video decoder such as an HEVC decoder. Atlas data can be decoded through an atlas decoder. Geometry data can be encoded using arithmetic coding in addition to a video codec, in which case the geometry data can be decoded through an arithmetic decoder. Each decoded data can go through a processing process such as format conversion, where it is converted into a buffer format according to the type of each data, and then stored in each buffer in each data bitstream format. Afterwards, the PE can access each buffer and perform 3D reconstruction to render the V-DMC data.
[0706] Each component data that constitutes V-DMC content passes through each decoder and is stored in the corresponding buffer in the form of raw data for each component. Since the 3D reconstruction process is not performed within the pipeline, the PE (Presentation Engine) can access the buffer for each component, and 3D reconstruction or partial rendering can be performed depending on the receiver operation according to the use case, etc.
[0707] Figure 20 shows a case where 3D restoration is located at the decoder stage before the buffer, and Figure 21 shows a case where 3D restoration is located at the PE stage after the buffer.
[0708] A decoder according to embodiments may refer to an EB decoder, an attribute decoder, a geometry decoder, a displacement decoder, an arithmetic encoding decoder, an atlas decoder, etc. In addition, it may be interpreted as a decoder including a decapsulator that parses a track of a file, decoding, and processing, and may be interpreted as a decoder including a buffer to a PE.
[0709] Figure 22 shows a buffer format for V-DMC according to embodiments.
[0710] A buffer according to embodiments may further include MPEG_primitive_VDMC_extension information.
[0711] To support VDMC compression objects in MPEG-I scene descriptions, the MPEG_media extension is used to reference a VDMC-encoded bitstream. The PE can support operations that perform 3D reconstruction of decoded VDMC components, as described in FIG. 21. The PE accesses the decoded VDMC data through a buffer. The syntax of a VDMC object is provided as an extension to mesh.primitive in the scene description format. This extension references the decoded data of the VDMC object. Each decoded VDMC component is signaled using the properties defined in the MPEG_primitive_VDMC extension.
[0712] Extension usage can be listed in the extensionsUsed or extensions top-level glTF properties.
[0713] "extensionsUsed": [
[0714] "MPEG_primitive_VDMC"
[0715] ]
[0716] or
[0717] "extensions": [
[0718] "MPEG_primitive_VDMC"
[0719] ]
[0720] The scene description information structure and properties for VDMC mesh compression extension can be hierarchically organized and stored in a buffer as shown in Fig. 22.
[0721] As shown in Fig. 22, the scene description can represent the relationship between the decoded V-DMC component data and each buffer using the MPEG_primitive_VDMC extension within the primitive of the mesh object of glTF2.0.
[0722] In the case of V-DMC, displacement data can be encoded using video coding or arithmetic coding, but since the displacement data itself is not data in video format, it can reference the buffer through an accessor rather than referencing the video texture object defined in the MPEG_texture_video extension in Table 61, like attribute data.
[0723] The respective semantics of Figure 22 are as follows:
[0724] The MPEG_primitive_VDMC extension references several VDMC components containing decoded base meshes, displacements, attributes, and metadata for the 3D reconstruction process.
[0725] MPEG_primitive_VDMC properties:
[0726] NameTypeDefaultUsageDescription_MPEG_VDMC_CONFIGintegerThis component provides a reference to a timed accessor that contains configuration information that is applicable to a sequence of frames of the VDMC decoded mesh primitive._MPEG_VDMC_ADobjectthis component shall reference a timed accessor that provides the VDMC atlas data buffer. The atlas buffer format is defined in section 4.6.3. Future specifications of the atlas data buffer format shall use a different version.Exactly one atlas component shall be present, irrespective of the version._MPEG_VDMC_BMDobjectthis component shall reference a timed accessor that provides the basemesh data buffer._MPEG_VDMC_GVDobjectThis component shall reference a timed accessor that provides the decoded displacement data buffer._MPEG_VDMC_AVDintegerThis component shall provide a video texture reference, which corresponds to the decoded attribute video data.
[0727] _MPEG_VDMC_CONFIG: This component provides a reference to a timed accessor containing configuration information applicable to a VDMC decoded mesh elementary frame sequence.
[0728] _MPEG_VDMC_AD: This element references a timed accessor that provides a VDMC atlas data buffer. The atlas buffer format is defined in Table 72.
[0729] There can be exactly one atlas component, regardless of version.
[0730] _MPEG_VDMC_BMD: This component references a timed accessor that provides a basemesh data buffer.
[0731] _MPEG_VDMC_GVD: This component references a timed accessor that provides a buffer of decoded displacement data.
[0732] _MPEG_VDMC_AVD: This component provides a video texture reference corresponding to the decoded attribute video data.
[0733] Properties of the MPEG_VDMC_AD object:
[0734] NameTypeDefaultUsageDescriptionbuffer_formatstringprovides an identifier of the associated atlas data buffer format. A list of supported atlas data buffer formats is provided in Table G.4 in ISO / IEC 23090-14.accessorintegerThis provides the index of the timed accessor that provides access to the atlas data buffer.
[0735] buffer_format: Provides an identifier for the associated atlas data buffer format. For a list of supported atlas data buffer formats, see Table G.4 of ISO / IEC 23090-14.
[0736] accessor: Provides the index of the timed accessor that provides access to the atlas data buffer.
[0737] Properties of the MPEG_VDMC_BMD object:
[0738] NameTypeDefaultUsageDescriptionbuffer_formatsringprovides an identifier of the associated basemesh data buffer format. Supported basemesh data buffer format is provided in section 4.6.6.accessorintegerThis provides the index of the timed accessor that provides access to the basemesh data buffer.
[0739] buffer_format: Provides an identifier for the associated basemesh data buffer format. Supported basemesh data buffer formats are listed in Table 75.
[0740] accessor: This provides the index of the timed accessor that provides access to the basemesh data buffer.
[0741] Properties of the MPEG_VDMC_GVD object:
[0742] NameTypeDefaultUsageDescriptionbuffer_formatsringprovides an identifier of the associated displacement data buffer format. Supported displacement data buffer format is provided in section 4.6.7.accessorintegerThis provides the index of the timed accessor that provides access to the displacement data buffer.
[0743] buffer_format: Provides an identifier for the associated displacement data buffer format. See Table 76 for supported displacement data buffer formats.
[0744] accessor: Provides the index of the timed accessor that provides access to the displacement data buffer.
[0745]
[0746] The basemesh data buffer formats in the buffers of FIGS. 17, 19, and 21 may be as follows.
[0747] The Basemesh data buffer format can be the meshpatch data unit syntax format described above, or it can refer to the format defined in the future V-DMC specification (ISO / IEC 23090-29) document.
[0748] FieldTypeDescriptionpatch_countuint16provides the total number of patches.for( i=0; i <patch_count; i++ ) {submesh_iduint16specifies the associated submesh ID specified in the current meshpatch.vertex_countuint16specifies the number of vertices associated with the current meshpatch.face_countuint16specifies the number of faces associated with the current meshpatch.2d_pos_xfloatspecifies the x-coordinate of the top-left corner of the meshpatch's bounding box for the current meshpatch.2d_pos_yfloatspecifies the y-coordinate of the top-left corner of the meshpatch's bounding box for the current meshpatch.2d_size_xfloatspecifies the width value of the current meshpatch's bounding box.2d_size_yfloatspecifies the height value of the current meshpatch's bounding box..........}
[0749] patch_count: Indicates the total number of patches.
[0750] submesh_id: Specifies the associated submesh ID assigned to the current meshpatch.
[0751] vertex_count: Indicates the number of vertices associated with the current meshpatch.
[0752] face_count: Specifies the number of faces associated with the current meshpatch.
[0753] 2d_pos_x: Specifies the x-coordinate of the upper left corner of the meshpatch bounding box of the current meshpatch.
[0754] 2d_pos_y: Specifies the y-coordinate of the upper left corner of the meshpatch bounding box of the current meshpatch.
[0755] 2d_size_x: Specifies the width value of the bounding box of the current meshpatch.
[0756] 2d_size_y: Specifies the height value of the bounding box of the current meshpatch.
[0757]
[0758] The displacement data buffer format is as follows:
[0759] Displacement data buffer format This can be in the form of the displacement sub-bit stream syntax described above, and may refer to the format defined in the future V-DMC specification (ISO / IEC 23090-29) document. For other data formats, the format defined in ISO / IEC 23090-14 may be followed.
[0760]
[0761] For example, if you define the buffer format as glTF2.0 JSON type, it could be as follows:
[0762] The content below is an example of glTF2.0 JSON format with properties related to the VDMC primitive extension applied.
[0763] "meshes": [
[0764] {
[0765] "name": "vdmc_mesh",
[0766] "primitives": [
[0767] {
[0768] "attributes": {
[0769] "POSITION": 0,
[0770] "COLOR_0": 1
[0771] },
[0772] "mode": 0,
[0773] "extensions": {
[0774] "MPEG_primitive_VDMC": {
[0775] "_MPEG_VDMC_CONFIG": 2,
[0776] "_MPEG_VDMC_AD": {
[0777] "buffer_format": "baseline",
[0778] "accessor": 10
[0779] }
[0780] "_MPEG_VDMC_BMD": 3,
[0781] "_MPEG_VDMC_GVD": 4,
[0782] "_MPEG_VDMC_AVD": 5
[0783] }}} ]} ]
[0784] Referring to Figure 17, the operation of each component of the scene description system is described as follows.
[0785] PE (Presentation Engine):
[0786] 1) The presentation engine receives and analyzes scene description information and the next scene description update.
[0787] 2) The presentation engine identifies media that has time constraints and identifies the required presentation time.
[0788] 3) The presentation engine uses the MAF API to request media and provides the following information: a) where MAF can find the requested media, b) what part of the media and at what level of detail, c) when the requested media should be provided, and d) in what format the data is desired and how to pass it to the presentation engine.
[0789] MAF (Media Access Function):
[0790] 1) MAF instantiates a media fetching and decoding pipeline for the requested media at the appropriate time.
[0791] a) Ensure that the requested media is available in an appropriate buffer so that the presentation engine can access it at the appropriate time.
[0792] b) Ensure that the presentation engine decodes the media and reformats it to match the expected format, as described in the scene description information.
[0793] Data (media and metadata) exchange is performed through buffers (both circular and static). Buffer management is controlled through the Buffer API. Each buffer contains sufficient header information to describe its content and timing.
[0794] The information the presentation engine provides to the media access function allows it to: Select an appropriate source for the media (multiple sources can be specified), and MAF can select a source based on preferences and capabilities. Capabilities can be, for example, decoding capabilities or supported formats. Defaults can be, for example, user-defined.
[0795] - For each source selected:
[0796] 2) Access the media using the media access protocol.
[0797] 3) Set up a media pipeline to provide information in the correct buffer format.
[0798] MAF can obtain additional information from the presentation engine to optimize delivery, such as the required quality for each buffer and precise timing information.
[0799] The Media Access function establishes and manages a pipeline for each requested media or metadata. The pipeline takes one or more media or metadata tracks as input and outputs one or more buffers. The pipeline performs all necessary processing, including streaming, demultiplexing, decoding, decryption, and format conversion, to match the expected buffer format. It then uses the final buffer or set of buffers to exchange data with the presentation engine.
[0800] Figure 23 shows an encoding method according to embodiments.
[0801] The encoding method according to the embodiments may include a step of encoding mesh data (S2300); and / or a step of encapsulating a file including a bitstream including mesh data (S2310).
[0802] The encoding step (S2300) may include a step of encoding a base mesh, displacement (or geometry), attributes, etc., as shown in FIGS. 7 to 8, FIG. 13, etc.
[0803] The encapsulating step (S2310) can encapsulate a bitstream (e.g., FIG. 15) including parameter syntax such as an encoded base mesh, displacement (or geometry), attribute, and / or atlas data into a file including a single track or into files including multiple tracks. When encapsulating a file into multiple tracks, a file including a track including a base mesh, a track including displacement (or geometry), a track including an attribute, and a track including an atlas as an entry point can be created.
[0804] The encoding method according to the embodiments can efficiently decode mesh data as in FIG. 18 and additionally generate scene description information to enhance immersion. The scene description information can include information about a scene, a camera, lights, animation, skin, material, and texture (source or image). The scene description information can be generated according to the data file format of the buffer in which the scene description is stored. The scene description can be generated and stored according to various file formats. As in FIG. 18, the scene description can have a hierarchical structure and have a reference relationship between related extended information during decoding.
[0805] The step of encoding mesh data (S2300) may include: encoding base mesh data, encoding attributes, encoding geometry or displacement, and encoding atlas data.
[0806] The step of encapsulating a file (S2310) may include: creating a basemesh track, an attribute track, a geometry track or a displacement track, and an atlas track of the file.
[0807] The encoding method of FIG. 23 may further include generating scene description information for mesh data.
[0808] The encoding method is performed by an encoding device (encoder). The encoding device includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: encode mesh data; and encapsulate a file including a bitstream including mesh data.
[0809] The embodiments further include a computer-readable storage medium storing a bitstream generated by the method according to FIG. 23.
[0810] Embodiments further include a method comprising: obtaining a bitstream for mesh data, the bitstream being generated based on encoding basemesh data, encoding attributes, encoding geometry or displacement, and encoding atlas data; and transmitting data including the bitstream.
[0811] Figure 24 shows a decryption method according to embodiments.
[0812] The decryption method according to the embodiments may include a step of decapsulating a file including a bitstream including mesh data (S2410); and / or a step of decoding the mesh data (S2410). The operation of the decryption method of FIG. 24 may follow the reverse process of the encoding method of FIG. 23.
[0813] Referring to FIGS. 19 to 21, the decapsulating step (S2410) can parse a single track of a file or parse multiple tracks of a file.
[0814] The decoding step (S2410) can decode parameter syntax such as base mesh, displacement (or geometry), attribute, and / or atlas data, as shown in FIG. 11, FIG. 12, FIG. 14, etc.
[0815] The decoding method according to the embodiments can efficiently decode mesh data based on a buffer in which a scene description is stored, a PE that parses scene description information in the buffer, and a MAF that accesses the media data, as shown in FIGS. 17 and 19 to 21.
[0816] With respect to decapsulating and decoding in FIGS. 19 to 21, the step of decapsulating the file (S2410) may include: parsing a basemesh track, an attribute track, a geometry track or a displacement track, and an atlas track of the file, and the step of decoding mesh data (S2410) may include: decoding basemesh data of the basemesh track, decoding an attribute of the attribute track, decoding geometry or displacement of the geometry track or the displacement track, and decoding atlas data of the atlas track.
[0817] With respect to the buffers illustrated in FIGS. 19 to 21, the decoder is connected to the buffer, the buffer includes scene description information for mesh data, and the method may further include a step of rendering the mesh data based on the scene description information of the buffer.
[0818] With respect to MPEG_media.alternative, scene description information includes media extension information, and mesh data is referenced as an item of media extension information within the scene description information, and the referenced item may include at least one of a basemesh track or an atlas track of the file.
[0819] With respect to MPEG_primitive_VDMC, scene description information includes component extension information, and the component extension information may include at least one of a reference to an accessor for configuration information of mesh data, a reference to an accessor for an atlas data buffer for mesh data, a component referencing an accessor providing a basemesh data buffer, a component referencing an accessor providing a displacement data buffer, and a component providing a video texture reference for attribute video data.
[0820] With respect to MPEG_VDMC_AD, a reference to an accessor for an atlas data buffer may include at least one of an identifier of the format of the atlas data buffer, or an index of an accessor that provides access to the atlas data buffer.
[0821] With respect to MPEG_VDMC_BMD and Basemesh data buffer formats, a component referencing an accessor providing a basemesh data buffer includes at least one of an identifier of a format of the basemesh data buffer, or an index of an accessor providing access to the basemesh data buffer, and the format of the buffer of the basemesh data may include at least one of a patch count indicating the number of patches for the mesh data, an ID of a submesh within a mesh patch based on the patch count, a number of vertices associated with the mesh patch, a number of faces associated with the mesh patch, or position information of a bounding box for the mesh patch, or size information of a bounding box.
[0822] With respect to MPEG_VDMC_GVD, a component referencing an accessor that provides a displacement data buffer may include at least one of an identifier of the format of the displacement data buffer, or an index of an accessor that provides access to the displacement data buffer.
[0823] The decryption method is performed by a decryption device (decoder). The decryption device includes a memory; and at least one processor connected to the memory; and the at least one processor can be configured to: decapsulate a file including a bitstream including mesh data; and decode the mesh data.
[0824] The method and device according to the embodiments provide the following technical effects.
[0825] The method according to the embodiments includes a method of configuring a pipeline in a Media Access Function of a scene description reference architecture to support decoding and rendering of V-DMC content in an existing MPEG-related extension content of a glTF2.0 file format, which is a 3D delivery format.
[0826] The method according to the embodiments includes a method for configuring a buffer format so that data decoded through a pipeline in a Media Access Function of a scene description reference architecture can be effectively accessed by a Presentation Engine.
[0827] The bitstream and metadata for each component type that constitutes V-DMC content can be effectively decoded by a receiver that supports glTF2.0, a scene description format.
[0828] The data representation method according to the embodiments provides the effect of efficiently accessing the V-DMC bitstream.
[0829] A transmitter or receiver according to embodiments can efficiently decode and render a file of a V-DMC bitstream through a pipeline and buffer configuration of a scene description reference architecture.
[0830] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.
[0831] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[0832] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may be implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium may be distributed across network-connected computer systems, allowing the processor-readable code to be stored and executed in a distributed manner.
[0833] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.
[0834] Terms such as "first," "second," etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.
[0835] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0836] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.
[0837] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.
[0838] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.
[0839]
[0840] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[0841]
[0842] As described above, the embodiments may be applied in whole or in part to a point cloud data transmission and reception device and system.
[0843] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.
[0844] Embodiments may include modifications / changes, which do not depart from the scope of the claims and their equivalents.
Claims
1. A step of decapsulating a file containing a bitstream including mesh data by a decoder; and A step of decoding the above mesh data; comprising: How to decrypt.
2. In paragraph 1, The steps to decapsulate the above file are: Parsing the basemesh track, attribute track, geometry track or displacement track, and atlas track of the above file, The steps for decoding the above mesh data are: Decode the basemesh data of the above basemesh track, Decode the attributes of the above attribute track, Decoding the geometry or displacement of the above geometry track or the above displacement track, comprising decoding atlas data of the above atlas track; How to decrypt.
3. In paragraph 1, The decoder is connected to a buffer, the buffer containing scene description information for the mesh data, The method further includes a step of rendering the mesh data based on the scene description information of the buffer. How to decrypt.
4. In paragraph 3, The above scene description information includes media extension information, The above mesh data is referenced as an item of the media extension information within the scene description information, The above referenced item comprises at least one of the basemesh track or the atlas track of the above file, How to decrypt.
5. In paragraph 3, The above scene description information includes component extension information, The component extension information includes at least one of a reference to an accessor for configuration information of the mesh data, a reference to an accessor for an atlas data buffer for the mesh data, a component referencing an accessor providing a basemesh data buffer, a component referencing an accessor providing a displacement data buffer, and a component providing a video texture reference for attribute video data. How to decrypt.
6. In paragraph 5, A reference to an accessor for the atlas data buffer comprises at least one of an identifier of the format of the atlas data buffer, or an index of an accessor that provides access to the atlas data buffer. How to decrypt.
7. In paragraph 5, A component referencing an accessor that provides the basemesh data buffer includes at least one of an identifier of the format of the basemesh data buffer, or an index of an accessor that provides access to the basemesh data buffer, The format of the buffer of the above base mesh data is a patch count indicating the number of patches for the above mesh data, Based on the above patch count, the ID of the submesh within the mesh patch, The number of vertices associated with the above mesh patch, The number of faces associated with the above mesh patch, or Containing at least one of position information of the bounding box for the above mesh patch and size information of the bounding box, How to decrypt.
8. In paragraph 5, A component referencing an accessor that provides the displacement data buffer includes at least one of an identifier of the format of the displacement data buffer, or an index of an accessor that provides access to the displacement data buffer. How to decrypt.
9. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Decapsulate a file containing a bitstream containing mesh data; configured to decode the above mesh data; Decryption device.
10. A step of encoding mesh data by an encoder; and A step of encapsulating a file including a bitstream including the above mesh data; comprising; Encoding method.
11. In paragraph 10, The steps for encoding the above mesh data are: Encode the basemesh data, Encode the attributes, Encode geometry or displacement, Including encoding the atlas data, The steps to encapsulate the above file are: Including creating a basemesh track, an attribute track, a geometry track or a displacement track, and an atlas track of the above file, Encoding method.
12. In paragraph 10, The method further comprises generating scene description information for the mesh data. Encoding method.
13. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Encode mesh data; and Encapsulating a file including a bitstream containing the above mesh data; Encoding device.
14. A computer-readable storage medium storing a bitstream generated by the method according to Article 10.
15. Step of obtaining bitstream for mesh data, The above bitstream is generated based on encoding basemesh data, encoding attributes, encoding geometry or displacement, and encoding atlas data; and A method comprising the step of transmitting data including the bitstream.
Citation Information
Patent Citations
Information processing device and method
US20240046562A1
Displacement coding for mesh compression
US20240089499A1
Dynamic re-lighting of volumetric video
WO2022189702A2
An apparatus, a method and a computer program for volumetric video
WO2023041838A1
3D data transmission device, 3D data transmission method, 3D data reception device, and 3D data reception method
WO2024063544A1
Cited By
Video transmission system and method, equipment and medium
CN121442119A