Parameterized guided displacement packing for dynamic grid coding

Through the V-DMC standard and parameterized guided displacement packaging technology, the problem of low encoding and decoding efficiency of volume video data is solved, and efficient encoding and decoding of high-resolution dynamic grids is realized.

CN120035845APending Publication Date: 2025-05-23NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072688.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-10-11
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is inefficient and difficult to achieve high resolution dynamic grid codec when compressing and decoding volume video data.

Method used

The video-based dynamic grid codec (V-DMC) standard is adopted to generate compressed dynamic grid sequences through parameterized guided displacement packaging, combined with multi-resolution grid analysis and codec technology.

Benefits of technology

It improves the encoding and decoding efficiency of volume video data, supports high-resolution dynamic mesh reconstruction, and reduces the cost of storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035845A_ABST
    Figure CN120035845A_ABST
Patent Text Reader

Abstract

An example method includes receiving, with an encoder for a frame of an input volume video grid, a simplified base grid having texture coordinates, a set of pre-computed displacements for a selected subdivision method, and a texture frame; generating a static base mesh code, the static base mesh code outputting a base mesh stream comprising the encoded and quantized texture coordinates, and generating an encoded static base mesh; decoding the encoded static base grid, and generating a reconstructed inverse quantization base grid and reconstructed texture coordinates; adapting the pre-calculated displacement to the reconstructed inverse quantized base grid by applying the subdivision to the reconstructed inverse quantized base grid; on the basis of a remapping technology, remapping of the reconstructed texture coordinates to the regular lattices is calculated, so that a displacement packaging image is generated; metadata describing the remapping technology through signal transmission; and signaling the presence of conflicting vertices in the shifted packed image along the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example and non-limiting embodiments relate generally to volumetric video codecs, and more particularly, to parameterized guided displacement packing for dynamic mesh codecs. Background Art

[0002] It is known to perform encoding and decoding of images and videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The foregoing aspects and other features are explained in the following description in conjunction with the accompanying drawings, in which:

[0004] Figure 1A is a diagram illustrating volumetric media conversion on the encoder side.

[0005] Figure 1B is a diagram illustrating volumetric media reconstruction at the decoder side.

[0006] Figure 2 An example of block-to-patch mapping is shown.

[0007] Figure 3A An example of an atlas coordinate system is shown.

[0008] Figure 3B An example of a local 3D patch coordinate system is shown.

[0009] Figure 3C An example of a final target 3D coordinate system is shown.

[0010] Figure 4 Elements of a grid are shown.

[0011] Figure 5 An example V-PCC extension for trellis coding based on embodiments described herein is shown.

[0012] Figure 6 An example V-PCC extension for trellis decoding based on embodiments described herein is shown.

[0013] Figure 7 A subdivision step is shown for subdividing a triangle into four triangles by connecting the midpoints of the sides of the initial triangle.

[0014] Figure 8 Multiresolution analysis of the mesh is depicted.

[0015] Fig. 9 is a block diagram of an encoder consisting of a preprocessing module.

[0016] Fig.10 Depicted are the preprocessing steps at the encoder.

[0017] Fig.11 is a block diagram of an intra-frame encoder scheme.

[0018] Fig.12 is a block diagram of an inter-frame encoder scheme.

[0019] Fig.13 A decoder scheme consisting of a decoder module that demultiplexes and decodes all substreams and a post-processing module that reconstructs the dynamic grid sequence is depicted.

[0020] Fig.14 The decoding process in intra mode is shown.

[0021] Fig.15 The decoding process in inter-frame mode is shown.

[0022] Fig.16 Shown is the basic trellis encoder in the VDMC encoder.

[0023] Fig.17 An example base grid encoder is shown.

[0024] Fig.18 An exemplary base grid decoder is shown.

[0025] Fig.19 Figure 1 shows an overview of the underlying mesh data sub-flow structure.

[0026] Fig. 20 Depicts the segmentation from a grid to subgrids.

[0027] Fig.21 The diagram shows 2 examples of sub-grids.

[0028] Fig.22A A visualization of the UV parameterization of the base mesh's vertices and triangles.

[0029] Fig. 22B This is a visualization of the texture coordinate map after midpoint subdivision.

[0030] Fig.23 A modified encoder scheme embodiment including an image packing module is depicted.

[0031] Fig.24 A modified decoder scheme embodiment including an image unpacking module is depicted.

[0032] Fig.25 Example remapping techniques are depicted.

[0033] Fig.26 The diagram shows different blocking strategies for different LODs.

[0034] Fig. 27 is an example apparatus for implementing the examples described herein.

[0035] Fig.28 A representation of an example of a non-volatile storage medium is shown.

[0036] Fig.29 are example methods of implementing the examples described herein.

[0037] Fig.30 are example methods of implementing the examples described herein. DETAILED DESCRIPTION

[0038] The examples described herein relate to a new standardization activity called Video-based Dynamic Mesh Codec (V-DMC) ISO / IEC 23090-29, which is a new application of the Visual Volumetric Video Codec (V3C) standards series ISO / IEC 23090-5.

[0039] Volumetric video data

[0040] Volumetric video data represents a three-dimensional scene or object and can be used as input for AR, VR and MR applications. Such data describes the geometry (shape, size, position in 3D space) and corresponding attributes (e.g., color, opacity, reflectivity...), plus any possible time transformation of the geometry and attributes at a given time instance (such as a frame in a 2D video). Volumetric video is generated from a 3D model (e.g., CGI) or captured from a real-world scene using various capture schemes (e.g., a combination of multiple cameras, laser scanning, video and dedicated depth sensors, etc.). Moreover, a combination of CGI and real-world data is possible. Typical representation formats for such volumetric data are triangular meshes, point clouds or voxels. Temporal information about the scene can be included in the form of individual capture instances (e.g., "frames" in 2D video) or other means (e.g., the position of an object as a function of time).

[0041] Because volumetric video describes a 3D scene (or object), the data can be viewed from any viewpoint. Therefore, volumetric video is an important format for AR, VR or MR applications, especially for providing 6DOF viewing capabilities.

[0042] Increased computing resources and advances in 3D data acquisition devices have made it possible to reconstruct high-detail volumetric video representations of natural scenes. Infrared, laser, time-of-flight, and structured light are all examples of devices that can be used to construct 3D video data. The representation of 3D data depends on how the 3D data is used. Dense voxel arrays have been used to represent volumetric medical data. In 3D graphics, polygonal meshes are widely used. On the other hand, point clouds are very suitable for applications such as capturing real-world 3D scenes, where the topology does not need to be a 2D manifold. Another way to represent 3D data is to encode and decode the 3D data into a set of textures and depth maps, as in the multi-view plus depth framework. Closely related to the techniques used in multi-view plus depth is the use of elevation maps and multi-level surface maps.

[0043] MPEG Visual Volumetric Video Codec (V3C)

[0044] This article mentions excerpts from the ISO / IEC 23090-5 Visual volumetric video codec and video-based point cloud compression, second edition standard.

[0045] Visual volumetric video (a sequence of visual volumetric frames) when uncompressed can be represented by a large amount of data, which can be costly in terms of storage and transmission. This leads to the need for high codec efficiency standards for compressing visual volumetric data.

[0046] The V3C specification implements the encoding and decoding process of various volumetric media by using video and image codec technologies. This is achieved by first converting the volumetric media from the corresponding 3D representation into multiple 2D representations (also called V3C components) before encoding and decoding such information. Such representations may include occupancy, geometry, and attribute components. The occupancy component can inform the V3C decoding and / or rendering system which samples in the 2D component are associated with the data in the final 3D representation. The geometry component includes information about the precise location of the 3D data in space, while the attribute component can provide additional attributes of such 3D data, such as texture or material information. Figure 1A and Figure 1B An example is shown in .

[0047] Figure 1A shows volumetric media conversion at the encoder, and Figure 1B Volumetric media conversion at the decoder side is shown. 3D media 102 is converted to a series of 2D representations: occupancy 118, geometry 120 and attributes 122. Additional atlas information 108 is also included in the bitstream to enable inverse reconstruction. Reference is made to ISO / IEC 23090-5.

[0048] like Figure 1AAs further shown, the volume capture operation 104 generates projections 106 from the input 3D media 102. In some examples, the projections 106 are projection operations. From the projections 106, the occupancy operation 110 generates an occupancy 2D representation 118, the geometry operation 112 generates a geometry 2D representation 120, and the attribute operation 114 generates an attribute 2D representation 122. Additional atlas information 108 is included in the bitstream 116. The atlas information 108, the occupancy 2D representation 118, the geometry 2D representation 120, and the attribute 2D representation 122 are encoded into a V3C bitstream 124 to encode a compressed version of the 3D media 102. Based on the examples described herein, V-DMC packetization signaling 129 may also be signaled in the V3C bitstream 124, or directly signaled to a decoder. The V-DMC packetization signaling 129 may be used on the decoder side, such as Figure 1B as shown in .

[0049] like Figure 1B As shown, a decoder using the V3C bitstream 124 derives a 2D representation using an occupancy operation 128, a geometry operation 130, and an attribute operation 132. An atlas information operation 126 provides atlas information to the bitstream 134. The occupancy operation 128 derives an occupancy 2D representation 136, the geometry operation 130 derives a geometry 2D representation 138, and the attribute operation 132 derives an attribute 2D representation 140. A 3D reconstruction operation 142 uses the atlas information 126 / 134, the occupancy 2D representation 136, the geometry 2D representation 138, and the attribute 2D representation 140 to generate a decompressed reconstruction 144 of the 3D media 102.

[0050] Additional information that allows associating all these subcomponents and enabling inverse reconstruction from a 2D representation back to a 3D representation is also included in a special component, referred to in this article as an atlas. The atlas consists of multiple elements, patches. Each patch identifies a region among all available 2D components and includes the information required to perform the appropriate inverse projection of that region back to 3D space. The shape of this region is determined by the 2D bounding box associated with each patch and their encoding and decoding order. The shapes of these regions are further refined after taking into account the occupancy information.

[0051] The atlas is divided into patches of equal size. Figure 2 202 in which Figure 2 An example of block to patch mapping is shown. The 2D bounding box of the patch and its encoding and decoding order determine the mapping between blocks of the atlas image and patch indices. Figure 2An example of a block-to-patch mapping of four projected patches (204, 204-2, 204-3, 204-4) projected onto atlas 201 is shown when asps_patch_precedence_order_flag is equal to 0. Projection points are represented in dark gray. Areas that do not include any projection points are represented in light gray. Patch packing blocks 202 are represented by dashed lines. The number within each patch packing block 202 represents the patch index of the patch (204, 204-2, 204-3, 204-4) to which it is mapped.

[0052] The axis orientation is specified for internal operations. For example, the origin of the atlas coordinates is located at the upper left corner of the atlas frame. For the reconstruction step, the intermediate axis definition for the local 3D patch coordinate system is used. The 3D local patch coordinate system is then transformed to the final target 3D coordinate system using appropriate transformation steps.

[0053] Figure 3A shows an example of an atlas coordinate system, Figure 3B shows an example of a local 3D patch coordinate system, and Figure 3C An example of a final target 3D coordinate system is shown. Refer to ISO / IEC 23090-5.

[0054] Figure 3A An example of a single patch 302 packed onto an atlas image 304 is shown. Figure 3B , the patch 302 is then transformed to a local 3D patch coordinate system (U, V, D) defined by a projection plane with origin O′, tangent axis (U), vice tangent axis (V), and normal axis (D). For orthogonal projection, the projection plane is equal to the side of the axis-aligned 3D bounding box 306, as Figure 3B The position of the bounding box 306 in the 3D model coordinate system (defined by a left-handed system with axes (X, Y, Z)) can be obtained by adding the offsets TilePatch3dOffsetU 308, TilePatch3DOffsetV 310, and TilePatch3DOffsetD 312, as shown in FIG. Figure 3C As shown in the picture.

[0055] V3C Advanced Syntax

[0056] The encoded and decoded V3C video component is referred to herein as a video bitstream, and the atlas component is referred to as an atlas bitstream. The video bitstream and atlas bitstream can be further split into smaller units, referred to herein as video sub-bitstreams and atlas sub-bitstreams, respectively, and can be interleaved together to construct a V3C bitstream after adding appropriate delimiters.

[0057] The V3C patch information is included in the atlas bitstream atlas_sub_bitstream(), which includes a sequence of NAL units. The NAL unit is specified to format the data and provide header information in a manner suitable for transmission on various communication channels or storage media. All data is included in the NAL unit, and each unit in the NAL unit includes an integer number of bytes. The NAL unit specifies a common format for both packet-oriented systems and bitstream systems. The format of the NAL unit for both packet-oriented transmission and sample streams is the same, except that in the sample stream format specified in Appendix D of ISO / IEC 23090-5, each NAL unit may be preceded by an additional element that specifies the size of the NAL unit.

[0058] NAL units in an atlas bitstream can be divided into atlas codec layer (ACL) units and non-atlas codec layer (non-ACL) units. The former is dedicated to carrying patch data, while the latter is dedicated to carrying data required for correct parsing of ACL units or any additional auxiliary data.

[0059] In the nal_unit_header() syntax, nal_unit_type specifies the type of RBSP data structure included in the NAL unit, as specified in Table 4 of ISO / IEC 23090-5. nal_layer_id specifies the identifier of the layer to which the ACL NAL unit belongs, or the identifier of the layer to which a non-ACL NAL unit applies. The value of nal_layer_id should be in the range of 0 to 62 (inclusive). The value 63 may be specified by ISO / IEC in the future. Decoders conforming to the profiles specified in Annex A of ISO / IEC 23090-5 should ignore (e.g., remove or discard from the bitstream) all NAL units whose value of nal_layer_id is not 0.

[0060] V3C extension mechanism

[0061] When designing the V3C specification, it was envisaged that modifications or new versions could be created in the future. To ensure that the first implementation of the V3C decoder is compatible with any future extensions, a number of fields are reserved for future extensions to the parameter sets.

[0062] For example, the second version of V3C introduced extensions in VPS related to MIV and packaged video components.

[0063] Rendering and Meshes

[0064] A polygon mesh is a collection of vertices, edges, and faces that defines the shape of a polyhedral object in 3D computer graphics and solid modeling. Faces typically consist of triangles (triangle meshes), quadrilaterals (quads), or other simple convex polygons (n-gons), as this simplifies rendering, but can also more generally consist of concave polygons or even polygons with holes.

[0065] refer to Figure 4 , an object 400 created using a polygonal mesh is represented by different types of elements. These different types of elements include vertices 402, edges 404, faces 406, polygons 408, and surfaces 410, such as Figure 4 Therefore, Figure 4 Elements of a grid are shown.

[0066] A polygonal mesh is defined by the following elements:

[0067] Vertex (402): A position in 3D space, defined as (x, y, z), along with other information such as color (r, g, b), normal vectors, and texture coordinates.

[0068] Edge (404): A connection between two vertices.

[0069] Face (406): A closed set of edges 404, where a triangular face has three edges and a quadrilateral face has four edges. A polygon 408 is a set of coplanar faces 406. In systems that support multiple sides, polygons and faces are equivalent. Mathematically, a polygon mesh can be considered an unstructured lattice or undirected graph with the additional properties of geometry, shape, and topology.

[0070] Surfaces (410): or smoothing groups, are useful for grouping smooth areas, but are not required.

[0071] Groups: Some mesh formats include groups, which define separate elements of a mesh and can be used to identify separate sub-objects for skeletal animation or separate characters for non-skeletal animation.

[0072] Materials: Defined to allow different parts of a mesh to use different shaders when rendered.

[0073] UV Coordinates: Most mesh formats also support some form of UV coordinates, which are a separate 2D representation of the mesh "unwrapped" to show what parts of a 2D texture map are applied to different polygons of the mesh. Meshes can also include other vertex attribute information such as color, tangent vectors, weight maps to control animation, etc. (sometimes also called channels).

[0074] V-PCC Trellis Codec Extension (MPEG M49588)

[0075] Figure 5 and Figure 6 Extensions to the V-PCC encoder and decoder are shown to support trellis encoding and trellis decoding, respectively, as proposed in the MPEG input document [MPEG M47608].

[0076] In encoder extension 500, input mesh data 502 is demultiplexed into vertex coordinates + attributes 506 and vertex connectivity 508 using demultiplexer 504. Vertex coordinates + attributes data 506 are encoded and decoded using MPEG-IV-PCC (such as using MPEG-I VPCC encoder 510), while vertex connectivity data 508 is encoded and decoded (using vertex connectivity encoder 516) into auxiliary data 518. The two (encoded vertex coordinates and vertex attributes 517 and auxiliary data 518) are multiplexed using multiplexer 520 to create a final compressed output bitstream 522. Vertex sorting 514 is performed on the reconstructed vertex coordinates 512 at the output of MPEG-IV-PCC 510 to reorder the vertices for optimal vertex connectivity encoding 516.

[0077] Based on the examples described in this article, Figure 5 As shown in Figure 5 The encoding process / apparatus 500 may be extended so that the encoding process / apparatus 500 transmits packetization signaling 530 (eg, V-DMC packetization signaling) within the output bitstream 522. Alternatively, the packetization signaling 530 may be provided separately from the output bitstream 522 and transmitted by signaling.

[0078] like Figure 6 As shown, in the decoder 600, the input bitstream 602 is demultiplexed using the demultiplexer 604 to generate compressed bitstreams for vertex coordinates + attributes 605 and vertex connectivity 606. The input / compressed bitstream 602 may include or may be the output from the encoder 500, i.e. Figure 5 The output bitstream 522 of the encoder 500 is shown in FIG. 5 . The vertex coordinates + attribute data 605 are decompressed using an MPEG-IV-PCC decoder 608 to generate vertex attributes 612. Vertex reordering 616 is performed on the reconstructed vertex coordinates 614 at the output of the MPEG-IV-PCC decoder 608 to match the vertex order at the encoder 500. The vertex connectivity data 606 is also decompressed using a vertex connectivity decoder 610 to generate vertex connectivity information 618, and all information (including vertex attributes 612, the output of vertex reordering 616, and vertex connectivity information 618) is multiplexed using a multiplexer 620 to generate a reconstructed mesh 622.

[0079] Based on the examples described in this article, Figure 6 As shown in Figure 6The decoding process / apparatus 600 may be extended so that the decoding process / apparatus 600 receives and decodes packetization signaling 630 (eg, V-DMC packetization signaling), which may be part of the compressed bitstream 602. Figure 6 The packetized signaling 630 may include or be associated with Figure 5 Alternatively, the packetization signaling 630 may be received and signaled separately from the compressed bitstream 602 or the output bitstream 522 (eg, signaled separately from the compressed bitstream 602 to the demultiplexer 604).

[0080] General Mesh Compression

[0081] Mesh data can be compressed directly without projecting it into a 2D plane as in V-PCC based mesh codecs. In fact, the anchor for the V-PCC Mesh Compression Proposal (CfP) uses an off-the-shelf mesh compression technology, Draco (https: / / google.github.io / draco / ), to compress mesh data without textures. Draco is used to compress vertex positions, connectivity data (faces), and UV coordinates in 3D. Additional per-vertex attributes can also be compressed using Draco. The actual UV texture can be compressed using traditional video compression techniques, such as H.265 or H.264.

[0082] Draco uses the Edgebreaker algorithm at its core to compress 3D mesh information. Draco offers a good balance between simplicity and efficiency and is part of the supporting Khronos extensions for the glTF specification. The main idea of ​​the algorithm is to traverse the mesh triangles in a deterministic way so that each new triangle is encoded next to an already encoded triangle. This enables vertex specific information to be predicted from previously encoded data by simply adding deltas to the previous data. Edgebreaker utilizes symbols to signal how each new triangle is connected to previously encoded parts of the mesh. Connecting triangles in this way results in an average of 1 to 2 bits per triangle when combined with existing binary encoding techniques.

[0083] V-DC

[0084] The V-DMC standardization work started after the completion of the Call for Proposals (CfP) issued by MPEG 3DG (ISO / IEC SC29 WG2) on integrating mesh compression into the V3C standard series (ISO / IEC 23090-5). The technology retained after the analysis of the CfP results is based on multi-resolution mesh analysis and encoding and decoding. The method includes:

[0085] 1. Generate a base mesh which is a simplified (low resolution) mesh approximation of the original mesh, called the base mesh (this is done for all frames of the dynamic mesh sequence) i .

[0086] 2. Perform several mesh subdivision iteration steps on the generated base mesh (e.g. Figure 7 As shown, each triangle 700 is converted into four triangles (701, 702, 703, 704) by connecting the midpoints of the triangle edges to generate other approximate meshes m n i , where n represents the number of iterations, and m i =m 0 i .

[0087] 3. Define the displacement vector d i , also called the error vector, for each grid approximation m n i Each vertex of has n>0, denoted by d n i .

[0088] 4. For each subdivision level, use m n i +d n i The deformed mesh obtained (eg, by adding displacement vectors to the subdivided mesh vertices) generates the best approximation of the original mesh at that resolution given the base mesh and the previous subdivision level.

[0089] 5. The displacement vector may undergo a lazy wavelet transform before compression.

[0090] 6. The attribute map of the original mesh is transferred to the deformed mesh at the highest resolution (eg, subdivision level) so that the texture coordinates are obtained for the deformed mesh, and a new attribute map is generated.

[0091] The program Figure 8 Medium picture. Figure 7 A subdivision step is shown of subdividing triangle 700 into four triangles (701, 702, 703, 704) by connecting the midpoints of the initial triangle edges. Figure 8 A multi-resolution analysis of a mesh is shown. The base mesh (left, 802) undergoes a first step of subdivision and an error vector is added to each vertex (arrows at 804), and after a series of iterative subdivisions and displacements, a highest resolution mesh (right, 808 from 806) is generated. The connectivity of the highest resolution deformed mesh 808 is typically different from the original mesh 802, however the geometry of the deformed mesh 808 is a good approximation of the geometry of the original mesh.

[0092] The encoding process 900 can be divided into two main modules: a pre-processing module 902 and an actual encoder module 904, such as Fig. 9 shown. Fig. 9 The encoder 901 is shown to be composed of a preprocessing module 902, which generates a base grid 906 and a displacement vector 908 given an input grid sequence 903 and its attribute map 905. The encoder module 904 generates a compressed bit stream 910 by taking the input and output of the preprocessing module 902. The encoder 904 provides feedback 912 to the preprocessing 902.

[0093] Preprocessing 902 mainly includes three steps: decimation 1002 (reducing the original mesh resolution to produce a base mesh 906 and a decimated mesh 1004), uv-atlas isocharting 1006 (creating a parameterization of the base mesh and generating a parameterized decimated mesh 1008), and subdivision surface fitting 1010, as shown in FIG. Fig.10 Therefore, Fig.10 A pre-processing step 902 at the codec 901 is shown.

[0094] exist Fig.11 and Fig.12 , the encoder is illustrated for the intra-frame (INTRA) case (1101) and the inter-frame (INTER) case (1201). In the latter 1201, the base grid connectivity of the first frame in a group of frames is applied to the base grids of subsequent frames to improve compression performance. Fig.11 An intra-frame encoder scheme 1100 is shown.

[0095] Fig.11The encoder process 1100 for INTRA frame encoding is shown. The input to this module is the base mesh 906 (i.e., an approximation of the input mesh 903, but including fewer faces and vertices), patch information 1102 associated with the input base mesh 906, displacements 908, static / dynamic input mesh frames 903, and attribute maps 905. The output of this module is a compressed bitstream 910, which includes a V3C extended signaling sub-bitstream, which includes patch data information 1102, a compressed base mesh substream 1104, a compressed displacement video component substream 1106, and a compressed attribute video component substream 1108. Module 1101 takes the input base mesh 906 and first quantizes its data in a quantization module 1110, which can be dynamically tuned by a control module 1112. The quantized base grid is then encoded using a static grid encoder module 1114, which outputs a compressed base grid sub-bitstream 1104, which is multiplexed 1116 in the output bitstream 910. The encoded base grid is decoded in a static grid decoder module 1118, generating a reconstructed quantized base grid 1120. An update displacement module 1122 takes the reconstructed quantized base grid 1120, the original base grid 906, and the input displacement 908 as input to generate a new updated displacement 1124, which is remapped to the reconstructed base grid data to avoid precision errors due to the static grid encoding and decoding process. The updated displacement 1124 is filtered using a wavelet transform in a wavelet transform module 1126 (also taking the reconstructed base grid 1120 as input) and then quantized in a quantization module 1128. The quantized wavelet coefficients 1130 generated from the updated displacement 1124 are then packed into a video component in an image packing module 1132. The video component 1133 is then encoded using a 2D video encoder such as HEVC, VVC, etc. in a video encoder module 1134, and the output compressed displacement video component sub-bitstream 1106 is multiplexed 1116 into the output compressed bitstream 910 along with the V3C signaling information sub-bitstream 1102. The compressed displacement video component is then first decoded and reconstructed and then unpacked into encoded and quantized wavelet coefficients 1138 in an image unpacking module 1136. These wavelet coefficients 1138 are then dequantized in an inverse quantization module 1140 and reconstructed using an inverse wavelet transform module 1142 that generates a reconstructed displacement 1144. The reconstructed base grid 1146 is dequantized in an inverse quantization module 1148 , and the dequantized base grid 1146 is combined with the reconstructed displacements 1144 in a deformed grid reconstruction module 1150 to obtain a reconstructed deformed grid 1152 .The reconstructed deformed mesh 1152 is then fed into the attribute transfer module 1154 along with the attribute map 905 produced by the pre-processing 902 and the input static / dynamic mesh frame 903. The output of the attribute transfer module is an updated attribute map 1156, which now corresponds to the reconstructed deformed mesh frame 1152. The updated attribute map 1154 is then filled 1156, undergoes color conversion 1158, and is encoded as a video component using a 2D video codec (such as HEVC or VVC) in the filling module 1156, the color conversion module 1158, and the video encoder module 1160, respectively. The output compressed attribute map bitstream 1108 is multiplexed 1116 into the encoder output bitstream 910.

[0096] V3C signaling information sub-bit stream and Fig.11 The long horizontal arrow at the top corresponds to the input patch information 1102 (derived in the pre-processing step). Additional metadata is added to the patch data information during encoding, but not in Fig.11 Indicated in.

[0097] Fig.12 An inter-frame encoder scheme 1200 is shown, which is similar to the intra-frame case 1100, but the base grid connectivity is restricted for all frames in a set of frames. A motion encoder 1202 is used to efficiently encode the displacement between the base grid compared to the base grid of the first frame in a set of frames.

[0098] The inter-coding process 1200 is similar to the intra-coding process with the following changes. The reconstructed reference base grid 1146 is the input to the inter-coding process. A new module called motion encoder 1202 takes as input the quantized input base grid 906 and the reconstructed quantized reference base grid 1146 to produce compressed motion information encoded as a compressed motion bitstream 1204, which is multiplexed 1206 into the encoder output compressed bitstream 910. All other modules and processes are similar to the intra-coding case (1100, 1101).

[0099] The compressed bitstream 910 generated by the encoder 1201 multiplexes 1206 a sub-bitstream with a base mesh encoded using a static mesh codec, a sub-bitstream 1204 with motion data encoded using an animation codec for the base mesh if the INTER codec is enabled, a sub-bitstream 1208 with wavelet coefficients of displacement vectors packed in an image and encoded using a video codec 1209, a sub-bitstream 1210 with property maps encoded using a video codec 1212, and a sub-bitstream including all metadata needed to decode and reconstruct a mesh sequence based on the aforementioned sub-bitstreams. The signaling of the metadata is based on the V3C syntax and includes necessary extensions specific to the mesh.

[0100] The decoding process 1300 Fig.13 First, the compressed bitstream 910 is demultiplexed into reconstructed sub-bitstreams, such as metadata 1302, reconstructed base mesh 1304, reconstructed displacement 1306, and reconstructed property map data 1308, using a decoder 1301. Reconstruction of the mesh sequence is performed based on the data in a post-processing module 1310.

[0101] therefore, Fig.13 A decoder scheme 1300 consisting of a decoder module 1201 that demuxes and decodes all sub-bitstreams and a post-processing module 1310 is described, where the post-processing module 1310 reconstructs the dynamic grid sequence to generate an output grid 1312 and output attributes 1314.

[0102] Fig.14 and Fig.15 The decoding process in the INTRA mode 1400 and the decoding process in the INTER mode 1500 are illustrated respectively.

[0103] Fig.14The decoding process 1400 in intra mode using intra frame decoding 1401 is depicted. The intra frame decoding process includes the following modules and processes. First, the input compressed bitstream is demultiplexed 1402 into V3C extended atlas data information (or patch information) 1403, compressed static mesh bitstream 1405, compressed displacement video component 1407 and compressed attribute map bitstream 1409. The static mesh decoding module 1404 converts the compressed static mesh bitstream 1405 into a reconstructed quantized static mesh 1411 representing a base mesh. The reconstructed quantized static mesh 1411 undergoes inverse quantization in the inverse quantization module 1406 to produce a decoded reconstructed base mesh 1413. The compressed displacement video component bitstream 1407 is decoded in the video decoding module 1408 to generate a reconstructed displacement video component 1415. The displacement video component 1415 is unpacked into reconstructed quantized wavelet coefficients in the image unpacking module 1410. The reconstructed quantized wavelet coefficients are inverse quantized in the inverse quantization module 1412 and then undergo an inverse wavelet transform in the inverse wavelet transform module 1414, which produces a decoded displacement vector 1416. The deformed grid reconstruction module 1418 takes into account the patch information and takes the decoded reconstructed base grid 1413 and the decoded displacement vector 1416 as input to produce an output decoded grid frame 1312. The compressed attribute map video component 1409 is decoded using a video decoder 1420 and may undergo a color conversion 1422 to produce a decoded attribute map frame 1314 corresponding to the decoded grid frame 1312.

[0104] Fig.15 Depicted is a decoding process 1500 in inter mode using inter frame decoding 1501. The inter decoding process 1500 is similar to the intra decoding process module 1401 with the following changes. The decoder also demultiplexes the compressed information bitstream 1503. The decoded reference base grid 1502 is taken as input to the motion decoder module 1504 along with the compressed motion information sub-bitstream 1503. The decoded reference base grid 1502 is selected from a buffer of previously decoded base grid frames (by the intra decoder process 1401 for the first frame in a group of frames). The reconstruction module 1506 of the base grid takes as input the decoded reference base grid 1502 and the decoded motion information 1505 to produce a decoded reconstructed quantized base grid 1507. All other processes are similar to the intra decoding process 1401.

[0105] The signaling of metadata and substreams generated by encoder 901 and obtained by decoder 1301 is proposed as an extension of V3C in the technical proposal for dynamic grid coding and decoding CfP, and should be regarded as purely indicative. The signaling is as follows and mainly includes additional V3C unit header syntax, additional V3C unit payload syntax and grid frame intra patch data unit.

[0106] V3C Unit Header Syntax

[0107] V3C Unit Load Syntax

[0108] Patch data unit in grid frame

[0109] The refinement of metadata and sub-stream signaling is discussed as follows.

[0110] The base grid is the output of the base grid substream decoder.

[0111] A submesh is a set of vertices, their connectivity and associated attributes that can be decoded completely independently in a mesh frame. Each base mesh can have one or more submeshes.

[0112] The resampled base mesh is the output of the mesh subdivision process. The input to this process is the base mesh (or groups of sub-meshes) and information from the atlas data substream on how to subdivide / resampling the mesh (sub-meshes).

[0113] The displacement video is the output of the displacement decoder. The input to this process is the decoded geometry video and information from the atlas data substream on how to interpret / process this video. The displacement video includes the displacement values ​​to be added to the corresponding vertices.

[0114] The facegroup identifier (facegroupId) is one of the attribute types assigned to each triangular face in the resampled base mesh. The FacebookId can be compared with the id of the subpart in the patch to determine the facegroup (facegroup) corresponding to the patch. If the facegroupId is not transmitted by the base mesh substream decoder, it is derived from information in the atlas data substream.

[0115] V3C Unit

[0116] The compressed base mesh is signaled in a new substream named base mesh data substream (unit type V3C_MD). Like other v3c units, the unit type and its associated v3c parameter set id and atlas id are signaled in v3c_unit_header().

[0117] V3c parameter set extension

[0118] A new extension needs to be introduced in the v3c_parameter_set syntax structure to handle V-DMC. Several new parameters are introduced in this extension, including the following:

[0119] vps_ext_mesh_data_facegroup_id_attribute_present_flag equal to 1 indicates that one of the attribute types present in the base mesh data stream is facegroup Id.

[0120] vps_ext_mesh_data_attribute_count indicates the number of total attributes in the base mesh, including both attributes signaled by the base mesh data substream and attributes signaled in the video substream (using ai_attribute_count). When vps_ext_mesh_data_facegroup_id_attribute_present_flag is equal to 1, this value shall be greater than or equal to ai_attribute_count+1. This may be constrained by profile / level.

[0121] The type of attributes signaled by the base mesh substream but not by the video substream are signaled as vps_ext_mesh_attribute_type data type.When vps_ext_mesh_data_facegroup_id_attribute_present_flag is equal to 1, one of the vps_ext_mesh_attribute_type must be facegroup_id.

[0122] vps_ext_mesh_data_substream_codec_id indicates the identifier of the codec used to compress the base mesh data. This codec may be identified by a profile, a component codec mapping SEI message, or by means outside of this document.

[0123] vps_ext_attribute_frame_width[i] and vps_ext_attribute_frame_height[i] indicate the corresponding width and height of video data corresponding to the i-th attribute among the signaled attributes in the video substream.

[0124] AFPS Sequence Parameter Set Extension

[0125] The information included in this extension may be overridden by the same information in an AFPS extension or patch data unit. The following parameters are introduced:

[0126] asps_vmc_ext_prevent_geometry_video_conversion_flag Prevents the output of the geometry video sub-stream decoder from being converted. When the flag is true, the output is used as is without any conversion process from Annex B of ISO / IEC 23090-5. When the flag is true, the size of the geometry video shall be the same as the nominal video size indicated in the bitstream.

[0127] asps_vmc_ext_prevent_attribute_video_conversion_flag Prevents the output of the attribute video sub-stream decoder from being converted. When the flag is true, the output is used as is without any conversion process from Annex B of ISO / IEC 23090-5. When the flag is true, the size of the attribute video shall be the same as the nominal video size indicated in the bitstream.

[0128] asps_vmc_ext_subdivision_method and asps_vmc_ext_subdivision_iteration_count Signals information about the subdivision method.

[0129] asps_vmc_ext_transform_index Indicates the transform applied to the displacement. The transform index may indicate that no transform is applied. When the transform is LINEAR_LIFTING, the necessary parameters are signaled as vmc_lift_transform_parameters. asps_vmc_ext_transform_index The name of the transformation method 0 NONE 1 LINEAR_LIFTING

[0130] asps_vmc_ext_patch_mapping_method indicates how subparts of a submesh are mapped to patches. When asps_vmc_ext_patch_mapping_method is equal to 0, all triangles in the corresponding submesh are associated with the current patch. In this case, there is only one patch associated with the submesh. When asps_vmc_ext_patch_mapping_method is equal to 1, subpart_ids are explicitly signaled in the mesh patch data unit to indicate the associated subparts. In other cases, the triangular faces in the corresponding submesh are divided into subparts by the method indicated by asps_vmc_ext_patch_mapping_method.

[0131] asps_vmc_ext_tjunction_removing_method indicates the method for removing T-junctions created by different subdivision methods or by different subdivision iterations of two triangles that share an edge.

[0132] asps_vmc_ext_num_attribute indicates the total number of attributes corresponding to the mesh. Its value should be less than or equal to vps_ext_mesh_data_attribute_count.

[0133] asps_vmc_ext_attribute_type is the type of the i-th attribute, which should be one of ai_attribute_type_ids or vps_ext_mesh_attribute_type.

[0134] asps_vmc_ext_direct_attribute_projection_enabled_flag indicates that the 2d position to which the attribute is projected is explicitly signaled in the mesh patch data unit. Therefore, the projection id and orientation index in V3CV-PCC ISO / IEC 23090-5:2021 can also be used as in ISO / IEC 23090-5:2021.

[0135] asps_vmc_extension() can be as follows:

[0136] The vmc_lifting_transform_parameters syntax element can be as follows:

[0137] Atlas Frame Parameter Set Extension

[0138] afps_vmc_ext_single_submesh_in_frame_flag indicates that there is only one submesh for a mesh frame.

[0139] When affs_vmc_ext_overriden_flag in afps_vmc_extension() is true, the subdivision method, displacement coordinate system, transform index, transform parameters and attribute transform parameters can also be transmitted through signals, and the information is overwritten on the information transmitted through signals in asps_vmc_extension().

[0140] afps_vmc_ext_single_attribute_tile_in_frame_flag indicates that there is only one tile for each attribute signaled in the video stream.

[0141] The afps_vmc_extension syntax element can be as follows:

[0142] afps_ext_vmc_attribute_tile_information() includes tile information for attributes signaled by the video substream.

[0143] Atlas Block Header

[0144] A tile can be associated with one or more submeshes with id ath_submesh_id.

[0145] Patch Data Unit

[0146] Like V-PCC patch data units, mesh patch data units are signaled in the atlas data substream.Mesh intra-frame patch data units, mesh inter-frame patch data units, mesh merge patch data units and mesh skip patch data units may be used.

[0147] The patch_information_data syntax element may be as follows:

[0148] The mdu_submesh_id indicates which submesh the patch is associated with among those indicated in the atlas tile header.

[0149] mdu_vertex_count_minus1 and mdu_triangle_count_minus1 indicate the number of vertices and triangles associated with the current patch.

[0150] The syntax elements mdu_num_subparts and mdu_subparts_id are signaled when asps_vmc_ext_patch_mapping_method is not 0. When asps_vmc_ext_patch_mapping_method is 1, the associated triangular faces are the union of the triangular faces whose facegroupId is equal to mdu_subpart_id.

[0151] When mdu_patch_parameters_enable_flag is true, the subdivision method, displacement coordinate system, transform index, transform parameters and attribute transform parameters may be signaled again, and the information overwrites the corresponding information signaled in asps_vmc_extension().

[0152] The mesh_inter_data_unit syntax element can be as follows:

[0153] The mesh merge data unit can be as follows:

[0154] The mesh_skip_data_unit syntax element may be as follows: mesh_skip_data_unit(tileID,patchIdx){ Descriptors }

[0155] The mesh_raw_data_unit syntax element can be as follows:

[0156] The signaling of the underlying mesh subflows is also under investigation and can be Fig.16 (output bit stream 1602), Fig.17 (base grid bitstream 1702) and Fig.18 (Input base grid bitstream 1802) as shown.

[0157] One of the main features of the current V-DMC specification design is to support a base mesh signal that can be encoded using any current or future specified static mesh codec. For example, this type of information can be encoded and decoded using Draco three-dimensional graphics compression. This representation can provide a basis for applying other decoded information to reconstruct the output mesh frame within the context of V-DMC.

[0158] Furthermore, for encoding and decoding dynamic grid frames, it is highly desirable to be able to exploit the temporal correlation that may exist with previously encoded base grid frames. Fig.16 , Fig.17 and Fig.18 ), which is achieved by encoding (using 1604, 1704) the grid motion field instead of directly encoding the base grid (1601, 1701), and using this information and the previously encoded base grid to reconstruct the base grid of the current frame (1606, 1608, 1706, 1708, 1806, 1808). This method can be considered equivalent to inter-frame prediction in video codecs.

[0159] It is also highly desirable to associate all encoded base grid frames or motion fields with information that can help determine their decoded output order and their reference relationship. For example, better codec efficiency may be achieved if the encoding and decoding order of all frames does not follow the display order, or by using any previously encoded motion field or base grid instead of the immediately previously encoded motion field or base grid as a reference for generating the motion field for frame N. It is also highly desirable to detect random access points on the fly and to independently decode multiple sub-grids that can together form a single grid, similar to sub-pictures in video compression.

[0160] For the above reasons, a new base grid substream format is introduced. This new format is very similar to video codec formats such as the atlas subbitstream used in HEVC or V3C, where the base grid subbitstream is also constructed using NAL units. Fig.19 , high-level syntax (HLS) structures such as base grid sequence parameter set 1902, base grid frame parameter set 1904 and sub-grid layer 1906 are also specified. Fig.19 An overview of this bitstream 1900 with its different subcomponents is shown in FIG. Fig.19It is an overview of the base grid data sub-stream structure 1900.

[0161] One of the desired features of this design is the ability to divide the grid into multiple smaller partitions, which are referred to as sub-grids in this document ( Fig. 20 ). Fig. 20 shows the division of grid 2002 into sub-grids (2004 and 2006). These sub-grids (2004, 2006) can be decoded completely independently, which can contribute to partial decoding and spatial random access. Although it may be a requirement for all applications, some applications may require the division in the sub-grids to be consistent and fixed in time. The sub-grids do not need to use the same codec type. For example, for one frame, one sub-grid can use intra-frame coding, while for another frame, inter-frame coding can be used at the same decoding moment. However, it is usually required to use the same coding order, and the same reference can be used for all sub-grids corresponding at a specific moment. Such a restriction can help ensure the proper random access ability for the entire stream. Fig.21 An example of using two sub-grids is shown in.

[0162] Fig.21 shows picture order count 0 with sub-grid 2102 and sub-grid 2104, picture order count 1 with sub-grid 2112 and sub-grid 2114, and picture order count 2 with sub-grid 2120 and sub-grid 2122. The base grid frame parameter set 2130 and the base grid sequence parameter set 2140 are also shown.

[0163] NAL unit syntax

[0164] As mentioned before, the new bitstream is also based on NAL units and is similar to those bitstreams of the atlas sub-stream in V3C. The syntax is provided below.

[0165] General NAL unit syntax

[0166] The bmesh_nal_unit_header syntax element can be as follows:

[0167] NAL unit header syntax

[0168] The bmesh_nal_unit_header syntax element can be as follows:

[0169] NAL unit semantics

[0170] This section includes some semantics corresponding to the above syntax structures. More details of the syntax elements that have not been fully defined in detail will be provided.

[0171] 1. General NAL unit semantics

[0172] NumBytesInNalUnit specifies the size of the NAL unit in bytes. This value is required to decode the NAL unit. Some form of demarcation of NAL unit boundaries is necessary to implement reasoning about NumBytesInNalUnit. One such demarcation method is specified in AnnexTBD for the sample stream format. Other demarcation methods may be specified outside of this document.

[0173] The trellis codec layer (MCL) is specified to effectively represent the content of the trellis data. The NAL is specified to format the data and provide header information in a manner suitable for transmission on various communication channels or storage media. All data is included in NAL units, each unit in which includes an integer number of bytes. The NAL unit specifies a common format for both packet-oriented systems and bitstream systems. The format of the NAL unit for both the transport and sample streams of packet-oriented systems is the same, except that in the sample stream format specified in Appendix TBD, each NAL unit may be preceded by an additional element that specifies the size of the NAL unit.

[0174] rbsp_byte[i] is the i-th byte of the RBSP. The RBSP is specified as the following ordered sequence of bytes:

[0175] The RBSP includes the following string of data bits (SODB): If the SODB is empty (e.g., the length is zero bits), the RBSP is also empty; otherwise, the RBSP includes the following SODB:

[0176] The first byte of the RBSP includes the first (most significant, leftmost) eight bits of SODB; the next byte of the RBSP includes the next eight bits of SODB, and so on, until fewer than eight bits of SODB remain.

[0177] The rbsp_trailing_bits() syntax structure is as follows after the SODB: i) the first (most significant, leftmost) bit of the final RBSP byte includes the remaining bits of the SODB (if any); ii) the next bit includes a single bit equal to 1 (e.g., rbsp_stop_one_bit); iii) when the rbsp_stop_one_bit is not the last bit of the byte-aligned byte, there are one or more bits equal to 0 (e.g., an instance of rbsp_alignment_zero_bit) to enable byte alignment.

[0178] Syntax structures with these RBSP characteristics are indicated in the syntax tables using the "_rbsp" suffix. These structures are carried within NAL units as the contents of the rbsp_byte[i] data bytes. The association of the RBSP syntax structure with the NAL unit is specified in Table 4.

[0179] When the boundaries of the RBSP are known, the decoder can extract the SODB from the RBSP by concatenating the bits of the bytes of the RBSP and discarding the rbsp_stop_one_bit (the last (least significant, rightmost) bit equal to 1) and discarding any subsequent (less significant, more right) bits following it that are equal to 0. The data required for the decoding process is included in the SODB portion of the RBSP.

[0180] 2. NAL unit header semantics

[0181] Similar NAL unit types (for the atlas case) are defined for the base grid, enabling similar functionality for random access and partitioning of the grid. Unlike atlases that are partitioned into tiles, in this document we define the concept of subgrids and define specific NAL units corresponding to the coded grid data. In addition, NAL units that can include metadata such as SEI messages are defined.

[0182] Specifically, the supported base grid NAL unit types are specified as follows:

[0183] The name of the specific type can be changed to avoid confusion with the atlas type.

[0184] Raw byte sequence payload, post-bit, and byte alignment syntax

[0185] Basic grid sequence parameter set RBSP syntax

[0186] As with similar bitstreams, the main syntax structure defined for the base grid bitstream is the sequence parameter set, which includes basic information about the bitstream, identification features of the codecs supported for the intra-codec grid and the inter-codec grid, and information about references.

[0187] 1.1 General Base Grid Sequence Parameter Set RBSP Syntax

[0188] bmsps_log2_max_mesh_frame_order_cnt_lsb_minus4, bmsps_max_dec_mesh_frame_bufferin g_minus1, bmsps_long_term_ref_mesh_frames_flag, bmsps_num_ref_mesh_frame_lists_in_bmsps, bmesh_ref_list_struct(i) are equivalent to those in ASPS.

[0189] bmsps_intra_mesh_codec_id indicates the static mesh codec used to encode the base mesh in this base mesh substream. It can be associated with a specific mesh or motion mesh codec by a profile specified in the corresponding specification, or it can be explicitly indicated using SEI messages, as done in the V3C specification for video sub-bitstreams.

[0190] bmsps_intra_mesh_data_size_precision_bytes_minus1(+1) specifies the precision of the size of the encoded and decoded mesh data in bytes.

[0191] bmsps_inter_mesh_codec_present_flag indicates when a specific codec indicated by bmsps_inter_mesh_codec_id is used to encode an inter-prediction sub-mesh.

[0192] bmsps_inter_mesh_data_size_precision_bytes_minus1(+1) specifies the precision of the size of the inter prediction mesh data in bytes. This precision is signaled considering the size of the coded mesh data, and the inter prediction mesh data (eg motion field) may be significantly different.

[0193] bmsps_facegroup_segmentation_method indicates how facegroups can be derived for a mesh. A facegroup is a collection of triangular faces in a submesh. Each triangular face is associated with a FacegroupId indicating the facegroup to which it belongs. When bmsps_facegroup_segmentation_method is 0, the FacegroupId exists directly in the encoded submesh. Other values ​​indicate that facegroups can be derived using different methods based on the characteristics of the stream. For example, a value of 1 means that there is no FacebookId associated with any face. A value of 2 means that all faces are identified using a single ID, a value of 3 means that facegroups are identified based on a connected component method, and a value of 4 indicates that each individual face has its own unique ID. Currently ue(v) is used to indicate bmsps_facegroup_segmentation_method, but fixed-length encoding or division into more elements may be used instead.

[0194] 1.2 Basic Grid Class, Layer and Level Syntax

[0195] The bmptl_extended_sub_profile_flag provides support for sub profiles, which is useful when further restricting the base mesh profile according to usage and application.

[0196] 1.3 Basic Grid Frame Parameter Set RBSP Syntax

[0197] The base mesh frame parameter set has frame level information such as the number of sub-grids in a frame corresponding to one mfh_mesh_frm_order_cnt_lsb. Sub-grids are encoded and decoded in one mesh_data_submesh_layer() and are decodable independently of other sub-grids. In the case of inter-frame prediction, a sub-grid can only refer to a sub-grid with the same smh id in its associated reference frame. This mechanism is equivalent to the mechanism specified in 8.3.6.2.2 in V3C.

[0198] This mechanism is equivalent to the Atlas frame tile information syntax (8.3.6.2.2 in V3C).

[0199] The bmesh_sub_mesh_information syntax element can be as follows:

[0200] 1.4 Base Grid Subgrid Layer RBSP Syntax

[0201] 1.4.1bmesh sub-mesh layer RBSP syntax

[0202] bmesh_submesh_layer includes submesh information. One or more bmesh_submesh_layer_rbsp may correspond to one mesh frame indicated by mfh_mesh_frm_order_cnt_lsb.

[0203] 1.4.2 Subgrid Header Syntax

[0204] This mechanism is equivalent to the atlas tile header (8.3.6.11 in the standard).

[0205] smh_id is the id of the current sub-mesh included in the mesh data sub-mesh data.

[0206] smh_type indicates how the mesh is encoded and decoded. When smh_type is I_SUBMESH, the mesh data is encoded and decoded using the indicated static mesh codec. mfh_submesh_type Name of mfh_submesh_type 1 0 I_SUBMESH 2 1 P_SUBMESH 3 2 SKIP_SUBMESH

[0207] 1.4.3 Subgrid Data Unit

[0208] smdu_intra_sub_mesh_unit(unitSize) contains a sub-mesh unit stream of size unitSize (in bytes) as an ordered byte stream or bit stream where the location of the unit boundaries can be identified from the pattern in the data. The format of this sub-mesh unit stream is identified by the 4CC code defined by bmptl_profile_codec_group_idc or by the component codec mapped to the SEI message.

[0209] smdu_inter_sub_mesh_unit(unitSize) contains a sub-mesh unit stream of size unitSize (in bytes) as an ordered byte stream or bit stream where the location of the unit boundaries can be identified from the pattern in the data. The format of this sub-mesh unit stream is identified by the 4CC code defined by bmptl_profile_codec_group_idc or by the component codec mapped to the SEI message.

[0210] The V-DMC method uses a suboptimal image-based packing of displacement vectors. Although the displacements are arranged in Morton order (see Background section), this does not ensure that neighboring pixels of the packed displacements correspond to the displacements of neighboring vertices at any level of detail (LOD). Since video codecs use blocks of connected pixels (macroblocks for AVC, codec units for HEVC, etc.) to perform intra- or inter-frame prediction and DCT transforms, neighboring pixels may correspond to displacements that are not similar to each other, resulting in poor compression performance.

[0211] It can be observed that there is redundancy between adjacent displacements on the surface with or without wavelet transform (the residual after wavelet transform still exhibits local smoothness because it is derived from local curvature). Therefore, it is important to design a packing strategy that leads to the minimum possible image resolution while maintaining the correspondence between the displacement neighborhood (global or at LOD level) and the packed pixel neighborhood. The same is true for the temporal aspect, the position of the pixel neighborhood should be similar from frame to frame to achieve efficient compression performance.

[0212] Suboptimal packing of displacement vectors not only has a negative impact on the bit rate as described above, but also increases reconstruction artifacts after lossy video decoding operations.

[0213] In the V-DMC framework, the UV coordinates of the base mesh vertices are encoded in the base mesh substream, and their reconstructed values ​​are available at both the encoder and the decoder (e.g., in a closed loop).

[0214] Fig.22A and Fig. 22B Parameterization and UV texture coordinates are shown. Fig.22A is a visualization of the UV parameterization of the base mesh vertices and triangles 2201. Fig. 22B 2202 is a visualization of the texture coordinate map after midpoint subdivision (three iterations). Starting from the base mesh to the highest LOD level, a displacement vector is assigned to each vertex (position in the UV map) of each LOD. The UV coordinates of the vertices of a given LOD level are obtained by midpoint subdivision (or another subdivision as indicated in the V-DMC bitstream).

[0215] As depicted in FIG. 22 , the uv coordinates represent a parameterization in 2D of the base mesh, which is a mapping in 2D that preserves the mesh connectivity of the area and shape of the 3D triangles.

[0216] This mapping has the interesting advantage that it preserves local neighborhoods inside sub-meshes or patches. However, it does not directly map multiple vertices / vertex positions onto a regular grid that can be compressed using a video codec. We define a method to use uv mapping to define displacement vectors packed into an image, via a process that is performed in the same way at the decoder and encoder, and that maps irregularly sampled parameterizations to a regular pixel grid.

[0217] Several embodiments are described herein that cover the following: an encoding process for generating displacement image packing based on a reconstructed parameterization for each LOD; a decoding process for assigning a decoded packed displacement vector to its corresponding vertex for each LOD; and the signaling required for unpacking the displacement vectors, including two categories of pixels: empty pixels (when no displacement vector is attached to the pixel) and duplicate pixels (when more than one displacement vector is attached to the pixel).

[0218] The following embodiments are based on the fact that a static mesh decoder, included in both the V-DMC encoder and the V-DMC decoder, outputs a reconstructed base mesh including reconstructed texture coordinates.

[0219] In the V-DM framework and for the examples described herein, any parameterization method can be used to generate texture coordinates. Nevertheless, local isometric mapping (e.g., isometric maps) is preferred because it minimizes distortion once the texture map (or patch) is mapped on the mesh based on the uv coordinates. Other possible parameterization techniques include, for example, harmonic mapping, conformal mapping, or signal-specific parameterization.

[0220] The modified encoder module and decoder module are respectively Fig.23 and Fig.24 Shown in. Fig.23 Modification of the intra-frame encoder 2301 Fig.11 The encoder module 1101 shown in FIG. Fig.24 Modification of the intra-frame decoder 2401 Fig.14 The decoder module 1401 shown in FIG. Although described for the intra-frame encoding and decoding case, these figures are applicable to the inter-frame case ( Fig.12 The encoder 1201 and Fig.15 The decoder 1501) remains valid. The modified image packing module at the encoder and the modified image unpacking at the decoder are described below.

[0221] Fig.23A modified encoder scheme 2300 embodiment is shown, including an intra frame encoder 2301. The image packing module 2332 takes the reconstructed texture coordinates (2320) from the static mesh decoder 2318 and performs a reparameterization to calculate the displaced pixel positions. The same applies to the inter case. Signaling is provided to instruct the decoder how to handle unused pixels or vertex conflicts in the packed image.

[0222] Fig.24 A modified decoder scheme 2400 embodiment is depicted, including intra frame decoding 2401. The image unpacking module (2410) obtains the reconstructed texture coordinates (2411) from the static mesh decoder 2404, extracts the corresponding signaling information about unused pixels and conflicts, and performs remapping to calculate the vertex index corresponding to the decoded pixel position for each displacement vector. The same applies to the inter case.

[0223] Fig.25 A simple example is used to illustrate the principles behind the remapping technique applied in both the encoder and the decoder.

[0224] Fig.25 Example remapping techniques are shown. At the top left, an example original parameterization 2502 is overlaid with a grid (3×3 pixels 2503) representing possible image packing, which is suboptimal because many vertices 2501 (and their displacement vectors) are mapped to the same pixel 2503. At the top right (2504), a pixel distance-based approach uses the shortest distance between the optimal two vertices to define a regular grid (10×8 pixels 2503 in this example), which avoids the problem of multiple vertices 2501 being mapped to the same vertex, but generates many unused vertices. At the bottom left (2506), a relaxation with a boundary mapping technique is shown, which performs a harmonic reparameterization constrained by a 3×4 pixel grid (each pixel is 2503), the size of which is selected based on the sub-mesh boundary. This method still generates conflicts (multiple vertices are mapped to the same pixel illustrated in pixel 2503-1). At the bottom right (2508), a variation based on the relaxation approach is shown, where the boundary constraints are relaxed to allow a grid with sufficient resolution to avoid vertex conflicts, at the expense of generating unused pixels (unused pixels 2503-2 and 2503-3 are shown).

[0225] Encoder Example:

[0226] An encoder that receives, for each frame of an input volumetric video mesh: a simplified base mesh with texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; generates a static base mesh encoding that outputs a base mesh substream including encoded and quantized texture coordinates, decodes the encoded static base mesh, and generates a reconstructed dequantized base mesh and texture coordinates, adapts the precomputed displacements to the reconstructed base mesh by applying subdivision to the reconstructed base mesh, optionally filters the adapted precomputed displacements using a wavelet filter, computes a mapping of the reconstructed texture coordinates to a regular grid according to a remapping technique to generate a displacement packed image, signals metadata describing the remapping technique, and signals the presence of conflicting vertices in the displacement packed image along a bitstream, encodes the displacement video frame to generate a displacement substream, remaps the texture mapping to generate an attribute video frame, encodes the attribute video frame as a substream, signals V-DMC metadata along the substream, and multiplexes the substream into an output stream.

[0227] In another embodiment for the encoder: as above, but wherein the static mesh encoder and decoder generate a reconstructed reference base mesh for one or more frames of the sequence, and wherein a motion encoder is used to compress the static base mesh relative to the reference base mesh.

[0228] Decoder Embodiment

[0229] The decoder receives and demultiplexes the extended V-DMC bitstream into substreams, the substreams including: a base grid substream, an attribute displacement substream, an attribute texture substream and a V-DMC signaling substream.

[0230] The decoder performs operations including: extracting V-DMC metadata; decoding and dequantizing a base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame including reconstructed texture coordinates; decoding an attribute texture substream into a texture frame; decoding a displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame according to the V-DMC metadata; extracting subset metadata required to remap displaced pixel positions to corresponding vertex indices, and optionally handling conflicts as signaled using the subset metadata; applying the decoded displacements to the subdivided mesh, optionally including an inverse wavelet transform specified by the V-DMC metadata; mapping the reconstructed and color converted texture to the subdivided mesh; storing or rendering the subdivided deformed mesh as an output mesh; signaling metadata describing the remapping technique; and signaling the presence of conflicting vertices in a packed displacement image along the bitstream.

[0231] In another embodiment for the decoder: as above, but wherein the static grid decoder generates a reconstructed reference base grid for one or more frames of the sequence, and wherein the motion decoder is used to decode and reconstruct the static base grid relative to the reconstructed reference base grid.

[0232] The remapping algorithm used in the encoding and decoding embodiments is Fig.25 and as follows.

[0233] Regular grid size based on edge distance

[0234] In the remapping technique, each sub-mesh of texture coordinates is analyzed to obtain the minimum edge length. This minimum edge length can be set per sub-mesh or for the entire mesh. Once the minimum edge length for each patch is obtained, a regular grid is set so that the pixel distance is equal to the edge length. The UV coordinates are then quantized and the vertices are mapped to their nearest pixels. This method does not generate conflicts, but may introduce many unoccupied pixels. For example, these unused pixels can be repaired by a push-pull blending algorithm to optimize encoding by reducing edges. The decoder can identify unused pixels by recalculating the remapping in exactly the same way.

[0235] In another embodiment, a regular grid is set per sub-grid, and the regular grids are stitched together in the atlas by taking into account the relative positions of the sub-grids in the UV map. This remapping is performed in the same way by the decoder, which does not require an offset to convert pixel positions to vertex indices. In one embodiment, the sub-grid positions in the fused regular grid (considered as an atlas) can be labeled as patches with their coordinates, width and height.

[0236] Harmonic Reparameterization

[0237] In the harmonic reparameterization method, the encoder and decoder first extract the boundary of each subgrid and map it to a regular grid boundary. Given a subgrid of M vertices with a boundary of N vertices, define Nf as the first larger even integer, and define integers W and H such that 2*(W+H)=Nf. These are the weights W and height H of the regular grid set of this subgrid that includes W*H pixels. Depending on the subgrid connectivity, W*H can be different from M.

[0238] The encoder and decoder do the following:

[0239] Map the border vertices to the regular grid border. This can be done by starting with the vertex with the smallest index and mapping it to the top left pixel of the regular grid, then selecting its first border vertex neighbor and mapping it to the next regular grid border pixel in a clockwise manner.

[0240] Map interior vertices. Once all boundary vertices are mapped, interior vertices (non-boundary vertices) are iteratively positioned at the centroid of their neighbors in sub-pixel coordinates (boundary vertex positions remain unchanged) and quantized to the nearest 1-connected available pixel position (e.g., the delta in two pixel coordinates is -1, 0, or 1). When there is no available pixel for a vertex, the conflict score is incremented by 1.

[0241] This operation is repeated until there are no more unoccupied pixels if M is greater than W*H, or until the conflict score does not decrease after additional iterations if M is less than W*H.

[0242] At the end of this step, all vertices are mapped to pixels and either a conflict list or unused pixels are generated. When patches are defined to correspond to submeshes, the conflict list per submesh can be signaled in the atlas frame parameter set or at the patch level.

[0243] Harmonic Relaxation Reparameterization

[0244] The reparameterization of harmonic relaxation is similar to the harmonic parameterization, but instead of choosing a boundary where 2*(W+H)=Nf, W and H are increased so that W*H is greater than M and 2*(W+H) is greater than Nf. This enables collisions to be avoided. The iterative process is also modified for the boundaries. Although it is initially set as in step 1 of the harmonic reparameterization (leaving several regular grid boundary pixels unoccupied due to the larger grid), the iterative process of mapping interior vertices includes a shift of the boundary mapping whenever the collision score cannot be reduced. The boundary pixels closest to the pixel with the most collisions (or the pixel with the smallest coordinate in the case of equality) are released by shifting the adjacent boundary vertices to their adjacent positions in a clockwise manner. Once this operation is performed, the iterative process is restarted until no collisions remain.

[0245] Conflict avoidance context filling

[0246] In this remapping method, any of the above methods can be used, but instead of listing and signaling conflicts, conflicting vertices are copied with their pixel neighborhood (3x3 pixels) and filled, for example, at the bottom of a regular grid.

[0247] This conflict-avoiding remapping has the following motivation. Conflicts of (u,v) positions may occur because after remapping or in case of subdivision of base mesh triangles, more than one base vertex may be mapped to the same pixel position, with some vertices obtained by subdivision being mapped to pixel positions already used by some other vertices. Based on the fact that the metadata and remapping algorithm are the same at the encoder and decoder, the decoder can detect the exact conflicts generated by the encoder.

[0248] Some conflicts do not pose a critical issue when the displacement values ​​are similar to each other and are still within the threshold range of the encoder options. The decoder can use the stored displacement values ​​on pixels with vertex conflicts and map the available values ​​to the corresponding conflicting vertex indices. Otherwise, when the values ​​are too different for vertices that conflict on the same pixel, these values ​​are signaled, for example, in the atlas frame parameter set (see later in the document).

[0249] In another embodiment, instead of signaling this value in the asps metadata, the additional padding of the displacement attribute frame may be set to paste the conflicting value with its pixel neighbors, which may be seen as a rectangular context at the bottom of the frame, for example.

[0250] In order to store those replicated contexts with proper displacement values ​​for conflicting points, the displacement atlas resolution has to be extended on the encoder side.

[0251] In encoder embodiments, the duplication of contexts can be treated as a coding mode like other remapping methods. The encoder can choose one method or the other by performing rate-distortion optimization and marking the remapping mode in the afps metadata.

[0252] If the generated displaced images include empty pixels, they are inpainted using, for example, the same push-pull algorithm used in V3C V-PCC and MIV to inpaint the attribute or geometry component atlas.

[0253] It should be noted that in V-DMC, when inter-frame prediction is used, a reference base grid including its uv coordinates is used for the complete set of frames by encoding the motion parameters, for example using the MPEG FAMC standard. In one embodiment, when the remapping is based on a base grid, a parameterization of the reference base grid is used instead of being recalculated at each frame.

[0254] Since the remapping technique may generate conflicting or unused pixels, it is necessary to mark them (if the value is not filled in the displacement video), for example in the atlas frame parameter set.

[0255] Several image packing is possible, since the remapping can be applied to either the reconstructed base mesh texture coordinates (left side of FIG. 22 ) or the subdivided reconstructed base mesh texture coordinates (right side of FIG. 22 ). In the first case, the mapping obtained for the base mesh texture coordinates can be reused for each subdivision step, for example by mapping each new vertex obtained by subdividing an edge to the pixel position of the edge vertex with the smallest index. The replicated remapping for each LOD can be stacked into a single image in several ways, for example by using vertical stacking, which can enable random access to the LODs.

[0256] Alternatively, the remapping operation may be performed on the full resolution of the reconstructed subdivided base mesh, which maximally preserves vertex neighborhoods but reduces the ease of extracting LODs.

[0257] Fig.26 Different blocking strategies for different LODs are shown. On the left, a simple example of a base mesh 2602 before subdivision and its image packing for displacement data. On the right, the same base mesh after one iteration of subdivision (LOD1) has three blocking examples: INTERLEAVED 2610, where the packing is based on the subdivided resolution, HIERARCHICAL 2620, where the result of the base mesh packing is hierarchically blocked to create a 2x larger and taller image at each LOD level, VERTICAL_STACKING 2630, which stacks the packing achieved by the base mesh packing for each LOD, thus retaining the width of the image, but multiplying the height by 4. Unused pixels are represented as 2604. It should be noted that on the large mesh, the total number of vertices is multiplied by 4 after each subdivision step.

[0258] Signaling Embodiments.

[0259] The remapping technique may be signaled in the atlas sequence parameter set asps for the V-DMC extension.

[0260] asps_vmc_ext_attribute_packing_method indicates the attribute (displacement or wavelet coefficient) packing method, with the values ​​defined in the table below.TRAVERSAL corresponds to the raster and Morton order used in the V-DMC test model, which converts a 1D traversal of attributes or wavelet coefficients into an image.PARAMETERIZATION_GUIDED corresponds to a packing method that is based on a reconstructed base mesh 2D parameterization (texture coordinates).

[0261] asps_vmc_ext_ap_remapping_method indicates the remapping method to be used when asps_vmc_ext_attribute_packing_method is equal to PARAMETERIZATION_GUIDED. The following table defines the following methods. EDGE_DISTANCE describes a remapping method in which the minimum edge length is used to calculate the packing grid resolution. HARMONIC describes a relaxation method based on boundary constraints, namely harmonic reparameterization. HARMONIC_RELAXED indicates a remapping method using harmonic reparameterization with relaxed boundary constraints.

[0262] The width and height of the attribute corresponding to the displacement or wavelet coefficient are provided by asps_vmc_ext_attribute_frame_width[i] and asps_vmc_ext_attribute_frame_height[i] respectively, where i represents the index of the attribute index corresponding to the displacement or wavelet coefficient. asps_vmc_ext_ap_remapping_method The name of the remapping method 0 EDGE_DISTANCE 1 HARMONIC 2 HARMONIC_RELAXED 3 RESERVED

[0263] asps_vmc_ext_ap_tiling_method indicates the blocking method used after packing when asps_vmc_ext_attribute_packing_method is equal to PARAMETERIZATION_GUIDED. These values ​​are described in the table below. VERTICAL_STACKING indicates that packing is obtained from the base mesh vertex remapping, and is copied and stacked at the bottom, keeping the width of the first map for other LOD iterations. HIERARCHICAL indicates that each LOD map is copied from the LOD map derived for the base mesh, and the LOD map is blocked to the right, bottom, and bottom right of the previous LOD level in a similar manner to that done for the 2D discrete wavelet transform of the image. INTERLEAVED indicates that remapping is performed at the full resolution of the subdivided mesh, resulting in interleaved packing of displacements for all LODs.

[0264] The asps_vmc_ext_ap_tiling_method syntax element may be as follows: asps_vmc_ext_ap_tiling_method The name of the chunking method 0 VERTICAL_STACKING 1 HIERARCHICAL 2 INTERLEAVED 3 RESERVED

[0265] Conflicting or unused pixels may be signaled to the decoder in different ways. In the case of unused pixels, the decoder may identify such pixels and avoid considering them when remapping displacements to reconstructed mesh vertices, so no specific signaling is required, but the number of unused pixels may be sent for error detection at the decoder.

[0266] Conflicts require displacements to be marked in the following way. By convention, the displacement corresponding to the vertex with the smallest index is encoded and decoded in the packed image. For other values, the following signaling is used. They can be signaled in the atlas frame parameter set.

[0267] afps_vmc_ext_ap_remapping_collisions_number_minus1 indicates the number of pixels that have multiple values ​​mapped and require additional signaling of their unpacked values.

[0268] afps_vmc_ext_ap_collision_first_index indicates the index in the attribute frame of the first collision.

[0269] afps_vmc_ext_ap_collision_attribute_number_minus1[i] indicates the number of values ​​encoded for this attribute frame pixel position.

[0270] afps_vmc_ext_ap_collision_attribute_vector[i][j] indicates the value of the vector.

[0271] afps_vmc_ext_ap_collision_attribute_scalar[i][j] indicates the value of the scalar attribute if afps_vmc_ext_displacement_coordinate_system_enable_flag is set to true.

[0272] afps_vmc_ext_ap_collision_offset[i] indicates the pixel offset towards the next collision.

[0273] The afps_vmc_extension syntax element can be as follows:

[0274] Fig. 27The device 2700 can be implemented in hardware and is configured to implement parameterized guided packing of displacements for dynamic mesh encoding and decoding based on any of the examples described herein. The device includes a processor 2702, at least one memory 2704 including computer program code 2705 (the memory 2704 can be non-transient, transient, non-volatile or volatile), wherein the at least one memory 1404 and the computer program code 2705 are configured to, together with the at least one processor 2702, cause the device to implement circuit systems, processes, components, modules, functions, encoding and decoding and / or decoding (collectively referred to as 2706) to implement hierarchical encoding and decoding of texture and geometry data for dynamic meshes based on the examples described herein. The device 2700 is also configured to provide or receive signaling 2707 based on the signaling embodiments described herein. The device 2700 optionally includes a display and / or I / O interface 2708, which can be used to display the output (e.g., image or volumetric video) of the result of the encoding and decoding 2706. The display and / or I / O interface 2708 may also be configured to receive input such as user input (e.g., using a keypad, a touch screen, a touch area, a microphone, biometrics, one or more sensors, etc.). The device 2700 also includes one or more communication interfaces ((multiple) I / Fs) 2710, such as a network (N / W) interface. (Multiple) communication I / Fs 2710 may be wired and / or wireless and communicate via any communication technology through a channel or the Internet / other network. (Multiple) communication I / Fs 2710 may include one or more transmitters and one or more receivers. (Multiple) communication I / Fs 2710 may include standard well-known components such as amplifiers, filters, frequency converters, (de) modulators, and (multiple) encoder / decoder circuit systems and one or more antennas. In some examples, the processor 2702 is configured to implement item 2706 and / or item 2707 without using memory 2704.

[0275] Device 2700 may be a remote, virtual or cloud device. Device 2700 may be a writer or a reader (e.g., a parser), or both a writer and a reader (e.g., a parser). Device 2700 may be a codec or a decoder, or both a codec and a decoder (codec). Device 2700 may be a user equipment (UE), a head mounted display (HMD), or any other fixed or mobile device.

[0276] Memory 2704 may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. Memory 2704 may include a database for storing data. Interface 2712 enables data communication between various items of apparatus 2700, such as Fig. 27 As shown. Interface 2712 may be one or more buses. For example, interface 2712 may be one or more buses, such as an address, data, or control bus, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, optical fiber or other optical communication device, etc. Computer program code 2705 may include object-oriented software. Apparatus 2700 need not include every feature mentioned, or may include other features. Apparatus 2700 may be Figure 1A , Figure 1B , Figure 5 , Figure 6 , Fig.11 , Fig.14 , Fig.23 , Fig.24 or an embodiment of any device shown in any other figure described and shown herein and having features thereof.

[0277] Fig.28 A schematic diagram of non-volatile memory media 2800a (e.g., a computer / compact disk (CD) or digital versatile disk (DVD)) and 2800b (e.g., a universal serial bus (USB) memory stick) storing instructions and / or parameters 2802 that, when executed by a processor, enable the processor to perform one or more steps of the methods described herein.

[0278] Fig.29 The method 2900 for implementing the examples described herein is provided. At 2910, the method includes, with an encoder for a frame of an input volumetric video mesh, receiving: a simplified base mesh with texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame. At 2920, the method includes generating a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh. At 2930, the method includes decoding the encoded static base mesh, and generating a reconstructed dequantized base mesh and reconstructed texture coordinates. At 2940, the method includes adapting the precomputed displacements to the reconstructed dequantized base mesh by applying subdivision to the reconstructed dequantized base mesh. At 2950, ​​the method includes calculating a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image. At 2960, the method includes signaling metadata describing the remapping technique. At 2970, the method includes signaling the presence of conflicting vertices in the displacement packed image along the bitstream. The method 1500 may be performed using an encoder device (eg, 1101 , 1201 , 2301 , 2700 ).

[0279] Fig.30The method 3000 for implementing the examples described herein. At 3010, the method includes receiving and demultiplexing an extended video-based dynamic mesh codec bitstream into substreams using a decoder, the substreams including: a base mesh substream, a displacement substream, an attribute substream, and a video-based dynamic mesh codec signaling substream. At 3020, the method includes extracting video-based dynamic mesh codec metadata. At 3030, the method includes decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame including reconstructed texture coordinates. At 3040, the method includes decoding the attribute substream into a texture frame. At 3050, the method includes decoding the displacement substream into a displacement image. At 3060, the method includes performing a subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh. At 3070, the method includes extracting subset metadata for remapping the displacement pixel positions of the displacement image to corresponding vertex indices. At 3080, the method includes applying the decoded displacement to the subdivided mesh as specified using the video-based dynamic mesh codec metadata. The method 3000 may be performed using a decoder device (eg, 1401, 1501, 2401, 2700).

[0280] The following examples are provided and described in this article.

[0281] Example 1. A device comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the device to at least: receive, using an encoder for a frame of an input volumetric video mesh: a simplified base mesh with texture coordinates, a set of pre-computed displacements for a selected subdivision method, and a texture frame; generate a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; decode the encoded static base mesh, and generate a reconstructed dequantized base mesh and reconstructed texture coordinates; adapt the pre-computed displacements to the reconstructed dequantized base mesh by applying subdivision to the reconstructed dequantized base mesh; calculate a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image; transmit metadata describing the remapping technique by a signal; and transmit the presence of conflicting vertices in the displacement packed image by a signal along a bitstream.

[0282] Example 2. The apparatus of example 1, wherein the apparatus is caused to filter the adapted pre-computed displacement using a wavelet filter.

[0283] Example 3. An apparatus according to any of Examples 1 to 2, wherein the displacement packed image comprises a displacement video frame.

[0284] Example 4. The apparatus of Example 3, wherein the apparatus is caused to: encode the displaced video frame to generate a displaced substream.

[0285] Example 5. An apparatus according to any one of Examples 1 to 4, wherein the apparatus is caused to: remap the texture map to generate the attribute video frame.

[0286] Example 6. The apparatus of Example 5, wherein the apparatus is caused to: encode the attribute video frame as a first substream.

[0287] Example 7. The apparatus of Example 6, wherein the apparatus is caused to: signal video-based dynamic lattice coding (V-DMC) metadata along the second substream.

[0288] Example 8. The apparatus of Example 7, wherein the apparatus is caused to: multiplex the first sub-stream and the second sub-stream into an output stream.

[0289] Example 9. An apparatus according to any one of Examples 1 to 8, wherein the static base grid encoding is performed using a static grid encoder, and the decoding of the encoded static base grid is performed using a static grid decoder.

[0290] Example 10. An apparatus according to Example 9, wherein the apparatus is caused to: generate a reconstructed reference base mesh for one or more frames of a sequence of input volumetric video meshes using a static mesh encoder and a static mesh decoder; and compress the static base mesh relative to the reconstructed reference base mesh using a motion encoder.

[0291] Example 11. A device comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the device to at least: utilize a decoder to receive and demultiplex an extended video-based dynamic mesh codec bitstream into substreams, the substreams comprising: a base mesh substream, a displacement substream, an attribute substream, and a video-based dynamic mesh codec signaling substream; extract video-based dynamic mesh codec metadata; decode and dequantize the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame comprising reconstructed texture coordinates; decode the attribute substream into a texture frame; decode the displacement substream into a displacement image; perform subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extract subset metadata for remapping displaced pixel positions of the displacement image to corresponding vertex indices; and apply the decoded displacement to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

[0292] Example 12. The apparatus of Example 11, wherein the apparatus is caused to: handle conflicts as signaled using the subset metadata.

[0293] Example 13. An apparatus according to any one of Examples 11 to 12, wherein the apparatus is caused to apply an inverse wavelet transform to the subdivided mesh as specified using video-based dynamic mesh codec (V-DMC) metadata.

[0294] Example 14. An apparatus according to any of Examples 11 to 13, wherein the apparatus is caused to: generate a reconstructed and color converted texture; and map the reconstructed and color converted texture to the subdivided mesh.

[0295] Example 15. The apparatus of any one of Examples 11 to 14, wherein the apparatus is caused to: store or render the subdivided deformed mesh as an output mesh.

[0296] Example 16. An apparatus according to any one of Examples 11 to 15, wherein the apparatus is caused to: transmit, by signaling, metadata describing a remapping technique, the remapping technique being used for remapping of displaced pixel positions.

[0297] Example 17. An apparatus according to any one of Examples 11 to 16, wherein the apparatus is caused to: signal along a bitstream the presence of conflicting vertices in the packed displacement image.

[0298] Example 18. An apparatus according to any one of Examples 11 to 17, wherein the apparatus is caused to: generate a reconstructed reference base grid for one or more frames of a sequence of volumetric video grids using a static grid decoder; and decode and reconstruct the static base grid relative to the reconstructed reference base grid using a motion encoder.

[0299] Example 19. A method comprising: utilizing an encoder for a frame of an input volumetric video mesh, receiving: a simplified base mesh having texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; generating a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; decoding the encoded static base mesh, and generating a reconstructed dequantized base mesh and reconstructed texture coordinates; adapting the precomputed displacements to the reconstructed dequantized base mesh by applying subdivision to the reconstructed dequantized base mesh; computing a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image; signaling metadata describing the remapping technique; and signaling the presence of conflicting vertices in the displacement packed image along a bitstream.

[0300] Example 20. The method of Example 19, further comprising filtering the adapted pre-computed displacement using a wavelet filter.

[0301] Example 21. A method according to any one of Examples 19 to 20, wherein the displacement packed image includes a displacement video frame.

[0302] Example 22. The method of Example 21 further comprises encoding the displaced video frame to generate a displaced sub-stream.

[0303] Example 23. A method according to any one of Examples 19 to 22, further comprising remapping the texture map to generate an attribute video frame.

[0304] Example 24. The method of Example 23 further comprises: encoding the attribute video frame as a first substream.

[0305] Example 25. The method of Example 24, further comprising: signaling video-based dynamic lattice codec (V-DMC) metadata along the second substream.

[0306] Example 26. The method of Example 25 further comprises multiplexing the first substream and the second substream into an output stream.

[0307] Example 27. A method according to any one of Examples 19 to 26, wherein static base grid encoding is performed using a static grid encoder, and decoding of the encoded static base grid is performed using a static grid decoder.

[0308] Example 28. The method of Example 27 also includes: generating a reconstructed reference base mesh for one or more frames of a sequence of input volumetric video meshes using a static mesh encoder and a static mesh decoder; and compressing the static base mesh relative to the reconstructed reference base mesh using a motion encoder.

[0309] Example 29. A method comprising: using a decoder, receiving and demultiplexing an extended video-based dynamic mesh codec bitstream into substreams, the substreams including: a base mesh substream, a displacement substream, an attribute substream and a video-based dynamic mesh codec signaling substream; extracting video-based dynamic mesh codec metadata; decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame including reconstructed texture coordinates; decoding the attribute substream into a texture frame; decoding the displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extracting subset metadata for remapping displaced pixel positions of the displacement image to corresponding vertex indices; and applying the decoded displacement to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

[0310] Example 30. The method of Example 29 further comprising handling conflicts as signaled using the subset metadata.

[0311] Example 31. The method of any one of Examples 29 to 31, further comprising: applying an inverse wavelet transform to the subdivided mesh as specified using video-based dynamic mesh codec (V-DMC) metadata.

[0312] Example 32. The method of any one of Examples 29 to 31, further comprising: generating a reconstructed and color-converted texture; and mapping the reconstructed and color-converted texture to the subdivided mesh.

[0313] Example 33. The method of any one of Examples 29 to 32, further comprising storing or rendering the subdivided deformed mesh as an output mesh.

[0314] Example 34. The method according to any one of Examples 29 to 33 further includes: transmitting metadata describing a remapping technique via a signal, the remapping technique being used to remap displaced pixel positions.

[0315] Example 35. The method of any one of Examples 29 to 34, further comprising: transmitting the presence of conflicting vertices in the packed displacement image via a signal along a bitstream.

[0316] Example 36. The method according to any one of Examples 29 to 35 further includes: using a static mesh decoder to generate a reconstructed reference base mesh for one or more frames of a sequence of volumetric video meshes; and using a motion encoder to decode and reconstruct the static base mesh relative to the reconstructed reference base mesh.

[0317] Example 37. An apparatus comprising: a component for receiving, with an encoder for a frame of an input volumetric video mesh: a simplified base mesh having texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; a component for generating a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; a component for decoding the encoded static base mesh and generating a reconstructed dequantized base mesh and reconstructed texture coordinates; a component for adapting the precomputed displacements to the reconstructed dequantized base mesh by applying subdivision to the reconstructed dequantized base mesh; a component for calculating a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image; a component for transmitting metadata describing the remapping technique via a signal; and a component for transmitting the presence of conflicting vertices in the displacement packed image via a signal along a bitstream.

[0318] Example 38. A device comprising: a component for receiving and demultiplexing an extended video-based dynamic mesh codec bitstream into substreams using a decoder, the substreams including: a base mesh substream, a displacement substream, an attribute substream and a video-based dynamic mesh codec signaling substream; a component for extracting video-based dynamic mesh codec metadata; a component for decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame including reconstructed texture coordinates; a component for decoding the attribute substream into a texture frame; a component for decoding the displacement substream into a displacement image; a component for performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; a component for extracting subset metadata for remapping displaced pixel positions of the displacement image to corresponding vertex indices; and a component for applying the decoded displacement to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

[0319] Example 39. A non-volatile program storage device is provided, readable by a machine, the non-volatile program storage device tangibly embodying a program of instructions, the instructions executable together with the machine, for performing operations, the operations comprising: receiving, using an encoder for a frame of an input volumetric video mesh: a simplified base mesh having texture coordinates, a set of pre-computed displacements for a selected subdivision method, and a texture frame; generating a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; decoding the encoded static base mesh, and generating a reconstructed dequantized base mesh and reconstructed texture coordinates; adapting the pre-computed displacements to the reconstructed dequantized base mesh by applying subdivision to the reconstructed dequantized base mesh; calculating a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image; transmitting metadata describing the remapping technique by signaling; and transmitting the presence of conflicting vertices in the displacement packed image by signaling along a bitstream.

[0320] Example 40. A non-volatile program storage device is provided, which is readable by a machine, and the non-volatile program storage device tangibly embodies a program of instructions, which are executable together with the machine to perform operations, the operations including: using a decoder, receiving and demultiplexing an extended video-based dynamic mesh codec bitstream into substreams, the substreams including: a base mesh substream, a displacement substream, an attribute substream and a video-based dynamic mesh codec signaling substream; extracting video-based dynamic mesh codec metadata; decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame including reconstructed texture coordinates; decoding the attribute substream into a texture frame; decoding the displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extracting subset metadata for remapping the displacement pixel positions of the displacement image to corresponding vertex indices; and applying the decoded displacement to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

[0321] References to "computers," "processors," and the like should be understood to cover not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures, but also special purpose circuits such as field programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices, and other processing circuit systems. References to computer programs, instructions, code, and the like should be understood to cover software for a programmable processor or firmware, such as, for example, programmable content of a hardware device, such as instructions for a processor, or configuration settings for a fixed function device, gate array, or programmable logic device, and the like.

[0322] As used herein, the term "circuitry" may refer to any of the following: (a) a hardware circuit implementation, such as an implementation in analog and / or digital circuitry, and (b) a combination of circuitry and software (and / or firmware), such as (where applicable): (i) a combination of (multiple) processors or (ii) portions of (multiple) processors / software, including (multiple) digital signal processors, software and (multiple) memories that work together to enable the device to perform various functions, and (c) circuitry, such as (multiple) microprocessors or portions of (multiple) microprocessors, which require software or firmware for operation, even when the software or firmware is not physically present. As another example, as used herein, the term "circuitry" would also cover an implementation of only a processor (or multiple processors) or a portion of a processor and its (or its) accompanying software and / or firmware. The term "circuitry" would also cover, for example and when applicable to the particular element, a baseband integrated circuit or an application processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device or another network device. Circuitry may also be used to mean a function or process, such as a function or process implemented by an encoder or decoder or codec.

[0323] In the figures, arrows between blocks indicate the operational couplings therebetween and the direction of data flow over those couplings.

[0324] It should be understood that the foregoing description is illustrative only. Those skilled in the art may design various alternatives and modifications. For example, the features described in the various dependent claims may be combined with each other in any suitable (multiple) combinations. In addition, the features from the above-mentioned different embodiments may be selectively combined into new embodiments. Therefore, this specification is intended to cover all such alternatives, modifications and variations that fall within the scope of the appended claims.

[0325] The following abbreviations and abbreviations that may be found in the specification and / or drawings are defined as follows: 2D and variant 2D 3D and Variant 3D 3DG 3D Graphics Codec Group 4CC four-character codec 6DOF Six Degrees of Freedom ACL Atlas Codec Layer AD Atlas Data afoc atlas frame order count AFPS and Variant Atlas Frame Parameter Sets ai attribute index ap attribute packaging AR Augmented Reality ASIC Application-Specific Integrated Circuit ASPS and Variant Atlas Sequence Parameter Sets ath atlas block header AUD Access Unit Delimiter Aux AVD properties video data BFPS and variant base grid frame parameter sets BLA Broken Link Access BMCL Basic Mesh Codec Layer bmesh base mesh BMFPS Base Grid Frame Parameter Set bmptl Basic grid levels, layers and classes BMSPS and variant base grid sequence parameter sets b(n) n bits bsmi base grid subgrid information BSPS Base Grid Sequence Parameter Set CD CfP Call for Proposals CGI Computer Generated Imagery cnt count CRA Completely Random Access DCT Discrete Cosine Transform DVD Digital Versatile Disc EOB End of bitstream EOS sequence ends ESEI Essential Supplementary Enhancement Information Exp Index ext extension FAMC Frame-based animation mesh compression FD fill data Final draft of FDIS international standard f(n) is a floating point with n bits, e.g. f(1) glTF Graphics Language Transmission Format GVD Geometric Video Data H.264 Advanced Video Codec Video Compression Standard H.265 High-efficiency video codec video compression standard HEVC High Efficiency Video Codec HLS Advanced Syntax HMD Head-mounted Display ID and variant identifiers IDC Indication Idx Index IDR Instant Decode Refresh IEC International Electrotechnical Commission I / F Interface I / O Input / Output IRAP Intra-frame random access point ISO International Organization for Standardization len length LOD and Variant Level of Detail LP lead image lsb least significant bit lt Lifting Transform ltp lifting transformation parameters MCL mesh codec layer MD Mesh Data mdu intra-mesh patch data unit mfoc grid frame order count midu inter-grid data unit miv and variants MPEG immersive video mmdu Grid Merge Data Unit mpdu Grid Patch Data Unit MPEG Moving Picture Experts Group MPEG-I MPEG Immersive MR Mixed Reality mrdu grid raw data unit msh Grid MUX multiplexing NAL and variants of the Network Abstraction Layer NBMCL Non-BMCL NSEI Non-essential Supplementary Enhanced Information N / W Network OVD Occupied Video Data poc picture sequence count Pos Quant RADL Random Access Decodable Preamble RASL Random Access Skip Preamble rbsp and variant raw byte sequence payloads ref RSV Retention SC Subcommittee SEI Supplemental Enhancement Information se(v) Signed integer order 0 exponential-Golomb encoding with left-bit priority (i.e., most significant bit first) smdu subgrid data unit smh and variant subgrids SODB Data bit string STSA Stepped Temporal Sublayer Access TBD To be determined TSA Time Sublayer Access ti Tile Information u(n) uses an n-bit unsigned integer, such as u(1), u(2) UE User Equipment ue(v) unsigned integer exponent with left-first digit Golomb-coded syntax element UNSPEC Unspecified USB Universal Serial Bus uv and variant coordinate textures, where "U" and "V" are the axes of the 2D texture. V3C Visual Volumetric Video Codec V-DMC or VDMC Video-based dynamic mesh codec vmc Volume Mesh Compression VPCC or V-PCC Video-based point cloud encoding / decoding / compression VPS V3C Parameter Set VR Virtual Reality vuh volume unit head VVC Multi-functional Video Codec WG Working Group

Claims

1. A device, include: at least one processor; as well as at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: With an encoder for a frame of an input volumetric video mesh, receiving: a simplified base mesh having texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; Generate a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; Decoding the encoded static base mesh and generating a reconstructed inverse quantized base mesh and reconstructed texture coordinates; adapting the pre-computed displacements to the reconstructed dequantized base grid by applying a subdivision to the reconstructed dequantized base grid; Based on a remapping technique, calculating a remapping of the reconstructed texture coordinates to a regular grid to generate a displacement packed image; signaling metadata describing the remapping technique; as well as The presence of conflicting vertices is signaled in the displacement packed image along a bitstream.

2. The apparatus according to claim 1, wherein the apparatus is caused to: The adapted pre-computed displacements are filtered using a wavelet filter.

3. The apparatus of any one of claims 1 to 2, wherein the displacement packed image comprises a displacement video frame.

4. The apparatus of claim 3, wherein the apparatus is caused to: The displaced video frame is encoded to generate a displaced sub-stream.

5. The device according to any one of claims 1 to 4, wherein the device is caused to: Remap texture maps to generate attributed video frames.

6. The apparatus of claim 5, wherein the apparatus is caused to: The attribute video frame is encoded as a first substream.

7. The apparatus of claim 6, wherein the apparatus is caused to: Video-based dynamic trellis codec (V-DMC) metadata is signaled along the second substream.

8. The apparatus of claim 7, wherein the apparatus is caused to: The first substream and the second substream are multiplexed into an output stream.

9. The apparatus according to any one of claims 1 to 8, wherein the static base grid encoding is performed using a static grid encoder, and the decoding of the encoded static base grid is performed using a static grid decoder.

10. The apparatus of claim 9, wherein the apparatus is caused to: generating, using the static mesh encoder and the static mesh decoder, a reconstructed reference base mesh for one or more frames of the sequence of input volumetric video meshes; and The static base mesh is compressed relative to the reconstructed reference base mesh using a motion encoder.

11. A device, include: at least one processor; as well as at least one memory storing instructions, which, when executed by the at least one processor, cause the apparatus to at least: Using a decoder, receiving and demultiplexing an extended video-based dynamic grid codec bit stream into substreams, wherein the substreams include: a basic grid substream, a displacement substream, an attribute substream, and a video-based dynamic grid codec signaling substream; Extract video-based dynamic mesh codec metadata; Decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, wherein the reconstructed base mesh frame includes reconstructed texture coordinates; Decoding the attribute substream into a texture frame; Decoding the displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extracting subset metadata for remapping displaced pixel positions of the displaced image to corresponding vertex indices; and The decoded displacements are applied to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

12. The apparatus of claim 11, wherein the apparatus is caused to: Conflicts are handled as signaled using the subset metadata.

13. The device according to any one of claims 11 to 12, wherein the device is caused to: An inverse wavelet transform is applied to the subdivided mesh as specified using the video-based dynamic mesh codec (V-DMC) metadata.

14. The device according to any one of claims 11 to 13, wherein the device is caused to: generating a reconstructed and color-converted texture; and The reconstructed and color converted texture is mapped to the subdivided mesh.

15. The device according to any one of claims 11 to 14, wherein the device is caused to: The tessellated deformed mesh is stored or rendered as an output mesh.

16. The device according to any one of claims 11 to 15, wherein the device is caused to: Metadata describing a remapping technique used for said remapping of said displaced pixel positions is signaled.

17. The device according to any one of claims 11 to 16, wherein the device is caused to: The presence of conflicting vertices in the packed displacement image is signaled along the bitstream.

18. The device according to any one of claims 11 to 17, wherein the device is caused to: generating, using a static mesh decoder, a reconstructed reference base mesh for one or more frames of a sequence of volumetric video meshes; and Using a motion encoder, a static base grid is decoded and reconstructed relative to the reconstructed reference base grid.

19. A method, include: With an encoder for a frame of an input volumetric video mesh, receiving: a simplified base mesh having texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; Generate a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; Decoding the encoded static base mesh and generating a reconstructed inverse quantized base mesh and reconstructed texture coordinates; adapting the pre-computed displacements to the reconstructed dequantized base grid by applying a subdivision to the reconstructed dequantized base grid; Based on a remapping technique, calculating a remapping of the reconstructed texture coordinates to a regular grid to generate a displacement packed image; signaling metadata describing the remapping technique; as well as The presence of conflicting vertices in the displacement packed image is signaled along the bitstream.

20. A method, include: Using a decoder, receiving and demultiplexing an extended video-based dynamic grid codec bit stream into substreams, wherein the substreams include: a basic grid substream, a displacement substream, an attribute substream, and a video-based dynamic grid codec signaling substream; Extract video-based dynamic mesh codec metadata; Decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, wherein the reconstructed base mesh frame includes reconstructed texture coordinates; Decoding the attribute substream into a texture frame; Decoding the displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extracting subset metadata for remapping displaced pixel positions of the displaced image to corresponding vertex indices; and The decoded displacements are applied to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

21. A device, include: An encoder for utilizing a frame for an input volumetric video mesh receives: a simplified base mesh with texture coordinates, a set of precomputed displacements for a selected subdivision method, and components of a texture frame; means for generating a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates and generating an encoded static base mesh; Means for decoding the encoded static base mesh and generating a reconstructed inverse quantized base mesh and reconstructed texture coordinates; means for adapting said pre-computed displacements to said reconstructed dequantized base grid by applying a subdivision to said reconstructed dequantized base grid; A component for calculating a remapping of the reconstructed texture coordinates to a regular grid based on a remapping technique to generate a displacement packed image; means for signaling metadata describing the remapping technique; as well as Means for signaling along a bitstream the presence of conflicting vertices in said displacement packed image.

22. A device, include: A component for receiving and demultiplexing the extended video-based dynamic grid codec bitstream into substreams using a decoder, wherein the substreams include: a base grid substream, a displacement substream, an attribute substream, and a video-based dynamic grid codec signaling substream; A component for extracting video-based dynamic mesh codec metadata; means for decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, the reconstructed base mesh frame comprising reconstructed texture coordinates; A component for decoding the attribute substream into a texture frame; A component for decoding the displacement substream into a displacement image; means for performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; means for extracting subset metadata for remapping displaced pixel positions of the displaced image to corresponding vertex indices; and Means for applying the decoded displacements to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.

23. A non-transitory program storage device readable by a machine, the non-transitory program storage device tangibly embodying a program of instructions executable with the machine for performing operations, the operations include: With an encoder for a frame of an input volumetric video mesh, receiving: a simplified base mesh having texture coordinates, a set of precomputed displacements for a selected subdivision method, and a texture frame; Generate a static base mesh encoding, the static base mesh encoding outputting a base mesh substream including encoded and quantized texture coordinates, and generating an encoded static base mesh; Decoding the encoded static base mesh and generating a reconstructed inverse quantized base mesh and reconstructed texture coordinates; adapting the pre-computed displacements to the reconstructed dequantized base grid by applying a subdivision to the reconstructed dequantized base grid; Based on a remapping technique, calculating a remapping of the reconstructed texture coordinates to a regular grid to generate a displacement packed image; signaling metadata describing the remapping technique; as well as The presence of conflicting vertices in the displacement packed image is signaled along the bitstream.

24. A non-transitory program storage device readable by a machine, the non-transitory program storage device tangibly embodying a program of instructions executable with the machine for performing operations, the operations include: Using a decoder, receiving and demultiplexing an extended video-based dynamic grid codec bit stream into substreams, wherein the substreams include: a basic grid substream, a displacement substream, an attribute substream, and a video-based dynamic grid codec signaling substream; Extract video-based dynamic mesh codec metadata; Decoding and dequantizing the base mesh substream into a reconstructed base mesh frame, wherein the reconstructed base mesh frame includes reconstructed texture coordinates; Decoding the attribute substream into a texture frame; Decoding the displacement substream into a displacement image; performing subdivision of the reconstructed base mesh frame based on the video-based dynamic mesh codec metadata to generate a subdivided mesh; extracting subset metadata for remapping displaced pixel positions of the displaced image to corresponding vertex indices; and The decoded displacements are applied to the subdivided mesh as specified using the video-based dynamic mesh codec metadata.