Signaling displacement data for video-based grid coding
By compressing the attribute information and spatial information of the three-dimensional grid, a compressed bit stream is generated, which solves the high cost and time-consuming storage and transmission of three-dimensional visual content, and achieves the effect of rapid storage and transmission.
Patent Information
- Application Number
- CN202380074937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2023-10-26
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is costly and time-consuming when storing and transmitting large amounts of three-dimensional visual content, especially the texture or attribute information of the three-dimensional grid occupies a large amount of data.
The attribute information and spatial information of the three-dimensional grid are compressed through the encoder system to generate a compressed bit stream, including compressed basic grid, displacement value and attribute information. Use the decoding unit to block and slice the grid to improve compression efficiency.
It realizes the rapid storage and transmission of three-dimensional visual content, reduces the use of storage space and improves the transmission speed.
Smart Images

Figure CN120112950A_ABST
Abstract
Description
Background Art Technical Field
[0001] The present disclosure generally relates to compression and decompression of three-dimensional meshes with associated textures or attributes.
[0002] Related technical description
[0003] Various types of sensors (such as light detection and ranging (LIDAR) systems, 3D cameras, 3D scanners, etc.) can capture data indicating the location of points in three-dimensional space (e.g., locations in the X, Y, and Z planes). In addition, such systems can capture attribute information, such as color information (e.g., RGB values), texture attributes, intensity attributes, reflectivity attributes, motion-related attributes, modal attributes, or various other attributes, in addition to spatial information about the corresponding points. In some cases, additional attributes can be assigned to the corresponding points, such as a timestamp when the point was captured. The points captured by such sensors can constitute a "point cloud", which includes a set of points each having associated spatial information and one or more associated attributes. In some cases, a point cloud can include thousands of points, hundreds of thousands of points, millions of points, or even more points. In addition, in some cases, a point cloud can be generated, for example, in software, as opposed to a point cloud being captured by one or more sensors. In either case, such a point cloud can include a large amount of data, and storing and transmitting these point clouds can be costly and time-consuming. Furthermore, three-dimensional visual content may also be captured in other ways, such as via 2D images of a scene captured from multiple viewing positions relative to the scene.
[0004] Such three-dimensional visual content can be represented by a three-dimensional mesh, which includes a plurality of polygons with connected vertices, which model the surface of the three-dimensional visual content (such as the surface of a point cloud). In addition, when modeled as a three-dimensional mesh, the texture or attribute values of the points of the three-dimensional visual content can be overlaid on the mesh to represent the attributes or texture of the three-dimensional visual content.
[0005] In addition, a 3D mesh can be generated in the software, for example, without first modeling it as a point cloud or other type of 3D visual content. For example, the software can directly generate a 3D mesh and apply textures or attribute values to characterize the object. Summary of the invention
[0006] In some embodiments, a system includes one or more sensors configured to capture a point representing an object in the view of the sensor and to capture a texture or attribute value associated with the point of the object. The system also includes one or more computing devices storing program instructions that, when executed, cause the one or more computing devices to generate a three-dimensional mesh that uses the vertices of the polygons defining the three-dimensional mesh and the connections between the vertices to model the point of the object. In addition, in some embodiments, the three-dimensional mesh can be generated without first being captured by one or more sensors. For example, a computer graphics program can generate a three-dimensional mesh with an associated texture or associated attribute value to represent an object in a scene without having to generate a point cloud representing the object.
[0007] In some embodiments, an encoder system includes one or more computing devices storing program instructions that, when executed by the one or more computing devices, also cause the one or more computing devices to determine multiple tiles of attributes for a three-dimensional mesh and a corresponding attribute map that maps the attribute tiles to the geometry of the mesh.
[0008] The encoder system can also encode the geometry of the mesh by encoding the displacement of the base mesh and the vertices relative to the base mesh. The compressed bitstream may include a compressed base mesh, a compressed displacement value, and compressed attribute information. In order to improve compression efficiency, a decoding unit may be used to encode parts of the mesh. For example, the encoding unit may include a block of the mesh, each block including an independently encoded fragment of the mesh. For another example, a decoding unit may include a slice composed of several sub-grids of the mesh, wherein the sub-grid utilizes the dependency between the sub-grids and is therefore not independently encoded. In addition, a higher-level decoding unit, such as a slice group, may be used. Different encoding parameters may be defined in the bitstream to be applied to different decoding units. For example, it is not necessary to repeatedly signal the encoding parameter, but the commonly signaled decoding parameter may be applied to members of a given decoding unit, such as a sub-grid of a slice or a mesh portion constituting a block. Some example encoding parameters that can be used include: entropy decoding parameters, intra-frame prediction parameters, inter-frame prediction parameters, local or sub-grid indexes, and the like.
[0009] In some embodiments, the displacement value may be signaled in a sub-bitstream other than the video sub-bitstream (such as in the displacement value's own sub-bitstream). For example, a dynamic grid encoder may generate a bitstream that includes a displacement data sub-bitstream in addition to a base grid sub-bitstream, an atlas data sub-bitstream, etc. In some embodiments, signaling the displacement information at least partially outside the video sub-bitstream may enable out-of-order decoding. In addition, this may enable the use of different levels of detail to reconstruct the sub-grids, and allow displacements to be signaled in a relative manner. For example, the displacement of a given sub-grid may be signaled relative to a baseline displacement signaled at a higher level (such as in a sequence parameter set, a frame header, etc.). In such embodiments, a network abstraction layer unit (NAL unit) syntax may be used for a displacement data sub-bitstream, such as a NAL unit header, a sequence parameter set, a frame parameter set, a block information, etc. In such embodiments, a NAL unit may be used to define a displacement data unit. In addition, in some embodiments, the displacement value may be signaled in an atlas data sub-bitstream, for example, together with the slice data. In some embodiments, the displacement values may be signaled in the atlas data sub-bitstream using the slice data units, or in some embodiments, the displacement data units may be signaled in the atlas data sub-bitstream, e.g., together with the slice data units. Additionally, in some embodiments, the displacement values may be signaled in a hybrid manner, with some displacement values being signaled in the video sub-bitstream and other displacement values being signaled outside the video sub-bitstream, such as in the atlas data sub-bitstream, or in the displacement data sub-bitstream. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Example input information for defining a three-dimensional mesh according to some embodiments is illustrated.
[0011] Figure 2 Alternative examples of input information for defining a three-dimensional mesh according to some embodiments are illustrated, where the input information is formatted according to an object format.
[0012] Figure 3 An example pre-processor and encoder for encoding a three-dimensional mesh according to some embodiments are illustrated.
[0013] Figure 4 A more detailed view of an example intra encoder is illustrated according to some embodiments.
[0014] Figure 5 An example intra decoder for decoding a three-dimensional mesh according to some embodiments is illustrated.
[0015] Figure 6A more detailed view of an example inter-frame encoder is illustrated in accordance with some embodiments.
[0016] Figure 7 An example inter-frame decoder for decoding a three-dimensional mesh according to some embodiments is illustrated.
[0017] FIG. 8A to FIG. 8B The partitioning of a grid into multiple tiles is illustrated according to some embodiments.
[0018] Fig. 9 A grid is illustrated that is partitioned into two tiles, each tile comprising sub-grids, according to some embodiments.
[0019] FIG. 10A to FIG. 10D Adaptive subdivision based on edge subdivision rules of shared edges according to some embodiments is illustrated.
[0020] FIG. 11A to FIG. 11C Adaptive subdivision of adjacent tiles is illustrated according to some embodiments, where shared edges are subdivided in a manner that ensures that vertices of adjacent tiles are aligned with each other.
[0021] Fig.12 An example of displacement applied to subdivided positions to adjust vertex positions according to some embodiments is illustrated.
[0022] Fig.13 Example color planes of a video sub-bitstream used at least in part to signal displacement values according to some embodiments are illustrated.
[0023] Fig.14 Example slices that may be included in a video sub-bitstream to signal displacement values according to some embodiments are illustrated.
[0024] Fig.15 An example of how slice size and position in a video frame of a video sub-bitstream may be signaled according to some embodiments is illustrated.
[0025] Fig.16 An example computer system that may implement an encoder or decoder according to some embodiments is illustrated.
[0026] This specification includes references to "one embodiment" or "an embodiment." The appearance of the phrase "in one embodiment" or "in an embodiment" does not necessarily refer to the same embodiment. The particular features, structures or characteristics may be combined in any suitable manner consistent with the present disclosure.
[0027] The term "comprising" is open ended. As used in the appended claims, the term does not exclude additional structures or steps. Consider the following cited claim: "an apparatus comprising one or more processor units..." Such a claim does not exclude the apparatus from including additional components (e.g., a network interface unit, a graphics circuit, etc.).
[0028] "Configured to", various units, circuits or other components may be described or described as "configured to" perform one or more tasks. In such contexts, "configured to" is used to imply a structure (e.g., a circuit) that performs the one or more tasks during operation by indicating that the unit / circuit / component includes the structure. In this way, the unit / circuit / component is said to be configured to perform the task even when the specified unit / circuit / component is currently inoperable (e.g., not turned on). The units / circuits / components used with the "configured to" language include hardware-such as circuits, memories storing executable program instructions to implement operations, etc. Reference to a unit / circuit / component "configured to" perform one or more tasks is explicitly intended not to invoke 35 U.S.C. §112 (f) for the unit / circuit / component. In addition, "configured to" may include a general structure (e.g., a general circuit) manipulated by software and / or firmware (e.g., an FPGA or a general processor executing software) to operate in a manner capable of performing one or more tasks to be solved. "Configured to" may also include adapting a manufacturing process (eg, a semiconductor fabrication facility) to produce a device (eg, an integrated circuit) suitable for implementing or performing one or more tasks.
[0029] "First," "second," etc. As used herein, these terms act as labels for the nouns that precede them, and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing a write operation of a "first" value and a "second" value. The terms "first" and "second" do not necessarily imply that the first value must be written before the second value.
[0030] "Based on". As used herein, this term is used to describe one or more factors that influence a determination. This term does not exclude additional factors that influence a determination. That is, a determination may be based solely on these factors or at least in part on these factors. Consider the phrase "A is determined based on B". In this case, B is a factor that influences the determination of A, and such a phrase does not exclude that the determination of A may also be based on C. In other examples, A may be determined based solely on B. DETAILED DESCRIPTION
[0031] As data acquisition and display technologies become more advanced, the ability to capture volumetric content including thousands of points in 2D or 3D space has increased (such as via LIDAR systems). In addition, the development of advanced display technologies (such as virtual reality or augmented reality systems) has increased the potential uses of volumetric content. However, volumetric content files are typically very large, and storing and transmitting these volumetric content files can be costly and time consuming. For example, communication of volumetric content over a private network or a public network (such as the Internet) can require a considerable amount of time and / or network resources, such that some uses of the volumetric content (such as real-time uses) may be limited. In addition, the storage requirements of the volumetric content files may consume a significant amount of storage capacity of the device storing the volumetric content files, which may also limit potential applications for using the volumetric content data.
[0032] In some embodiments, an encoder may be used to generate compressed volumetric content to reduce the cost and time associated with storing and transmitting large volumetric content files. In some embodiments, a system may include an encoder that compresses attribute information and / or spatial information of volumetric content so that the volumetric content file can be stored and transmitted faster than non-compressed volumetric content and in a manner that the volumetric content file can occupy less storage space than non-compressed volumetric content.
[0033] In some embodiments, additionally or alternatively, such encoders and decoders or other encoders and decoders described herein may be adapted to encode three degrees of freedom plus (3DOF+) scenes, visual volumetric content (such as MPEG V3C scenes), immersive video scenes (such as MPEG MIV), etc.
[0034] In some embodiments, the static or dynamic meshes to be compressed and / or encoded may include a set of 3D meshes M(0), M(1), M(2) ... M(n). Each mesh M(i) may be defined by connectivity information C(i), geometry information G(i), texture coordinates T(i), and texture connectivity CT(i). For each mesh M(i), one or more 2D images A(i,0), A(i,1) ... A(i,D-1) may be included that describe the texture or attributes associated with the mesh. For example, Figure 1 An example static or dynamic mesh M(i) including connectivity information C(i), geometry information G(i), texture image A(i), texture connectivity information TC(i) and texture coordinate information T(i) is illustrated. In addition, Figure 2 illustrates an example of a textured mesh stored in object (OBJ) format.
[0035] For example, Figure 2The example texture mesh stored in the object format shown includes geometric information listed as X, Y, and Z coordinates of the vertices and texture coordinates listed as two-dimensional (2D) coordinates of the vertices, where the 2D coordinates identify the pixel position of the pixel storing the texture information for the given vertex. The example texture mesh stored in the object format also includes texture connectivity information indicating the mapping between geometric coordinates and texture coordinates to form polygons such as triangles. For example, a first triangle is formed by three vertices, where the first vertex (1 / 1) is defined as a first geometric coordinate (e.g., 64.062500, 1237.739990, 51.757801), which corresponds to a first texture coordinate (e.g., 0.0897381, 0.740830). The second vertex (2 / 2) of the triangle is defined as a second geometric coordinate (e.g., 59.570301, 1236.819946, 54.899700), which corresponds to a second texture coordinate (e.g., 0.899059, 0.741542). Finally, the third vertex of the triangle corresponds to the third listed geometric coordinate that matches the third listed texture coordinate. However, it should be noted that in some instances, the vertices of a polygon such as a triangle may be mapped to a set of geometric coordinates and texture coordinates that may have different index positions in the corresponding lists of geometric coordinates and texture coordinates. For example, the second triangle has a first vertex that corresponds to the fourth listed set of geometric coordinates and the seventh listed set of texture coordinates. The second vertex corresponds to the first listed set of geometric coordinates and the first listed set of texture coordinates, and the third vertex corresponds to the third listed set of geometric coordinates and the ninth listed set of texture coordinates.
[0036] In some embodiments, the geometric information G(i) may characterize the positions of the vertices of the mesh in 3D space, and the connectivity C(i) may indicate how the vertices are connected together to form polygons that constitute the mesh M(i). In addition, the texture coordinates T(i) may indicate the positions of pixels in the 2D image corresponding to the vertices of the corresponding sub-mesh. The attribute patch information may indicate how to map the texture coordinates defined relative to the 2D bounding box to the three-dimensional space of the 3D bounding box associated with the attribute patch based on the way the points are projected onto the projection plane of the attribute patch. In addition, the texture connectivity information TC(i) may indicate how the vertices represented by the texture coordinates T(i) are connected together to form the polygons of the sub-mesh. For example, each texture or attribute patch of the texture image A(i) may correspond to a corresponding sub-mesh defined using the texture coordinates T(i) and the texture connectivity TC(i).
[0037] In some embodiments, the mesh encoder may perform a tile generation process in which the mesh is subdivided into a set of sub-meshes. These sub-meshes may correspond to connected components of texture connectivity, or may be sub-meshes that have texture connectivity different from that of the mesh. In some embodiments, the number and size of sub-meshes to be determined may be adjusted to balance discontinuities and flexibility in updating the mesh, such as via inter-frame prediction. For example, smaller sub-meshes may allow for more fine-grained updates to change specific areas of the mesh, such as in Figure 3 A high-level block diagram of the encoding process in some embodiments is illustrated. Note that a feedback loop during the encoding process enables the encoder to guide the pre-processing steps and change their parameters to achieve the best possible compromise based on various criteria such as: rate distortion, encoding / decoding complexity, random access, reconstruction complexity, terminal capabilities, encoder / decoder power consumption, network bandwidth and latency, and / or other factors.
[0038] A mesh, which may be a static or dynamic mesh, is received at preprocessing module 302. In addition, a property map characterizing how to map a property image (e.g., a texture image) of the static / dynamic mesh to the mesh is received at preprocessing module 302. For example, the property map may include texture coordinates and texture connectivity of the texture image of the mesh. Preprocessing module 302 divides the static / dynamic mesh M(i) into a base mesh M(i) and a displacement d(i). The displacement characterizes how vertices will be displaced to recreate the original static / dynamic mesh from the base mesh. For example, in some embodiments, vertices included in the original static / dynamic mesh may be omitted from the base mesh (e.g., the base mesh may be a compressed version of the original static / dynamic mesh). As will be discussed in more detail below, the decoder may predict additional vertices to be added to the base mesh, for example, by subdividing the edges between the remaining vertices included in the base mesh. In such an example, the displacement may indicate how the additional vertices will be displaced, wherein the displacement of the added vertices modifies the base mesh to better represent the original static / dynamic mesh. For example, Figure 4 A detailed intra encoder 402 is illustrated that can be used to encode the base mesh m(i) and the displacements d(i) of the added vertices. For dynamic meshes, a Figure 5 The inter-frame encoder shown in FIG. Figure 5 As shown, instead of signaling a new base grid for each frame, the base grid of the current time frame can be compared to a reconstructed quantized reference base grid m'(i) (e.g., the base grid that the decoder will see from the previous time frame), and a motion vector representing how the current base grid has changed relative to the reference base grid can be encoded instead of encoding a new base grid for each frame. Note that the motion vectors may not be encoded directly, but may be further compressed to exploit the relationship between the motion vectors.
[0039] The separated base mesh m(i) and displacement d(i) separated by the pre-processing module 302 are provided to the encoder 304, which can be Figure 4 Intra-frame encoder as shown or as Figure 6 The inter-frame encoder shown. In addition, the attribute map A(i) is provided to the encoder 304. In some embodiments, in addition to the separated base grid m(i) and displacement d(i), the original static / dynamic grid M(i) can also be provided to the encoder 304. For example, the encoder 304 can compare the reconstructed version of the static / dynamic grid (which has been reconstructed from the base grid M(i) and displacement d(i)) to determine the geometric distortion. In some embodiments, an attribute transfer process can be performed to adjust the attribute values of the attribute image to address the slight geometric distortion. In some embodiments, feedback can be provided back to the pre-processing 302 by changing the way the original static / dynamic grid is extracted to generate the base grid, for example to reduce distortion. It should be noted that in some embodiments, the intra-frame encoder and the inter-frame encoder can be combined into a single encoder, which includes a logic component for switching between intra-frame encoding and inter-frame encoding. The output of the encoder 304 is a compressed bit stream representing the original static / dynamic grid and its associated attributes / textures.
[0040] With respect to mesh decimation, in some embodiments, a portion of the surface of a static / dynamic mesh may be considered to be an input 2D curve (represented by a 2D polyline), which is referred to as the "original" curve. The original curve may first be downsampled to generate a base curve / polyline, which is referred to as the "decimated" curve. A subdivision scheme (such as those described herein) may then be applied to the decimated polyline to generate the "subdivided" curve. For example, a subdivision scheme using an iterative interpolation scheme may be applied. The subdivision scheme may include inserting a new point in the middle of each edge of the polyline at each iteration. The inserted points represent additional vertices that may be moved by the displacement.
[0041] For example, the subdivided polyline is then deformed to better approximate the original curve. More precisely, a displacement vector is calculated for each vertex of the subdivided mesh so that the shape of the displaced curve approximates the shape of the original curve. The advantage of a subdivided curve is that it has a subdivision structure that allows efficient compression while ensuring a reliable approximation of the original curve. Compression efficiency is achieved through the following properties:
[0042] The decimated / base curve has a small number of vertices and requires a limited number of bits to encode / transmit.
[0043] The subdivision curves are automatically generated by the decoder once the base / decimated curves are decoded (eg, without signaling or hard-coding any information other than the subdivision scheme type and subdivision iteration count at the decoder).
[0044] The displacement curves are generated by decoding and applying the displacement vectors associated with the subdivision curve vertices.In addition to allowing spatial / quality scalability, the subdivision structure enables efficient wavelet decomposition, which provides high compression performance (eg, with respect to rate-distortion performance).
[0045] For example, Figure 4 A more detailed view of an example intra encoder is illustrated according to some embodiments.
[0046] In some embodiments, the intra-frame encoder 402 receives a base grid m(i), a displacement d(i), an original static / dynamic grid M(i), and an attribute map A(i). The base grid m(i) is provided to a quantization module 404, wherein various aspects of the base grid may be (optionally) further quantized. In some embodiments, various grid encoders may be used to encode the base grid. In addition, in some embodiments, the intra-frame encoder 402 may allow customization, wherein different corresponding grid encoding schemes may be used to encode the base grid. For example, the static grid encoder 406 may be a selected grid encoder selected from a group of feasible grid encoders, such as a DRACO encoder (or another suitable encoder). The encoded base grid encoded by the static grid encoder 406 is provided to a multiplexer (MUX) 438 to be included in the compressed bitstream b(i). In addition, the encoded base grid is provided to a static grid decoder to generate a reconstructed version of the base grid (which the decoder will see). This reconstructed version of the base grid is used to update the displacement d(i) to resolve any geometric distortion between the original base grid and the reconstructed version of the base grid (which the decoder will see). For example, the static grid decoder 408 generates a reconstructed quantized base grid m'(i) and provides the reconstructed quantized base grid m'(i) to a displacement update module 410, which also receives the original base grid and the original displacement d(i). The displacement update module 410 compares the reconstructed quantized base grid m'(i) (as seen by the decoder) with the base grid m(i) and adjusts the displacement d(i) to account for the difference between the base grid m(i) and the reconstructed quantized base grid m'(i). These updated displacements d'(i) are provided to a wavelet transform 412, which applies a wavelet transform to further compress the updated displacements d'(i) and outputs wavelet coefficients e(i), which are provided to a quantization module 414, which generates quantized wavelet coefficients e'(i). The quantized wavelet coefficients may then be packaged into 2D image frames via an image packaging module 416, where the packaged 2D image frames are further video encoded via video encoding 418. The encoded video images are also provided to a multiplexer (MUX) 438 for inclusion in the compressed bitstream b(i). Additionally, in some embodiments, displacement values (such as indicated in the generated quantized wavelet coefficients e'(i), or indicated using other compression schemes) may be at least partially encoded outside of the video sub-bitstream, such as in their own displacement data sub-bitstream, in a base grid sub-bitstream, or in an atlas data sub-bitstream.
[0047] Furthermore, to account for any geometric distortion introduced relative to the original static / dynamic mesh, an attribute transfer process 430 may be used to modify attributes to account for the differences between the reconstructed deformed mesh DM(i) and the original static / dynamic mesh.
[0048] For example, video encoding 418 may further perform video decoding (or a complementary video decoding module ( Figure 4 )). This produces reconstructed packed quantized wavelet coefficients, which are unpacked via the image unpacking module 420. In addition, inverse quantization can be applied via the inverse quantization module 422, and the inverse wavelet transform 424 can be applied to generate the reconstructed displacement d"(i). In some embodiments, other decoding techniques can be used to generate the reconstructed displacement d"(i), such as decoding the displacement signaled in the atlas data sub-bitstream, the displacement data sub-bitstream, or the base grid sub-bitstream. In addition, the reconstructed quantized base grid m'(i) generated by the static grid decoder 408 can be inverse quantized via the inverse quantization module 428 to generate a reconstructed base grid m"(i). The reconstructed deformed grid generation module 426 applies the reconstructed displacement d"(i) to the reconstructed base grid m"(i) to generate a reconstructed deformed grid DM(i). It should be noted that the reconstructed deformed grid DM(i) represents the reconstructed grid that the decoder will generate, and accounts for any geometric deformation caused by losses introduced during the encoding process.
[0049] The attribute transfer module 430 compares the geometry of the original static / dynamic mesh M(i) with the reconstructed deformed mesh DM(i) and updates the attribute map to account for any geometric deformations, which is output as an updated attribute map A'(i). The updated attribute map A'(i) is then filled, where the 2D image including the attribute image is filled so that the space not used to transmit the attribute image has been filled. In some embodiments, color space conversion is optionally applied at the color space conversion module 434. For example, the RGB color space used to characterize the color values of the attribute image can be converted to the YCbCr color space, and color space subsampling, such as 4:2:0, 4:0:0, etc., can also be applied. The updated attribute map A'(i), which has been filled and optionally color space converted, is then video encoded via the video encoding module 436 and provided to the multiplexer 438 for inclusion in the compressed bitstream b(i).
[0050] In some embodiments, the controller 400 can coordinate the various quantization and inverse quantization steps and the video encoding and decoding steps so that inverse quantization "cancels" quantization and so that video decoding "cancels" video encoding. In addition, the attribute transfer module 430 can take into account the quantization level being applied based on communications from the controller 400.
[0051] Figure 5 An example intra decoder for decoding a three-dimensional mesh according to some embodiments is illustrated.
[0052] The intra-frame decoder 502 receives a compressed bit stream b(i), such as that obtained by Figure 4 The compressed bitstream generated by the intra encoder 402 is shown. The demultiplexer (DEMUX) 504 parses the bitstream into a base grid subcomponent, a displacement subcomponent, and an attribute map subcomponent. In some embodiments, the displacement subcomponent may be signaled in the displacement data subbitstream, and may also be at least partially signaled in other subbitstreams such as an atlas data subbitstream, a base grid subbitstream, or a video subbitstream. In this case, the displacement decoder 522 decodes the displacement subbitstream, and / or the atlas decoder 524 decodes the atlas subbitstream.
[0053] The static grid decoder 506 decodes the base grid sub-components to generate a reconstructed quantized base grid m'(i), which is provided to the inverse quantization module 518, which in turn outputs the decoded base grid m"(i) and provides it to the reconstructed deformed grid generator 520.
[0054] In some embodiments, a portion of the displacement subcomponent of the bitstream is provided to video decoding 508, where video-encoded image frames are video-decoded and provided to image unpacking 510. Image unpacking 510 extracts the packed displacements from the video-decoded image frames and provides them to inverse quantization 512, where the displacements are inversely quantized. In addition, the inverse quantized displacements are provided to an inverse wavelet transform 514, which outputs decoded displacements d"(i). A reconstructed deformed grid generator 520 applies the decoded displacements d'(i) to a decoded base grid m"(i) to generate a decoded static / dynamic grid M"(i). The decoded displacements may come from any combination of a video sub-bitstream, an atlas data sub-bitstream, a base grid sub-bitstream, and / or a displacement data sub-bitstream. In addition, the attribute map subcomponent is provided to video decoding 516, which outputs a decoded attribute map A"(i). The decoded grid M"(i) and the decoded attribute map A"(i) may then be used to present a reconstructed version of the three-dimensional visual content at a device associated with the decoder.
[0055] like Figure 5 As shown, the bit stream is demultiplexed into three or more separate sub-streams:
[0056] Grid subflow;
[0057] a displacement substream for positions and possibly for each vertex attribute; and
[0058] A property graph subflow for each property graph.
[0059] The trellis substream is fed to a trellis decoder to generate a reconstructed quantized base trellis m'(i). The decoded base trellis m"(i) is then obtained by applying inverse quantization on m'(i). The proposed scheme is not restricted by which trellis codec is used. The trellis codec used may be explicitly specified in the bitstream or may be implicitly defined / fixed by the specification or the application.
[0060] The displacement substream can be decoded by a video / image decoder. The generated image / video is then unpacked and inverse quantization is applied to the wavelet coefficients. In an alternative embodiment, the displacements can be decoded by a dedicated displacement data decoder or an atlas decoder. The proposed scheme is not restricted by which codec / standard is used. Image / video codecs such as [HEVC][AVC][AV1][AV2][JPEG][JPEG2000] can be used. Dictionary-based decoders such as ZIP or motion decoders for decoding grid motion information can be used, for example, as dedicated displacement data decoders. The decoded displacements d"(i) are then generated by applying an inverse wavelet transform to the unquantized wavelet coefficients. The final decoded grid is generated by applying a reconstruction process to the decoded base grid m"(i) and adding the decoded displacement field d"(i).
[0061] The attribute substream is directly decoded by a video decoder and a decoded attribute map A"(i) is generated as output. The proposed scheme is not restricted by which codec / standard is used. Image / video codecs such as [HEVC] [AVC] [AV1] [AV2] [JPEG] [JPEG2000] can be used. Alternatively, the attribute substream can be decoded by a non-image / video decoder (e.g., using a dictionary-based decoder such as ZIP). Multiple substreams can be decoded, each of which is associated with a different attribute map. Each substream can use a different codec.
[0062] Figure 6 A more detailed view of an example inter-frame encoder is illustrated in accordance with some embodiments.
[0063] In some embodiments, the inter-frame encoder 602 may include similar components as the intra-frame encoder 402, but instead of encoding a base grid, the inter-frame encoder may encode motion vectors that can be applied to a reference grid to generate a base grid at a decoder.
[0064] For example, in the case of dynamic meshes, a temporally consistent remeshing process is used, which can produce the same subdivision structure shared by the current mesh M'(i) and the reference mesh M'(j). Such a coherent temporal remeshing process makes it possible to skip the encoding of the base mesh m(i) and reuse the base mesh m(j) associated with the reference frame M(j). This can also enable better temporal prediction for both attribute and geometry information. More precisely, a motion field f(i) that describes how to move the vertices of m(j) to match the positions of m(i) can be calculated and encoded. Figure 6 Such a process is described in . For example, motion encoder 406 may generate a motion field f(i) that describes how to move the vertices of m(j) to match the positions of m(i).
[0065] In some embodiments, the base grid m(i) associated with the current frame is first quantized (e.g., using uniform quantization) and encoded by a static grid encoder. The proposed scheme is not restricted by which grid codec is used. The grid codec used can be explicitly specified in the bitstream by encoding the grid codec ID, or can be implicitly defined / fixed by the specification or application.
[0066] Depending on the application and target bitrate / visual quality, the encoder may optionally encode a set of displacement vectors associated with the subdivided mesh vertices, referred to as a displacement field d(i).
[0067] The displacement field d(i) is then updated (at update displacement module 410) using the reconstructed quantized base grid m'(i) (e.g., the reconstructed output of base grid 408) to generate an updated displacement field d'(i), thereby processing the difference between the reconstructed base grid m'(i) and the original base grid m(i). A wavelet transform is then applied to d'(i) at wavelet transform 412 by utilizing the subdivision surface grid structure, and a set of wavelet coefficients is generated. The wavelet coefficients are then quantized at quantization 414, packed into a 2D image / video (at image packing 416), and compressed by an image / video encoder (at video encoding 418). The encoding of the wavelet coefficients can be lossless or lossy. A reconstructed version of the wavelet coefficients is obtained by applying image unpacking and inverse quantization to the reconstructed wavelet coefficient video generated during the video encoding process (e.g., at 420, 422, and 424). The reconstructed displacements d”(i) are then calculated by applying the inverse wavelet transform to the reconstructed wavelet coefficients. The reconstructed base mesh m”(i) is obtained by applying inverse quantization to the reconstructed quantized base mesh m'(i). The reconstructed deformed mesh DM(i) is obtained by subdividing m”(i) and applying the reconstructed displacements d”(i) to its vertices.
[0068] Since the quantization step and / or the grid compression module may be lossy, a reconstructed quantized version of m(i) is calculated (denoted as m'(i)). If the grid information is losslessly encoded and the quantization step is skipped, then m(i) will exactly match m'(i).
[0069] like Figure 6 As shown, the reconstructed quantized reference base grid m'(j) is used to predict the current frame base grid m(i). Figure 3 The pre-processing module 302 described in can be configured so that m(i) and m(j) share the same:
[0070] The number of vertices;
[0071] Connectivity;
[0072] Texture coordinates; and
[0073] Texture connectivity.
[0074] The motion field f(i) is computed by considering the quantized version of m(i) and the reconstructed quantized base mesh m'(j). Since m'(j) may have a different number of vertices than m(j) (e.g., vertices may be merged / removed), the encoder keeps track of the transform applied to m(j) to obtain m'(j) and applies it to m(i) to ensure a 1-to-1 correspondence between m'(j) and the transformed quantized version of m(i), denoted as m*(i). The motion field f(i) is computed by subtracting the quantized position p(i,v) of vertex v of m*(i) from the position p(j,v) of vertex v of m'(j):
[0075] f(i,v)=p(i,v)-p(j,v)
[0076] The motion field is then further predicted by the connectivity information of m'(j) and entropy coded (eg, context adaptive binary arithmetic coding may be used).
[0077] Since the motion field compression process may be lossy, a reconstructed motion field (denoted as f'(i)) is calculated by applying the motion decoder module 408. The reconstructed quantized basis grid m'(i) is then calculated by adding the motion field to the position of m'(j). The rest of the encoding process is similar to intra-frame coding.
[0078] Figure 7 An example inter-frame decoder for decoding a three-dimensional mesh according to some embodiments is illustrated.
[0079] The interframe decoder 702 includes Figure 5The inter-frame decoder 702 is similar to the intra-frame decoder 502 shown in FIG. However, instead of receiving a directly encoded base grid, the inter-frame decoder 702 reconstructs the base grid of the current frame based on the motion vector of the displacement field relative to the reference frame. For example, the inter-frame decoder 702 includes a motion field / vector decoder 704 and a reconstruction of the base grid module 706.
[0080] In a similar manner to the intra decoder, the inter decoder 702 splits the bitstream into three separate substreams:
[0081] Motion sub-flow;
[0082] displacement sub-stream; and
[0083] Attribute subflow.
[0084] The motion substream is decoded by applying a motion decoder 704. The proposed scheme is not restricted by which codec / standard is used to decode the motion information. For example, any motion decoding scheme can be used. Optionally, the decoded motion is then added to the decoded reference quantization base grid m'(j) to generate a reconstructed quantization base grid m'(i), i.e., the already decoded grid at instance j can be used for prediction of the grid at instance i. The decoded base grid m"(i) is then generated by applying inverse quantization to m'(i).
[0085] About Figure 5 The displacement substream and attribute substream are decoded in a similar manner to the intra-frame decoding process described above. The decoded grid M" (i) is also reconstructed in a similar manner.
[0086] The inverse quantization and reconstruction process is not standardized and may be implemented in various ways and / or combined with the rendering process.
[0087] Grid division and control coding to avoid cracks
[0088] In some embodiments, the grid can be subdivided into sets of tiles (e.g., sub-parts), and these tiles can potentially be grouped into multiple groups of tiles, such as sets of tile groups / blocks. In such embodiments, different encoding parameters (e.g., subdivision, quantization, wavelet transform, coordinate system, etc.) can be used to compress each tile or tile group. In some embodiments, to avoid cracks at tile boundaries, lossless coding can be used for boundary vertices. In addition, quantization of wavelet coefficients for boundary vertices can be disabled, and local coordinate systems for boundary vertices are not used.
[0089] In some embodiments, scalability can be supported at different levels. For example, temporal scalability can be achieved by temporal subsampling and frame reordering. In addition, different mechanisms can be used to achieve quality and spatial scalability for geometry / vertex attribute data and attribute map data. In addition, region of interest (ROI) reconstruction can be supported. For example, the encoding process described in the previous section can be configured to encode ROI with higher resolution and / or higher quality for geometry, vertex attributes and / or attribute map data. This is particularly useful for providing higher visual quality content (e.g., higher quality of the face relative to the rest of the body) under strict bandwidth and complexity constraints. Priority / importance / space / bounding box information can allow the decoder to adaptively decode a subset of the grid based on the view volume, power budget or terminal capability in association with slices, slice groups, blocks, network abstraction layer (NAL) units and / or sub-bitstreams. It should be noted that any combination of such decoding units can be used together to implement such functionality. For example, NAL units and sub-bitstreams can be used together.
[0090] In some embodiments, temporal and / or spatial random access may be supported. Temporal random access may be achieved by introducing IRAPs (intra-frame random access points) in different substreams (e.g., attribute atlas, video, grid, motion, and displacement substreams). Spatial random access may be supported by defining and using blocks, sub-pictures, slice groups, and / or slices, or any combination of these decoding units. Metadata describing the layout and relationships between different units may also need to be generated and included in the bitstream to help the decoder determine the units that need to be decoded.
[0091] As discussed above, various functions can be supported, such as:
[0092] Random access to space;
[0093] Adaptive mass allocation (e.g., bump compression, such as assigning higher mass to the face compared to the body of a human model);
[0094] Region of interest (ROI) access;
[0095] Decoding unit-level metadata (e.g., object description, bounding box information);
[0096] Spatial and quality scalability; and
[0097] Adaptive streaming and decoding (e.g., high priority regions are streamed / decoded first).
[0098] The disclosed compression scheme allows various coding units (e.g., slices, slice groups, and blocks) to be compressed using different encoding parameters (e.g., subdivision schemes, subdivision iteration counts, quantization parameters, etc.), which may introduce compression artifacts (e.g., cracks between slice boundaries). In some embodiments, as further discussed below, efficient strategies (e.g., efficient strategies in terms of computational complexity, compression efficiency, power consumption, etc.) can be used, which allow the scheme to handle different coding unit parameters without introducing artifacts.
[0099] Grid Blocks
[0100] The grid can be divided into a set of blocks (e.g., parts / segments) that can be encoded and decoded independently. Fig. 8A / Figure 8B As illustrated in , vertices / edges shared by more than one tile are replicated. Note that the mesh 800 is divided into two tiles 850 and 852 by replicating three vertices {V0, V1, V2} and two shared edges {(V0, V1), (V1, V2)}. In other words, when the mesh is divided into two tiles, each of the two tiles includes vertices and edges that are the previous set of vertices and edges in the combined mesh, so these vertices and edges are repeated in the tiles.
[0101] Subgrid
[0102] Each tile can be further split into a set of sub-grids that can exploit dependencies between them during encoding. For example, Fig. 9 An example of a grid 900 is shown, which is divided into two blocks (902 and 904), which contain three sub-grids and four sub-grids, respectively. For example, block 902 includes sub-grids 0, 1, and 2; and block 904 includes sub-grids 0, 1, 2, and 3.
[0103] The subgrid structure can be defined in any of the following ways:
[0104] explicitly encode a per-face integer attribute indicating the index of the submesh to which each face of the mesh belongs, or
[0105] The CCs of a mesh are detected implicitly by treating each connected component (CC) as a sub-mesh, either with respect to the connectivity of the position or the connectivity of the texture coordinates, or both. The mesh vertices are traversed from neighbor to neighbor, which enables the detection of CCs in a deterministic manner. The index assigned to the CCs starts at 0 and increases by one each time a new CC is detected.
[0106] Sharding
[0107] A tile is a group of sub-grids. The encoder can explicitly store the index of the sub-grids belonging to it for each tile. In one specific embodiment, a sub-grid can belong to one or more tiles (e.g., associating metadata with overlapping parts of the grid). In another embodiment, a sub-grid can belong to only a single tile. Vertices on the boundaries between tiles are not repeated. Tiles are also encoded using correlations between them, so they cannot be encoded / decoded independently.
[0108] The list of submeshes associated with a shard can be encoded using various strategies, such as:
[0109] Entropy decoding
[0110] Intra prediction
[0111] Inter prediction
[0112] Local subgrid index (i.e., smaller extent)
[0113] Shard Group
[0114] A shard group is a group of shards. Shard groups are particularly useful for storing parameters shared by the shards within the group or allowing unique handles that can be used to associate metadata with those shards.
[0115] In some embodiments, to support using different encoding parameters for each slice, the following may be used:
[0116] The relationship between submeshes and tiles can be exploited to assign a tile ID to each face of the base mesh.
[0117] Vertices and edges belonging to a single submesh are assigned the shard ID to which the submesh belongs.
[0118] Vertices and edges that lie on the border of two or more shards are assigned to all corresponding shards.
[0119] When a subdivision scheme is applied, the decision whether to subdivide an edge depends on the subdivision parameters of all shards to which the edge belongs.
[0120] For example, FIG. 11A to FIG. 11C Demonstrates adaptive subdivision based on shared edges. FIG. 11A to FIG. 11C The subdivision decision determined in (for example, suppose there is an edge belonging to two shards Patch0 and Patch1, the subdivision iteration count of Patch0 is 0, and the subdivision iteration count of Patch1 is 2. The shared edge will be subdivided 2 times (i.e., taking the maximum subdivision count)), using FIG. 10A to FIG. 10D The subdivision scheme shown in Figure 2 subdivides the edges. For example, based on the subdivision decisions associated with different edges, apply FIG. 10A to FIG. 10DGiven an adaptive subdivision scheme among the adaptive subdivision schemes described in . FIG. 10A to FIG. 10D It shows how to subdivide a triangle when 3, 2, 1, or 0 of its edges need to be subdivided, respectively. The vertices created after each subdivision iteration are assigned to the fragments of their parent edges.
[0121] When applying quantization to wavelet coefficients, the quantization parameter for the vertex is selected based on the quantization parameters of all the tiles to which the vertex belongs. For example, suppose a tile belongs to two tiles Patch0 and Patch1. Patch0 has quantization parameter QP1. Patch1 has quantization parameter QP2. The wavelet coefficient associated with the vertex will be quantized using quantization parameter QP=min(QP1, QP2).
[0122] Grid Blocks
[0123] To support trellis blocking, the encoder can divide the trellis into a set of tiles and then encode / decode these tiles independently by applying any trellis codec, or use a trellis codec that natively supports blocked trellis decoding.
[0124] In both cases, the set of shared vertices that lie on the tile boundary is repeated. To avoid gaps between tiles, the encoder can do either of the following:
[0125] Encodes the stitching information indicating the mapping between repeated vertices as follows:
[0126] ○ Encode per-vertex labels identifying duplicate vertices (by encoding vertex attributes using the mesh codec)
[0127] ○ Encode for each duplicate vertex the index of the vertex it should be merged with.
[0128] Ensure that the decoded positions and vertex attributes associated with duplicate vertices match exactly
[0129] ○ Each vertex tag identifies duplicate vertices, and the mesh codec is used to decode this information into vertex attributes
[0130] ○ Apply the adaptive subdivision scheme described in the previous section to ensure consistent subdivision behavior across tile boundaries
[0131] ○ The encoder needs to maintain the mapping between duplicate vertices and adjust the encoding parameters to ensure matching values
[0132] ■ Encode repeated vertex positions and vertex attribute values in a lossless manner
[0133] ■Disable wavelet transform (transform bypass mode)
[0134] ■ Perform a search in the coded parameter space to determine a set of parameters encoded in the bitstream and ensure that the positions and attribute values match
[0135] ■Do nothing
[0136] ○ The decoder shall decompress each vertex label information to be able to identify repeated vertices. If the encoder signals different encoding parameters for repeated vertices compared to non-repeated vertices, the decoder shall adaptively switch between the two sets of parameters based on the vertex type (e.g., repeated vs. non-repeated).
[0137] ○ The decoder may apply smoothing and automatic stitching as post-processing based on signaling provided by the encoder or based on analysis of the decoded mesh. In a specific embodiment, each vertex or patch includes a flag for enabling / disabling such post-processing. Among other things, the flag may:
[0138] ■Signaled as a SEI message,
[0139] ■ Encode using the trellis codec, or
[0140] ■Signaled in the atlas sub-bitstream.
[0141] In another embodiment, the encoder may replicate regions of the grid and store them in multiple tiles. This may be done to:
[0142] Fault tolerance,
[0143] Protective tape, and
[0144] Seamless / adaptive streaming.
[0145] Example Scheme for Signaling Displacement Information
[0146] As discussed above, the displacement information indicates the spatial differences between vertices in the target mesh and the predicted mesh, such as may be predicted by subdividing the base mesh. As further described herein, the number of vertices in the target mesh and the predicted mesh is the same, and the vertices match exactly. The displacement data may be signaled as is, i.e., as the difference between the target mesh and the predicted mesh. Alternatively, the displacement data may be mathematically transformed. The transformed coefficients may be decoded and signaled in a manner similar to that described above. Whether transformed or not, the difference data may be signaled using a video sub-bitstream. Alternatively, the information may be signaled via another separate and dedicated data sub-bitstream, referred to as a displacement sub-bitstream. In another scenario, the information may be signaled using an atlas data sub-bitstream, for example, as part of the mesh tile data information. Additionally, the displacement information may be signaled in a base mesh sub-bitstream. For example, Fig.12Applying displacements (such as signaled in displacement information) to predicted vertex positions, such as may be predicted by subdividing a base mesh, to reposition vertices to target positions that better match an initial version of volumetric visual content that is (or has been) compressed is illustrated.
[0147] When the total number of target vertices and predicted vertices is N, the vertex position in the target mesh is in And the vertex positions in the predicted mesh are in The displacement value of vertex k is as follows:
[0148] D vk ={D k,0 ,D k,1 ,D k,2}={O k,0 -P k,0 ,O k,1 -P k,1 ,O k,2 -P k,2}
[0149] On the encoder side, the displacement data can be transformed by a transformation method such as wavelet transform. Transformed to In the case where the displacement data is not transformed, The transformed data may then be quantized.
[0150] After quantization (if applicable), Can be placed on the image. Corresponds to the {X,Y,Z} coordinates in 3D space Each entry of k,0 ,C k,1 ,C k,2} can be placed on different image planes. For example, the first entry {C 0,0 ,C 1,0 ,C 2,0 ,...,C N-1,0} is placed on the Y plane, and the second entry {C 0,1 ,C 1,1 ,C 2,1 ...C N-1,1} and the third entry {C 0,2 ,C 1,2 ,C 2,2 ...C N-1,2} are placed on the U plane and V plane respectively.
[0151] In another embodiment, The entries of can be placed on the three planes of the image at different granularities. For example, each value of the first entry {C 0,0 ,C 1,0 ,C 2,0 ,...,C N-1,0} is placed multiple times on the Y plane, while the values of the other entries are placed once on the corresponding planes. This multiple placement takes into account the following chroma plane scaling for chroma format changes. In another embodiment, 0 can be used instead of duplicating the first entry. For example, Fig.13 Three color planes of a video image frame, such as a Y plane, a U plane, and a V plane, and subsampling (e.g., a 4:2:0 chroma format) are illustrated, wherein coefficients (resulting from applying a wavelet transform to a displacement) are signaled in the corresponding color planes to signal coefficients of displacement motion in multiple directions (such as X, Y, and Z). For example, a coefficient of an X displacement may be signaled in the Y plane, a coefficient of a Y displacement may be signaled in the U plane, and a coefficient of a Z displacement may be signaled in the V plane. In the notation used, the first digit after "C" may indicate an index value of the displacement, such as a first displacement, a second displacement, etc. The last digit may indicate a displacement component, such as 0 may indicate movement in the X direction, 1 may indicate movement in the Y direction, and 2 may indicate movement in the Z direction, etc.
[0152] In another embodiment, the values may be placed contiguously on a plane. In this case, after all first entries are placed on the plane, all second entries are then placed. All third entries are then placed on the same plane. Conceptually, when the image resolution is WxH, the first entry (C k,0 ) can be placed on (k / W, k%W). The second entry (C k,1 ) can be placed on ((N+k) / W, (N+k)%W), and then the first entry of the coefficient value of vertex k (C k,2 ) can be placed on ((2N+k) / W, (2N+k)%W).
[0153] In another embodiment, the values may be placed in an interleaved manner on a plane. In this case, the 3 values of a vertex and then the 3 values of the next vertex are placed consecutively on the same plane. Conceptually, when the image resolution is WxH, the first entry of the coefficient value of vertex k (C k,0 ) can be placed on (3k / W, 3k%W). The second entry of the coefficient value of vertex k (C k,1 ) can be placed on ((3k+1) / W, (3k+1)%W), and then the first entry of the coefficient value of vertex k (C k,2) can be placed on ((3k+2) / W, (3k+2)%W).
[0154] In another embodiment, the coefficients may be placed in various orders. Instead of being placed in a grid order, they may be placed in a zigzag order, vertical first order, horizontal first order, diagonal up order, or diagonal down order. The order may be signaled in the bitstream.
[0155] In another embodiment, these values may be placed on a restricted area, such as a block in an image. In this case, the W value corresponds to the width of the block. Fig.14 shows a block-based Example of placement.
[0156] The concepts of sub-grids and tiles and how to signal them are also described above. A sub-grid is an independently decodable sub-portion of a grid. A tile is a group of triangles in a sub-portion that share the same transform / subdivision information and whose coefficient data is signaled together. In some embodiments, the number of coefficients of a tile is the same as the vertex represented by the tile.
[0157] In some embodiments, the coefficient values of each tile may be placed in different areas of the video image frame, such as Fig.15 shown.
[0158] Displacement data in atlas sub-bitstream
[0159] Since coefficient data can be bound to a slice, the corresponding coefficients can be signaled in the slice. This guarantees partial decoding and also enables random access to the decoded mesh. The slice structure also gives the functionality of inter-frame prediction. For example, at the end of the slice data for V-DMC, within the structures mesh_intra_data_unit, mesh_inter_data_unit and mesh_merge_data_unit, the size of the coefficient data and a set of coefficients can be signaled. An example of a signaling mechanism is as follows:
[0160]
[0161] In another embodiment, the signaling of coefficients in the tile data may be controlled by a flag in the corresponding atlas frame parameter set. In the following example, the flag afps_patch_coefficient_signal_enable_flag signaled in the corresponding atlas frame parameter set indicates that the size of the coefficient data (mdu_coefficient_size) is signaled or inferred to be 0. For example:
[0162]
[0163] In another embodiment, the signaling of coefficients in the tile data may be controlled by a flag in the corresponding atlas sequence parameter set.
[0164] In another embodiment, a flag in the atlas frame parameter set may be controlled by a flag in the atlas sequence parameter set.
[0165] In another embodiment, mdu_coefficient_size is not signaled, but the size of the coefficient is inferred to be equal to mdu_vertex_count_minus1 + 1. In this case, an example of a syntax table is as follows:
[0166]
[0167] If flags from the atlas frame parameter set or atlas sequence parameter set control the signaling mechanism, the flags apply directly to the signaling loop as follows:
[0168]
[0169] When the current slice is predicted from its reference slice, mdu_coefficient_size and mdu_coefficient_byte may be explicitly signaled.
[0170] In another embodiment, mdu_coefficient_size and mdu_coefficient_byte may be set to be the same as the mdu_coefficient_size and mdu_coefficient_byte of the reference slice. In this case, the elements mdu_coefficient_size and mdu_coefficient_byte may not be signaled at all.
[0171] In another embodiment, mdu_coefficient_size may be explicitly signaled. Then, the difference between the coefficient value of the current slice and the coefficient value of the reference slice may be signaled. When the current coefficient size is greater than the coefficient size of the reference, the coefficient value of the reference slice is inferred to be equal to 0. In another embodiment, the coefficient value may be inferred to be equal to the value of the last coefficient in the current slice.
[0172] In another embodiment, mdu_coefficient_size is inferred from the size signaled in the reference slice. Then, the coefficient value is signaled. In another case, the difference between the coefficient value of the current slice and the coefficient value of the reference slice may be signaled. When the current coefficient size is greater than the coefficient size of the reference, the coefficient value of the reference slice is inferred to be 0. When the coefficient value of the current slice is When the number of coefficients is Nc, the coefficient value of the reference slice is And the number of coefficients is Nr,C k It can be described as follows:
[0173]
[0174] In another embodiment, the coefficient value may be inferred to be equal to the value of the last coefficient in the current slice, as follows:
[0175]
[0176] The same approach as described above can be applied when some information of the current slice is predicted from its reference slices and some information is explicitly signaled in the mesh merged data unit.
[0177] In another embodiment, additional coefficients to the coefficients from the reference slice may be signaled. The final coefficients will be a continuous series of coefficients from the reference slice and the signaled coefficients. For example:
[0178]
[0179]
[0180] In another embodiment, the size may indicate the difference between the current coefficient size and the referenced coefficient size. When the size is less than 0, mmdu_coefficient_byte may not be signaled at all. The current coefficient size is derived as (mmdu_vertex_count_minu1+1+mmdu_additional_coefficient_size), and only this size of the referenced coefficient value will be used.
[0181]
[0182] In some implementations, mdu_coefficient_size and mdu_coefficient_byte may be explicitly signaled.
[0183] In another embodiment, mdu_coefficient_size and mdu_coefficient_byte are the same as the mdu_coefficient_size and mdu_coefficient_byte of the reference slice. In this case, mdu_coefficient_size and mdu_coefficient_byte may not be signaled at all.
[0184] In another embodiment, mdu_coefficient_size is explicitly signaled. Then, the difference between the coefficient value of the current slice and the coefficient value of the reference slice may be signaled. When the current coefficient size is greater than the coefficient size of the reference, the coefficient value of the reference slice is inferred to be 0. In another embodiment, the coefficient value may be inferred to be the value of the last coefficient in the current slice.
[0185] In another embodiment, mdu_coefficient_size is inferred from the size signaled in the reference slice. Then, the coefficient value is signaled. In other cases, the difference between the coefficient value of the current slice and the coefficient value of the reference slice may be signaled. When the current coefficient size is greater than the coefficient size of the reference, the coefficients with indices greater than the last index of the reference are inferred to be 0. In another embodiment, the coefficient value may be inferred to be the value of the last coefficient in the current slice.
[0186] For skip slice mode, no information is signaled. When decoding a coefficient, all information including the coefficient in the reference slice is copied to use as itself.
[0187] When coefficient data can be signaled in the slice data, each slice data can indicate whether its corresponding coefficient data is in the slice unit (atlas data sub-bitstream) or in the video (geometry video sub-bitstream). In this case, each slice will have a flag indicating each slice as follows. When a flag (e.g., mdu_coefgender_in_outstream_flag) is equal to 1 (true), which indicates that the coefficient value is in the video stream, the location and size of the geometry information is signaled. In this case, the coefficient value is only signaled when the flag is equal to 0 (false).
[0188]
[0189] In another embodiment, slice data units dedicated to signaling coefficient data may be defined. That is, coefficient intra data units, coefficient inter data units, coefficient merge data units, and coefficient skip data units. The syntax table of coefficient intra data units may be as follows:
[0190]
[0191] In some implementations, cdu_patch_id indicates the index of the tile to which the current coefficient data unit corresponds. In addition, in some implementations, cdu_coefficient_size indicates the size of the coefficient data.
[0192] For coefficient inter data units, the index of the reference coefficient data unit is signaled. When the number of coefficients (or the size of the coefficient data) is different from the reference, the newly added coefficients are signaled. In the following example, when cdu_coefficient_size_diff is positive, cduCoefficientDiffSize is the same as cdu_coefficient_size_diff. Otherwise, it is set to 0.
[0193]
[0194] In another embodiment, cdu_remaining_coefficient_byte is not signaled, but when the current coefficient size is greater than the reference coefficient size, coefficients with indices greater than the last index of the reference are inferred to be 0. In another embodiment, the coefficient value may be inferred to be the value of the last coefficient in the current slice.
[0195] The coefficient skip data unit can be described as follows:
[0196] coefficient skip data unit(tileID,patchIdx){ Descriptors }
[0197] When coefficient data may be signaled in the slice data and a new sub-bitstream for displacement data (coefficient data) only is also defined as further described below, each slice data unit may indicate whether its corresponding coefficient data is in the slice unit (atlas data sub-bitstream) or in the new sub-bitstream (displacement sub-bitstream). When the flag mdu_coefficient_in_outstream_flag is true, it indicates that the coefficient value is in the displacement data sub-bitstream. The coefficient value is signaled only when the flag is false. Since the video stream may not be used at all, there is no need to signal the slice information of the geometry video, such as mdu_geometry_2d_pos_x, mdu_geometry_2d_pos_y, mdu_geometry_2d_size_x_minus1, and mdu_geometry_2d_size_y_minus1.
[0198] In another embodiment, each slice can be in the video sub-bitstream, in the atlas data sub-bitstream or in the displacement data sub-bitstream to signal the corresponding coefficient. In this case, the index mdu_coefgender_stream_index can be signaled, rather than a flag. For example, when mdu_coefgender_stream_index is 0, the coefficient is signaled in the video, and 2d position and size information is needed in the slice. When mdu_coefgender_stream_index is 1, the coefficient is signaled in the displacement data sub-bitstream. When mdu_coefgender_stream_index is 2, the coefficient is signaled in the slice data unit (e.g., the atlas data sub-bitstream). When mdu_coefgender_stream_index is 1 or 2, 2d position and size information is not needed in the slice data unit. When mdu_coefgender_stream_index is 0 or 1, the coefficient value is not signaled in the slice data unit (e.g., the atlas data sub-bitstream). The index can be different from 0, 1 or 2. In another embodiment, there may be more than 3 cases, namely, signaling coefficients only in the video sub-bitstream, signaling coefficients only in the displacement sub-bitstream, and signaling coefficients only in the atlas data sub-bitstream. The coefficients of a tile may be signaled in both the tile data unit (atlas data sub-bitstream) and in the other sub-bitstream. In this case, the coefficients in the tile data unit appear first.
[0199] Displacement data sub-bitstream
[0200] In some embodiments, the displacement data (transformed into wavelet coefficients or other) may be signaled independently in a sub-bitstream separate from other sub-bitstreams (such as an atlas data sub-bitstream, a base grid sub-bitstream, or a video sub-bitstream). In this case, the displacement data sub-bitstream structure has its own high-level syntax to ensure that partial decoding, random accessibility, inter-frame prediction, etc. are enabled. The displacement data sub-bitstream may include a displacement sequence parameter set, a displacement frame parameter set, and a displacement block layer. Each block layer may have a header containing a frame parameter set id, a block type (intra-frame, inter-frame, merge, skip), the size of its payload, etc. The payload contains displacement data, such as wavelet transform coefficients. When using a displacement data sub-bitstream, signaling the transform information (such as a transform parameter set) in the atlas data sub-bitstream is replaced by signaling in the displacement data sub-bitstream. The signaling mechanism for how to signal the transform parameter set in the displacement data sub-bitstream is a similar signaling mechanism as described above for when the displacement data is signaled in the atlas data sub-bitstream.
[0201] NAL unit syntax for displacement data sub-bitstream
[0202] As discussed above, the displacement data sub-bitstream is also based on NAL units, and the NAL units are similar to the NAL units of the atlas sub-bitstream. As a NAL sample stream, the NAL unit size precision can be signaled at the beginning of the displacement data sub-bitstream, and the NAL unit size (interpreted as NumBytesInalUnit) can be signaled per NAL unit. An example syntax is provided below.
[0203] Generic NAL unit syntax
[0204]
[0205] NAL unit header syntax
[0206]
[0207] NAL unit semantics for displacement data sub-bitstream
[0208] This part contains some semantics corresponding to the above grammatical structure.
[0209] 1. General NAL unit semantics
[0210] NumBytesInNalUnit specifies the size of the NAL unit in bytes. This value is used for decoding of the NAL unit. Some form of demarcation of NAL unit boundaries is necessary to implement the inference of NumBytesInNalUnit.
[0211] rbsp_byte[i] is the i-th byte of the RBSP (raw byte sequence payload). The RBSP is specified as an ordered sequence of bytes as follows:
[0212] RBSP contains the following data bit string (SODB):
[0213] If SODB is empty (ie, zero bits in length), then RBSP is also empty.
[0214] Otherwise, the RBSP contains the following SODB:
[0215] 1) The first byte of the RBSP contains the first (most significant, leftmost) eight bits of SODB; the next byte of the RBSP contains the next eight bits of SODB, and so on, until less than eight bits of SODB remain.
[0216] 2) The syntax structure of rbsp_trailing_bits() is as follows and exists after SODB:
[0217] i) The first (most significant, leftmost) bit of the final RBSP byte contains the remaining bits of SODB (if any).
[0218] ii) The next bit consists of a single bit equal to 1 (eg, rbsp_stop_one_bit).
[0219] iii) When rbsp_stop_one_bit is not the last bit of a byte that is byte-aligned, there are one or more bits equal to 0 (eg, an instance of rbsp_alignment_zero_bit) to cause byte alignment.
[0220] The "_rbsp" suffix is used in the syntax tables to indicate syntax structures that have these RBSP properties. These structures are carried within NAL units as the contents of the rbsp_byte[i] data bytes. The association of RBSP syntax structures with NAL units is specified in the following table.
[0221] In some embodiments, when the boundaries of the RBSP are known, the decoder can extract the SODB from the RBSP by concatenating the bits of the bytes of the RBSP and discarding the rbsp_stop_one_bit, which is the last (least significant, rightmost) bit that is equal to 1, and discarding any subsequent (less significant, more rightward) bits that are equal to 0. The data required for the decoding process is contained in the SODB portion of the RBSP.
[0222] NAL unit header semantics for displacement data sub-bitstream
[0223] Similar NAL unit types as for the atlas data sub-bitstream may be used, such as defined for coefficients that enable functionality for random access and segmentation of the grid. In the displacement data sub-bitstream, the concept of displacement partitioning and specific NAL units are defined to correspond to the decoded grid data. Additionally, NAL units are defined that may include metadata, such as SEI messages.
[0224] Specifically, the supported coefficient NAL unit types are specified as follows:
[0225]
[0226]
[0227] The main syntax structure defined for a bitstream is the sequence parameter set. This syntax structure contains basic information about the bitstream, features that identify the codecs supported for intra- and inter-coding trellises, and information about references.
[0228] Generic Shift Sequence Parameter Set RBSP Syntax
[0229]
[0230]
[0231] dsps_displacement_data_size_precision_bytes_minus1(+1) specifies the precision of the size of the displacement data (in bytes).
[0232] dsps_transform_index indicates which transform is used for the displacement in the sequence. 0 may indicate that it is not applied.
[0233] When the transformation is LINEAR_LIFTING, the transformation parameter may be signaled as vmc_lifting_transform_parameter.
[0234]
[0235] Coefficient distribution, layers, and level syntax
[0236]
[0237]
[0238] The dptl_extended_sub_profile_flag providing support for sub-profiles may be very useful to further restrict coefficient profiles according to usage and application.
[0239] Displaced frame parameter set RBSP syntax for the displaced data sub-bitstream
[0240] The displacement frame parameter set has frame level information such as the number of displacement tiles (dispTile) in a frame corresponding to one disp_frm_order_cnt_lsb. The displacement data of a grid data tile is decoded in one displacement_data_tile_layer() and can be decoded independently of other dispTiles. In case of inter prediction, a dispTile can only refer to dispTiles with the same dh_id in its associated reference frame.
[0241]
[0242]
[0243] Displacement Block Layer RBSP Syntax in the Displacement Data Sub-Bitstream
[0244] disp_tile_layer contains the tile information. One or more disp_tile_layer_rbsp may correspond to one frame indicated by dh_frm_order_cnt_lsb. DispUnitSize may be derived from NumBytesInNalUnit and the size of disp_header(). In another embodiment, the size may be explicitly signaled in disp_header().
[0245]
[0246] Displacement Header Syntax
[0247]
[0248]
[0249] dh_id is the id of the current displacement contained in the mesh data displacement data.
[0250] dh_type indicates how the displacement is decoded. If dh_type is I_DISP, the data is not predicted from any other frame or partition. If dh_type is P_DISP or M_DISP, the displacement data is decoded using inter-frame prediction.
[0251]
[0252] Displacement Data Unit
[0253]
[0254] In some embodiments, disp_data_unit(unitSize) contains a stream of displacement units of size unitSize in bytes, as an ordered stream of bytes or bits within which the locations of unit boundaries can be identified based on patterns in the data. coefficient_byte may be interpreted as described in Section 5 using dh_type as the slice type and dispUnitSize as the coefficient_size. In this case, coefficient_size is not predicted at all.
[0255] In some embodiments, the index of the tile may be signaled in both the atlas data sub-bitstream and the displacement data sub-bitstream. For example, where a tile in the atlas data sub-bitstream needs to find a corresponding tile in the displacement data sub-bitstream, the corresponding displacement tile index may be signaled in a grid tile data unit. In another embodiment, the corresponding grid tile data unit may be signaled in a displacement tile data unit.
[0256] Example Computer System
[0257] Fig.16 An example computer system 1600 is shown that can implement an encoder or decoder or any other of the components described herein (e.g., as described above with reference to FIG. 1 ). Figures 1 to 15 1600). The computer system 1600 may be configured to perform any or all of the embodiments described above. In various embodiments, the computer system 1600 may be any of various types of devices, including, but not limited to, a personal computer system, a desktop computer, a laptop computer, a notebook computer, a tablet computer, an all-in-one computer, a tablet or netbook computer, a mainframe computer system, a handheld computer, a workstation, a network computer, a camera, a set-top box, a mobile device, a consumer device, a video game controller, a handheld video game device, an application server, a storage device, a television, a video recording device, a peripheral device (such as a switch, a modem, a router), or generally any type of computing or electronic device.
[0258] The various embodiments of the point cloud encoder or decoder described herein may be executed on one or more computer systems 1600, which may interact with various other devices. Figures 1 to 15 Any component, action or functionality described may be implemented in a configuration as Fig.161600. In the illustrated embodiment, the computer system 1600 includes one or more processors 1610 coupled to a system memory 1620 via an input / output (I / O) interface 1630. The computer system 1600 also includes a network interface 1640 coupled to the I / O interface 1630, and one or more input / output devices 1650, such as a cursor control device 1660, a keyboard 1670, and one or more displays 1680. In some cases, it is contemplated that the embodiments may be implemented using a single instance of the computer system 1600, while in other embodiments, multiple such systems or multiple nodes making up the computer system 1600 may be configured to host different portions or instances of the embodiments. For example, in one embodiment, some elements may be implemented via one or more nodes of the computer system 1600 that are different from those nodes that implement other elements.
[0259] In various embodiments, computer system 1600 may be a uniprocessor system including one processor 1610, or a multiprocessor system including a number of processors 1610 (e.g., two, four, eight, or another suitable number). Processor 1610 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1610 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as x86, PowerPC, SPARC, or MIPS ISAs, or any other suitable ISAs. In a multiprocessor system, each of processors 1610 may typically, but not necessarily, implement the same ISA.
[0260] The system memory 1620 may be configured to store point cloud compression or point cloud decompression program instructions 1622 and / or sensor data accessible by the processor 1610. In various embodiments, the system memory 1620 may be implemented using any suitable memory technology, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash type memory, or any other type of memory. In the illustrated embodiment, the program instructions 1622 may be configured to implement an image sensor control application incorporating any of the above functionalities. In some embodiments, the program instructions and / or data may be received, sent, or stored on different types of computer accessible media or similar media separate from the system memory 1620 or the computer system 1600. Although the computer system 1600 is described as implementing the functionality of the functional blocks of the preceding figures, any functionality described herein may be implemented via such a computer system.
[0261] In one embodiment, the I / O interface 1630 may be configured to coordinate I / O communications between the processor 1610, the system memory 1620, and any peripheral devices in the device (including the network interface 1640 or other peripheral device interfaces, such as the input / output device 1650). In some embodiments, the I / O interface 1630 may perform any necessary protocol, timing, or other data conversion to convert data signals from one component (e.g., the system memory 1620) into a format suitable for use by another component (e.g., the processor 1610). In some embodiments, the I / O interface 1630 may include support for devices attached, for example, via various types of peripheral buses (such as a variation of the peripheral component interconnect (PCI) bus standard or the universal serial bus (USB) standard). In some embodiments, the functionality of the I / O interface 1630 may be divided into two or more separate components, such as a north bridge and a south bridge, for example. In addition, in some embodiments, some or all of the functionality of the I / O interface 1630 (such as an interface to the system memory 1620) may be directly incorporated into the processor 1610.
[0262] The network interface 1640 may be configured to allow data to be exchanged between the computer system 1600 and other devices (e.g., carriers or proxy devices) attached to the network 1685 or between nodes of the computer system 1600. In various embodiments, the network 1685 may include one or more networks, including, but not limited to, a local area network (LAN) (e.g., an Ethernet or an enterprise network), a wide area network (WAN) (e.g., the Internet), a wireless data network, some other electronic data network, or some combination thereof. In various embodiments, the network interface 1640 may support communication via a wired or wireless general data network (such as any suitable type of Ethernet network), for example; communication via a telecommunications / telephone network (such as an analog voice network or a digital fiber optic communication network); communication via a storage area network (such as a Fiber Channel SAN), or communication via any other suitable type of network and / or protocol.
[0263] In some embodiments, input / output devices 1650 may include one or more display terminals, keyboards, keypads, touchpads, scanning devices, voice or optical recognition devices, or any other device suitable for inputting or accessing data by one or more computer systems 1600. Multiple input / output devices 1650 may be present in computer system 1600, or may be distributed across various nodes of computer system 1600. In some embodiments, similar input / output devices may be separate from computer system 1600 and may interact with one or more nodes of computer system 1600 via a wired or wireless connection (such as via network interface 1640).
[0264] like Fig.16 As shown, memory 1620 may include program instructions 1622, which may be executable by a processor to implement any of the elements or actions described above. In one embodiment, the program instructions may execute the method described above. In other embodiments, different elements and data may be included. It should be noted that the data may include any data or information described above.
[0265] Those skilled in the art will appreciate that computer system 1600 is merely illustrative, and is not intended to limit the scope of the embodiments. Specifically, computer systems and devices may include any combination of hardware or software that can perform the indicated functions, including computers, network equipment, Internet equipment, personal digital assistants, wireless telephones, pagers, etc. Computer system 1600 may also be connected to other devices not shown, or may otherwise be operated as an independent system. In addition, the functions provided by the illustrated components may be combined in fewer components or distributed in additional components in some embodiments. Similarly, in some embodiments, the functions of some components in the illustrated components may not be provided, and / or other additional functions may be available.
[0266] Those skilled in the art will also recognize that, although various items are illustrated as being stored in memory or on storage devices during use, for the purpose of memory management and data integrity, these items or parts thereof may be transmitted between memory and other storage devices. Alternatively, in other embodiments, some or all of these software components may be executed in a memory on another device, and communicate with the illustrated computer system via inter-computer communication. Some or all of the system components or data structures may also be stored (e.g., as instructions or structured data) on a computer accessible medium or portable article to be read by a suitable driver, and various examples thereof are described above. In some embodiments, instructions stored on a computer accessible medium separated from computer system 1600 may be transmitted to computer system 1600 via a transmission medium or signal (such as an electrical signal, an electromagnetic signal, or a digital signal transmitted by a communication medium such as a network and / or a wireless link). Various embodiments may also include receiving, sending, or storing instructions and / or data implemented according to the above description on a computer accessible medium. Generally speaking, computer-accessible media may include non-transitory computer-readable storage media or memory media, such as magnetic or optical media, for example, disks or DVD / CD-ROMs, volatile or non-volatile media, such as RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), ROM, etc. In some embodiments, computer-accessible media may include transmission media or signals, such as electrical, electromagnetic, or digital signals transmitted via a communication medium such as a network and / or a wireless link.
[0267] In different embodiments, the methods described herein can be implemented in software, hardware, or a combination thereof. In addition, the order of the frames of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc. For those skilled in the art who benefit from the present disclosure, various modifications and changes can obviously be made. The various embodiments described herein are intended to be illustrative and not restrictive. Many variations, modifications, additions, and improvements are possible. Therefore, multiple examples can be provided for the components described herein as a single example. The boundaries between various components, operations, and data repositories are arbitrary to a certain extent, and specific operations are shown in the context of a specific example configuration. Other allocations of functions are contemplated, and they may fall within the scope of the appended claims. Finally, the structure and function presented as discrete components in the example configuration may be implemented as a combined structure or component. These and other variations, modifications, additions, and improvements may fall within the scope of the embodiments defined in the following claims.
Claims
1. A non-transitory computer-readable storage medium storing program instructions that, when executed using one or more computing devices, cause the one or more computing devices to: The visual volume content is compressed using a dynamic mesh compression algorithm, wherein the visual volume is compressed include: compressing the base mesh to generate a base mesh sub-bitstream to be included in a bitstream for compressed visual volumetric content; Determining displacement information for displacements to be applied to subdivided locations of the base mesh; and compressing the attribute information, wherein the compressed attribute information is to be included in a video sub-bitstream of the bitstream for the compressed visual volumetric content; as well as The bitstream for the compressed visual volumetric content is provided, wherein the displacement information is at least partially signaled in the bitstream in a sub-bitstream other than the video sub-bitstream.
2. The non-transitory computer-readable storage medium of claim 1 , wherein the displacement information is signaled in a displacement sub-bitstream, the displacement sub-bitstream being a sub-bitstream separate from the video sub-bitstream, the base grid sub-bitstream, and the atlas data sub-bitstream of the bitstream for the compressed visual volumetric content. 3 . The non-transitory computer-readable storage medium of claim 1 , wherein the displacement information is signaled in an atlas data sub-bitstream of the bitstream for the compressed visual volumetric content.
4. The non-transitory computer-readable storage medium of claim 3, wherein the displacement information is signaled at least in part in a network abstraction layer unit (NAL unit) of a slice data unit syntax used in the atlas data sub-bitstream.
5. The non-transitory computer-readable storage medium of claim 3, wherein the displacement information is signaled at least in part in a sequence parameter set header of the slice data unit syntax used in the atlas data sub-bitstream.
6. The non-transitory computer-readable storage medium of claim 5, wherein the displacement information is signaled at least in part in a frame parameter set header of the slice data unit syntax used in the atlas data sub-bitstream. 7 . The non-transitory computer-readable storage medium of claim 1 , wherein the displacement information is signaled in the base grid sub-bitstream of the bitstream for the compressed visual volumetric content.
8. The non-transitory computer-readable storage medium of claim 1, wherein the displacement information is further signaled at least in part in the video sub-bitstream, wherein a flag in a portion of the displacement information signaled in a sub-bitstream other than the video sub-bitstream is used to signal the portions of the displacement information signaled in the video sub-bitstream.
9. The non-transitory computer-readable storage medium of claim 8, wherein the portion of the displacement information signaled in the video sub-bitstream is signaled such that: signaling a first displacement component in a first color plane of the video sub-bitstream; signaling a second displacement component in a second color plane of the video sub-bitstream; and A third displacement component is signaled in a third color plane of the video sub-bitstream, wherein one or more resolutions of the first component, the second component, and the third component signaling the displacement information are adjusted to account for subsampling between the first color plane, the second color plane, and the third color plane of the video sub-bitstream.
10. The non-transitory computer-readable storage medium of claim 8, wherein the portion of the displacement information signaled in the video sub-bitstream is signaled using a single color plane of the video sub-bitstream.
11. The non-transitory computer-readable storage medium of claim 8, wherein the portion of the displacement information signaled in the video sub-bitstream is signaled such that: Starting points in different blocks of the image frames of the video sub-bitstream are used to signal displacement information of different sub-grids of the reconstructed grid.
12. A non-transitory computer-readable storage medium storing program instructions that, when executed using one or more computing devices, cause the one or more computing devices to: receiving a bitstream representing a compressed version of a visual volume content, the bitstream include: Base grid sub-bitstream; Video sub-bitstream; and displacement information to be applied to displacements of subdivision positions of a base grid signaled in said base grid sub-bitstream, wherein the displacement information is at least partially signaled in the bitstream in a sub-bitstream other than the video sub-bitstream; and reconstructing a mesh of the visual volumetric content, wherein to reconstruct the mesh, the program instructions cause the one or more computing devices to: subdividing the edge of the base mesh to generate the subdivided position; parsing the bit stream to identify the displacement information; as well as The displacement indicated in the displacement information is applied to the subdivided position of the base mesh.
13. The non-transitory computer-readable storage medium of claim 12, wherein the displacement information is signaled in a displacement sub-bitstream, the displacement sub-bitstream being a sub-bitstream separate from the video sub-bitstream, the base grid sub-bitstream, and the atlas data sub-bitstream.
14. The non-transitory computer-readable storage medium of claim 12, wherein the displacement information is signaled in an atlas data sub-bitstream of the bitstream.
15. The non-transitory computer-readable storage medium of claim 12, wherein the displacement information is signaled in the base grid sub-bitstream of the bitstream.
16. The non-transitory computer-readable storage medium of claim 12, wherein the displacement information is further signaled at least in part in the video sub-bitstream.
17. The non-transitory computer-readable storage medium of claim 16, wherein a flag in a portion of the displacement information signaled in a sub-bitstream other than the video sub-bitstream is used to signal the portions of the displacement information signaled in the video sub-bitstream.
18. A device, the device include: a memory storing program instructions; and one or more processors, wherein the program instructions, when executed on or across the one or more processors, cause the one or more processors to: Receiving a bitstream representing a compressed version of visual volumetric content, the bitstream comprising: Base grid sub-bitstream; a video sub-bitstream; and displacement information to be applied to displacements of subdivision positions of a base grid signaled in said base grid sub-bitstream, wherein the displacement information is at least partially signaled in the bitstream in a sub-bitstream other than the video sub-bitstream; Reconstructing a mesh of the visual volume content, wherein reconstructing the mesh comprises: subdividing the edge of the base mesh to generate the subdivided position; parsing the bitstream to identify the displacement information; and The displacement indicated in the displacement information is applied to the subdivided position of the base mesh.
19. The apparatus of claim 18, wherein the displacement information is signaled at least in part in an atlas data sub-bitstream of the bitstream.
20. The apparatus of claim 18, wherein the displacement information is signaled at least in part in a displacement sub-bitstream, the displacement sub-bitstream being a sub-bitstream separate from the video sub-bitstream, the base grid sub-bitstream, and the atlas data sub-bitstream of the bitstream.