An apparatus, a method and a computer program for volumetric video
Patent Information
- Application Number
- EP2022860687
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-26
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-02
AI Technical Summary
Current volumetric video coding standards, such as V3C, do not allow for the mixing of different representations like MIV and V-PCC in a single bitstream, limiting flexibility and adoption due to decoder complexity and compatibility issues.
A method and apparatus that combine bitstreams of different 3D volumetric representations encoded in various formats by analyzing syntax elements, determining signaling elements to indicate format differences, and merging them into a multi-format volumetric representation, enabling decoding by a single decoder.
Enables the combination of different V3C technologies and profiles into a single bitstream, enhancing flexibility and adoption by providing signaling information for correct interpretation and decoding, thus improving coding efficiency and 6DOF capabilities.
Smart Images

Figure 1.1
Abstract
Description
AN APPARATUS, A METHOD AND A COMPUTER PROGRAM FORVOEUMETRIC VIDEOTECHNICAE FIEED
[0001] The present invention relates to an apparatus, a method and a computer program for volumetric video coding.BACKGROUND
[0002] Visual volumetric video-based coding (V3C; defined in ISO / IEC DIS 23090-5) provides a generic syntax and mechanism for volumetric video coding. The generic syntax can be used by applications targeting volumetric content, such as point clouds, immersive video with depth, and mesh representations of volumetric frames. The purpose of the specification is to define how to decode and interpret the associated data (atlas data in ISO / IEC 23090-5) which tells a Tenderer how to interpret 2D frames for reconstructing volumetric frames.
[0003] The two applications of V3C (ISO / IEC 23090-5), i.e. video-based point cloud compression (V-PCC; defined in ISO / IEC 23090-5) and MPEG immersive video (MIV; defined in ISO / IEC 23090-12), use a number of V3C syntax elements with a slightly modified semantics. Moreover, MPEG 3DG (ISO SC29 WG7) group has started a work on a third volumetric video coding application, i.e. V3C mesh compression.
[0004] Thus, a volumetric frame can be created and represented by a mix of technologies. Some type of content, materials and surfaces could benefit from one representation over another. However, the current V3C does not allow to mix different representations (i.e. MIV and V-PCC) in one V3C bitstream in a way that a client would be able to decode and consume the content as one. Such inflexibility in the usage of V3C family of standards may slow down the adoption of V3C technologies.Therefore, there is a need to find an enhanced solution to allow combining different technologies and profiles into a single bitstream.SUMMARY
[0005] Now, an improved method and technical equipment implementing the method has been invented, by which the above problems are alleviated. Various aspects include a method, an apparatus and a computer readable medium comprising a computer program, or a signal stored therein, which are characterized by what is stated in the independent claims. Various details of the embodiments are disclosed in the dependent claims and in the corresponding images and description.
[0006] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
[0007] According to a first aspect, there is provided a method comprising obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyzing content and semantics of syntax elements of said at least two 3D volumetric representations; determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and merging the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
[0008] An apparatus according to a second aspect comprises means for obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for analyzing content and semantics of syntax elements of said at least two 3D volumetric representations; means for determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and means for merging the bitstreams of said at least two 3D volumetric representations into a bitstream of a multi-format volumetric representation.
[0009] According to an embodiment, the apparatus comprises means for providing said bitstream of the multi-format volumetric representation to a decoder; and means forproviding said one or more signaling elements, in or along said bitstream of the multiformat volumetric representation to the decoder.
[0010] According to an embodiment, the apparatus comprises means for receiving volumetric frames segmented into said at least two 3D volumetric representations; means for mapping the first 3D volumetric representation to a first format-specific sub-encoder and the second 3D volumetric representation to a second format-specific sub-encoder, wherein said first and second format-specific sub-encoders are configured to encode the content of the respective 3D volumetric representation into video and atlas data according to the format.
[0011] According to an embodiment, the apparatus comprises means for adjusting said first and second format-specific sub-encoders to use a same coordinate space.
[0012] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to merge at least atlas components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0013] According to an embodiment, a signaling of differences in the semantics of the syntax elements of the atlas components of the first format and the second format is configured to be carried out by at least one syntax element included in a visual volumetric video-based coding (V3C) parameter set extension or V3C common atlas sequence parameter set extension data syntax structure.
[0014] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least video components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0015] According to an embodiment, a signaling of differences in the semantics of the syntax elements of the video components of the first format and the second format is configured to be carried out by at least one syntax element included in an atlas frame parameter set extension data syntax structure.
[0016] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least patch dataof the first format and the second format into the bitstream of the multi-format volumetric representation.
[0017] According to an embodiment, said first and second formats are one of the following: V3C MPEG Immersive Video (MIV) format, V3C Video-based Point Cloud Compression (V-PCC) format, or V3C mesh format.
[0018] An apparatus according to a third aspect comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyze content and semantics of syntax elements of said at least two 3D volumetric representations; determine one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and merge the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
[0019] A method according to a fourth aspect comprises: receiving a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receiving, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; separating, from said bitstream, the encoded first 3D volumetric representation to a first formatspecific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and decoding the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
[0020] An apparatus according to a fifth aspect comprises at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at leastone processor, cause the apparatus at least to perform: receive a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receive, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; separate, from said bitstream, the encoded first 3D volumetric representation to a first formatspecific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and decode the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
[0021] An apparatus according to a sixth aspect comprises means for receiving a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for receiving, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; means for separating, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and means for decoding the encoded first 3D volumetric representation with the first format-specific subdecoder and the encoded second 3D volumetric representation with the second formatspecific sub-decoder at least partly based on said one or more syntax elements.
[0022] Computer readable storage media according to further aspects comprise code for use by an apparatus, which when executed by a processor, causes the apparatus to perform the above methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] For a more complete understanding of the example embodiments, reference is now made to the following descriptions taken in connection with the accompanying drawings in which:
[0024] Figs, la and lb show an encoder and decoder for encoding and decoding 2D pictures;
[0025] Figs. 2a and 2b show a compression and a decompression process for 3D volumetric video;
[0026] Fig. 3 shows an example of block-to-patch mapping with 4 projected patches onto an atlas;
[0027] Figs. 4a - 4c show an illustrative example of a patch projection into 2D domain for atlas data;
[0028] Fig. 5 shows a flow chart for an encoding method according to an embodiment;
[0029] Fig. 6 shows an exemplified block chart of an apparatus according to an embodiment;
[0030] Fig. 7 shows an exemplified block chart of an apparatus according to another embodiment;
[0001] Fig. 8 shows an example of input V3C bitstreams and multi-format V3C bitstream according to an embodiment;
[0032] Fig. 9 shows an exemplified block chart of an apparatus according to yet another embodiment; and
[0033] Fig. 10 shows a flow chart for decoding method according to an embodiment.DETAILED DESCRIPTON OF SOME EXAMPLE EMBODIMENTS
[0034] In the following, several embodiments of the invention will be described in the context of point cloud models for volumetric video coding. It is to be noted, however, that the invention is not limited to specific scene models or specific coding technologies. In fact, the different embodiments have applications in any environment where coding of volumetric scene data is required.
[0035] A video codec comprises an encoder that transforms the input video into a compressed representation suited for storage / transmission, and a decoder that can un-compress the compressed video representation back into a viewable form. An encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e. at lower bitrate).
[0036] Volumetric video may be captured using one or more three-dimensional (3D) cameras. When multiple cameras are in use, the captured footage is synchronized so that the cameras provide different viewpoints to the same world. In contrast to traditional 2D / 3D video, volumetric video describes a 3D model of the world where the viewer is free to move and observer different parts of the world.
[0037] Volumetric video enables the viewer to move in six degrees of freedom (6DOF): in contrast to common 360° video, where the user has from 2 to 3 degrees of freedom (yaw, pitch, and possibly roll), a volumetric video represents a 3D volume of space rather than a flat image plane. Volumetric video frames contain a large amount of data because they model the contents of a 3D volume instead of just a two-dimensional (2D) plane. However, only a relatively small part of the volume changes over time. Therefore, it may be possible to reduce the total amount of data by only coding information about an initial state and changes which may occur between frames. Volumetric video can be rendered from synthetic 3D animations, reconstructed from multi-view video using 3D reconstruction techniques such as structure from motion, or captured with a combination of cameras and depth sensors such as LiDAR (Light Detection and Ranging), for example.
[0038] Volumetric video data represents a three-dimensional scene or object, and thus such data can be viewed from any viewpoint. Volumetric video data can be used as an input for augmented reality (AR), virtual reality (VR) and mixed reality (MR) applications. Such data describes geometry (shape, size, position in 3D-space) and respective attributes (e.g. color, opacity, reflectance, . ..), together with any possible temporal changes of the geometry and attributes at given time instances (e.g. frames in 2D video). Volumetric video is either generated from 3D models, i.e. computer-generated imagery (CGI), or captured from real-world scenes using a variety of capture solutions, e.g. a multi-camera, a laser scan, a combination of video and dedicated depths sensors, etc. Also, a combination of CGI and real-world data is possible. Examples of representation formats for such volumetric data are triangle meshes, point clouds, or voxel. Temporal information aboutthe scene can be included in the form of individual capture instances, i.e. “frames” in 2D video, or other means, e.g. position of an object as a function of time.
[0039] Increasing computational resources and advances in 3D data acquisition devices has enabled reconstruction of highly detailed volumetric video representations of natural scenes. Infrared, lasers, time-of-flight and structured light are all examples of devices that can be used to construct 3D video data. Representation of the 3D data depends on how the 3D data is used. Dense voxel arrays have been used to represent volumetric medical data. In 3D graphics, polygonal meshes are extensively used. Point clouds on the other hand are well suited for applications, such as capturing real world 3D scenes where the topology is not necessarily a 2D manifold. Another way to represent 3D data is coding this 3D data as a set of texture and depth map as is the case in the multi-view plus depth. Closely related to the techniques used in multi-view plus depth is the use of elevation maps, and multi-level surface maps.
[0040] In 3D point clouds, each point of each 3D surface is described as a 3D point with color and / or other attribute information such as surface normal or material reflectance. Point cloud is a set of data points in a coordinate system, for example in a three- dimensional coordinate system being defined by X, Y, and Z coordinates. The points may represent an external surface of an object in the screen space, e.g. in a three-dimensional space.
[0041] In dense point clouds or voxel arrays, the reconstructed 3D scene may contain tens or even hundreds of millions of points. If such representations are to be stored or interchanged between entities, then efficient compression of the presentations becomes fundamental. Standard volumetric video representation formats, such as point clouds, meshes, voxel, suffer from poor temporal compression performance. Identifying correspondences for motion-compensation in 3D-space is an ill-defined problem, as both, geometry and respective attributes may change. For example, temporal successive “frames” do not necessarily have the same number of meshes, points or voxel. Therefore, compression of dynamic 3D scenes is inefficient. 2D-video based approaches for compressing volumetric data, i.e. multiview with depth, have much better compression efficiency, but rarely cover the full scene. Therefore, they provide only limited 6DOF capabilities.
[0042] Instead of the above-mentioned approach, a 3D scene, represented as meshes, points, and / or voxel, can be projected onto one, or more, geometries. These geometries may be “unfolded” or packed onto 2D planes (two planes per geometry: one for texture, one for depth), which are then encoded using standard 2D video compression technologies. Relevant projection geometry information may be transmitted alongside the encoded video files to the decoder. The decoder decodes the video and performs the inverse projection to regenerate the 3D scene in any desired representation format (not necessarily the starting format).
[0043] Projecting volumetric models onto 2D planes allows for using standard 2D video coding tools with highly efficient temporal compression. Thus, coding efficiency can be increased greatly. Using geometry-projections instead of 2D-video based approaches based on multiview and depth, provides a better coverage of the scene (or object). Thus, 6DOF capabilities are improved. Using several geometries for individual objects improves the coverage of the scene further. Furthermore, standard video encoding hardware can be utilized for real-time compression / decompression of the projected planes. The projection and the reverse projection steps are of low complexity.
[0044] Figs, la and lb show an encoder and decoder for encoding and decoding the 2D texture pictures, geometry pictures and / or auxiliary pictures. A video codec consists of an encoder that transforms an input video into a compressed representation suited for storage / transmission and a decoder that can uncompress the compressed video representation back into a viewable form. Typically, the encoder discards and / or loses some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate). An example of an encoding process is illustrated in Figure la. Figure la illustrates an image to be encoded (In); a predicted representation of an image block (P'n); a prediction error signal (Dn); a reconstructed prediction error signal (D'n); a preliminary reconstructed image (I'n); a final reconstructed image (R'n); a transform (T) and inverse transform (T1); a quantization (Q) and inverse quantization (Q1); entropy encoding (E); a reference frame memory (RFM); inter prediction (Pinter); intra prediction (Pintra); mode selection (MS) and filtering (F).
[0045] An example of a decoding process is illustrated in Figure lb. Figure lb illustrates a predicted representation of an image block (P'n); a reconstructed predictionerror signal (D'n); a preliminary reconstructed imagea final reconstructed image (R'n); an inverse transform (T1); an inverse quantization (Q1); an entropy decoding (E1); a reference frame memory (RFM); a prediction (either inter or intra) (P); and filtering (F).
[0046] Many hybrid video encoders encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate). Video codecs may also provide a transform skip mode, which the encoders may choose to use. In the transform skip mode, the prediction error is coded in a sample domain, for example by deriving a sample-wise difference value relative to certain adjacent samples and coding the sample-wise difference value with an entropy coder.
[0047] Many video encoders partition a picture into blocks along a block grid. For example, in the High Efficiency Video Coding (HEVC) standard, the following partitioning and definitions are used. A coding block may be defined as an NxN block of samples for some value of N such that the division of a coding tree block into coding blocks is a partitioning. A coding tree block (CTB) may be defined as an NxN block of samples for some value of N such that the division of a component into coding tree blocks is a partitioning. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture that has three sample arrays, or a coding tree block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture that has three sample arrays,or a coding block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A CU with the maximum allowed size may be named as LCU (largest coding unit) or coding tree unit (CTU) and the video picture is divided into non-overlapping LCUs.
[0048] In HEVC, a picture can be partitioned in tiles, which are rectangular and contain an integer number of LCUs. In HEVC, the partitioning to tiles forms a regular grid, where heights and widths of tiles differ from each other by one LCU at the maximum. In HEVC, a slice is defined to be an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined to be an integer number of coding tree units ordered consecutively in the tile scan and contained in a single NAL unit. The division of each picture into slice segments is a partitioning. In HEVC, an independent slice segment is defined to be a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for a preceding slice segment, and a dependent slice segment is defined to be a slice segment for which the values of some syntax elements of the slice segment header are inferred from the values for the preceding independent slice segment in decoding order. In HEVC, a slice header is defined to be the slice segment header of the independent slice segment that is a current slice segment or is the independent slice segment that precedes a current dependent slice segment, and a slice segment header is defined to be a part of a coded slice segment containing the data elements pertaining to the first or all coding tree units represented in the slice segment. The CUs are scanned in the raster scan order of LCUs within tiles or within a picture, if tiles are not in use. Within an LCU, the CUs have a specific scan order.
[0049] Entropy coding / decoding may be performed in many ways. For example, context-based coding / decoding may be applied, where in both the encoder and the decoder modify the context state of a coding parameter based on previously coded / decoded coding parameters. Context-based coding may for example be context adaptive binary arithmetic coding (CABAC) or context-adaptive variable length coding (CAVLC) or any similar entropy coding. Entropy coding / decoding may alternatively or additionally be performed using a variable length coding scheme, such as Huffman coding / decoding or Exp-Golombcoding / decoding. Decoding of coding parameters from an entropy-coded bitstream or codewords may be referred to as parsing.
[0050] The phrase along the bitstream (e.g. indicating along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out- of-band data is associated with the bitstream. The phrase decoding along the bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, an indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
[0051] A first texture picture may be encoded into a bitstream, and the first texture picture may comprise a first projection of texture data of a first source volume of a scene model onto a first projection surface. The scene model may comprise a number of further source volumes.
[0052] In the projection, data on the position of the originating geometry primitive may also be determined, and based on this determination, a geometry picture may be formed. This may happen for example so that depth data is determined for each or some of the texture pixels of the texture picture. Depth data is formed such that the distance from the originating geometry primitive such as a point to the projection surface is determined for the pixels. Such depth data may be represented as a depth picture, and similarly to the texture picture, such geometry picture (such as a depth picture) may be encoded and decoded with a video codec. This first geometry picture may be seen to represent a mapping of the first projection surface to the first source volume, and the decoder may use this information to determine the location of geometry primitives in the model to be reconstructed. In order to determine the position of the first source volume and / or the first projection surface and / or the first projection in the scene model, there may be first geometry information encoded into or along the bitstream. It is noted that encoding a geometry (or depth) picture into or along the bitstream with the texture picture is only optional and arbitrary for example in the cases where the distance of all texture pixels to the projection surface is the same or there is no change in said distance between a plurality of texture pictures. Thus, a geometry (or depth) picture may be encoded into or along thebitstream with the texture picture, for example, only when there is a change in the distance of texture pixels to the projection surface.
[0053] An attribute picture may be defined as a picture that comprises additional information related to an associated texture picture. An attribute picture may for example comprise surface normal, opacity, or reflectance information for a texture picture. A geometry picture may be regarded as one type of an attribute picture, although a geometry picture may be treated as its own picture type, separate from an attribute picture.
[0054] Texture picture(s) and the respective geometry picture(s), if any, and the respective attribute picture(s) may have the same or different chroma format.
[0055] Terms texture (component) image and texture (component) picture may be used interchangeably. Terms geometry (component) image and geometry (component) picture may be used interchangeably. A specific type of a geometry image is a depth image. Embodiments described in relation to a geometry (component) image equally apply to a depth (component) image, and embodiments described in relation to a depth (component) image equally apply to a geometry (component) image. Terms attribute image and attribute picture may be used interchangeably. A geometry picture and / or an attribute picture may be treated as an auxiliary picture in video / image encoding and / or decoding.
[0056] Figures 2a and 2b illustrate an overview of exemplified compression / decompression processes. The processes may be applied, for example, in MPEG visual volumetric video-based coding (V3C), defined currently in ISO / IEC DIS 23090-5: “Visual Volumetric Video-based Coding and Video-based Point Cloud Compression”, 2nd Edition.
[0057] Visual volumetric video, a sequence of visual volumetric frames, if uncompressed, may be represented by a large amount of data, which can be costly in terms of storage and transmission. This has led to the need for a high coding efficiency standard for the compression of visual volumetric data.
[0058] V3C specification enables the encoding and decoding processes of a variety of volumetric media by using video and image coding technologies. This is achieved through first a conversion of such media from their corresponding 3D representation to multiple 2D representations, also referred to as V3C components, before coding such information. Such representations may include occupancy, geometry, and attribute components. The occupancy component can inform a V3C decoding and / or rendering system of whichsamples in the 2D components are associated with data in the final 3D representation. The geometry component contains information about the precise location of 3D data in space, while attribute components can provide additional properties, e.g. texture or material information, of such 3D data. An example of volumetric media conversion at an encoder is shown in Figure 2a and an example of a 3D reconstruction at a decoder is shown in Figure 2b.
[0059] Additional information that allows associating all these subcomponents and enables the inverse reconstruction, from a 2D representation back to a 3D representation is also included in a special component, referred to in this document as the atlas. An atlas consists of multiple elements, named as patches. Each patch identifies a region in all available 2D components and contains information necessary to perform the appropriate inverse projection of this region back to the 3D space. The shape of such regions is determined through a 2D bounding box associated with each patch as well as their coding order. The shape of these regions is also further refined after the consideration of the occupancy information.
[0060] Atlases are partitioned into patch packing blocks of equal size. The 2D bounding boxes of patches and their coding order determine the mapping between the blocks of the atlas image and the patch indices. Figure 3 shows an example of block-to-patch mapping with 4 projected patches onto an atlas when asps_patch_precedence_order_flag is equal to 0. Projected points are represented with dark grey. The area that does not contain any projected points is represented with light grey. Patch packing blocks are represented with dashed lines. The number inside each patch packing block represents the patch index of the patch to which it is mapped.
[0061] Axes orientations are specified for internal operations. For instance, the origin of the atlas coordinates is located on the top-left comer of the atlas frame. For the reconstruction step, an intermediate axes definition for a local 3D patch coordinate system is used. The 3D local patch coordinate system is then converted to the final target 3D coordinate system using appropriate transformation steps.
[0062] Figure 4a shows an example of a single patch packed onto an atlas image. This patch is then converted to a local 3D patch coordinate system (U, V, D) defined by the projection plane with origin O’, tangent (U), bi-tangent (V), and normal (D) axes. For anorthographic projection, the projection plane is equal to the sides of an axis-aligned 3D bounding box, as shown in Figure 4b. The location of the bounding box in the 3D model coordinate system, defined by a left-handed system with axes (X, Y, Z), can be obtained by adding offsets TilePatch3dOffsetU, TilePatch3DOffsetV, and TilePatch3DOffsetD, as illustrated in Figure 4c.
[0063] The generic mechanism of V3C may be used by applications targeting volumetric content. One of such applications is MPEG immersive video (MIV; defined in ISO / IEC 23090-12).
[0064] MIV enables volumetric video coding for applications in which a scene is recorded with multiple RGB(D) (red, green, blue, and optionally depth) cameras with overlapping fields of view (FoVs). One example setup is a linear array of cameras pointing towards a scene. This multi-scopic view of the scene allows a 3D reconstruction and therefore 6DoF / 3DoF+ consumption.
[0065] MIV uses the patch data unit concept from V3C and extends it by using camera views for reprojection.
[0066] Coded V3C video components are referred to in this document as video bitstreams, while an atlas component is referred to as the atlas bitstream. Video bitstreams and atlas bitstreams may be further split into smaller units, referred to here as video and atlas sub-bitstreams, respectively, and may be interleaved together, after the addition of appropriate delimiters, to construct a V3C bitstream.
[0067] V3C patch information is contained in atlas bitstream, atlas_sub_bistream(), which contains a sequence of NAL units. A NAL unit is specified to format data and provides header information in a manner appropriate for conveyance on a variety of communication channels or storage media. All data are contained in NAL units, each of which contains an integer number of bytes. A NAL unit specifies a generic format for use in both packet-oriented and bitstream systems. The format of NAL units for both packet- oriented transport and sample streams is identical except that in the sample stream format specified in Annex D of ISO / IEC 23090-5 each NAL unit can be preceded by an additional element that specifies the size of the NAL unit.
[0068] NAL units in atlas bitstream can be divided to atlas coding layer (ACL) and nonatlas coding layer (non-ACL) units. The former dedicated to carry patch data while thelatter to carry data necessary to properly parse the ACL units or any additional auxiliary data.
[0069] In the nal_unit_header() syntax nal unit type specifies the type of the RBSP data structure contained in the NAL unit as specified in Table 4 of ISO / IEC 23090-5. nal layer id specifies the identifier of the layer to which an ACL NAL unit belongs or the identifier of a layer to which a non-ACL NAL unit applies. The value of nal layer id shall be in the range of 0 to 62, inclusive. The value of 63 may be specified in the future by ISO / IEC. Decoders conforming to a profile specified in Annex A of ISO / IEC 23090-5 shall ignore (i.e., remove from the bitstream and discard) all NAL units with values of nal layer id not equal to 0.
[0070] Thus, the visual volumetric video-based coding (V3C; ISO / IEC DIS 23090-5) as described above specifies a generic syntax and mechanism for volumetric video coding. The generic syntax can be used by applications targeting volumetric content, such as point clouds, immersive video with depth, and mesh representations of volumetric frames. The purpose of the specification is to define how to decode and interpret the associated data (atlas data in ISO / IEC 23090-5) which tells a Tenderer how to interpret 2D frames for reconstructing volumetric frames.
[0071] The two applications of V3C (ISO / IEC 23090-5), i.e. V-PCC (ISO / IEC 23090- 5) and MIV (ISO / IEC 23090-12), use a number of V3C syntax elements with a slightly modified semantics. An example on how the generic syntax element can be differently interpreted by the application is pdu_projection_id.In case of V-PCC the syntax element specifies the index of the projection plane for the patch. There can be 6 or 18 projections planes in V-PCC, and they are implicit, i.e. pre-determined.In case of MIV pdu_projection_id corresponds to a view ID, i.e. identifies which view the patch originated from. View IDs and their related information is explicitly provided in MIV view parameters list and may be tailored for each content.
[0072] Moreover, MPEG 3DG (ISO SC29 WG7) group has started a work on a third application, i.e. V3C mesh compression. It is also envisaged that Mesh coding will re-use V3C syntax as much as possible and can also slightly modify the semantics.
[0073] As described above, a volumetric frame can be created and represented by a mix of technologies. Some type of content, materials and surfaces could benefit from one representation over another.
[0074] For example, in a 3D representation of human, the body and clothes could be represented by a mesh while hair and their fluid motion could be represented by a point cloud (i.e. particles). Another example could be a fire pit combining mesh information with point cloud coded flames and ember.
[0075] However, the current V3C does not allow to mix different representations (i.e. MIV and V-PCC) in one V3C bitstream in a way that a client would be able to decode and consume the content as one. The underlying reason for discouraging combinations of different profiles and technologies in a single V3C bitstream stems from an effort to limit the complexity of the V3C decoder and guarantee similar behaviour of implementations. For example:A content provided as mix representation would have to be split before encoding and be encoded as two separate V3C bitstreamHigher level information (e.g. scene description) would need to be provided to properly present the two decoded V3C bitstreamsEncoding as different V3C bitstreams would in practice mean increasing the amount of overlapping configuration related data as well as number of separate video bitstreamsIncreased amount of separate video bitstreams could result in significant problems for the decoder as supporting large number of parallel instances of video decoders is typically limited client systems.
[0076] Looking forward, for wider adoption of V3C technologies and to increase flexibility of V3C family of standards, it is required to find an enhanced solution to allow combining different technologies and profiles into a single bitstream.
[0077] In the following, an enhanced method for combining different V3C technologies and profiles into a single bitstream will be described in more detail, in accordance with various embodiments.
[0078] The method, which is disclosed in Figure 5, comprises obtaining (500) bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetricrepresentation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyzing (502) content and semantics of syntax elements of said at least two 3D volumetric representations; determining (504) one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and merging (506) the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
[0079] Thus, the method enables to provide signalling information that allows to differentiate V3C bitstream elements encoded in multi-format V3C bitstream based on based on the differences in the semantics of some syntax elements between different V3C formats.
[0080] An apparatus suitable for implementing the method comprises means for obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for analyzing content and semantics of syntax elements of said at least two 3D volumetric representations;means for determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and means for merging the bitstreams of said at least two 3D volumetric representations into a bitstream of a multi-format volumetric representation.
[0081] Figure 6 shows an exemplified block chart of such an apparatus. In the example of Figure 6, the apparatus obtains three different 3D volumetric representation bitstreams: a first V3C bitstream according to V3C mesh format, a second V3C bitstream according to V3C MIV format and a third V3C bitstream according to V3C V-PCC format. The bitstreams are fed into a merging unit, which analyzes the content and the semantics of the syntax elements used in the different formats of the 3D volumetric representations. The analysis reveals the differences in the semantics of the syntax elements between the three formats, and one or more signaling elements are determined to indicate said differences to a decoder. The merging unit then merges the three bitstreams of said 3D volumetric representations into a single multi-format V3C bitstream.
[0082] According to an embodiment, the apparatus comprises means for providing said bitstream of the multi-format volumetric representation to a decoder; and means for providing said one or more signaling elements, in or along said bitstream of the multiformat volumetric representation, to the decoder. The single multi-format V3C bitstream and the one or more signaling elements may be stored in memory of a buffer. However, at least in real-time use cases, the single multi-format V3C bitstream is provided to a decoder. For enabling the decoder to differentiate the semantics of the syntax element of the different formats and how to interpret the multi-format V3C bitstream, said one or more signaling elements are also provided, in or along said bitstream of the multi-format volumetric representation, to the decoder.
[0083] According to an embodiment, the apparatus comprises means for receiving volumetric frames segmented into said at least two 3D volumetric representations; means for mapping the first 3D volumetric representation to a first format-specific sub-encoder and the second 3D volumetric representation to a second format-specific sub-encoder, wherein said first and second format-specific sub-encoders are configured to encode the content of the respective 3D volumetric representation into video and atlas data according to the format.
[0084] Figure 7 shows an exemplified block chart of an apparatus according to such embodiment. In the example of Figure 7, the apparatus obtains the three different 3D volumetric representations in one or more volumetric frames with mixed content. The apparatus analyses the content at least to the extent that the three different 3D volumetric representations can be split and mapped to their format-specific sub-encoders. The example of Figure 7 discloses a first sub-encoder for V3C mesh format, a second subencoder for V3C MIV format and a third sub-encoder for V3C V-PCC format.
[0085] According to an embodiment, the apparatus comprises means for adjusting said first and second format-specific sub-encoders to use a same coordinate space. Thus, in order to enable a meaningful processing of 3D volumetric representations of different formats, it is ensured that each sub-V3C encoder share the same coordinate space.
[0086] Mixing of the formats may be performed on multiple levels of the V3C bitstream and a new signalling can be provided accordingly.
[0087] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to merge at least atlas components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0088] Thus, each format may have dedicated one or more atlases (each with own atlas component and video components) that contain a data of one type (i.e. V-PCC, MIV, Mesh etc.). The atlases representing the different formats may be mixed into a single V3C bitstream, i.e., the multi-format volumetric representation.
[0089] According to an embodiment, a signaling of differences in the semantics of the syntax elements of the atlas components of the first format and the second format is configured to be carried out by at least one syntax element included in a visual volumetric video-based coding (V3C) parameter set extension or V3C common atlas sequence parameter set extension data syntax structure.
[0090] For example, to differentiate the atlases, the merging unit may set ptl_profile_toolset_idc to indicate a multi-format format and number of ptl_syb_profile_idc to reflect the constraints of each separate format. Moreover, new signalling information would be added to V3C extension that indicates the type / format of a content of a given atlas. Such extension could as well provide a priority of a given format.
[0091] Table 1 shows an example of including syntax elements of the new signalling information into VPS extension data syntax structure.Table 1. (An ISO / IEC 23090-5 example)
[0092] In the above syntax, vps multi format present flag equal to 1 specifies that the vps_multi_format_extension( ) syntax structure is present in the v3c_parameter_set( ) syntax structure. vps_multi_format_present flag equal to 0 specifies that this syntax structure is not present. When not present, the value of vps_multi_format_present flag is inferred to be equal to 0.Table 2. (An ISO / IEC 23090-5 example)
[0093] In Table 2, vmf_content_format_id[ atlasID ] indicates format of content in the atlas with atlas ID. Error! Reference source not found, describes the list of defined format ids and their relationship with vmf_content_format_id[ atlasID ].
[0094] vmf_content_format_id[ atlasID ] equal to 0 indicates atlas contains multiple content formats and the information on how to interpret the content is provided on the atlas level signalling (e.g. Atlas Frame Parameter Set (AFPS))
[0095] vmf_prioroity_level[ atlasID ] indicates the degradation priority of a format.Table 3.
[0096] In one embodiment, a vmf content format id may instead be used to signal predefined values. Therein, the index of a sub-profile ptl_sub_profile_idc syntax element could be linked in for loop of profile_tier_level( ) syntax structure.
[0097] In one embodiment, the extension presented in the above embodiments may be alternatively recorded in Common Atlas Sequence Parameter Set (CASPS).
[0098] In one embodiment to differentiate the atlases, a merging unit may link the V3C units containing the atlases of one format to one unique V3C parameter set id, rewrite a ptl_profile_toolset_idc to indicate the multiple formats and set a ptl_sub_profile_idc to indicate the format of the content that link to this parameter set. An example of input V3C bitstreams and multi-format V3C bitstream is presented in Figure 8.
[0099] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least video components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0100] Hence, the new signalling and merging may be performed on a lower level, and the mixed bitstreams may be generated, besides on atlas component, but also on video component, i.e. on tile level. Accordingly, multiple component bitstreams may be merged into single multi-format bitstreams, whereupon the decoding process is simplified, since all geometry video components could be carried in the same video frame.
[0101] Thus, in this approach, the mixing may be performed on tile level, wherein formats may share atlas data and video component. For example, an atlas may contain two tiles, each tile containing patches describing data of a different format. This would allowmerging, for example, attribute tiles containing point cloud data with mesh attribute tiles in a single video frame.
[0102] Figure 9 shows an exemplified block chart of an apparatus according to such embodiment. In the example of Figure 9, the merging unit is configured to extract the components of the V3C bitstreams and perform the merging on the component level separately for each V3C component, whereas the parameter sets are processed in their common process. Each of the multi-format V3C components and the parameter sets are then combined into a single bitstream of the multi-format volumetric representation.
[0103] According to an embodiment, a signaling of differences in the semantics of the syntax elements of the video components of the first format and the second format is configured to be carried out by at least one syntax element included in an atlas frame parameter set extension data syntax structure.
[0104] Table 4 shows an example of including syntax elements of the new signalling information into AFPS data syntax structure.Table 4. (An ISO / IEC 23090-5 example)
[0105] In the above syntax, afps multi format extension present flag equal to 1 specifies that the afps_multi_format_extension( ) syntax structure is present in the atlas_frame_parameter_set( ) syntax structure. afps_multi_format_extension_present_flag equal to 0 specifies that this syntax structure is not present. When not present, the value of afps_multi_format_extension_present_flag is inferred to be equal to 0.
[0106] afps_tile_id[ k ] specifies the tile ID of the k-th tile.Table 5. (An ISO / IEC 23090-5 example)
[0107] amf_content_format_id[ tilelD ] indicates format of content in the tile with tileID. Table 5 describes the list of defined format ids and their relationship with vmf_content_format_id[ atlasID ].
[0108] vmf_content_format_id[ atlasID ] equal to 0 indicates that an atlas contains multiple content formats and the information on how to interpret the content is provided on the atlas level signalling (e.g. AFPS).Table 6.
[0109] amf_priority_level[ tilelD ] indicates the degradation priority of a format.
[0110] In one embodiment, an amf content format id may instead be used to signal pre-defined values. Therein, the index of a sub-profile ptl_sub_profile_idc syntax element could be linked in for loop of profile_tier_level( ) syntax structure.
[0111] According to an embodiment, said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least patch data of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0112] Hence, the signalling of multiple formats on atlas level or on tile level may be extended to a further lower level such that a similar functionality may be performed on patch level. Linking patches explicitly to given format increases the flexibility of the implementation significantly, since it would allow packing of patches representing different formats inside the same atlas and even the same tile group.
[0113] Another aspect relates to the operation of a decoder. Figure 10 shows an example of a decoding method comprising receiving (1000) a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receiving (1002), either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; separating (1004), from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and decoding (1006) the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
[0114] Thus, the decoder receives the multi-format composite V3C bitstream, as well as the signaling elements indicating the differences in the semantics of the syntax elements between the V3C formats used in the multi-format composite V3C bitstream. The decoder splits the 3D volumetric representations of different formats to their format-specific sub-decoders, which utilize the information about the differences in the semantics of the syntax elements between the V3C formats so as to render the content of their format-specific 3D volumetric representations correctly.
[0115] Consequently, the embodiments as described herein enable to combine different V3C technologies and profiles into a single bitstream, thereby fostering the encoding where some type of content, materials and surfaces benefit from encoding with mixed V3C formats. The embodiments may facilitate wider adoption of V3C technologies and increase flexibility of V3C family of standards. The signaling of the information about the differences in the semantics of the syntax elements between the V3C formats enable the decoder to differentiate the semantics of the syntax element and interpret the multi-format composite V3C bitstream correctly.
[0116] The embodiments relating to the encoding aspects may be implemented in an apparatus comprising: means for obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for analyzing content and semantics of syntax elements of said at least two 3D volumetric representations; means for determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and means for merging the bitstreams of said at least two 3D volumetric representations into a bitstream of a multi-format volumetric representation.
[0117] The embodiments relating to the encoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyze content and semantics of syntax elements of said at least two 3D volumetric representations; determine one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format andthe second format; and merge the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
[0118] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to provide said bitstream of the multi-format volumetric representation to a decoder; and provide said one or more signaling elements, in or along said bitstream of the multi-format volumetric representation to the decoder.
[0119] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to receive volumetric frames segmented into said at least two 3D volumetric representations; and map the first 3D volumetric representation to a first format-specific sub-encoder and the second 3D volumetric representation to a second format-specific sub-encoder, wherein said first and second format-specific sub-encoders are configured to encode the content of the respective 3D volumetric representation into video and atlas data according to the format.
[0120] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to adjust said first and second format-specific subencoders to use a same coordinate space.
[0121] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to merge at least atlas components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0122] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to carry out a signaling of differences in the semantics of the syntax elements of the atlas components of the first format and the second format by at least one syntax element included in a visual volumetric video-based coding (V3C) parameter set extension or V3C common atlas sequence parameter set extension data syntax structure.
[0123] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to further merge at least video components of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0124] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to carry out a signaling of differences in the semantics ofthe syntax elements of the video components of the first format and the second by at least one syntax element included in an atlas frame parameter set extension data syntax structure.
[0125] According to an embodiment, the apparatus comprises computer code configured to cause the apparatus to further merge at least patch data of the first format and the second format into the bitstream of the multi-format volumetric representation.
[0126] According to an embodiment, said first and second formats are one of the following: V3C MPEG Immersive Video (MIV) format, V3C Video-based Point Cloud Compression (V-PCC) format, V3C mesh format.
[0127] The embodiments relating to the decoding aspects may be implemented in an apparatus comprising means for receiving a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for receiving, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; means for separating, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and means for decoding the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
[0128] The embodiments relating to the decoding aspects may likewise be implemented in an apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receive, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the firstformat and the second format; separate, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and decode the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
[0129] Such apparatuses may comprise e.g. the functional units disclosed in any of the Figures la, lb, 2a and 2b for implementing the embodiments.
[0130] In the above, some embodiments have been described with reference to encoding. It needs to be understood that said encoding may comprise one or more of the following: encoding source image data into a bitstream, encapsulating the encoded bitstream in a container file and / or in packet(s) or stream(s) of a communication protocol, and announcing or describing the bitstream in a content description, such as the Media Presentation Description (MPD) of ISO / IEC 23009-1 (known as MPEG-DASH) or the IETF Session Description Protocol (SDP). Similarly, some embodiments have been described with reference to decoding. It needs to be understood that said decoding may comprise one or more of the following: decoding image data from a bitstream, decapsulating the bitstream from a container file and / or from packet(s) or stream(s) of a communication protocol, and parsing a content description of the bitstream,
[0131] In the above, where the example embodiments have been described with reference to an encoder or an encoding method, it needs to be understood that the resulting bitstream and the decoder or the decoding method may have corresponding elements in them. Likewise, where the example embodiments have been described with reference to a decoder, it needs to be understood that the encoder may have structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0132] In general, the various embodiments of the invention may be implemented in hardware or special purpose circuits or any combination thereof. While various aspects of the invention may be illustrated and described as block diagrams or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples,hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0133] Embodiments of the inventions may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.
[0134] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well established rules of design as well as libraries of pre stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) may be transmitted to a semiconductor fabrication facility or "fab" for fabrication.
[0135] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiment of this invention. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended examples. However, all such and similar modifications of the teachings of this invention will still fall within the scope of this invention.
Claims
CLAIMS1. An apparatus comprising: means for obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for analyzing content and semantics of syntax elements of said at least two 3D volumetric representations; means for determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and means for merging the bitstreams of said at least two 3D volumetric representations into a bitstream of a multi-format volumetric representation.
2. The apparatus according to claim 1, comprising: means for providing said bitstream of the multi-format volumetric representation to a decoder; and means for providing said one or more signaling elements, in or along said bitstream of the multi-format volumetric representation to the decoder.
3. The apparatus according to claim 1 or 2, comprising: means for receiving volumetric frames segmented into said at least two 3D volumetric representations; and means for mapping the first 3D volumetric representation to a first formatspecific sub-encoder and the second 3D volumetric representation to a second formatspecific sub-encoder, wherein said first and second format-specific sub-encoders are configured to encode the content of the respective 3D volumetric representation into video and atlas data according to the format.
4. The apparatus according to claim 3, comprising:means for adjusting said first and second format-specific sub-encoders to use a same coordinate space.
5. The apparatus according to any of claims 1 - 4, wherein said means for merging the bitstreams of said at least two 3D volumetric representations is configured to merge at least atlas components of the first format and the second format into the bitstream of the multi-format volumetric representation.
6. The apparatus according claim 5, wherein a signaling of differences in the semantics of the syntax elements of the atlas components of the first format and the second format is configured to be carried out by at least one syntax element included in a visual volumetric video-based coding (V3C) parameter set extension or V3C common atlas sequence parameter set extension data syntax structure.
7. The apparatus according to claim 5 or 6, wherein said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least video components of the first format and the second format into the bitstream of the multi-format volumetric representation.
8. The apparatus according claim 7, wherein a signaling of differences in the semantics of the syntax elements of the video components of the first format and the second format is configured to be carried out by at least one syntax element included in an atlas frame parameter set extension data syntax structure.
9. The apparatus according to claim 7 or 8, wherein said means for merging the bitstreams of said at least two 3D volumetric representations is configured to further merge at least patch data of the first format and the second format into the bitstream of the multi-format volumetric representation.
10. The apparatus according to any preceding claim, whereinsaid first and second formats are one of the following: V3C MPEG Immersive Video (MIV) format, V3C Video-based Point Cloud Compression (V-PCC) format, or V3C mesh format.
11. A method comprising: obtaining bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyzing content and semantics of syntax elements of said at least two 3D volumetric representations; determining one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and merging the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
12. An apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: obtain bitstreams of at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; analyze content and semantics of syntax elements of said at least two 3D volumetric representations; determine one or more signaling elements, based on the analysis, to indicate differences in the semantics of the syntax elements between the first format and the second format; and merge the bitstreams of said at least two 3D volumetric representations into a multi-format volumetric representation.
13. A method comprising: receiving a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receiving, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; separating, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and decoding the encoded first 3D volumetric representation with the first formatspecific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
14. An apparatus comprising at least one processor and at least one memory, said at least one memory stored with computer program code thereon, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: receive a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; receive, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; separate, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; anddecode the encoded first 3D volumetric representation with the first formatspecific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
15. An apparatus comprising: means for receiving a bitstream in a decoder, said bitstream comprising at least two 3D volumetric representations, wherein a first 3D volumetric representation is encoded according to a first format and a second 3D volumetric representation is encoded according to a second format; means for receiving, either in said bitstream or in a further bitstream, one or more signaling elements indicating differences in semantics of syntax elements between the first format and the second format; means for separating, from said bitstream, the encoded first 3D volumetric representation to a first format-specific sub-decoder and the encoded second 3D volumetric representation to a second format-specific sub-decoder; and means for decoding the encoded first 3D volumetric representation with the first format-specific sub-decoder and the encoded second 3D volumetric representation with the second format-specific sub-decoder at least partly based on said one or more syntax elements.
Citation Information
Patent Citations
Unified coding of 3D objects and scenes
US20210092345A1