Return of coded base mesh and displacement data for video-based dynamic mesh coding in ISOBMFF media containers

KR1020260140253APending Publication Date: 2026-09-22INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020267026647
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-17
Publication Date
2026-09-22

Smart Images

  • Figure P1020267026647_ABST
    Figure P1020267026647_ABST
Patent Text Reader

Abstract

Examples of systems and methods discussed herein provide an encapsulation scheme that may include a base mesh bitstream, sub-mesh(s), and / or displacement information, which provides efficient storage and transmission of content in an extensible container such as an ISO Base Media File Format (ISOBMFF) container. In particular, examples of systems and methods discussed herein provide a general-purpose and scalable design that supports the transport of sub-components of a Video-Based Dynamic Mesh Coding (V-DMC) or similar bitstream, such as a base mesh, sub-mesh(s), and displacement information, in an ISOBMFF media container.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference regarding related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 622,977 filed January 19, 2024, the contents of which are incorporated herein by reference.

[0003] The systems, devices, and methods generally relate to video encoding and transmission. In particular, the systems, devices, and methods relate to the transport of a coded base mesh and displacement data for video-based dynamic mesh coding. Background Technology

[0004] New 3D video encoding standards include Video-based Dynamic Mesh Coding (V-DMC), published as ISO / IEC 23090-29 by Working Group 7 of the Motion Picture Experts Group (MPEG), which is an extension of the Visual Volumetric Video-based Coding (V3C) specification (ISO / IEC 23090-5). Communicating V-DMC and / or V3C data between encoders and decoders or other such devices presents problems due to the large bandwidth and throughput typically required, as well as the potential for data diversity and different data subtypes (e.g., 2D data, 3D data, mesh data, attribute data, etc.). Existing file types may be unsuitable for describing these various forms of data and / or may be bloated and inefficient.

[0005] In some embodiments, the present disclosure relates to examples of systems and methods for extensible and scalable encapsulation of mesh and sub-mesh information for V-DMC or V3C video data in an extensible container, such as an ISO Base Media File Format (ISOBMFF) container. In particular, the examples of systems and methods discussed herein provide a general-purpose and scalable design that supports the carrying of sub-components of a V-DMC or similar bitstream, such as a base mesh, sub-mesh, and displacement information, in an ISOBMFF media container. Brief explanation of the drawing

[0006] A more detailed understanding can be obtained from the following description, which is provided as an example in relation to the attached drawings, and identical reference numbers within the drawings indicate identical elements. FIG. 1 is a schematic diagram of an implementation of an ISOBMFF media container for V-DMC bitstreams according to some examples. FIG. 2 is a schematic diagram of another implementation of an ISOBMFF media container for V-DMC bitstreams according to some examples. FIG. 3 illustrates a block diagram of a system in which the embodiments of the examples can be implemented. Figure 4 illustrates a block diagram of an implementation of a video-based point cloud compression (V-PCC) encoder. Figure 5 illustrates a block diagram of the implementation of a V-PCC decoder. Specific details for implementing the invention

[0007] The following ISO / IEC specification(s) and standard(s) (including any draft versions of such standard(s)) are incorporated herein by reference in their entirety and, for all purposes, become part of the disclosure: ISO / IEC 23090-5 "Information Technology -- Coded Representation of Immersive Media -- Part 5: Visual Volume Video-Based Coding (V3C) and Video-Based Point Cloud Compression (V-PCC)" ISO / IEC 23090-10, "Information Technology -- Coded Representation of Immersive Media -- Part 10: Return of Visual Volume Video-Based Coding Data"; ISO / IEC 23090-29 "Information Technology -- Coded Representation of Immersive Media -- Part 29: Video-Based Dynamic Mesh Coding (V-DMC); and ISO / IEC 14496-12, "Information Technology -- Coding of Audiovisual Objects -- Part 12: ISO Base Media File Format". Although the present disclosure may refer to aspects of these standards(s), the present disclosure is by no means limited by these standards(s).

[0008] Video-based dynamic mesh coding (V-DMC), which extends the Visual Volume Video-based Coding (V3C) specification, provides encoding, communication, and decoding of three-dimensional video bitstreams. These bitstreams may include subcomponents containing sequences of dynamic meshes that can represent entities such as objects, environmental features, people, animals, machines, etc. These dynamic meshes may be represented by a number of V3C components including a base mesh, a set of displacements, a 2D representation of attributes, and an atlas.

[0009] In some examples, the base mesh component may contain a simplified, low-resolution approximation of the original mesh. The base mesh component may be encoded using any mesh codec. For instance, in some examples, a static mesh codec may be used to code static mesh instances, and a motion decoder may be used to temporally transform the static mesh over a set period (e.g., when an entity moves over time). In some examples, the base mesh codec may delegate the compression process to the static mesh codec and the motion codec. In many examples, the V3C codec may be codec-agnostic with respect to the base mesh, meaning that any static mesh / motion codec may be used to code the base mesh bitstream.

[0010] In some examples, the displacement component provides displacement vectors that can be encoded as a V3C geometry video component using any suitable video codec. In some examples, the profile may indicate that the displacement component is encoded using arithmetic coding, video coding, or any other suitable type of coding.

[0011] In some examples, the atlas component provides the V3C decoding and / or rendering system with information on how to perform inverse reconstruction. For example, the atlas component can specify how to subdivide the base mesh, how to apply displacement vectors to the subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.

[0012] In some examples, a submesh can be represented by a set of vertices, their connectivity, and associated mesh properties. In many examples, submeshes can be decoded completely independently of each other. Each base mesh may contain one or more submeshes, each of which may be associated with a unique identifier within the base mesh subbitstream frame parameter set. Similarly, in many examples, each patch within an atlas tile may be associated with a corresponding submesh identifier.

[0013] Communicating V-DMC and / or V3C data between encoders and decoders or other such devices presents problems due to the diversity of data and the potential for different data subtypes (e.g., 2D data, 3D data, mesh data, attribute data, etc.), as well as the typically required large bandwidth and throughput. Existing file types may be unsuitable for describing these various forms of data and may lack scalability or extensibility.

[0014] This description provides examples of encapsulation schemes that include a base mesh bitstream and expose its high-level features, such as sub-meshes, to provide efficient storage and transmission of base mesh content. Examples of encapsulation schemes may also include arithmetic-coded displacement information, which can satisfy the requirements of V-DMC profiles where displacement information is arithmetic-coded rather than video-based. By using flexible encapsulation, access to sub-meshes within the bitstream is enabled, and the bitstream can be multiplexed together with displacement bitstreams, atlases, and base mesh bitstreams in an extensible container, such as an ISO Base Media File Format (ISOBMFF) container. In particular, examples of systems and methods discussed herein provide a general-purpose and scalable design that supports the transport of V-DMC or similar bitstream subcomponents, such as base meshes, sub-meshes, and displacement information, in an ISOBMFF media container.

[0015] For clarity, sub-mesh information is part of the base mesh substream. Subsequently, the substreams are encapsulated by structuring them in a media container as separate tracks (in the case of timed media) (on the encoder / packager side). Generally, descriptions of the base mesh and displacement substreams are provided. In the case of a base mesh substream containing multiple sub-meshes, methods are provided for storing these sub-meshes in individual sub-mesh tracks, which enable the player to select only the tracks for the sub-meshes it needs or wants.

[0016] A system and method for communicating mesh and displacement data are provided. The system and method comprises the steps of: receiving a plurality of substreams of items of media content by an encoder of a device — said plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream —; generating a track reference by the encoder that identifies an association between the base mesh substream and the one or more sub-mesh substreams — said track reference is encapsulated within a header of a first substream among the plurality of substreams —; and generating the plurality of substreams into a container file format by the encoder, wherein the header of the first substream generated into the container file format includes a reference identifier for each of the one or more video substreams, the base mesh substream, the one or more sub-mesh substreams, and the displacement substream. The container file format may be the International Standards Organization Base Media File Format (ISOBMFF). The base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be data timed in synchronization with the atlas substream and one or more video substreams. The base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be untimed data.At least one of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be data timed in synchronization with the atlas substream and one or more video substreams, and at least one of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be untimed data. The substreams may correspond to tracks of the generated container file format. The header of the generated base mesh substream includes an identifier for each sub-mesh substream. The base mesh substream may identify multiple base meshes of an item of the media content, and each base mesh is associated with a different time period within the item of the media content. The one or more generated sub-mesh substreams may be identified as members of a track group in the header of the atlas substream.

[0017] A device for communicating mesh and displacement data is also disclosed. The device includes an encoder; and an input / output (I / O) device operatively connected to the encoder. The device may be configured to receive a plurality of substreams of items of media content, wherein the plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream. The device may be further configured to generate a track reference identifying an association between the base mesh substream and the one or more sub-mesh substreams, wherein the track reference is encapsulated in the header of the first substream among the plurality of substreams. The device may be further configured to generate the plurality of substreams in a container file format, wherein the header of the first substream generated in the container file format includes a reference identifier for each of the one or more video substreams, the base mesh substream, the one or more sub-mesh substreams, and the displacement substream. The container file format may be the International Organization for Standardization Base Media File Format (ISOBMFF). The base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be data timed in synchronization with the atlas substream and one or more video substreams. The base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be untimed data.At least one of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be data timed in synchronization with the atlas substream and one or more video substreams, and at least another of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream may be untimed data. The substreams may correspond to tracks of the generated container file format. The header of the generated base mesh substream includes an identifier for each sub-mesh substream. The header of a single sub-mesh track includes an identifier for each sub-mesh substream. The base mesh substream may identify multiple base meshes of an item of the media content, and each base mesh is associated with a different time period within the item of the media content. The generated one or more sub-mesh substreams may be identified as members of a track group in the header of the atlas substream.

[0018] Additional methods for communicating mesh and displacement data are also described. The method comprises the steps of: receiving information including a container file format by a decoder of a device, wherein the container format includes a plurality of substreams of items of media content, and the plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream; decoding the information including the container file format including the plurality of substreams by the decoder; parsing the plurality of substreams from the container file format, wherein the header of the first substream generated in the container file format includes reference identifiers for each of the one or more video substreams, the base mesh substream, the one or more sub-mesh substreams, and the displacement substream; parsing a track reference encapsulated in the header of the first substream among the plurality of substreams to identify an association between the base mesh substream and the one or more sub-mesh substreams; and rendering the media content using the track reference. Rendering of the above media content can be performed on an external display.

[0019] Within the ISO / IEC 14496 (MPEG-4) standard, there are several parts defining file formats for storing time-based media. These parts are all based on and derived from the ISO Base Media File Format (ISOBMFF), which is a structural and media-independent definition. FIG. 1 illustrates a schematic diagram of an implementation of an ISOBMFF media container (100) for timed data of V-DMC bitstreams. ISOBMFF (100) contains structure and media data information for timed presentations of media data, such as audio and video. There is also support for untimed data, such as metadata at different levels within the file structure. The logical structure of the file is a logical structure of a movie containing a set of time-parallel tracks (110) (e.g., Track 1 (1101), Track 2 (1102), Track 3 (1103), collectively referred to as Tracks (110)). The temporal structure of the file is that tracks (110) contain sequences of samples at time, and these sequences are mapped to the timeline of the entire movie. ISOBMFF (100) is based on the concept of box structure files. A box structure file consists of a series of boxes (sometimes referred to as atoms) having a size and a type. The types are 32-bit values ​​and are typically selected from four printable characters, also known as 4-character codes (4CC) (e.g., in the illustrated implementation, "v3vg" for the V-DMC geometry bitstream, "v3va" for the V-DMC attribute bitstream). Untimed data may be included in a metadata box at the file level, in a movie box, or attached to one of the streams of timed data, called tracks, within the movie.

[0020] Among the top-level boxes within the ISOBMFF container (100) is a MovieBox ('moov') containing metadata for consecutive media streams existing in a file. This metadata is signaled within the hierarchy of boxes within the MovieBox, for example, within a TrackBox ('trak'). A track represents a consecutive media stream existing in a file. The media stream itself consists of a sequence of samples, such as audio or video access units of the underlying media stream, and is contained within a MediaDataBox ('mdat') existing at the top level of the container (100). The metadata for each track includes a list of sample description entries, each providing the coding or encapsulation format used in the track and initialization data for processing that format. Each sample may be associated with one of the sample description entries of the track. ISO / IEC 14496-12 provides a tool for defining a clear timeline map for each track. This is known as an edit list and is signaled using an EditListBox with the following syntax, where each entry defines a part of the track timeline: by mapping a part of the composition timeline, or by indicating 'empty' time (parts of the presentation timeline not mapped to any media, 'empty' edits):

[0021]

[0022] The ISO / IEC 23090-10 specification also includes extensions to the ISO / IEC 14496-12 specification that enable the encapsulation (return) of V3C bitstreams within ISOBMFF media containers.

[0023] FIG. 2 is a schematic diagram of another implementation of an ISOBMFF media container (200) for V-DMC bitstreams having additional tracks (210) defined for a base mesh bitstream, a sub-mesh bitstream, and a displacement bitstream according to some examples (e.g., track 1 (2101), track 2 (2102), track 3 (2103), track 4 (2104), track 5 (2105), track 6 (2106), collectively referred to as tracks (210)). Although each of the base mesh (2104), sub-mesh (2105), and displacement bitstream (2106) is illustrated as one, in some examples, there may be more or fewer of them. For example, in one implementation, the media container (200) may include a base mesh bitstream and a plurality of sub-mesh bitstreams (e.g., for different sub-meshs of the base mesh).

[0024] In some examples, a base mesh track (2104) is included to carry data of a base mesh sub-bitstream of a V-DMC bitstream. The base mesh track (2104) may have a sample entry of type BaseMeshSampleEntry derived from VolumetricVisualSampleEntry and has an ISOBMFF box type identified by a 4-character code (4CC) such as 'vbm1', 'vdmb', or 'vbmg' or a similar identifier. In some examples, BaseMeshSampleEntry may include a BaseMeshConfigurationBox containing configuration information necessary to initialize a base mesh decoder. The decoder initialization information may be signaled in a BaseMeshDecoderConfigurationRecord structure within the BaseMeshConfigurationBox.

[0025] In some examples, BaseMeshSampleEntry and BaseMeshConfigurationBox can be defined as follows:

[0026]

[0027] The semantics of the fields in BaseMeshConfigurationBox are as follows:

[0028] unit_size_precision_bytes_minus1 / plus1 specifies the precision in bytes of the sample stream Network Abstraction Layer (NAL) unit to which this configuration record applies. The value of this field may depend on the 4CC-code of the sample entry. For example, in some examples, for base mesh tracks, unit_size_precision_bytes_minus1 may be equivalent to or equal to ssnh_unit_size_precision_bytes_minus1 in sample_stream_nal_header().

[0029] num_of_setup_unit_arrays indicates the number of arrays of base mesh NAL units of the specified type(s).

[0030] array_completeness, when equivalent to 1, may indicate that all base mesh NAL units of the given type are in the following array and there is nothing in the stream; and when equivalent to 0, may indicate that additional base mesh NAL units of the specified type may be in the stream. Default and allowed values ​​may be restricted by the sample entry name.

[0031] nal_unit_type indicates the type of base mesh NAL units in the following array, which may all be of that type or, in some examples, a mixture of types (in which case there may be multiple nal_unit_type entries or the types may be predefined as a mixture). The types may be encoded as predetermined values, such as those defined in ISO / IEC 23090-29. In some examples, the field may be restricted to taking one of the values ​​indicating NAL_BMSPS, NAL_BMFPS, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI base mesh NAL units.

[0032] num_nal_units indicates the number of base mesh NAL units of type nal_unit_type included in the configuration record for the stream to which this configuration record applies.

[0033] setup_unit_length specifies the size of the setup_unit field in bytes. The length field includes the sizes of both the NAL unit header and the NAL unit payload, but in many examples, it does not include the length of the length field itself.

[0034] The setup_unit may contain NAL units corresponding to the associated nal_unit_type. When present in the setup_unit, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI may contain SEI (supplemental enhancement information) messages of a 'declarative' nature, that is, those providing information about the entire stream. An example of such an SEI may be a user-data SEI.

[0035] In some examples, the meanings of the fields inherited from VolumetricVisualSampleEntry are as follows.

[0036] compressorname within the base class VolumetricVisualSampleEntry indicates the name of the compressor used and may include padding or length indicators. For example, given the value "\012VBM coding", the first byte is the count of the remaining bytes, represented here as \012, which is 10 (decimal) and is the number of bytes in the remainder of the string.

[0037] In some examples, the base mesh (2104) includes one or more sub-meshes (2105) that can be decoded independently. When the base mesh subbitstream (2105) includes two or more sub-meshes, data for each sub-mesh may be carried out on separate tracks within the ISOBMFF container (200). For example, FIG. 2 illustrates only one sub-mesh track (2105), but in some examples, multiple sub-mesh tracks may be included (e.g., one for each sub-mesh). This is particularly useful when the spatial partitioning of the base mesh into sub-meshes does not change during the duration of the dynamic mesh sequence. These tracks may be referred to as sub-mesh tracks and may include a SubMeshSampleEntry having 4CC 'vbs1' (and, in some examples, 'vbs2' for the second sub-mesh track, 'vbs3' for the third sub-mesh track, etc.—although other identifiers may be used). The sub-mesh track (2105) returns sub-mesh data for one or more sub-meshes within the base mesh sub-bitstream. Identifiers for all sub-meshes returned by the sub-mesh track (e.g., when multiple sub-meshes are used or when sub-meshes change over the course of a dynamic mesh sequence) can be signaled in the SubMeshConfigurationBox within the SubMeshSampleEntry.

[0038] In some examples, a base mesh track (2104) containing decoder configuration information necessary to instantiate and initialize a base mesh decoder can be linked to associated sub-mesh tracks (2105) using a track reference type such as 'vlvb'. In some examples, a TrackReferenceTypeBox having the reference type 'vlvb' can be added to the TrackReferenceBox within the TrackBox of the base mesh track (2104). The TrackReferenceTypeBox may contain an array of track_IDs or a string containing identifiers for the referenced sub-mesh tracks. In other embodiments, another track reference may be added from the sub-mesh track (2105) to the base mesh track (2104).

[0039] In some examples, SubMeshSampleEntry can be defined as follows:

[0040]

[0041] In some examples, the meanings of the fields of SubMeshConfigurationBox can be defined as follows:

[0042] unit_size_precision_bytes_minus1 / plus1 can specify the precision in bytes of the sample stream NAL units to which the sample entry containing this configuration box is applied. The value of this field can be equivalent to ssnh_unit_size_precision_bytes_minus1 in sample_stream_nal_header() for the base mesh component bitstream.

[0043] num_submeshes can indicate the number of submeshes returned from this track.

[0044] submesh_id can specify the submesh ID of a submesh existing on the track. The value of submesh_id can be equivalent to the value of the corresponding bmsi_submesh_id syntax element within bmesh_sub_mesh_information() defined in ISO / IEC 23090-29.

[0045] compressorname within the base class VolumetricVisualSampleEntry may specify the name of the compressor used (along with padding or other specifiers if necessary). For example, given the value "\014VBM submesh", the first byte is the count of the remaining bytes, represented here as \014, which is 12 (decimal) (14 in octal) and is the number of bytes in the remainder of the string.

[0046] A sync sample within a base mesh track (2104) or a sub-mesh track (2105) is a sample containing an intra-random access point (IRAP) coded base mesh access unit as defined in ISO / IEC 23090-29. Base mesh parameter sets and SEI messages may be repeated in the sync sample to allow random access if necessary.

[0047] A V-DMC decoder typically requires information from an atlas subbitstream (2101) to reconstruct a base mesh. To efficiently access and decode specific sub-meshes from the base mesh subbitstream (2104), in some examples, atlas information (2101) regarding these sub-meshes may be placed on independent atlas tiles. To signal the association between a sub-mesh track (2105) and a V3C atlas tile track containing atlas information (2101) for the corresponding sub-mesh, a new track group VDMCSubMeshTrackGroupBox extending the TrackGroupTypeBox defined in ISO / IEC 14496-12 may be defined as follows:

[0048]

[0049] The meanings of the fields in VDMCSubMeshTrackGroupBox can be as follows:

[0050] submesh_id may contain an identifier for a submesh for this track group instance. The value of submesh_id may be equivalent to the value of the corresponding bmsi_submesh_id syntax element within bmesh_sub_mesh_information() defined in ISO / IEC 23090-29.

[0051] num_tiles can indicate the number of atlas tiles associated with a sub-mesh for this track group instance.

[0052] tile_id can specify the atlas tile ID of the atlas tile that returns patches associated with the submesh for this track group instance.

[0053] In another embodiment, VDMCSubMeshGroupBox may represent an entity group containing each group of V-DMC tracks belonging to the same sub-mesh. The definition of VDMCSubMeshGroupBox is as follows:

[0054]

[0055] The meanings of the fields in VDMCSubMeshGroupBox can be as follows:

[0056] submesh_id may contain an identifier for a submesh for this entity group instance. The value of submesh_id may be equivalent to the value of the corresponding bmsi_submesh_id syntax element within bmesh_sub_mesh_information() defined in ISO / IEC 23090-29.

[0057] num_tiles can indicate the number of atlas tiles associated with a submesh for this entity group instance.

[0058] tile_id can specify the atlas tile ID of the atlas tile that returns patches associated with the submesh for this entity group instance.

[0059] The V-DMC displacement track (2106) can return data from the V-DMC displacement subbitstream. If the V-DMC displacement component is coded using a traditional 2D video codec, the displacement track (2106) may include a VisualSampleEntry with 4CC 'resv' and a restricted video track including a RestrictedSchemeInfoBox.

[0060] In some examples, arithmetic-coded displacement tracks (2106) may be returned as restricted volume visual tracks. These tracks use a general-purpose restricted VolumetricVisualSampleEntry having 4CC 'res3' containing RestrictedSchemeInfoBox (e.g., as illustrated in the implementation of FIG. 2).

[0061] In many examples, the following additional requirements apply to both video-coded and arithmetic-coded displacement tracks (2106):

[0062] A SchemeTypeBox exists in the RestrictedSchemeInfoBox, and the scheme_type of that box is set to 'vvbm'; and

[0063] schemeInformationBox exists in RestrictedSchemeInfoBox and includes V3CUnitHeaderBox.

[0064] In the track header of the displacement track (2106), the track_in_movie flag may be set to 0 to indicate that this track should not be presented alone.

[0065] In some examples, the displacement track sample entry may also include a VDMCDisplacementConfigurationBox defined as follows:

[0066]

[0067] The meanings of the fields in VDMCDisplacementConfigurationBox can be as follows:

[0068] unit_size_precition_bytes_minus1 + 1 specifies the precision in bytes of the sample stream NAL unit to which this configuration record applies. The value of this field will depend on the 4CC-code of the sample entry. For bass mesh tracks, unit_size_precision_bytes_minus1 will be equivalent to ssnh_unit_size_precision_bytes_minus1 in sample_stream_nal_header().

[0069] num_of_setup_unit_arrays can indicate the number of arrays of displacement NAL units of the specified type(s).

[0070] array_completeness, when equivalent to 1, indicates that all displacement NAL units of the given type are in the following array and there is nothing in the stream; when equivalent to 0, indicates that additional displacement NAL units of the specified type may be in the stream. Default and allowed values ​​may be restricted by the sample entry name.

[0071] nal_unit_type may indicate the type of displacement NAL units in the following array, which may all be of that type or, in some examples, a mixture of types (in which case, there may be multiple nal_unit_type entries or the types may be predefined as a mixture). nal_unit_type may include a value that may be predetermined to represent the type of displacement NAL units, such as those defined in ISO / IEC 23090-29; in some examples, it may be limited to taking one of the values ​​indicating NAL_DSPS, NAL_DFPS, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI displacement NAL units.

[0072] num_nal_units can indicate the number of base mesh NAL units of type nal_unit_type included in the configuration record for the stream to which this configuration record is applied.

[0073] setup_unit_length can specify the size of the setup_unit field in bytes. The length field can include the sizes of both the NAL unit header and the NAL unit payload, but in many examples, the length field itself is not included.

[0074] The setup_unit contains NAL units corresponding to the associated nal_unit_type. When present in the setup_unit, NAL_PREFIX_ESEI, NAL_PREFIX_NSEI, NAL_SUFFIX_ESEI, or NAL_SUFFIX_NSEI contain SEI messages of a 'declarative' nature, that is, those that provide information about the entire stream. An example of such an SEI may be a user-data SEI.

[0075] Each sample within the V-DMC displacement track (2106) may correspond to a single coded displacement access unit. In some examples, the V-DMC displacement sample may be defined as follows:

[0076]

[0077] Accordingly, examples of systems and methods discussed herein provide an encapsulation scheme that may include a base mesh bitstream (2104), sub-mesh(s) (2105), and / or displacement information (2106), which provides efficient storage and transmission of content in a scalable container, such as an ISOBMFF container (200). In particular, examples of systems and methods discussed herein provide a general-purpose and scalable design for supporting the transport of sub-components of a V-DMC or similar bitstream, such as a base mesh (2104), sub-mesh(s) (2105), and displacement information (2106), in an ISOBMFF media container (200).

[0078] FIG. 3 illustrates a block diagram of an example of a system (300) in which various embodiments and examples may be implemented. The system (300) may be implemented as a device comprising various components described below and configured to perform one or more of the embodiments described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and servers. The elements of the system (300) may be implemented, alone or in combination, as a single integrated circuit, a plurality of ICs, and / or individual components. For example, in at least one example, the processing and encoder / decoder elements of the system (300) are distributed across a plurality of ICs and / or individual components. In various examples, the system (300) is communicated to other systems or other electronic devices, for example, via a communication bus or via dedicated input and / or output ports. In various examples, the system (300) is configured to implement one or more of the embodiments described in this application.

[0079] The system (300) includes, for example, at least one processor (310) configured to execute loaded instructions to implement the various embodiments described in this application. The processor (310) may include an embedded memory, an input / output interface, and various other circuits as known in the art. The system (300) includes at least one memory (320) (e.g., a volatile memory device and / or a non-volatile memory device). The system (300) includes a storage device (340) that may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive and / or optical disk drive. The storage device (340) may include, as non-limiting examples, an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0080] The system (300) includes, for example, an encoder / decoder module (330) configured to process data to provide an encoded video / 3D object or a decoded video / 3D object, and the encoder / decoder module (330) may include its own processor and memory. The encoder / decoder module (330) represents a module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, the device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module (330) may be implemented as a separate element of the system (300) or may be integrated within the processor (310) as a combination of hardware and software, as is known to those skilled in the art.

[0081] Program code to be loaded onto a processor (310) or an encoder / decoder (330) to perform the various embodiments described in this application may be stored in a storage device (340) and subsequently loaded onto a memory (320) for execution by the processor (310). According to various examples, one or more of the processor (310), memory (320), storage device (340), and encoder / decoder module (330) may store one or more of the various items during the execution of the processes described in this application. These stored items may include, but are not limited to, input video / 3D objects, decoded video / 3D objects or parts of decoded video / 3D objects, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0082] In some examples, memory within the processor (310) and / or encoder / decoder module (330) is used to store instructions and provide working memory for processing required during encoding or decoding. In other examples, memory outside the processing device (e.g., the processing device may be either the processor (310) or the encoder / decoder module (330)) is used for one or more of these functions. The external memory may be memory (320) and / or storage device (340), e.g., dynamic volatile memory and / or non-volatile flash memory. In some examples, the external non-volatile flash memory is used to store the operating system of the television. In at least one example, high-speed external dynamic volatile memory, such as RAM, is used as working memory for coding and decoding operations, such as MPEG-2, HEVC, or VVC.

[0083] Inputs to the elements of the system (300) may be provided through various input devices as indicated in block (305). These input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted wirelessly by, for example, a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0084] In various examples, the input devices of block (305) have their respective associated input processing elements as known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting it again to a narrower band of frequencies to select a signal frequency band (e.g., which may be referred to as a channel in certain examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a stream of desired data packets. The RF section of various examples includes one or more elements for performing these functions, e.g., a frequency selector, a signal selector, a band-limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, for example, including down-converting the received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or to the baseband. In one set-top box example, the RF section and its associated input processing element receive an RF signal transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering it to a desired frequency band. Various examples rearrange the order of the elements described above (and others), remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0085] Additionally, USB and / or HDMI terminals may include their own interface processors for connecting the system (300) to other electronic devices via the USB and / or HDMI connections. It should be understood that various modes of input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor (310) as needed. Similarly, modes of USB or HDMI interface processing may be implemented within separate interface ICs or within the processor (310) as needed. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor (310) and the encoder / decoder (330), which operate in combination with memory and storage elements to process the data streams as needed for presentation on the output device.

[0086] Various elements of the system (300) may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected using an internal bus as known in the art, which includes a suitable connection arrangement, e.g., an I2C bus, wiring, and printed circuit boards, and may transmit data between them.

[0087] The system (300) includes a communication interface (350) that enables communication with other devices through a communication channel (390). The communication interface (350) may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel (390). The communication interface (350) may include, but is not limited to, a modem or a network card, and the communication channel (390) may be implemented, for example, within a wired and / or wireless medium.

[0088] In various examples, data is streamed to the system (300) using a Wi-Fi network such as IEEE 802.11. In these examples, the Wi-Fi signal is received through a communication interface (350) and a communication channel (390) adapted for Wi-Fi communication. The communication channel (390) in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples provide the streamed data to the system (300) using a set-top box that transmits data via the HDMI connection of the input block (305). Still other examples provide the streamed data to the system (300) using the RF connection of the input block (305).

[0089] The system (300) may provide output signals to various output devices, including a set of audio outputs such as a display (365) and speakers (375), and other peripheral devices (385). Other peripheral devices (385) include, in various examples, one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system (300). In various examples, control signals are communicated between the system (300) and the display (365), speakers (375), or other peripheral devices (385) using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicated to the system (300) through dedicated connections via their respective interfaces, such as a display interface (360), an audio interface (370), and a peripheral device interface (380). Output devices can be connected to the system (300) using a communication channel (390) through a communication interface (350). The display (365) and speakers (375) can be integrated into a single unit with other components of the system (300) in an electronic device, for example, a television. In various examples, the display interface (360) includes a display driver, for example, a timing controller (T Con) chip.

[0090] For example, if the RF part of the input (305) is part of a separate set-top box, the display (365) and speaker (375) may alternatively be separated from one or more of the other components. In various examples where the display (365) and speaker (375) are external components, the output signal may be provided through dedicated output connections, such as an HDMI port, a USB port, or a COMP output.

[0091] In many examples of two-dimensional or three-dimensional environments, environmental details and / or objects or entities can be represented by point clouds or meshes of vertices and / or edges forming a very large number of triangles or other shapes. For example, in three-dimensional virtual reality, augmented reality, or extended reality (VR, AR, or XR) environments, virtual objects within the environment can be represented as a mesh of triangles to represent the geometry of the objects. To provide fine detail, the triangles or other shapes used in the mesh can be very small, and accordingly, the mesh has a very large number of vertices and edges, potentially millions or billions of vertices for complex objects. This can correspondingly require a large amount of memory for storage and bandwidth for transmission. When used for three-dimensional video or animation, the problems mentioned above can become much more severe: even a relatively simple environment with one million vertices can require a total bandwidth of 3.6 gigabits per second at 30 frames per second if uncompressed.

[0092] Different encoding and decoding standards and examples have been described, including Video-based Point Cloud Compression (V-PCC), Geometry-based Point Cloud Compression (GPCC), and Video-based Dynamic Mesh Coding (V-DMC) published by the Motion Picture Experts Group (MPEG); P11 Dynamic Mesh Coding developed by Apple Inc.; and others. A mesh may include one or more of the following features: a list of vertex locations; a topology defining connections between vertices, e.g., a list of faces; and optionally photometric data, such as texture maps or color values ​​associated with the vertices or faces. A 3D mesh may be derived from the point cloud of a 3D object. Faces defined by connected vertices may be triangles or any other possible polygons or combinations of polygons. In many examples, photometric data may be projected onto a texture map so that the texture map can be encoded as a video image.

[0093] FIG. 4 illustrates a block diagram of the implementation of a V-PCC encoder (400). In a brief overview, in some examples, volume 3D data or a mesh may be divided into a plurality of sub-meshes or patches. In some examples, a frame image may be separated into a set of 3D projections within different components, such as far and near components for geometry and corresponding attribute components. In some examples, a 2D occupancy map may be created to indicate parts of an image to be used as textures or projections on a mesh. The 2D projection may include a plurality of independent patches based on the geometric characteristics of the input point cloud frame. After the patches are created and 2D projection frames for video encoding are created, the occupancy map, geometry information, attribute information, and auxiliary information may be compressed and, in many examples, post-processed (e.g., smoothing, normalization, scaling, etc.). Separate bitstreams (e.g., patch substream (410), attribute substream (420), geometry substream (430) and occupancy substream (440)) can be multiplexed into an output compressed bitstream (450).

[0094] FIG. 5 illustrates a block diagram of an implementation of a V-PCC decoder (500). An incoming bitstream (510) may be demultiplexed to recover patch (520), geometry (540), attribute (550), and occupancy (530) substreams. In some examples, additional auxiliary information (560), such as a sequence parameter set (SPS) substream (560), may be included in the multiplexed bitstream. This auxiliary information stream (560) may be entropy-coded and compressed. The occupancy map may be compressed using video compression (535), and in some examples, may be upscaled to nominal resolution using any suitable technique (e.g., nearest neighbor, etc.). The geometry stream (540) may be decoded (545) and, in combination with the occupancy map and auxiliary information, smoothed or interpolated to reconstruct point cloud geometry information (590). Based on the decoded attribute video stream (550), reconstructed geometry information (540), occupancy map (530), and auxiliary information (560), the point cloud topology and its attributes (e.g., texture, color, lighting, etc.) can be reconstructed.

[0095] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "encoded" or "coded" may be used interchangeably, the terms "pixel" or "sample" may be used interchangeably, and the terms "image," "picture," and "frame" may be used interchangeably. Typically, though not necessarily, the term "reconstructed" is used on the encoder side, whereas "decoded" is used on the decoder side.

[0096] Although examples using V-PCC codecs have been discussed above, the systems and methods discussed herein are not limited to these codecs and may be used to encode and decode any suitable 2D or 3D mesh or point cloud, and accordingly, the encoder and decoder of FIGS. 4 and 5 are provided merely as examples.

[0097] Various methods are described herein, each of which comprises one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first,” “second,” etc., may be used to modify elements, components, steps, actions, etc., in various examples, such as “first decoding” and “second decoding.” The use of these terms does not imply a specific order of modified actions unless specifically required. Accordingly, in this example, the first decoding does not need to be performed before the second decoding, but may be performed, for example, before, during, or within an overlapping period with the second decoding.

[0098] Furthermore, the embodiments described herein are not limited to the V-PCC, G-PCC, P11, or V-DMC, but may be applied, for example, to other standards and recommendations, and any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the embodiments described herein may be used individually or in combination. Various numerical values ​​are used in this application. Specific values ​​are for illustrative purposes only, and the described embodiments are not limited to these specific values.

[0099] Various examples involve decoding. As used in this application, "decoding" may include all or part of the processes performed on a received encoded sequence, for example, to produce a final output suitable for display. In various examples, these processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transform, and difference decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader decoding process will be apparent based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.

[0100] Various examples involve encoding. In a manner similar to the above discussion regarding "decoding," the term "encoding" as used in this application may include all or part of the processes performed on, for example, an input video sequence to generate an encoded bitstream.

[0101] The examples and embodiments described herein may be implemented, for example, as methods or processes, devices, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), implementations of the features discussed may also be implemented in other forms (e.g., devices or programs). Devices may be implemented, for example, in appropriate hardware, software, and firmware. Methods may be implemented, for example, in devices, for example, in processors, where processors generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, for example, computers, mobile phones, portable / personal information terminals (“PDAs”), and other devices that facilitate communication of information between end users.

[0102] References to "one embodiment," "an embodiment," "one implementation," or "an implementation," as well as other variations thereof, mean that specific features, structures, characteristics, etc. described in relation to the embodiment are included in at least one embodiment. Accordingly, the appearance of phrases such as "in one embodiment," "in an embodiment," "in one implementation," or "in an implementation," as well as any other variations, appearing in various places throughout this application, do not necessarily all refer to the same embodiment. Additionally, this application may refer to "determining" various information. Determining information may include, for example, estimating information, calculating information, predicting information, or retrieving information from memory.

[0103] Additionally, the present application may refer to "accessing" various information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0104] Additionally, the present application may refer to "receiving" various information. Receiving is intended to be a broad term, just like "accessing." Receiving information may include, for example, accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" is typically involved in some way during operations, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0105] For example, as in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be noted that the use of any of the following " / ", "and / or", and "at least one of ~" is intended to include the selection of only the first enumerated option (A), or the selection of only the second enumerated option (B), or the selection of both options (A and B). As an additional example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to include the selection of only the first enumerated option (A), or the selection of only the second enumerated option (B), or the selection of only the third enumerated option (C), or the selection of only the first and second enumerated options (A and B), or the selection of only the first and third enumerated options (A and C), or the selection of only the second and third enumerated options (B and C), or the selection of all three options (A, B, and C). This can be extended to as many items as are listed, as is obvious to a person skilled in the art of this and related fields.

[0106] Additionally, as used herein, the word "signal" refers specifically to indicating something to the corresponding decoder. For example, in certain examples, the encoder signals a quantization matrix for inverse quantization. In this way, the same parameter is used on both the encoder side and the decoder side in the examples. Thus, for example, the encoder can transmit a specific parameter to the decoder (explicit signaling) so that the decoder can use the same specific parameter. Conversely, if the decoder already possesses not only the specific parameter but also others, signaling may be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameter. By avoiding the transmission of any actual functions, bit savings are realized in various examples. It should be noted that signaling can be achieved in various ways.

[0107] For example, in various examples, one or more syntactic elements, flags, etc. are used to signal information to a corresponding decoder. The foregoing relates to the verb form of the word "signal," but the word "signal" may also be used as a noun in this specification.

[0108] As will be obvious to those skilled in the art, examples may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for performing a method, or data generated by one of the described examples. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted via various different wired or wireless links as known. The signal may be stored on a processor-readable medium.

Claims

Claim 1 A method for communicating mesh and displacement data, comprising: receiving a plurality of substreams of items of media content by an encoder of a device — said plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream —; generating a track reference by the encoder that identifies an association between the base mesh substream and the one or more sub-mesh substreams — said track reference is encapsulated within a header of a first substream among the plurality of substreams —; and generating the plurality of substreams in a container file format by the encoder, wherein the header of the first substream generated in the container file format includes reference identifiers of the one or more video substreams, the base mesh substream, the one or more sub-mesh substreams, and the displacement substream. Claim 2 A method according to claim 1, wherein the container file format is the International Organization for Standardization Base Media File Format (ISOBMFF). Claim 3 A method according to claim 1, wherein the base mesh substream, the one or more sub-mesh substreams, or the displacement substream are data timed in synchronization with the atlas substream and one or more video substreams. Claim 4 A method according to claim 1, wherein the base mesh substream, the one or more sub-mesh substreams, or the displacement substream are untimed data. Claim 5 A method according to claim 1, wherein at least one of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream is data timed in synchronization with the atlas substream and one or more video substreams, and at least another of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream is data not timed. Claim 6 A method according to claim 1, wherein the substreams correspond to tracks of the generated container file format. Claim 7 A method according to claim 1, wherein the header of the generated base mesh substream includes an identifier of each submesh substream. Claim 8 A method according to claim 1, wherein the base mesh substream identifies a plurality of base meshes of an item of the media content, and each base mesh is associated with a different time period within the item of the media content. Claim 9 A method according to claim 1, wherein one or more generated sub-mesh sub-streams are identified as members of a track group in the header of the atlas sub-stream. Claim 10 A device for communicating mesh and displacement data, comprising: an encoder; and an input / output (I / O) device operatively connected to the encoder, wherein the device is configured to receive a plurality of substreams of items of media content, and the plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream; wherein the device is further configured to generate a track reference identifying an association between the base mesh substream and the one or more sub-mesh substreams, and the track reference is encapsulated in the header of a first substream among the plurality of substreams; wherein the device is further configured to generate the plurality of substreams in a container file format, and the header of the first substream generated in the container file format includes a reference identifier for each of the one or more video substreams, the base mesh substream, the one or more sub-mesh substreams, and the displacement substream. Claim 11 In paragraph 10, the above container file format is a device that is the International Organization for Standardization Base Media File Format (ISOBMFF). Claim 12 A device according to claim 10, wherein the base mesh substream, the one or more sub-mesh substreams, or the displacement substream are data timed in synchronization with the atlas substream and one or more video substreams. Claim 13 A device according to claim 10, wherein the base mesh substream, the one or more sub-mesh substreams, or the displacement substream are untimed data. Claim 14 A device according to claim 10, wherein at least one of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream is data timed in synchronization with the atlas substream and one or more video substreams, and at least another of the base mesh substream, the one or more sub-mesh substreams, or the displacement substream is data not timed. Claim 15 In paragraph 10, the device, wherein the header of the generated base mesh substream includes an identifier of each submesh substream. Claim 16 In paragraph 10, a device in which the header of one sub-mesh track includes an identifier of each sub-mesh sub-stream. Claim 17 In paragraph 10, the base mesh substream identifies a plurality of base meshes of an item of the media content, and each base mesh is associated with a different time period within the item of the media content, a device. Claim 18 In paragraph 10, the device, wherein one or more of the generated sub-mesh substreams are identified as members of a track group in the header of the atlas substream. Claim 19 A method for communicating mesh and displacement data, comprising: receiving information including a container file format by a decoder of a device — said container format includes a plurality of substreams of items of media content, said plurality of substreams include an atlas substream, one or more video substreams, a base mesh substream, one or more sub-mesh substreams, and a displacement substream —; decoding said information including the container file format including the plurality of substreams by the decoder; parsing said substreams from the container file format — said header of a first substream generated in said container file format includes a reference identifier for each of said one or more video substreams, said base mesh substream, said one or more sub-mesh substreams, and said displacement substream —; parsing a track reference encapsulated in the header of the first substream among said plurality of substreams to identify an association between said base mesh substream and said one or more sub-mesh substreams; and rendering said media content using said track reference. Claim 20 In paragraph 19, the method wherein the rendering of the media content is performed on an external display.