Bearing of dynamic system stream unit head length accuracy information for volumetric video

By defining the 'vsuh' sample group in the ISOBMFF file, the problem of storing the time-varying accuracy information of the NAL/V3C unit header in the system stream was solved, and efficient compression and transmission of dynamic point cloud data were achieved.

CN121986483APending Publication Date: 2026-05-05INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technical specifications do not discuss the storage method for the length precision information of the system stream NAL unit and V3C unit header when it changes over time in the V3C bitstream, which makes it impossible to effectively manage the compression and transmission of dynamic point cloud data.

Method used

By defining V3C sample stream unit header (SSUH) information sample groups with grouping type 'vsuh' in the ISOBMFF file, dynamically changing NAL or V3C unit header length precision information is stored and mapped to ensure association with the samples.

Benefits of technology

It realizes the storage and transmission of system stream NAL/V3C cell header length precision information in V3C bitstream, and supports efficient dynamic point cloud data compression and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121986483A_ABST
    Figure CN121986483A_ABST
Patent Text Reader

Abstract

A point cloud encoding and decoding system and method. In one example encoding method, at least a first track comprising a plurality of samples containing point cloud information is encoded in a container file, such as an ISOBMFF file. A first sample set description box associated with the first track is encoded, wherein the first sample set description box includes at least a first sample set description entry and a second sample set description entry. The first sample group description entry stores first cell head length precision information and the second sample group description entry stores second cell head length precision information different from the first cell head length precision information. A first sample-to-group box is encoded, and it associates a first plurality of samples with a first sample group description entry and a second plurality of samples with a second sample group description entry.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing This application claims priority to European Patent Application No. 23306718.0, filed on October 6, 2023, entitled “Carrying of Precision Information of Stream Unit Header Length for Volumetric Video,” the entire contents of which are incorporated herein by reference. Background Technology

[0002] High-quality 3D point clouds have recently become an advanced representation for immersive media. A point cloud consists of a set of points represented in three-dimensional space, where each point is represented by coordinates and one or more attributes, such as color, transparency, acquisition time, laser reflectivity, or material properties associated with each point. Point clouds can be acquired in several ways. For example, one technique for acquiring point clouds is using multiple cameras and depth sensors. LiDAR (Light Detection and Ranging) laser scanners are also commonly used for point cloud acquisition. The number of points required to realistically reconstruct objects and scenes from point clouds can reach millions (or even billions). Therefore, efficient representation and compression are ideal for storing and transmitting point cloud data.

[0003] In its uncompressed form, volumetric video is represented by massive amounts of data. The ISO / IEC 23090-5 specification for Visual Volumetric Video Coding (V3C) leverages the compression efficiency of existing 2D video codecs to reduce the amount of data required for storing and transmitting volumetric video. The V3C encoder converts volumetric frames into a set of 2D image sequences and associated metadata, which enables the reconstruction of the volumetric frames; this metadata is called atlas data. The resulting 2D image sequences are then encoded using conventional video or image encoders, such as those according to ISO / IEC 14496-10 (H.264 / AVC), ISO / IEC 23008-2 (HEVC), or ISO / IEC 23090-3 (VVC) international standards. The atlas data is encoded according to the mechanism specified in the ISO / IEC 23090-5 standard. V3C is a general-purpose volumetric video coding and decoding mechanism that can be used for various applications aimed at compressing volumetric content, such as point clouds, immersive video, and mesh representations of visual volumetric frames. Examples of such applications include Video-Based Point Cloud Compression (V-PCC) and MPEG Immersive Video (MIV).

[0004] The V3C specification employs a High-Level Syntax (HLS) design, commonly found in traditional 2D video codecs, to represent encoded atlas data. The encoded atlas data is represented by a series of Network Abstraction Layer (NAL) units.

[0005] Recent advancements in 3D point acquisition and rendering technologies have enabled new applications in fields such as telepresence, virtual reality, and large-scale dynamic 3D maps. The 3D Graphics subgroup of ISO / IEC JTC1 / SC29 / WG7 is currently developing two 3D point cloud compression (PCC) standards: one is a geometry-based compression standard for static point clouds and time-varying sparse point clouds acquired by LiDAR, and the other is a video-based compression standard for dynamic point clouds. These standards aim to support efficient and interoperable storage and transmission of 3D point clouds. One of the goals of these standards is to provide support for lossy and / or lossless encoding and decoding of point cloud geometric coordinates and attributes. WG7 has released the FDIS version of the ISO / IEC 23090-5 standard for video compression of dynamic point clouds. The MPEG Systems Working Group (ISO / IEC / SC29 / WG3) has developed the V3C data bearer standard (ISO / IEC 23090-10) and has released its FDIS version. Summary of the Invention

[0006] An example point cloud encoding method includes: encoding at least a first track in a container file (such as an ISOBMFF file), the first track including a plurality of samples containing point cloud information; encoding a first sample group description box associated with the first track in the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; wherein the first sample group description entry stores first unit header length precision information, and the second sample group description entry stores second unit header length precision information different from the first unit header length precision information; and encoding a first sample-to-group box in the container file, the first sample-to-group box associating a first plurality of samples with the first sample group description entry, and associating a second plurality of samples with the second sample group description entry.

[0007] The first track can be, for example, an atlas track, a V3C bitstream track, or an atlas tile track.

[0008] In some embodiments, the cell header length precision information is the system flow network abstraction layer (NAL) cell header length precision information.

[0009] In some embodiments, the cell header length accuracy information is the system flow V3C cell header length accuracy information.

[0010] Some embodiments further include: encoding at least a second track in a container file, the track including a plurality of samples containing point cloud information; encoding a second sample group description box associated with the second track in the container file, the second sample group description box including at least a third sample group description entry and a fourth sample group description entry; wherein the third sample group description entry stores third unit head length precision information, and the fourth sample group description entry stores fourth unit head length precision information different from the third unit head length information; and encoding a second sample-to-group box in the container file, the second sample-to-group box associating a third plurality of samples with the third sample group description entry, and associating a fourth plurality of samples with the fourth sample group description entry.

[0011] The decoding method according to some embodiments includes: decoding at least a first track from a container file, the first track including a plurality of samples containing point cloud information; decoding a first sample group description box associated with the first track from the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; wherein the first sample group description entry stores first unit header length precision information, and the second sample group description entry stores second unit header length precision information different from the first unit header length precision information; and decoding a first sample to group box from the container file, the first sample to group box associating a first plurality of samples with the first sample group description entry, and associating a second plurality of samples with the second sample group description entry.

[0012] The example embodiments further include the following: corresponding decoding techniques, means including one or more processors configured to perform the methods described herein, means including at least one processor and a computer-readable medium storing instructions for performing the methods described herein, a computer-readable medium storing point clouds encoded according to the methods described herein, and signals for transmitting point clouds encoded according to the methods described herein. Attached Figure Description

[0013] Figure 1 The V3C bitstream structure, which consists of a series of V3C units, is shown.

[0014] Figure 2 The atlas frame is shown, divided into 24 tiles and 9 tile groups.

[0015] Figure 3 The multi-track V3C container structure is shown.

[0016] Figure 4 An example of using the “vsuh” sample group in a multi-track V3C file is shown according to an example embodiment.

[0017] Figure 5An example of using the “vsuh” sample group in a single-track V3C file according to an example embodiment is shown.

[0018] Figure 6 An example of using the “vsuh” sample group in a V3C file with atlas tracks and multiple atlas tile tracks, according to an example embodiment, is shown.

[0019] Figure 7 It is a block diagram of an example system that implements various aspects and embodiments. Detailed Implementation

[0020] Overview of video-based point cloud compression Figure 1 The structure of a bitstream for video-based point cloud compression (V-PCC) is shown. In ISO / IEC 23090-5, a V3C bitstream consists of a series of V3C units, such as... Figure 1 As shown. Each V3C unit has a V3C unit header and a V3C unit payload. The V3C unit header describes the V3C unit type. Supported V3C unit types are given in Table 1. The attribute video data V3C unit header specifies the attribute type and its index, allowing multiple instances of the same attribute type to be supported. Table 2 specifies the supported V3C attribute types. Occupancy, geometry, and attribute video data unit payloads correspond to video data units (e.g., HEVC NAL units) that can be decoded by an appropriate video decoder. The video decoder corresponding to each video coded component sub-bitstream (i.e., occupancy, geometry, or attribute sub-stream) is signaled in the V3C parameter set. vuh_unit_type identifier V3C Unit Type describe 0 V3C_VPS V3C Parameter Set V3C level parameters 1 V3C_AD Image Data Image Gallery Information 2 V3C_OVD Occupy video data Occupancy information 3 V3C_GVD Geometric video data Geometric Information 4 V3C_AVD Attribute video data Attribute information 5……31 V3C_RSVD Reserved - Table 1. V3C Unit Types ai_attribute_type_id[j][i] identifier Attribute type 0 ATTR_TEXTURE texture 1 ATTR_MATERIAL_ID Material ID 2 ATTR_TRANSPARENCY <![CDATA[ transparency ]]> 3 ATTR_REFLECTANCE reflectivity 4 ATTR_NORMAL normal 5……14 ATTR_RESERVED Reserved 15 ATTR_UNSPECIFIED not specified Table 2. V3C Attribute Types.

[0021] V3C Bitstream High-Level Syntax (HLS) supports data that arranges tiles into tile groups within an atlas frame. An atlas frame is divided into one or more tile rows and one or more tile columns. A tile is a rectangular area within an atlas frame. A tile group contains multiple tiles from an atlas frame, and these tiles can be independently decoded. Figure 2 An example of tile grouping for an atlas frame is shown (specifically, the atlas frame is divided into 24 tiles and 9 tile groups). In the current implementation, only rectangular tile groups are supported. Therefore, a tile group contains multiple tiles of the atlas frame that together form a rectangular region of the atlas frame.

[0022] The ISO / IEC DIS 23090-5 specification also includes several Supplemental Enhancement Information (SEI) messages, which can be signaled in the V3C bitstream to associate patches or volume rectangles within an atlas frame with objects within the scene represented by a point cloud. These SEI messages also support annotation, tagging, and attribute addition to these objects. Objects can correspond to "real" objects (physical objects in the scene) or conceptual objects associated with physical or other attributes. Objects can be associated with different parameters or attributes, which may correspond to information provided during the creation or editing of the point cloud or scene graph. The specification supports the use of information defining dependencies between different objects (e.g., one object may be part of another object).

[0023] Objects in a point cloud can be persistent over time and can be updated at any time / frame, while the associated information can persist from that point in time. Multiple patches or 2D volume rectangles (which can contain multiple patches) can be associated with a single object or multiple objects.

[0024] Overview of ISO Basic Media File Format The ISO / IEC 14496 (MPEG-4) standard defines several time-based media storage file formats. These are based on and derived from the ISO Basic Media File Format (ISOBMFF), a structured, media-independent definition. ISOBMFF contains structured information and media data information, primarily used for time-based presentation of media data such as audio and video. It also supports non-time-based data, such as metadata at different levels within the file structure. The logical structure of a file is a movie, which contains a set of time-parallel tracks. The temporal structure of the file consists of tracks containing temporally sampled sequences that are mapped onto the timeline of the entire movie. ISOBMFF is based on the concept of a box-structured file. A box-structured file consists of a series of boxes (sometimes called atoms) with length and type. These types are 32-bit values, typically chosen as four printable characters, also known as four-character codes (4CC). Non-time-based data can be contained in metadata boxes at the file level or appended to movie boxes or one of the time-based data streams called tracks within the movie.

[0025] Within the top-level box of the ISOBMFF container are movie boxes (“moov”), which contain metadata about the continuous media streams present in the file. This metadata is signaled in a box hierarchy within the movie box, such as in the track box (“trak”). A track represents a continuous media stream present in the file. The media stream itself consists of a series of samples, such as audio or video access units of the basic media stream. Samples are encapsulated in a media data box (“mdat”) present at the top level of the container. The metadata for each track includes a list of sample description entries, each entry providing the codec or container format used in the track and initialization data for processing that format. Each sample is associated with one of the sample description entries of the track. ISO / IEC 14496-12 provides a tool for defining an explicit timeline mapping for each track. This is called an edit list and is signaled using an edit list box with the following syntax, where each entry defines a portion of the track timeline by mapping a part of the composite timeline, or by indicating “empty” time (i.e., a portion of the presentation timeline not mapped to the media, an “empty” edit).

[0026] ISOBMFF provides a mechanism for handling situations where file authors need to perform certain actions (such as post-processing activities) on the player or renderer. For video streams, this mechanism is implemented using restricted video scheme tracks. When a video track is a restricted video scheme track as defined in Section 8.15 of the ISO / IEC 14496-12 standard, it indicates that the track has decoding post-processing requirements. A track is converted to a restricted video scheme track by setting the track's sample entry code to the four-character code (4CC) "resv" and adding a restricted scheme information box to its sample description, while keeping all other boxes unchanged. The original sample entry type based on the video codec used to encode the stream is then stored in the original format box within the restricted scheme information box. The restricted scheme information box contains three boxes: the original format box, the scheme type box, and the scheme information box. The original format box stores the original sample entry type, which is based on the video codec used to encode the component stream. The nature of the restrictions is defined in the scheme type box.

[0027] V3C Container File Format Overview Figure 3 The diagram illustrates the structure of a multitrack ISOBMFF V3C container as defined in MPEG-I Part 10 (ISO / IEC 23090-10). The box in the diagram corresponds to the ISOBMFF box in ISO / IEC 14496-12. Based on this structure, the multitrack V3C ISOBMFF container includes the following: • The V3C track contains the V3C parameter set and atlas parameter set (in the sample entries) as well as samples of NAL units carrying atlas sub-bitstreams. This track also includes track references to other tracks carrying payloads of video compression V3C units (e.g., unit types V3C_OVD, V3C_GVD, and V3C_AVD). • Restricted video scheme tracks, where samples contain access units for video-coded elementary streams of occupancy graph data (e.g., payloads of V3C cells of type V3C_OVD). • One or more restricted video scheme tracks, wherein the samples contain access units for the video-coded elementary stream of geometric data (e.g., payloads of V3C units of type V3C_GVD). • Zero or more restricted video scheme tracks, wherein the samples contain access units for the video-coded elementary stream for attribute data (e.g., payloads of V3C units of type V3C_AVD).

[0028] The “restricted video track” mechanism is used to store V3C bitstreams and indicates the process that the receiver should perform on the corresponding track.

[0029] Problems solved in some embodiments The V3C codec specification (ISO / IEC 23090-5) defines the syntax and semantics of the V3C sample stream format and the NAL sample stream format. The V3C bearer specification (ISO / IEC 23090-10) specifies that the system stream NAL unit (SSNU) header length precision information and the system stream V3C unit (SSVU) header length precision information are stored in the V3C decoder configuration record. This specification assumes that this data does not change over time in the V3C bitstream.

[0030] The V3C codec specification (ISO / IEC 23005) defines the syntax and semantics of the sample stream format for use by applications that deliver some or all of the V3C cell streams as an ordered stream of bytes or bits.

[0031] The V3C sample stream format begins with a sample stream header, followed by a series of sample stream V3C unit syntax structures. Each sample stream V3C unit syntax structure contains an element named `ssvu_V3C_unit_size`, which specifies a length value, followed by a `V3C_unit` (`ssvu_V3C_unit_sizes`) syntax structure. The length of this `v3C_unit` is actually `ssvu_v3c_unit_size`. The number of bits used to encode the `ssvu_v3c_unit_size` element is specified in the sample stream header. The sample stream V3C header syntax is shown below: .

[0032] The sample stream V3C unit syntax is as follows: .

[0033] The syntactic elements shown above can have the following semantics: Increasing 1 to ssvh_unit_size_pecision_bytes_minus1 specifies the precision (in bytes) of the ssvu_v3c_unit_size element in all sample stream v3c units. The value of ssvh_unit_size_pecision_bytes_minus1 should be between 0 and 7.

[0034] `ssvu_v3c_unit_size` specifies the length (in bytes) of the subsequent `v3c_unit`. The number of bits used to represent `ssvu_v3c_unit_size` is equal to `(ssvh_unit_size_pecision_bytes_minu1+1)*8`.

[0035] The V3C codec specification (ISO / IEC 23005) defines the syntax and semantics of the sample stream format for use by applications that deliver some or all of the atlas NAL cell streams as an ordered stream of bytes or bits.

[0036] The NAL unit sample stream format begins with a sample stream header, followed by a series of sample stream NAL unit syntax structures. Each sample stream NAL unit syntax structure contains an element named `ssnu_nal_unit_size`, which specifies a length value, followed by a `nal_unit(ssnu_na_unit_size)` syntax structure. The length of this `nal_unit` is actually `ssnu_nal_unit_size`. The number of bits used to encode the element `ssnu_nal_unit_size` is specified in the sample stream header. The sample stream NAL header syntax is shown below: .

[0037] The syntax for the sample stream NAL unit is as follows: .

[0038] The syntactic elements shown above can have the following semantics: Increasing 1 to ssnhunit_size_pecision_bytes_minus1 specifies the precision (in bytes) of the ssnu_nal_unit_size element in all sample stream NAL units. The value of ssnhunit_size_pecision_bytes_minus1 should be between 0 and 7.

[0039] ssnu_nal_unit_size specifies the length (in bytes) of the subsequent NAL_unit. The number of bits used to represent ssnu_nal_unit_size is equal to (ssnhunit_size_pecision_bytes_minu1+1)*8.

[0040] Currently, existing specifications do not discuss how to store information such as the system stream NAL cell header length precision byte being decremented by one or the system stream V3C cell header length precision byte being decremented by one over time in the V3C bitstream. Some embodiments of this disclosure aim to address this shortcoming. Example embodiments provide details on how to store system stream NAL / V3C cell header length precision byte information in an ISOBMFF file as it changes over time. Example embodiments further describe the association between each sample present in the track and its associated system stream NAL / V3C cell header length precision byte information.

[0041] Overview of some example embodiments This disclosure provides systems and methods for storing NAL / V3C cell header length precision information in a system stream as it dynamically changes over time within the V3C bitstream. Example embodiments further provide systems and methods relating to, for example, providing samples in an ISOBMFF file and their associated NAL / V3C cell header length precision information.

[0042] In some embodiments, a V3C Sample Stream Unit Header (SSUH) information sample group with a grouping type equal to "vsuh" is defined in the V3C Atlas track, V3C Atlas Tile track, or V3C Bitstream track.

[0043] In some embodiments, a sample group description entry of type "vsuh" is defined in the V3C atlas track, V3C atlas tile track, or V3C bitstream track to store dynamically changing NAL or V3C cell header length precision byte information.

[0044] In some embodiments, sample-to-box mapping is used to map samples existing in V3C atlas tracks, V3C atlas tile tracks, or V3C bitstream tracks and their associated NAL or V3C cell header length precision byte information.

[0045] Some implementations allow multiple instances of sample entries in a V3C bitstream track, atlas track, or atlas tile track to store dynamically changing system stream V3C or NAL cell header length precision byte information.

[0046] Example signaling for dynamic precision information In one example embodiment, when a single track is used to carry V3C data and the SSVU header length precision information changes over time, a V3C SSUH information sample group with a group type equal to "vsuh" is used to signal the SSVU header length precision byte information associated with the sample present in the V3C bitstream track.

[0047] In one example embodiment, when multiple tracks are used to carry V3C data and the SSNU header length precision information changes over time, a V3C SSUH information sample group with a group type equal to "vsuh" is used to signal the SSNU header length precision byte information associated with the sample present in the atlas track or atlas tile track.

[0048] In some embodiments, a V3C SSUH information sample group with grouping type "vsuh" is used to group V3C atlas samples that use the same SSNU header length precision byte information in the V3C atlas track or V3C atlas tile track.

[0049] In some embodiments, a V3C SSUH information sample group with grouping type "vsuh" is used to group V3C bitstream samples that use the same SSVU header length precision byte information in the V3C bitstream track.

[0050] Using "vsuh" as the grouping type in sample grouping indicates that a sample in a V3C track is assigned to the corresponding SSNU or SSVU header length precision byte information carried in the sample group description entry box. When a sample-to-group box with a grouping type equal to "vsuh" exists, an accompanying sample group description box with the same grouping type should also exist, and the sample-to-group box should contain the index of the sample's sample group description entry.

[0051] In one example embodiment, the grouping type "vsuh" may have the following characteristics: Group type: "vsuh" Container: Sample group descriptor box (“sgpd”) Mandatory: No Quantity: 0 or 1 In one example embodiment, the V3C Sample Stream Unit Header (SSUH) information sample group entry defines the NAL unit header length precision information for all atlas samples using the same precision byte information. The V3C SSUH information sample group entry defines the V3C unit header length precision information for all V3C bitstream samples using the same precision byte information in the V3C bitstream track.

[0052] In some embodiments, the sample group type "vsuh" may be used only with tracks having sample entries "v3e1", "v3eg", "v3c1", "v3cg", "w3cb", "v3t1", "v3a1", or "v3ag". This sample group and its associated description entry may not exist in the V3C video component track. When the SSNU header length precision information changes over time, this sample group and its associated description entry may exist in the V3C atlas track or V3C atlas tile track. When the SSVU header length precision information changes over time, this sample group and its associated description entry may exist in the V3C bitstream track. When the sample stream NAL unit header precision length or the sample stream V3C unit header precision length does not change over time, this information may exist in the V3C decoder configuration record, and this sample group and its associated description entry may not exist in any V3C track.

[0053] In some embodiments, the "vsuh" sample grouping type has the following syntax:

[0054] In some embodiments, the value `unit_size_pecision_bytes_minus1` plus 1 specifies the precision (in bytes) of the NAL unit length of the sample stream. In V3C atlas tracks or V3C atlas tile tracks, `unit_size_pecision_bytes_minus1` can be constrained to be equal to `ssnh_unit_size_pecision_bytes_minus1` in `sample_stream_nal_header()`. For V3C bitstream tracks, `unit_size_pecision_bytes_minus1` can be constrained to be equal to `ssvh_unit_size_precision_bytes_minus1` in `sample_stream_V3C_header()`. In some embodiments, the bits in the value "Reserved" are set to zero.

[0055] Figure 4An example of a V3C ISOBMFF file structure with an atlas track is shown. In this example, track 1 is an atlas track containing a sample group description box “sgpd” and a sample-to-group box “sbgp”. The sample-to-group box “sbgp” has a grouping type “vsuh”. The sample group description boxes with grouping type “vsuh” present in track 1 contain distinct entries, which are values ​​for each SSNU header length precision byte minus one, and each SSNU header length precision byte minus one is used in different samples of that track. The sample-to-group box with grouping type “vsuh” contains an index of the associated sample group description entry for each sample in that track.

[0056] As the SSNU header length precision byte information changes over time, the new SSNU header length precision byte value is carried in the sample group description entry. Samples using the new SSNU header length precision byte value are indicated in the sample group box with a group type equal to "vsuh". Figure 4 In the example shown, samples 1 to 100 in track 1 use SSNU header length precision byte minus one value information described by the sample group entry at index 1. Samples 101 to 300 use SSNU header length precision byte minus one value information described by the sample group entry at index 2 in the sample group description box.

[0057] Figure 5 An example of a V3C ISOBMFF file structure with a single track is shown. Track 1 is a V3C bitstream track containing a sample group description box "sgpd" and a sample-to-group box "sbgp" with group type "vsuh". The sample group description boxes with group type "vsuh" in Track 1 contain different entries for each SSVU header length precision byte minus one used for different samples in that track. The sample-to-group boxes with group type "vsuh" contain an index of the associated sample group description entry for each sample present in that track.

[0058] As the SSVU header length precision byte information changes over time, the new SSVU header length precision byte value is carried in the sample group description entry. Samples using the new SSVU header length precision byte value are indicated in the sample group box with group type equal to "vsuh". Figure 5 In the example shown, samples 1 to 100 in track 1 use SSVU header length precision byte minus one value information from the sample group entry at index 1. Samples 101 to 300 use SSVU header length precision byte minus one value information from the sample group entry at index 2 in the sample group description box.

[0059] Figure 6An example of a V3C ISOBMFF file structure with one atlas track and two atlas tile tracks is shown. Track 1 is an atlas track containing a sample group description box “sgpd” and a sample-to-group box “sbgp” with group type equal to “vsuh”. The sample group description boxes with group type equal to “vsuh” in Track 1 contain different entries for each SSNU header length precision byte minus one used for different samples in that track. The sample-to-group boxes with group type equal to “vsuh” contain an index of the associated sample group description entry for each sample present in that track. Similarly, Tracks 2 and 3 (atlas tile tracks) also contain sample-to-group boxes with group type equal to “vsuh” and sample group description boxes.

[0060] As the SSNU header length precision byte information changes over time, the new SSNU header length precision byte value is carried in the sample group description entry, and samples using the new SSNU header length precision byte value are indicated in the sample-to-group box with a group type equal to "vsuh". Figure 6 In the examples shown, samples 1 to 100 in track 1 use the SSNU header length precision byte minus one value from the sample group entry at index 1. Samples 101 to 300 use the SSNU header length precision byte minus one value from the sample group entry at index 2 present in the sample group description box. Similarly, samples 1 to 200 in tracks 2 and 3 use the SSNU header length precision byte minus one value from the sample group entry at index 1. Samples 201 to 300 use the SSNU header length precision byte minus one value from the sample group entry at index 2.

[0061] refer to Figure 6 The example point cloud encoding method includes encoding at least a first track 602 in a container file (such as an ISOBMFF file), the track comprising multiple samples containing point cloud information. A first sample group description box 604 associated with the first track is encoded in the container file. The first sample group description box includes at least a first sample group description entry 606 and a second sample group description entry 608. Some embodiments may include more such sample group description entries. The first sample group description entry stores first unit header length precision information and the second sample group description entry stores second unit header length precision information different from the first unit header length precision information. A first sample-to-group box 610 is also encoded in the container file. The first sample-to-group box associates a first plurality of samples with the first sample group description entry and associates a second plurality of samples with the second sample group description entry.

[0062] Some embodiments further include encoding at least a second track 612 in the container file, the track comprising multiple samples containing point cloud information. A second sample group description box 614 associated with the second track is also encoded within the container file. The second sample group description box includes at least a third sample group description entry 616 and a fourth sample group description entry 618. The third sample group description entry stores third unit head length precision information and the fourth sample group description entry stores fourth unit head length precision information different from the third unit head length precision information. A second sample-to-group box 620 is also encoded in the container file. The second sample-to-group box associates a third plurality of samples with the third sample group description entry and associates a fourth plurality of samples with the fourth sample group description entry. Figure 6 As shown, in some embodiments, an additional track (e.g., "track 3") with additional unit head length accuracy information may also be provided.

[0063] In other embodiments, when using a single track to carry V3C data and the SSVU header length precision information changes over time, multiple sample entries can be used to store this information. The SSVU header length precision information associated with a set of samples in that V3C bitstream track is stored in the V3C configuration box of the sample entry instance. When using multiple tracks to carry V3C data and the SSNU header length precision information changes over time, multiple sample entries can be used to store this information. The SSNU header length precision information associated with a set of samples in that V3C atlas track is stored in the V3C configuration box of the sample entry instance. The SSNU header length precision information associated with a set of samples in that V3C atlas tile track is stored in the V3C atlas tile configuration box of the sample entry instance.

[0064] Example system hardware Example embodiments configured to implement the encoders and / or decoders (collectively, the encoders) described herein may use, for example... Figure 7 This is achieved through a system. Figure 7This is a block diagram of an example system in which various aspects and embodiments are implemented. System 1000 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0065] System 1000 includes at least one processor 1010 configured to execute instructions loaded thereon to implement various aspects, such as those described in this document. Processor 1010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives.

[0066] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide an encoded or decoded point cloud, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 1030 can be implemented as a separate element of system 1000, or it can be incorporated into processor 1010 as a combination of hardware and software known to those skilled in the art.

[0067] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to execute the various aspects described in this document may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input point clouds, decoded point clouds or portions of decoded point clouds, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0068] In some embodiments, the memory within the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. External memory may be memory 1020 and / or storage device 1040, such as dynamically volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Codec, also known as H.265 and MPEG-H Part 2), or VVC (Universal Video Codec, a new standard being developed by the Joint Video Experts Group JVET).

[0069] As indicated in box 1130, inputs can be provided to the components of system 1000 through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal.

[0070] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a certain band), (ii) down-converting the selected signal, (iii) again limiting the band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band. Various embodiments rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0071] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, within a separate input processing IC or processor 1010. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed within a separate interface IC or processor 1010. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, an encoder / decoder 1030 operating in conjunction with memory and storage elements, and processor 1010, to process the data streams as needed for presentation on the output device.

[0072] Various components of system 1000 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data therebetween using suitable connection means 1140, such as internal buses known in the art, including inter-IC (I2C) buses, lines, and printed circuit boards.

[0073] System 1000 includes a communication interface 1050, which enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0074] In various embodiments, a wireless network, such as Wi-Fi, or IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), is used to stream or otherwise provide data to system 1000. In these embodiments, Wi-Fi signals are received via a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 1000, delivering data via an HDMI connection to input box 1130. Still other embodiments use an RF connection to input box 1130 to provide streaming data to system 1000. As shown above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0075] System 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used in a television, tablet computer, laptop computer, mobile phone, or other device. The display 1100 can also be integrated with other components (e.g., in a smartphone) or separate (e.g., as an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, used for both terms), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functionality based on the output of system 1000. For example, an optical disc player performs the function of playing the output of system 1000.

[0076] In various embodiments, system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 transmit control signals using signaling protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols, enabling inter-device control regardless of user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. Display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000 in electronic devices such as, for example, televisions. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0077] Display 1100 and speaker 1110 may alternatively be separated from one or more other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output port.

[0078] The embodiments can be implemented by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 1010 can be of any type suitable for the technical environment and can include one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0079] Additional Examples An example point cloud encoding method includes: encoding at least a first track in a container file (such as an ISOBMFF file), the first track including a plurality of samples containing point cloud information; encoding a first sample group description box associated with the first track in the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; wherein the first sample group description entry stores first cell header length precision information, and the second sample group description entry stores second cell header length precision information different from the first cell header length precision information; and encoding a first sample to a group box in the container file, the box associating a first plurality of samples with the first sample group description entry and associating a second plurality of samples with the second sample group description entry.

[0080] The first track can be, for example, an atlas track, a V3C bitstream track, or an atlas tile track.

[0081] In some embodiments, the cell header length precision information is the system flow network abstraction layer (NAL) cell header length precision information.

[0082] In some embodiments, the cell header length accuracy information is the system flow V3C cell header length accuracy information.

[0083] Some embodiments further include: encoding at least a second track in a container file, the second track including a plurality of samples containing point cloud information; encoding a second sample group description box associated with the second track in the container file, the second sample group description box including at least a third sample group description entry and a fourth sample group description entry; wherein the third sample group description entry stores third unit head length precision information, and the fourth sample group description entry stores fourth unit head length precision information different from the third unit head length precision information; and encoding a second sample-to-group box in the container file, the box associating a third plurality of samples with the third sample group description entry, and associating a fourth plurality of samples with the fourth sample group description entry.

[0084] A decoding method according to some embodiments includes: decoding at least a first track from a container file, the first track including a plurality of samples containing point cloud information; decoding a first sample group description box associated with the first track from the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; wherein the first sample group description entry stores first unit header length precision information and the second sample group description entry stores second unit header length precision information different from the first unit header length precision information; and decoding a first sample to a group box from the container file, the box associating a first plurality of samples with the first sample group description entry and a second plurality of samples with the second sample group description entry.

[0085] The example embodiments further include the following: corresponding decoding techniques, means including one or more processors configured to perform the methods described herein, means including at least one processor and a computer-readable medium storing instructions for performing the methods described herein, a computer-readable medium storing point clouds encoded according to the methods described herein, and signals for transmitting point clouds encoded according to the methods described herein.

[0086] It is important to note that the various hardware elements in one or more of the described embodiments are referred to as “modules”, which implement (i.e., perform, execute, etc.) the various functions described herein in conjunction with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more storage devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and it is noteworthy that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or include hardware (i.e., hardwired), firmware, software instructions, and / or similar instructions, and may be stored in any suitable non-transitory computer-readable medium, such as media commonly referred to as RAM, ROM, etc.

[0087] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware contained in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multifunction disks (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host.

Claims

1. An encoding method, comprising: Encode at least a first track in the container file, the first track comprising multiple samples containing point cloud information; The container file encodes a first sample group description box associated with the first track, the first sample group description box including at least a first sample group description entry and a second sample group description entry; The first sample group description entry stores the first unit head length precision information, and the second sample group description entry stores the second unit head length precision information, which is different from the first unit head length precision information. as well as The container file encodes a first sample-to-group box, which associates a first plurality of samples with a first sample group description entry and associates a second plurality of samples with a second sample group description entry.

2. An encoding apparatus comprising one or more processors, said processors being configured to perform at least the following operations: Encode at least a first track in the container file, the first track comprising multiple samples containing point cloud information; The container file encodes a first sample group description box associated with the first track, the first sample group description box including at least a first sample group description entry and a second sample group description entry; in, The first sample group description entry stores the first unit header length precision information, and the second sample group description entry stores the second unit header length precision information, which is different from the first unit header length precision information. as well as The container file encodes a first sample-to-group box, which associates a first plurality of samples with a first sample group description entry and associates a second plurality of samples with a second sample group description entry.

3. The method according to claim 1 or the apparatus according to claim 2, wherein the first track is an atlas track, a V3C bitstream track, or an atlas tile track.

4. The method according to claim 1 or claim 3 which is a subset of claim 1, or the apparatus according to claim 2 or claim 3 which is a subset of claim 2, wherein the cell header length precision information is the system flow network abstraction layer (NAL) cell header length precision information.

5. The method according to claim 1 or claim 3 which is a subset of claim 1, or the apparatus according to claim 2 or claim 3 which is a subset of claim 2, wherein the unit header length accuracy information is system flow V3C unit header length accuracy information.

6. The method according to claim 1 or claims 3-5 which are dependent on claim 1, or the apparatus according to claim 2 or claims 3-5 which are dependent on claim 2, further comprising: At least a second track is encoded in the container file, the second track comprising multiple samples containing point cloud information; The container file encodes a second sample group description box associated with the second track, the second sample group description box including at least a third sample group description entry and a fourth sample group description entry; The third sample group description entry stores the third unit header length precision information, and the fourth sample group description entry stores the fourth unit header length precision information, which is different from the third unit header length precision information; and The container file encodes a second sample-to-group box, which associates a third plurality of samples with the third sample group description entry and associates a fourth plurality of samples with the fourth sample group description entry.

7. The method according to claim 1 or claims 3-6 which are dependent on claim 1, or the apparatus according to claim 2 or claims 3-6 which are dependent on claim 2, wherein the container file is an ISOBMFF file.

8. A decoding method, comprising: Decode at least a first track from the container file, the first track comprising multiple samples containing point cloud information; Decode the first sample group description box associated with the first track from the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; Wherein, the first sample group description entry provides first unit head length precision information, and the second sample group description entry provides second unit head length precision information that is different from the first unit head length precision information; as well as The first sample to group box is decoded from the container file, the first sample to group box associating a first plurality of samples with a first sample group description entry, and associating a second plurality of samples with a second sample group description entry.

9. A decoding apparatus comprising one or more processors, said processors being configured to perform at least the following operations: Decode at least a first track from the container file, the first track comprising multiple samples containing point cloud information; Decode the first sample group description box associated with the first track from the container file, the first sample group description box including at least a first sample group description entry and a second sample group description entry; in, The first sample group description entry provides first unit head length precision information, and the second sample group description entry provides second unit head length precision information that is different from the first unit head length precision information; as well as The first sample to group box is decoded from the container file, the first sample to group box associating a first plurality of samples with a first sample group description entry, and associating a second plurality of samples with a second sample group description entry.

10. The method according to claim 8 or the apparatus according to claim 9, further comprising processing the point cloud information based on the unit head length accuracy information.

11. The method according to claim 8 or claim 10 which is a subset of claim 8, or the apparatus according to claim 9 or claim 10 which is a subset of claim 9, wherein the first track is an atlas track, a V3C bitstream track, or an atlas tile track.

12. The method according to claim 8 or claims 10-11 which are dependent on claim 8, or the apparatus according to claim 9 or claims 10-11 which are dependent on claim 9, wherein the cell header length precision information is the system flow network abstraction layer (NAL) cell header length precision information.

13. The method according to claim 8 or claims 10-11 which are dependent on claim 8, or the apparatus according to claim 9 or claims 10-11 which are dependent on claim 9, wherein, The unit header length accuracy information is the system flow V3C unit header length accuracy information.

14. The method according to claim 8 or claims 10-13 which are dependent on claim 8, or the apparatus according to claim 9 or claims 10-13 which are dependent on claim 9, further comprising: Decode at least a second track from the container file, the second track comprising multiple samples containing point cloud information; Decode the second sample group description box associated with the second track from the container file, the second sample group description box including at least a third sample group description entry and a fourth sample group description entry; The third sample group description entry stores the third unit head length precision information, and the fourth sample group description entry stores the fourth unit head length precision information that is different from the third unit head length precision information. as well as The second sample is decoded from the container file into a group box, which associates a third plurality of samples with the third sample group description entry and associates a fourth plurality of samples with the fourth sample group description entry.

15. The method according to claim 8 or claims 10-14 which are dependent on claim 8, or the apparatus according to claim 9 or claims 10-14 which are dependent on claim 9, wherein the container file is an ISOBMFF file.