Operation point sample group in the encoded and decoded video

By using VVC standard video encoding and decoding technology, the problem of inefficient storage and decoding of multi-layer video files is solved, flexible storage and real-time adaptation of multi-layer videos are realized, and network adaptability and decoding efficiency are improved.

CN114205606BActive Publication Date: 2025-08-26FACE CUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111092639.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-09-17
Publication Date
2025-08-26
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have problems such as inefficiency and insufficient flexibility in multi-layer bitstream processing and file format storage, especially when network conditions change, they cannot effectively support real-time adaptation and storage of multi-layer video.

Method used

Using VVC standard video encoding and decoding technology, by storing multi-layer video bitstreams in the ISO basic media file format, using the format rules of operation point information and layer correlation syntax elements, it realizes efficient storage and analysis of multi-layer bitstreams, and supports flexible output and real-time adaptation of multi-layer video.

Benefits of technology

It improves the storage and decoding efficiency of multi-layer video files, supports flexible output at different levels, adapts to changes in network conditions, and reduces the complexity and storage requirements of codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114205606B_ABST
    Figure CN114205606B_ABST
Patent Text Reader

Abstract

The present application relates to sample groups of operation points in encoded and decoded video, and describes systems, methods, and apparatus for processing visual media data. The method includes performing conversion between visual media data and a visual media file, wherein the visual media file stores a bitstream of the visual media data in a plurality of tracks according to a format rule, wherein the format rule specifies that file-level information includes syntax elements, and the syntax elements identify one or more tracks from the plurality of tracks that contain sample groups of a particular type including operation point information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is filed to timely claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 079,946, filed on September 17, 2020, and U.S. Provisional Patent Application No. 63 / 088,786, filed on October 7, 2020, under applicable patent laws and / or under the rules of the Paris Convention. The entire disclosures of the foregoing applications are incorporated herein by reference as a part of the disclosure of this application for all purposes under law. Technical Field

[0003] This patent document relates to the generation, storage and consumption of digital audio and video media information in file format. Background Art

[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that may be used by a visual media processing device for writing to or parsing from visual media files according to a file format.

[0006] In one example aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in a plurality of tracks according to a format rule, the format rule specifying that file-level information includes syntax elements, the syntax elements identifying one or more tracks from the plurality of tracks that contain a sample group of a particular type including operation point information.

[0007] In another example aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule. The visual media file stores one or more tracks comprising one or more video layers. The format rule specifies that whether a first set of syntax elements indicating layer dependency information is stored in the visual media file depends on whether a second syntax element indicating that all layers in the visual media file are unrelated has a value of 1.

[0008] In another example aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in a plurality of tracks according to format rules, the format rules specifying a manner in which redundant access unit delimiter network access layer (AUD NAL) units stored in the plurality of tracks are processed during implicit reconstruction of the bitstream from the plurality of tracks.

[0009] In another example aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule. The visual media file stores one or more tracks comprising one or more video layers. The visual media file includes operation point (OP) information. The format rule specifies whether or how syntax elements are included in a sample group entry and an OP group box based on whether the OP contains a single video layer. The syntax element is configured to indicate an index of an output layer set of the OP.

[0010] In another example aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule, wherein the visual media file stores a plurality of tracks belonging to a specific type of entity group, and wherein the format rule specifies that, in response to the plurality of tracks having a specific type of track reference to a group identifier, the plurality of tracks (A) omit carrying a sample group of the specific type or (B) carry a sample group of the specific type such that information in the sample group of the specific type is consistent with information in the entity group of the specific type.

[0011] In another exemplary aspect, a visual media processing method is disclosed. The method includes performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data. The visual media file includes a plurality of tracks, and the visual media file stores an entity group that carries information about operation points in the visual media file and a track that carries each operation point. A format rule specifies attributes of the visual media file in response to the visual media file storing an entity group or a sample group that carries information about each operation point.

[0012] In yet another exemplary aspect, a visual media writing apparatus is disclosed, wherein the apparatus includes a processor configured to implement the above method.

[0013] In yet another exemplary aspect, a visual media parsing apparatus is disclosed, wherein the apparatus includes a processor configured to implement the above method.

[0014] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.

[0015] In yet another exemplary aspect, a computer-readable medium having a visual media file stored thereon is disclosed. The visual media is generated or parsed using the methods described herein.

[0016] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a block diagram of an example video processing system.

[0018] Figure 2 It is a block diagram of a video processing device.

[0019] Figure 3 is a flow chart of an example method of video processing.

[0020] Figure 4 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0021] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0022] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0023] Figure 7 An example of an encoder block diagram is shown.

[0024] Figure 8 An example of a bitstream with two OLSs is shown, where vps_max_tid_il_ref_pics_plus1[1][0] of OLS2 is equal to 0.

[0025] 9A to 9F A flow chart depicting an example method for visual media processing is depicted. DETAILED DESCRIPTION

[0026] Section headings are used in this document to facilitate understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is intended solely to facilitate understanding and is not intended to limit the scope of the disclosed techniques. As such, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to text relative to the current draft of the VVC specification or the ISOBMFF file format specification are shown with strikethrough indicating deleted text and highlighting indicating added text (including bold italics).

[0027] 1. Preliminary Discussion

[0028] This document is related to video file formats. Specifically, it relates to the storage of scalable generic video codec (VVC) video bitstreams in media files based on the ISO base media file format (ISOBMFF). The ideas can be applied individually or in various combinations to video bitstreams encoded or decoded by any codec (e.g., the VVC standard), and to any video file format, such as the VVC video file format under development.

[0029] 2. Abbreviation

[0030] ACT adaptive color transform adaptive color transform

[0031] ALF adaptive loop filter adaptive loop filter

[0032] AMVR adaptive motion vector resolution

[0033] APS adaptation parameter set

[0034] AU access unit access unit

[0035] AUD access unit delimiter access unit delimiter

[0036] AVC advanced video coding (Rec.ITU-T H.264|ISO / IEC14496-10)

[0037] B bi-predictive

[0038] BCW bi-prediction with CU-level weights

[0039] BDOF bi-directional optical flow

[0040] BDPCM block-based delta pulse code modulation

[0041] BP buffering period

[0042] CABAC context-based adaptive binary arithmetic coding

[0043] CB coding block

[0044] CBR constant bit rate

[0045] CCALF cross-component adaptive loop filter cross-component adaptive loop filter

[0046] CPB coded picture buffer encoded and decoded picture buffer

[0047] CRA clean random access

[0048] CRC cyclic redundancy check

[0049] CTB coding tree block

[0050] CTU coding tree unit

[0051] CU coding unit

[0052] CVS coded video sequence video sequence after encoding and decoding

[0053] DPB decoded picture buffer

[0054] DCI decoding capability information

[0055] DRAP dependent random access point

[0056] DU decoding unit

[0057] DUI decoding unit information

[0058] EG exponential-Golomb index-Columbus

[0059] EGk k-th order exponential-Golomb k-th order exponential-Golomb

[0060] EOB end of bitstream

[0061] EOS end of sequence

[0062] FD filler data filling data

[0063] FIFO first-in, first-out

[0064] FL fixed-length fixed length

[0065] GBR green,blue,and red

[0066] GCI general constraints information

[0067] GDR gradual decoding refresh

[0068] GPM geometric partitioning mode

[0069] HEVC high efficiency video coding (Rec.ITU-T H.265|ISO / IEC 23008-2)

[0070] HRD hypothetical reference decoder

[0071] HSS hypothetical stream scheduler

[0072] I intraframe

[0073] IBC intra block copy

[0074] IDR instantaneous decoding refresh

[0075] ILRP inter-layer reference picture

[0076] IRAP intra random access point intra frame random access point

[0077] LFNST low frequency non-separable transform low frequency non-separable transform

[0078] LPS least probable symbol

[0079] LSB least significant bit

[0080] LTRP long-term reference picture long-term reference picture

[0081] LMCS luma mapping with chroma scaling

[0082] MIP matrix-based intra prediction matrix-based intra prediction

[0083] MPS most probable symbol

[0084] MSB most significant bit

[0085] MTS multiple transform selection

[0086] MVP motion vector prediction motion vector prediction

[0087] NAL network abstraction layer

[0088] OLS output layer set output layer set

[0089] OP operation point

[0090] OPI operating point information

[0091] P predictive

[0092] PH picture header

[0093] POC picture order count picture order count

[0094] PPS picture parameter set

[0095] PROF prediction refinement with optical flow

[0096] PT picture timing

[0097] PU picture unit

[0098] QP quantization parameter quantization parameter

[0099] RADL random access decodable leading (picture)

[0100] RASL random access skipped leading(picture)

[0101] RBSP raw byte sequence payload

[0102] RGB red,green,and blue

[0103] RPL reference picture list

[0104] SAO sample adaptive offset

[0105] SAR sample aspect ratio

[0106] SEI supplemental enhancement information

[0107] SH slice header

[0108] SLI subpicture level information

[0109] SODB string of data bits

[0110] SPS sequence parameter set

[0111] STRP short-term reference picture

[0112] STSA step-wise temporal sublayer access

[0113] TR truncated rice

[0114] VBR variable bit rate

[0115] VCL video coding layer

[0116] VPS video parameter set video parameter set

[0117] VSEI versatile supplemental enhancement information (Rec.ITU-T H.274|ISO / IEC 23002-7)

[0118] VUI video usability information

[0119] VVC versatile video coding (Rec.ITU-T H.266|ISO / IEC23090-3)

[0120] 3. Video Codec Introduction

[0121] 3.1. Video Codec Standards

[0122] Video codec standards have evolved primarily through the well-known ITU-T and ISO / IEC developments. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and these two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new approaches have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which ended on July 1, 2020.

[0123] The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple coded video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360-degree immersive media.

[0124] 3.2. File format standards

[0125] Media streaming applications are typically based on IP, TCP, and HTTP transport methods and often rely on file formats such as the ISO Base Media File Format (ISOBMFF) [7]. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats that utilize ISOBMFF and DASH, file format specifications specific to the video format, such as the AVC file format and the HEVC file format, will be needed to encapsulate the video content in ISOBMFF tracks and in DASH representations and segments. Important information about the video bitstream (e.g., profile, tier, and level) and many other information will need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g., selection of appropriate media segments for both initialization at the start of a streaming session and stream adaptation during a streaming session.

[0126] Similarly, in order to use a picture format utilizing ISOBMFF, a file format specification specific to an image format, such as the AVC image file format and the HEVC image file format, would be required.

[0127] The VVC video file format, a file format for storing VVC video content based on ISOBMFF, is currently under development by MPEG. The latest draft specification is MPEG output document N19454 ("Information technology—Coding of audio-visual objects—Part 15: Carriage of network abstraction layer (NAL) unitstructured video in the ISO base media file format—Amendment 2: Carriage of VVC and EVC in ISOBMFF," July 2020).

[0128] The VVC image file format, a file format based on ISOBMFF for storing image content using the VVC codec, is currently under development by MPEG. The latest draft specification for the VVC image file format is included in MPEG output document N19460 ("Information technology—High efficiency coding and media delivery inheterogeneous environments—Part 12: Image File Format—Amendment 3: Support for VVC, EVC, slideshows and other improvements," July 2020).

[0129] 3.3. Temporal Scalability Support in VVC

[0130] Like HEVC, VVC includes similar support for temporal scalability. This support includes signaling of the temporal ID in the NAL unit header, the restriction that pictures of a particular temporal sub-layer cannot be used as inter-frame prediction references by pictures of lower temporal sub-layers, the sub-bitstream extraction process, and the requirement that each sub-bitstream extraction output of the appropriate input must be a consistent bitstream. Media-aware network elements (MANEs) can utilize the temporal ID in the NAL unit header for stream adaptation purposes based on temporal scalability.

[0131] 3.4. Image Resolution Changes Within a Sequence in VVC

[0132] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC implements the ability to change the resolution of a picture at a position within a sequence without encoding an IRAP picture that is always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference picture used for inter prediction when it has a different resolution than the current picture being decoded.

[0133] To allow reuse of motion compensation modules from existing implementations, the scaling ratio is constrained to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference and current pictures.

[0134] RPR allows resolution changes without encoding or decoding IRAP pictures, which can result in momentary bitrate spikes in streaming or video conferencing scenarios, for example, to cope with changing network conditions. RPR can also be used in applications where zooming in on the entire video area or a region of interest is required. Negative zoom window offsets are allowed to support a wider range of zoom-based applications. Negative zoom window offsets also enable the extraction of sub-picture sequences from a multi-layer bitstream while maintaining the same zoom window as in the original bitstream for the extracted sub-bitstream.

[0135] Unlike spatial scalability in the scalable extension of HEVC, where picture resampling and motion compensation are applied in two different stages, RPR in VVC is performed as part of the same process at the block level, where the derivation of sample positions and motion vector scaling are performed during motion compensation.

[0136] To limit implementation complexity, picture resolution changes within a CLVS are not allowed when the pictures in the CLVS each have multiple sub-pictures. Furthermore, decoder-side motion vector refinement, bidirectional optical flow, and optical flow prediction refinement are not applied when RPR is used between the current picture and a reference picture. The collocated pictures used to derive temporal motion vector candidates are also restricted to have the same picture size, scaling window offset, and CTU size as the current picture.

[0137] To support RPR, some other aspects of the VVC design differ from HEVC. First, the picture resolution and the corresponding consistency and scaling windows are signaled in the PPS rather than in the SPS, where the maximum picture resolution and the corresponding consistency window are signaled. In applications, the maximum picture resolution with the corresponding consistency window offset in the SPS can be used as the expected or desired picture output size after cropping. Second, for a single-layer bitstream, each picture store (a time slot in the DPB used to store one decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.

[0138] 3.5. Multi-layer scalability support in VVC

[0139] The ability to use RPR in the VVC core design to perform inter-frame prediction based on reference pictures of different sizes than the current picture allows VVC to easily support bitstreams containing multiple layers of different resolutions, for example, two layers with standard definition and high definition resolutions respectively. In the VVC decoder, such functionality can be integrated without any additional signal processing level codec tools, because the upsampling function required for spatial scalability support can be provided by reusing the RPR upsampling filter. However, additional high-level syntax design is required to implement scalability support for the bitstream.

[0140] Scalability is supported in VVC, but only in multi-layer profiles. Unlike scalability support in any earlier video codec standards, including extensions to AVC and HEVC, VVC scalability has been designed to be as friendly as possible to single-layer decoder implementations. The decoding capabilities of multi-layer bitstreams are specified as if there was only a single layer in the bitstream. For example, the decoding capabilities, such as the DPB size, are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for single-layer bitstreams can decode multi-layer bitstreams without significant changes.

[0141] Compared to the designs of multi-layer extensions of AVC and HEVC, HLS aspects are significantly simplified at the expense of some flexibility. For example, 1) the IRAP AU needs to contain pictures of each layer present in the CVS, which avoids the need to specify the layer-by-layer start decoding process, and 2) a simpler design for POC signaling is included in VVC, rather than a complex POC reset mechanism, to ensure that the derived POC value is the same for all pictures in the AU.

[0142] Like in HEVC, information about layers and layer dependencies is included in the VPS. OLS information is provided for signaling which layers are included in the OLS, which layers are output, and other information such as PTL and HRD parameters associated with each OLS. Similar to HEVC, there are three operating modes to output all layers, output only the highest layer, or output specific indicated layers in custom output mode.

[0143] There are some differences between the OLS designs in VVC and HEVC. First, in HEVC, layer sets are signaled, then OLSs are signaled based on the layer sets, and for each OLS, the output layers are signaled. The design in HEVC allows for layers belonging to an OLS that are neither output layers nor required for decoding the output layers. In VVC, the design requires that any layer in an OLS is either an output layer or a layer required for decoding the output layers. Therefore, in VVC, the OLS is signaled by indicating the output layer of the OLS, and then the other layers belonging to the OLS are simply inferred by the layer dependencies indicated in the VPS. In addition, VVC requires that each layer be included in at least one OLS.

[0144] Another difference in the VVC OLS design is that, contrary to HEVC where the OLS consists of all NAL units belonging to the set of identified layers mapped to the OLS, VVC can exclude some NAL units belonging to non-output layers mapped to the OLS. More specifically, VVC's OLS consists of a set of layers mapped to the OLS, where the non-output layers only include IRAP or GDR pictures with ph_recovery_poc_cnt equal to 0 or pictures from sub-layers used for inter-layer prediction. This allows the optimal level value of a multi-layer bitstream to be indicated by considering only all "necessary" pictures of all sub-layers within the layers forming the OLS, where "necessary" here means required for output or decoding. Figure 8 An example of a two-layer bitstream is shown with vps_max_tid_il_ref_pics_plus1[1][0] equal to 0, i.e. a sub-bitstream for which only IRAP pictures from layer L0 are kept when extracting OLS2.

[0145] Taking into account some scenarios where it is beneficial to allow different RAP periods at different layers, similar to ABC and HEVC, AUs are allowed to have layers with unaligned RAPs. In order to more quickly identify RAPs in multi-layer bitstreams, i.e. AUs with RAPs at all layers, the Access Unit Delimiter (AUD) is extended with a flag indicating whether the AU is an IRAP AU or a GDR AU, compared to HEVC. In addition, when a VPS indicates multiple layers, the AUD is mandatory to be present in such IRAP or GDR AU. However, for single-layer bitstreams indicated by a VPS or bitstreams that do not involve a VPS, the AUD is completely optional as in HEVC, because in this case the RAP can be easily detected from the first slice of the NAL unit type in the AU and the corresponding parameter set.

[0146] In order to enable sharing of SPS, PPS and APS by multiple layers and at the same time ensure that the bitstream extraction process does not discard parameter sets required by the decoding process, the VCL NAL unit of the first layer can refer to the SPS, PPS or APS with the same or lower layer ID value, as long as all OLSs including the first layer also include the layer identified by the lower layer ID value.

[0147] 3.6. Some details of the VVC video file format

[0148] 3.6.1. Overview of VVC Storage with Multiple Layers

[0149] Support for VVC bitstreams with multiple layers includes multiple tools, and there are various "models" for how they can be used. A VVC stream with multiple layers can be placed in a track in a number of ways, including the following:

[0150] 1. All layers are in one track, where all layers correspond to operating points;

[0151] 2. All layers are in one track, where there is no operation point that contains all layers;

[0152] 3. One or more layers or sublayers are in separate tracks, where the bitstream containing all samples of the indicated track or tracks corresponds to the operation point;

[0153] 4. One or more layers or sub-layers are in separate tracks, where there is no operation point that contains all NAL units of a set of one or more tracks.

[0154] The VVC file format allows one or more layers to be stored in a track. The storage of multiple layers per track can be used. For example, when a content provider wants to provide a multi-layer bitstream that is not intended for subsetting, or when a bitstream has been created for several predefined sets of output layers, where each layer corresponds to a view (e.g., a stereo pair), tracks can be created accordingly.

[0155] When a VVC bitstream is represented by multiple tracks and a player uses an operation point whose layers are stored in multiple tracks, the player has to reconstruct VVC access units before passing them to the VVC decoder.

[0156] A VVC operation point can be explicitly represented by a track, i.e., each sample in the track contains an access unit either locally or by parsing the "subp" track reference (when present) and by parsing the "vvcN" track reference (when present). An access unit contains NAL units from all layers and sublayers that are part of the operation point.

[0157] The storage of VVC bitstreams is supported by structures such as:

[0158] a) Sample entry,

[0159] b) Operation point information ("vopi") sample group,

[0160] c) Layer information ('linf') sample group,

[0161] d) Operation Point Entity Group ("opeg").

[0162] The structure within a SampleEntry provides information for the decoding or use of the sample, in this case the codec video and non-VCL data information associated with the SampleEntry.

[0163] The operation point information sample group records information about the operation point, such as the layers and sub-layers constituting the operation point, the dependencies between them (if any), the profile, layer and level parameters of the operation point, and other such operation point related information.

[0164] The layer information sample group lists all layers and sub-layers carried in the samples of the track.

[0165] The operation point entity group records information about the operation point, such as the layers and sub-layers that constitute the operation point, the dependencies between them (if any), the profile, layer and level parameters of the operation point, and other such operation point related information, as well as the track identifier that carries each operation point.

[0166] The information in these sample groups combined with the information in the track reference to look up the track or operation point entity group is sufficient for the reader to select an operation point according to its capabilities, identify the tracks containing the relevant layers and sub-layers required to decode the selected operation point, and extract them efficiently.

[0167] 3.6.2. Data Sharing and Reconstructing VVC Bitstream

[0168] 3.6.2.1. Overview

[0169] In order to reconstruct an access unit based on samples of multiple tracks carrying multi-layer VVC bitstreams, it is necessary to first determine the operating point.

[0170] NOTE: When a VVC bitstream is represented by multiple VVC tracks, the file parser can identify the required track for the selected operating point as follows:

[0171] Find all tracks with VVC sample entries.

[0172] If the track contains an "oref" track reference of the same ID, then that ID is resolved to a VVC track or an "opeg" entity group.

[0173] Select an operating point from the "opeg" entity group or the "vopi" sample group that is appropriate for the decoding capabilities and application purpose.

[0174] When the "opeg" entity group is present, it indicates that the track set exactly represents the selected operating point. Therefore, the VVC bitstream can be reconstructed and decoded based on the track set.

[0175] When the "opeg" entity group does not exist (ie when the "vopi" sample group exists), it is discovered from the "vopi" and "linf" sample groups which track set is needed to decode the selected operation point.

[0176] In order to reconstruct a bitstream based on multiple VVC tracks carrying VVC bitstreams, it may be necessary to first determine a target highest value TemporalId.

[0177] If several tracks contain data for an access unit, alignment of corresponding samples in the tracks is performed based on the sample decoding times, i.e. using time-sample tables without taking edit lists into account.

[0178] When a VVC bitstream is represented by multiple VVC tracks, the decoding times of the samples should be such that if the tracks are combined into a single stream ordered by increasing decoding time, the access unit order will be correct, as specified in ISO / IEC 23090-3.

[0179] The sequence of access units is reconstructed from the corresponding samples in the desired track according to the implicit reconstruction procedure described in 3.6.2.2.

[0180] 3.6.2.2. Implicit Reconstruction of VVC Bitstream

[0181] When an operation point information sample group is present, the required tracks are selected based on the layers they carry and their reference layers, as indicated by the operation point information and layer information sample groups.

[0182] When an operating point entity group exists, the required track is selected based on the information in the OperatingPointGroupBox.

[0183] When reconstructing a bitstream of a sub-layer containing VCL NAL units with TemporalId greater than 0, all lower sub-layers within the same layer (i.e., those layers with VCL NAL units with smaller TemporalId) are also included in the resulting bitstream and the required tracks are selected accordingly.

[0184] When reconstructing an access unit, picture units from samples with the same decoding time (as specified in ISO / IEC 23090-3) are placed into the access unit in increasing order of nuh_layer_id values.

[0185] When reconstructing an access unit with dependent layers and max_tid_il_ref_pics_plus1 is greater than 0, the sub-layers of the reference layers whose VCL NAL units within the same layer have a TemporalId less than or equal to max_tid_il_ref_pics_plus1-1 (as indicated in the operation point information sample group) are also included in the resulting bitstream, and the required tracks are selected accordingly.

[0186] When reconstructing an access unit with dependent layers and max_tid_il_ref_pics_plus1 is equal to 0, only IRAP picture units of the reference layer are included in the resulting bitstream, and the required track is selected accordingly.

[0187] If the VVC track contains a "subp" track reference, each picture unit is reconstructed as specified in clause 11.7.3, with additional constraints on EOS and EOB NAL units as specified below. The process in clause 11.7.3 is repeated for each layer of the target operation point in increasing nuh_layer_id order. Otherwise, each picture unit is reconstructed as described below.

[0188] The reconstructed access units are placed into the VVC bitstream in increasing order of decoding time, and duplication of end-of-bitstream (EOB) and end-of-sequence (EOS) NAL units is removed from the VVC bitstream, as described further below.

[0189] For access units within the same coded video sequence of a VVC bitstream and belonging to different sub-layers stored in multiple tracks, there may be more than one track containing an EOS NAL unit with a specific nuh_layer_id value in the corresponding sample. In this case, only one EOS NAL unit should be kept in the final reconstructed bitstream in the last access unit (the one with the longest decoding time) of these access units, placed after all NAL units of the last access unit of these access units except the EOB NAL unit (when present), and the other EOS NAL units are discarded. Similarly, there may be more than one such track containing EOB NAL units in the corresponding sample. In this case, only one EOB NAL unit should be kept in the final reconstructed bitstream, placed at the end of the last access unit of these access units, and the other EOB NAL units are discarded.

[0190] Since a particular layer or sub-layer may be represented by more than one track, when computing the required tracks for an operation point, it may be necessary to select among the set of tracks that all carry the particular layer or sub-layer.

[0191] When the operation point entity group does not exist, after selecting among the tracks carrying the same layers or sub-layers, the final required tracks may still jointly carry some layers or sub-layers that do not belong to the target operation point. The reconstructed bitstream of the target operation point shall not contain the layers or sub-layers carried by the final required tracks but not belonging to the target operation point.

[0192] NOTE: The VVC decoder implementation takes as input the bitstream corresponding to the target output layer set index and the highest TemporalId value of the target operation point, which correspond to the TargetOlsIdx and HighestTid variables in clause 8 of ISO / IEC 23090-3, respectively. The file parser needs to ensure that the reconstructed bitstream does not contain any other layers and sub-layers than those included in the target operation point before sending it to the VVC decoder.

[0193] 3.6.3. Operation point information sample group

[0194] 3.6.3.1. Definition

[0195] By using the Operation Point Information sample group ("vopi"), the application is informed of the different operation points provided by a given VVC bitstream and their composition. Each operation point is associated with a set of output layers, a maximum TemporalId, and profile, level, and layer signaling. All of this information is captured by the "vopi" sample group. In addition to this information, the sample group also provides information about the dependencies between layers.

[0196] When there is more than one VVC track for a VVC bitstream and no operation point entity groups exist for the VVC bitstream, both of the following apply:

[0197] - Among the VVC tracks in a VVC bitstream, there shall be one and only one track carrying the "vopi" sample group.

[0198] - All other VVC tracks of a VVC bitstream shall have a track reference of type "oref" to the track carrying the "vopi" sample group.

[0199] For any particular sample in a given track, a temporally collocated sample in another track is defined as a sample whose decoding time is the same as that of the particular sample. k The orbital reference of the "oref" orbital T N Each sample point S in N , the following apply:

[0200] - If on track T k There are time domain collocated samples S k , then the sample point S N and sample point S k The "vopi" sample group entry of the same "vopi" sample group entry is associated with the same "vopi" sample group entry.

[0201] - Otherwise, sample point S N and track T kIn the decoding time, the sample point S N The "vopi" sample group entry before the last sample point is associated with the same "vopi" sample group entry.

[0202] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the Sample Group Description box with grouping_type "vopi". For the more common case where there is a single VPS, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the operating point information sample group in the Sample Table box rather than in each track fragment.

[0203] The grouping_type_parameter is not defined for SampleToGroupBox with grouping type 'vopi'.

[0204] 3.6.3.2. Syntax

[0205]

[0206]

[0207]

[0208] 3.6.3.3. Semantics

[0209] num_profile_tier_level_minus1 plus 1 gives the number of profile, tier, and level combinations and related fields below.

[0210] ptl_max_temporal_id[i]: gives the maximum TemporalId of the NAL units of the associated bitstream of the specified i-th profile, layer and level structure.

[0211] NOTE: The semantics of ptl_max_temporal_id[i] and max_temporal_id for the operating points given below are different, although they may carry the same value.

[0212] ptl[i] specifies the i-th profile, layer, and level structure.

[0213] ISO / IEC 23090-3 defines all_independent_layers_flag, each_layer_is_an_ols_flag, ols_mode_idc, and max_tid_il_ref_pics_plus1.

[0214] num_operating_points: gives the number of operating points that follow the information.

[0215] output_layer_set_idx is the index of the output layer set that defines the operation point. The mapping between output_layer_set_idx and layer_id values ​​should be the same as the mapping specified by the VPS for the output layer set indexed by output_layer_set_idx.

[0216] ptl_idx: Signals the zero-based index of the listed profile, level and layer structure of the output layer set indexed by output_layer_set_idx.

[0217] max_temporal_id: gives the maximum TemporalId of the NAL units of this operation point.

[0218] NOTE: The maximum TemporalId value indicated in the Layer Information Sample Group has different semantics from the maximum TemporalId indicated here. However, they may carry the same literal value.

[0219] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.

[0220] layer_id: nuh_layer_id value of the layer providing the operation point.

[0221] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.

[0222] A frame_rate_info_flag equal to 0 indicates that no frame rate information is present for the operation point. A value of 1 indicates that frame rate information is present for the operation point.

[0223] bit_rate_info_flag equal to 0 indicates that no bit rate information is present for the operation point. A value of 1 indicates that bit rate information is present for the operation point.

[0224] avgFrameRate gives the average frame rate at the operating point in frames / (256 seconds). A value of 0 indicates that the average frame rate is not specified.

[0225] constantFrameRate equal to 1 indicates that the stream of the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream of the operation point has a constant frame rate. A value of 0 indicates that the stream of the operation point may or may not have a constant frame rate.

[0226] maxBitRate gives the maximum bit rate of the stream at any operating point within any window of one second, in bits per second.

[0227] avgBitRate gives the average bit rate of the stream at the operating point in bits per second.

[0228] max_layer_count: Count of all unique layers in all operation points related to this associated base track.

[0229] layerID: The nuh_layer_id of the layer for which all direct reference layers are given in the direct_ref_layerID loop below.

[0230] num_direct_ref_layers: The number of direct reference layers of the layer whose nuh_layer_id is equal to layerID.

[0231] direct_ref_layerID: nuh_layer_id of the direct reference layer.

[0232] 3.6.4. Layer Information Sample Group

[0233] The list of layers and sub-layers carried by a track is signaled in a layer information sample group. When there is more than one VVC track for the same VVC bitstream, each of these VVC tracks shall carry a "linf" sample group.

[0234] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the Sample Group Description box with grouping_type "linf". For the more common case where there is a single VPS, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the layer information sample groups in the Sample Table box rather than in each track fragment.

[0235] The grouping_type_parameter is not defined for SampleToGroupBox with grouping type 'linf'.

[0236] The syntax and semantics of the "linf" sample group are specified in clauses 9.6.3.2 and 9.6.3.3 respectively.

[0237] 3.6.5. Operation Point Entity Group

[0238] 3.6.5.1. Overview

[0239] An operation point entity group is defined to provide mapping of tracks to operation points and profile level information of the operation points.

[0240] When aggregating samples of tracks mapped to an operation point described in this entity group, the implicit reconstruction process does not require removing any further NAL units to produce a consistent VVC bitstream. Tracks belonging to an operation point entity group shall have a track reference of type "oref" to the group_id indicated in the operation point entity group.

[0241] All entity_id values ​​included in an operating point entity group shall belong to the same VVC bitstream.When present, the OperatingPointGroupBox shall be included in the GroupsListBox of the movie-level MetaBox, but shall not be included in the file-level or track-level MetaBox.

[0242] 3.6.5.2. Syntax

[0243]

[0244]

[0245] 3.6.5.3. Semantics

[0246] num_profile_tier_level_minus1 plus 1 gives the number of profile, tier, and level combinations and related fields below.

[0247] opeg_ptl[i] specifies the i-th profile, layer and level structure.

[0248] num_operating_points: gives the number of operating points that follow the information.

[0249] output_layer_set_idx is the index of the output layer set that defines the operation point. The mapping between output_layer_set_idx and layer_id values ​​should be the same as the mapping specified by the VPS for the output layer set indexed by output_layer_set_idx.

[0250] ptl_idx: Signals the zero-based index of the listed profile, level and layer structure of the output layer set indexed by output_layer_set_idx.

[0251] max_temporal_id: gives the maximum TemporalId of the NAL unit of this operation point.

[0252] NOTE: The maximum TemporalId value indicated in the Layer Information Sample Group has different semantics from the maximum TemporalId indicated here. However, they may carry the same literal value.

[0253] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.

[0254] layer_id: nuh_layer_id value of the layer providing the operation point.

[0255] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.

[0256] A frame_rate_info_flag equal to 0 indicates that no frame rate information is present for the operation point. A value of 1 indicates that frame rate information is present for the operation point.

[0257] bit_rate_info_flag equal to 0 indicates that no bit rate information is present for the operation point. A value of 1 indicates that bit rate information is present for the operation point.

[0258] avgFrameRate gives the average frame rate at the operating point in frames / (256 seconds). A value of 0 indicates that the average frame rate is not specified.

[0259] constantFrameRate equal to 1 indicates that the stream of the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream of the operation point has a constant frame rate. A value of 0 indicates that the stream of the operation point may or may not have a constant frame rate.

[0260] maxBitRate gives the maximum bit rate of the stream at any operating point within any window of one second, in bits per second.

[0261] avgBitRate gives the average bit rate of the stream at the operating point in bits per second.

[0262] entity_count specifies the number of tracks present in the operating point.

[0263] entity_idx specifies the index into the entity_id list in the entity group that belongs to the operation point.

[0264] 4. Examples of technical problems solved by the disclosed solutions

[0265] The latest design of the VVC video file format for storage of scalable VVC bitstreams has the following issues:

[0266] 1) When a VVC bitstream is represented by multiple VVC tracks, the file parser can identify the tracks required for the selected operation point by first finding all tracks with VVC sample entries, then finding all tracks containing the "vopi" sample group, and so on, to find information for all operation points provided in the file. However, finding all these tracks can be quite complex.

[0267] 2) "linf" sample group signaling information about which layers and / or sublayers are included in a track. When using the "vopi" sample group to select an OP, it is necessary to find all tracks containing the "linf" sample group, and the information carried in the "linf" sample group entries of these tracks is used together with the information in the "vopi" sample group entries on the OP's layers and / or sublayers to find the desired track. This can also be quite complex.

[0268] 3) In the "vopi" sample group entry, layer dependency information is signaled even when the value of all_independent_layers_flag is equal to 1. However, when all_independent_layers_flag is equal to 1, the layer dependency information is already known, so the bits used for signaling in this case are wasted.

[0269] 4) In the process of implicitly reconstructing the VVC bitstream from multiple tracks, the removal of redundant EOS and EOB NAL units is specified. However, in this process, it may be necessary to remove and / or rewrite AUD NAL units, but there is no corresponding process.

[0270] 5) It is specified that when the "opeg" entity group is present, the bitstream is reconstructed from multiple tracks by including all NAL units of the required tracks without removing any NAL units. However, for example, this will not allow NAL units (such as AUD, EOS, and EOB NAL units for a particular AU) to be included in more than one track carrying a VVC bitstream.

[0271] 6) The container of the "opeg" entity group box is specified as a movie-level MetaBox. However, the entity_id value of an entity group can refer to a track ID only when it is contained in a file-level MetaBox.

[0272] 7) In the "opeg" entity group box, the field output_layer_set_idx is always signaled for each OP. However, if the OP contains only one layer, it is usually not necessary to know the value of the OLS index, and even if it is useful to know the OLS index, it can be easily derived as the OLS index of the OLS containing only that layer.

[0273] 8) It is allowed that when there is an "opeg" entity group for a VVC bitstream, one of the tracks representing the VVC bitstream may have a "vopi" sample group. However, allowing both is unnecessary and would only increase the file size unnecessarily and confuse the file parser as to which one should be used.

[0274] 5. List of solutions

[0275] In order to solve the above problems, the following methods are disclosed. These items should be regarded as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these items can be applied alone or in any combination.

[0276] 1) To address issue 1, one or more of the following items are proposed:

[0277] a. Add signaling of file-level information about all OPs provided in the file, including the tracks required by each OP. The tracks required by an OP may carry layers or sub-layers not included in the OP.

[0278] b. Add signaling of file level information about which tracks contain the "vopi" sample groups.

[0279] i. In one example, a new box is specified, for example, named as the operation point information track box, with the file-level MetaBox as the container, for signaling the track carrying the "vopi" sample group.

[0280] A piece of file level information can be signaled in a file level box or a movie level box, or in a track level box, but the position of the track level box is identified in the file level box or the movie level box.

[0281] 2) To address issue 2, one or more of the following items are proposed:

[0282] a. Add information about the tracks required for each OP in the "vopi" sample group entry.

[0283] b. The use of the “linf” sample group is deprecated.

[0284] 3) To solve problem 3, when all_independent_layers_flag is equal to 1, signaling the layer dependency information in the "vopi" sample group entry is skipped.

[0285] 4) To solve problem 4, an operation for removing redundant AUD NAL units is added during the process of implicitly reconstructing the VVC bitstream based on multiple tracks.

[0286] a. Alternatively, when necessary, an operation for rewriting the AUD NAL unit is further added.

[0287] i. In one example, it is specified that, when an access unit is reconstructed from multiple picture units from different tracks, when the aud_irap_or_gdr_flag of the AUD NAL unit retained in the reconstructed access unit is equal to 1 and the reconstructed access unit is not an IRAP or GDR access unit, the value of aud_irap_or_gdr_flag of the AUD NAL unit is set equal to 0.

[0288] ii. The aud_irap_or_gdr_flag of the AUD NAL unit in the first PU may be equal to 1, while another PU of the same access unit but in a separate track has a picture that is not an IRAP or GDR picture. In this case, the value of aud_irap_or_gdr_flag of the AUD NAL unit in the reconstructed access unit is changed from 1 to 0.

[0289] b. In one example, additionally or alternatively, it is specified that when at least one picture unit among multiple picture units from different tracks of an access unit has an AUD NAL unit, the first picture unit (i.e., the picture unit with the smallest nuh_layer_id value) should have an AUD NAL unit.

[0290] c. In one example, it is specified that when multiple picture units from different tracks of an access unit have AUD NAL units present, only the AUD NAL unit in the first picture unit is kept in the reconstructed access unit.

[0291] 5) To address issue 5, one or more of the following items were proposed:

[0292] a. Specifies that when the "opeg" entity group is present and used, the required track provides the exact set of VCL NAL units required for each OP, but some non-VCL NAL units may become redundant in the reconstructed bitstream and therefore may need to be removed.

[0293] i. Alternatively, it is specified that when the "opeg" entity group is present and used, the required tracks provide the exact set of layers and sub-layers required for each OP, but some non-VCL NAL units may become redundant in the reconstructed bitstream and therefore may need to be removed.

[0294] b. During implicit reconstruction of a VVC bitstream from multiple tracks, the operation for removing redundant EOB and EOS NAL units is applied even when the “opeg” entity group exists and is used.

[0295] c. In the process of implicitly reconstructing a VVC bitstream from multiple tracks, even when the "opeg" entity group exists and is used, an operation for removing redundant AUD units and an operation for rewriting redundant AUD units are applied.

[0296] 6) To solve problem 6, specify the container of the "opeg" entity group box as the GroupsListBox in the file level MetaBox as follows: When present, the OperatingPointGroupBox should be contained in the GroupsListBox in the file level MetaBox and should not be contained in other level MetaBoxes.

[0297] 7) To address issue 7, when an OP contains only one layer, the signaling of the VvcOperatingPointsRecord present in the "vopi" point group entry and the output_layer_set_idx field in the "opeg" entity group box OperatingPointGroupBox is skipped for the OP.

[0298] a. In one example, the output_layer_set_idx field in VvcOperatingPointsRecord and / or OperatingPointGroupBox is moved after the loop that signals layer_id and is conditional on "if (layer_count>1)".

[0299] b. Furthermore, in one example, it is specified that when output_layer_set_idx does not exist for an OP, its value is inferred to be equal to the OLS index of the OLS containing only the layers in that OP.

[0300] 8) To address issue 8, it is specified that, when an "opeg" entity group exists for a VVC bitstream, none of the tracks representing the VVC bitstream should have a "vopi" sample group.

[0301] a. Alternatively, allow both to exist, but require that when both exist, they are consistent so that choosing either one makes no difference.

[0302] b. In one example, it is specified that tracks belonging to the "opeg" entity group (all of which have a track reference of type "oref" to the group_id indicated in the entity group) should not carry the "vopi" sample group.

[0303] 9) It may be specified that a VVC bitstream is not allowed to have an "opeg" entity group or a "vopi" sample group when the VVC bitstream is represented in only one track.

[0304] 6. Example of embodiment

[0305] Below are some example embodiments of some of the inventive aspects outlined in Section 5 above, which can be applied to the standard specification of the VVC video file format. The revised text is based on the latest draft specification of VVC. Most of the relevant parts that have been added or modified are Bold Underline Highlighted, and some deleted parts are marked with [[ Bold Italic ]] is highlighted. There may also be some other changes of an editorial nature, and therefore not highlighted.

[0306] 6.1. First embodiment

[0307] This embodiment targets items 2a, 3, 7, 7a, and 7b.

[0308] 6.1.1. Operation point information sample group

[0309] 6.1.1.1. Definition

[0310] By using the Operation Point Information sample group ("vopi"), the application is informed of the different operation points provided by a given VVC bitstream and their composition. Each operation point is associated with a set of output layers, a maximum TemporalId value, and profile, level, and layer signaling. All of this information is captured by the "vopi" sample group. In addition to this information, the sample group also provides information about the dependencies between layers.

[0311] When there is more than one VVC track for a VVC bitstream and no operation point entity groups exist for the VVC bitstream, both of the following apply:

[0312] - Among the VVC tracks in a VVC bitstream, there shall be one and only one track carrying the "vopi" sample group.

[0313] - All other VVC tracks of a VVC bitstream shall have a track reference of type "oref" to the track carrying the "vopi" sample group.

[0314] For any particular sample in a given track, a temporally collocated sample in another track is defined as a sample whose decoding time is the same as that of the particular sample. k The orbital reference of the "oref" orbital T N Each sample point S in N , the following apply:

[0315] - If on track T k There are time domain collocated samples S k , then the sample point S N and sample point S k The "vopi" sample group entry of the same "vopi" sample group entry is associated with the same "vopi" sample group entry.

[0316] - Otherwise, sample point S N and track T k In the decoding time, the sample point S N The "vopi" sample group entry before the last sample point is associated with the same "vopi" sample group entry.

[0317] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the Sample Group Description box with grouping_type "vopi". For the more common case where there is a single VPS, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the operating point information sample groups in the Sample Table box rather than in each track fragment.

[0318] The grouping_type_parameter is not defined for SampleToGroupBox with grouping type 'vopi'.

[0319] 6.1.1.2. Syntax

[0320]

[0321]

[0322]

[0323] 6.1.1.3. Semantics

[0324]

[0325] num_operating_points: gives the number of operating points that follow the information.

[0326] [[output_layer_set_idx is the index of the output layer set that defines the operation point. The mapping between output_layer_set_idx and layer_id values ​​shall be the same as that specified by the VPS for the output layer set indexed by output_layer_set_idx. ]]

[0327] ptl_idx: Signals the zero-based index of the listed profile, level and layer structure of the output layer set indexed by output_layer_set_idx.

[0328] max_temporal_id: gives the maximum TemporalId of the NAL unit of this operation point.

[0329] NOTE: The maximum TemporalId value indicated in the Layer Information Sample Group has different semantics from the maximum TemporalId indicated here. However, they may carry the same literal value.

[0330] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.

[0331] layer_id: nuh_layer_id value of the layer providing the operation point.

[0332] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.

[0333] output_layer_set_idx is the index of the output layer set that defines the operation point. The mapping between idx and layer_id values ​​should be consistent with the output layer set indexed by output_layer_set_idx in the VPS. When output_layer_set_idx does not exist for an OP, its value is inferred to be equal to the value containing only that OP. OLS index for the OLS of the layers in .

[0334] op_track_count specifies the number of tracks carrying VCL NAL units in this operation point.

[0335] op_track_id[j] specifies the jth track in the track that carries the VCL NAL unit at this operation point. track_ID value.

[0336] A frame_rate_info_flag equal to 0 indicates that no frame rate information is present for the operation point. A value of 1 indicates that frame rate information is present for the operation point.

[0337] bit_rate_info_flag equal to 0 indicates that no bit rate information is present for the operation point. A value of 1 indicates that bit rate information is present for the operation point.

[0338] avgFrameRate gives the average frame rate at the operating point in frames / (256 seconds). A value of 0 indicates that the average frame rate is not specified.

[0339] constantFrameRate equal to 1 indicates that the stream of the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream of the operation point has a constant frame rate. A value of 0 indicates that the stream of the operation point may or may not have a constant frame rate.

[0340] maxBitRate gives the maximum bit rate of the stream at any operating point within any window of one second, in bits per second.

[0341] avgBitRate gives the average bit rate of the stream at the operating point in bits per second.

[0342] max_layer_count: Count of all unique layers in all operation points related to this associated base track.

[0343] layerID: The nuh_layer_id of the layer for which all direct reference layers are given in the direct_ref_layerID loop below.

[0344] num_direct_ref_layers: The number of direct reference layers of the layer whose nuh_layer_id is equal to layerID.

[0345] direct_ref_layerID: nuh_layer_id of the direct reference layer.

[0346] 6.2. Second embodiment

[0347] This embodiment targets items 1.bi, 4, 4a, 4.ai, 4b, 4c, 5a, 6, 8, and 8b.

[0348] Implicit reconstruction of VVC bitstream:

[0349] When an operation point information sample group is present, the required tracks are selected based on the layers they carry and their reference layers, as indicated by the operation point information and layer information sample groups.

[0350] When an operating point entity group exists, the required track is selected based on the information in the OperatingPointGroupBox.

[0351] When reconstructing a bitstream of a sub-layer containing VCL NAL units with TemporalId greater than 0, all lower sub-layers within the same layer (i.e., those layers with smaller TemporalId of VCL NAL units) are also included in the resulting bitstream, and the required tracks are selected accordingly.

[0352] When reconstructing an access unit, picture units from samples with the same decoding time (as specified in ISO / IEC 23090-3) are placed into the access unit in increasing order of nuh_layer_id values. When accessing multiple units When at least one picture unit in the picture units has an AUD NAL unit, the first picture unit (i.e., the picture unit with the smallest nuh_ The picture unit of layer_id value shall have an AUDNAL unit, and only the AUDNAL unit in the first picture unit shall be The AUDNAL units are kept in the reconstructed access unit, while other AUDNAL units (when present) are discarded. In the unit, when the aud_irap_or_gdr_flag of the AUDNAL unit is equal to 1 and the reconstructed access unit is not an IRAP or When the GDR access unit is used, the value of aud_irap_or_gdr_flag of the AUDN A L unit is set to 0.

[0353] NOTE 1: The aud_irap_or_gdr_flag of the AUD NAL unit in the first PU may be equal to 1, while the same access unit Another PU of the same sub-unit but in a separate track has a picture that is not an IRAP or GDR picture. In this case, the reconstructed The value of aud_irap_or_gdr_flag of the AUD NAL unit in an access unit is changed from 1 to 0.

[0354]

[0355] Entity Group and other file-level information :

[0356] Sub-image entity group:

[0357]

[0358] Operation point entity group:

[0359] Overview:

[0360] An operation point entity group is defined to provide mapping of tracks to operation points and profile level information of the operation points.

[0361] When aggregating the samples of the track that map to the operating points described in this entity group, the implicit reconstruction process does not require removing any further VCL NAL units to produce a consistent VVC bitstream. Tracks belonging to an operation point entity group shall have a track reference of type "oref" to the group_id indicated in the operation point entity group [[track reference]] , and should not Carry "vopi" sample group .

[0362] All entity_id values ​​included in an operating point entity group shall belong to the same VVC bitstream. When present, the OperatingPointGroupBox shall be included in the [[movie]] document The GroupsListBox in the level MetaBox should not be included Other levels [[File level or track level]] MetaBox.

[0363] Operating point information track box

[0364] definition:

[0365] Box type: "topi"

[0366] Container: File-level MetaBox

[0367] Mandatory: No

[0368] Quantity: 0 or 1

[0369] The Operation Point Information Track Box contains the track ID of the track set that carries the "vopi" sample group. The absence of this box indicates There is no track with "vopi" sample group in the file.

[0370] grammar

[0371]

[0372] Semantics

[0373] num_tracks_with_vopi specifies the number of tracks with "vopi" sample groups in the file.

[0374] track_ID[i] specifies the track ID of the i-th track carrying the "vopi" sample group.

[0375] Figure 1 1 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, a passive optical network (PON), etc.) and wireless interfaces (such as a Wi-Fi interface or a cellular interface).

[0376] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or transmitted via a connected communication, as represented by component 1906. Component 1908 can use the bitstream (or codec) representation of the stored or transmitted video received at input 1902 to generate pixel values ​​or displayable video to be sent to display interface 1910. The process of generating user-viewable video based on the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the results of the codec will be performed by the decoder.

[0377] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0378] Figure 2 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuits. In some embodiments, video processing hardware 3606 can be at least partially included in processor 3602 (e.g., a graphics coprocessor).

[0379] Figure 4 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0380] like Figure 4 As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0381] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0382] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded and decoded pictures and associated data. The coded and decoded pictures are coded and decoded representations of the pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the network 130a via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0383] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0384] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 and configured to interface with an external display device.

[0385] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.

[0386] Figure 5 is a block diagram illustrating an example of a video encoder 200, which may be Figure 4 The video encoder 114 in the system 100 is shown.

[0387] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 5 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0388] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0389] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.

[0390] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are shown in FIG. Figure 5 are represented separately in the examples.

[0391] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0392] The mode selection unit 203 can select one of the coding modes (e.g., intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).

[0393] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.

[0394] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0395] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0396] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures containing the reference video blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0397] In some examples, motion estimation unit 204 may output a complete motion information set for use in a decoding process by a decoder.

[0398] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0399] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0400] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0401] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0402] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0403] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0404] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0405] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0406] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0407] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0408] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0409] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0410] Figure 6 is a block diagram illustrating an example of a video decoder 300, which may be Figure 4 The video decoder 114 in the system 100 is shown.

[0411] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 6 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0412] exist Figure 6 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations generally related to the video encoder 200 ( Figure 5 ) describes the decoding pass that is the opposite of the encoding pass.

[0413] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by implementing AMVP and Merge modes.

[0414] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0415] The motion compensation unit 302 may calculate interpolated values ​​of sub-integer pixels of a reference block using the interpolation filter used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.

[0416] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.

[0417] The intra prediction unit 303 can form a prediction block based on spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0418] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.

[0419] A list of solutions preferred by some embodiments is provided below.

[0420] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (eg, items 1, 2).

[0421] 1. A method for processing visual media (e.g., Figure 3 ), comprising: performing (702) a conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the file includes file-level information for all operation points included in the file, wherein the file-level information includes information on tracks required for each operation point.

[0422] 2. The method of claim 1, wherein the format rules allow a track to include layers and sub-layers that are not required at the corresponding operation point.

[0423] 3. The method according to any one of schemes 1-2, wherein information of tracks required for each operation point is included in a vopi sample group entry.

[0424] The following scenario illustrates an example embodiment of the technique discussed in the previous section (eg, item 3).

[0425] 4. A method of visual media processing, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the format rules specify skipping layer dependency information from a vopi sample group entry if all layers are unrelated.

[0426] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (eg, items 5, 6).

[0427] 5. A method for processing visual media, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the format rules define rules associated with the processing of operation point entity groups (OPEGs) in the bitstream representation.

[0428] 6. The method of claim 5, wherein the format rules specify that, if an opeg is present, each required track in the file provides an exact set of video codec network abstraction layers (VCL NALs) corresponding to each operation point in the opeg.

[0429] 7. The method of claim 6, wherein the format rules allow for inclusion of non-VCL cells in the track.

[0430] The following scenario illustrates an example embodiment of the technique discussed in the previous section (eg, item 4).

[0431] 8. A method for processing visual media, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the conversion includes performing implicit reconstruction of the bitstream representation according to multiple tracks, wherein redundant access unit delimiter network access units (AUD NALs) are processed according to the rules.

[0432] 9. The method of claim 8, wherein the rule specifies removal of AUD NAL units.

[0433] 10. The method of claim 8, wherein the rule specifies rewriting of AUD NAL units.

[0434] 11. The method according to any one of schemes 8-10, wherein the rule specifies that if at least one picture unit of a plurality of picture units from different tracks of an access unit has an AUD NAL unit, the first picture unit has another AUD NAL unit.

[0435] 12. A method according to any one of schemes 8-10, wherein the rule stipulates that when multiple picture units from different tracks of an access unit have AUD NAL units present, during decoding, only the AUD NAL unit in the first picture unit is kept in the reconstructed access unit.

[0436] 13. The method according to any one of schemes 1-12, wherein converting comprises generating a bitstream representation of the visual media data and storing the bitstream representation in a file according to format rules.

[0437] 14. The method of any one of claims 1 to 12, wherein converting comprises parsing the file according to format rules to recover the visual media data.

[0438] 15. A video decoding device comprising a processor, the processor being configured to implement the method according to one or more of schemes 1 to 14.

[0439] 16. A video encoding apparatus comprising a processor configured to implement the method according to one or more of schemes 1 to 14.

[0440] 17. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of schemes 1 to 14.

[0441] 18. A computer-readable medium having a bitstream representation thereon conforming to a file format generated according to any one of schemes 1 to 14.

[0442] 19. A method, apparatus or system as described in this document.

[0443] Some preferred embodiments of the above-listed solutions may include the following (eg, items 1, 2).

[0444] In some embodiments, a method of processing visual media (e.g., Figure 9A The method depicted in claim 910 includes performing (912) conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in a plurality of tracks according to a format rule, the format rule specifying that file-level information includes syntax elements, the syntax elements identifying one or more tracks from the plurality of tracks that contain a particular type of sample group including operation point information.

[0445] In the above embodiment, the format rule specifies that the visual media file includes file-level information of all operation points provided in the visual media file, wherein the format rule further specifies that for each operation point, the file-level information includes information of the corresponding track in the visual media file.

[0446] In some embodiments, format rules allow tracks required for a particular operation point to include layers and sub-layers that are not required for the particular operation point.

[0447] In some embodiments, wherein the syntax element comprises a box containing a file-level container.

[0448] In some embodiments, the format rules specify that file-level information be included in file-level boxes.

[0449] In some embodiments, the file level information is specified in the format rules to be included in a movie level box.

[0450] In some embodiments, the format rules specify that file-level information be included in a track-level box, which is identified within another track-level box or another file-level box.

[0451] In some embodiments, the format rules also specify that a particular type of sample point group includes information about the tracks required for each operation point.

[0452] In some embodiments, the format rules further specify that information about tracks required for each operation point be omitted from another specific type of sample group that includes layer information about the number of layers in the bitstream.

[0453] In some embodiments, the visual media data is processed by the Versatile Video Codec (VVC), and the plurality of tracks are VVC tracks.

[0454] Some preferred embodiments may include the following (eg, item 3).

[0455] In some embodiments, a method of visual media processing (e.g., Figure 9B The method depicted in claim 920 includes performing (922) conversion between visual media data and a visual media file that stores a bitstream of the visual media data according to a format rule. The visual media file stores one or more tracks comprising one or more video layers. The format rule specifies that whether a first set of syntax elements indicating layer dependency information is stored in the visual media file depends on whether a second syntax element indicating that all layers in the visual media file are unrelated has a value of 1.

[0456] In some embodiments, the first set of syntax elements is stored in sample point groups that indicate information about one or more operation points stored in the visual media file.

[0457] In some embodiments, the format rule specifies that, in response to the second syntax element having a value of 1, the first set of syntax elements is omitted from the visual media file.

[0458] Some preferred embodiments of the above-listed solutions may include the following aspects (eg, item 4).

[0459] In some embodiments, a method of processing visual media data (e.g., Figure 9C The method depicted in 930) includes: performing (932) conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in a plurality of tracks according to format rules, the format rules specifying a manner in which redundant access unit delimiter network access layer (AUD NAL) units stored in the plurality of tracks are handled during implicit reconstruction of the bitstream from the plurality of tracks.

[0460] In some embodiments, the format rules specify that redundant AUD NAL units are removed during implicit reconstruction.

[0461] In some embodiments, the format rules specify that redundant AUD NAL units are rewritten during implicit reconstruction.

[0462] In some embodiments, the format rules specify that, in response to implicit reconstruction including generating a specific access unit having a specific type different from an instantaneous random access point type or a gradual decoding refresh type based on multiple pictures of multiple tracks, the syntax field of the specific redundant AUD NAL included in the specific access unit is rewritten to a 0 value, indicating that the specific redundant AUD NAL does not represent an instantaneous random access point or a gradual decoding refresh type.

[0463] In some embodiments, the format rules further specify that the value of the syntax element in the AUD NAL unit in the first picture unit (PU) is rewritten to 0, indicating that the particular AUD NAL does not represent an instantaneous random access point or gradual decoding refresh type in the case where the second PU from a different track includes a picture that is not an intra random access point picture or a gradual decoding refresh picture.

[0464] In some embodiments, the format rule specifies that at least one picture unit of a plurality of picture units from different tracks responsive to an access unit has a first AUD NAL unit, and the first picture unit of the access unit generated according to the implicit reconstruction includes a second AUD NAL unit.

[0465] In some embodiments, the format rules specify that in response to multiple picture units from different tracks of an access unit including AUD NAL units, a single AUD NAL unit corresponding to the AUD NAL unit of the first picture unit is included in the access unit generated according to the implicit reconstruction.

[0466] Some preferred embodiments of the above-listed solutions may include the following aspects (eg, item 7).

[0467] In some embodiments, a visual media processing method (e.g., Figure 9D The method 940 depicted in claim 1 includes performing (942) conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule. The visual media file stores one or more tracks comprising one or more video layers. The visual media file includes information about an operation point (OP); wherein the format rule specifies whether or how syntax elements are included in a sample group entry and a group box of the OP based on whether the OP comprises a single video layer; wherein the syntax elements are configured to indicate an index of an output layer set of the OP.

[0468] In some embodiments, the format rules specify that, in response to the OP containing a single video layer, syntax elements are omitted from sample group entries and group boxes.

[0469] In some embodiments, the format rules specify that, in response to the OP including more than one video layer, a syntax element is included following the information indicating identification of more than one video layer.

[0470] In some embodiments, in response to omitting syntax elements from sample group entries and group bins, the index of the output layer set of the OP is inferred to be equal to the index of the output layer set comprising a single video layer.

[0471] Some preferred embodiments of the above-listed solutions may be combined with the following aspects (eg, item 8).

[0472] In some embodiments, a visual media processing method (e.g., Figure 9E The method 950 depicted in claim 1 includes performing (952) conversion between visual media data and a visual media file that stores a bitstream of the visual media data according to a format rule; wherein the visual media file stores a plurality of tracks that belong to a particular type of entity group, and wherein the format rule specifies that, in response to the plurality of tracks having a particular type of track reference to a group identifier, the plurality of tracks (A) omit carrying a sample group of the particular type or (B) carry a sample group of the particular type such that information in the sample group of the particular type is consistent with information in the entity group of the particular type.

[0473] In some embodiments, multiple tracks represent a bitstream.

[0474] In some embodiments, a particular type of entity group indicates that multiple tracks exactly correspond to an operation point.

[0475] In some embodiments, a particular type of sample point group includes information about which tracks of a plurality of tracks correspond to an operation point.

[0476] Some preferred embodiments of the above-listed solutions may be combined with the following aspects (eg, items 5, 6, 9).

[0477] In some embodiments, a visual media processing method (e.g., Figure 9F The method 960 depicted in includes: performing (962) conversion between visual media data and a visual media file storing a bitstream of the visual media data, wherein the visual media file includes a plurality of tracks; wherein the visual media file stores entity groups that carry information about operation points in the visual media file and tracks that carry each operation point; and wherein the format rule specifies properties of the visual media file responsive to the visual media file storing entity groups or sample groups that carry information about each operation point.

[0478] In some embodiments, the format rules specify that the entity group provide a set of tracks that carry an exact set of Video Codec Layer (VCL) Network Abstraction Layer (NAL) units for each operation point.

[0479] In some embodiments, the format rules further specify that redundant non-VCL NAL units in a track set are removed during reconstruction of the bitstream.

[0480] In some embodiments, the format rules specify that the entity group provide a set of tracks carrying an exact set of one or more layers and one or more sub-layers for each operation point.

[0481] In some embodiments, the format rules further specify that redundant non-Video Codec Layer (VCL) Network Abstraction Layer (NAL) units in the track set are removed during reconstruction of the bitstream.

[0482] In some embodiments, the format rules specify that redundant end-of-bitstream (EOB) or end-of-stream (EOS) network abstraction layer (NAL) units are removed during reconstruction of the bitstream from the multiple tracks.

[0483] In some embodiments, the format rules specify that access delimiter units (AUDs) are removed or overwritten during reconstruction of the bitstream from the multiple tracks.

[0484] In some embodiments, the format rules specify a property that does not allow containers of entity group boxes associated with an entity group to be stored in the visual media file at any level other than pre-specified file-level boxes.

[0485] In some embodiments, the pre-designated file level box is a group list box included in the file level metadata box.

[0486] In some embodiments, the format rules specify that, in response to the bitstream being stored in a single track in the visual media file, neither entity groups and / or sample groups are allowed to be stored for the bitstream.

[0487] In several of the above embodiments, converting includes storing the bitstream into a visual media file according to format rules.

[0488] In several of the above embodiments, the conversion includes parsing the visual media file according to the format rules to reconstruct the bitstream.

[0489] In some embodiments, the visual media file parsing apparatus may include a processor configured to implement the method disclosed in the above embodiments.

[0490] In some embodiments, the visual media file writing apparatus includes a processor configured to implement the method disclosed in the above embodiments.

[0491] Some embodiments may include a computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of the above embodiments.

[0492] Some embodiments may include a computer-readable medium having stored thereon a visual media file conforming to a file format generated according to any of the methods described above.

[0493] In the scheme described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the scheme described herein, a decoder can parse syntax elements in the codec representation using the format rules to produce decoded video, knowing the presence or absence of syntax elements according to the format rules.

[0494] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. As defined by the syntax, the bitstream representation of a current video block may, for example, correspond to bits located at different locations within the bitstream or spread across different locations within the bitstream. For example, a macroblock may be encoded based on a transformed and encoded error residual value and also using bits in the header and other fields in the bitstream. Furthermore, during conversion, a decoder may, based on this determination, parse the bitstream with the knowledge that some fields may or may not be present, as described in the above scheme. Similarly, an encoder may determine whether certain syntax fields are included and generate a codec representation accordingly by including or excluding syntax fields from the codec representation. The term "visual media" may refer to video or images, and the term visual media processing may refer to video processing or image processing.

[0495] The disclosed and other schemes, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0496] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form (including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0497] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0498] By way of example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data (e.g., magnetic, magneto-optical, or optical disks), to receive data from or send data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0499] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of the claims, but rather as descriptions of features that may be specific to a particular embodiment of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations, and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may point to a variant of a sub-combination or sub-combination.

[0500] Similarly, while operations may be described in a particular order in the drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in the sequential order shown, or that all illustrated operations be performed, in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0501] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method of processing visual media, comprising: performing conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in a plurality of tracks according to format rules; wherein the format rule specifies that the file-level information includes a syntax element, wherein the syntax element identifies one or more tracks containing a specific type of sample group from a plurality of tracks, wherein the specific type of sample group includes operation point information; The format rule specifies that whether a first set of syntax elements indicating layer dependency information is stored in the visual media file depends on whether a second syntax element indicating that all layers in the visual media file are unrelated has a value of 1.

2. The method according to claim 1, in, The format rule specifies that the visual media file includes file-level information for all operation points provided in the visual media file; The format rule further specifies that, for each operation point, the file-level information includes information of a corresponding track in the visual media file.

3. The method according to claim 2, wherein: The format rules allow tracks required for a particular operation point to include layers and sub-layers not required for the particular operation point.

4. The method according to claim 1, wherein The syntax element comprises a box containing a file-level container.

5. The method according to claim 1, wherein The format rules specify that the file level information is included in a file level box.

6. The method according to claim 1, wherein The format rules specify that the file level information is included in a movie level box.

7. The method according to claim 1, wherein The format rule specifies that the file-level information is included in a track-level box, which is identified within another track-level box or another file-level box.

8. The method according to claim 3, wherein: The format rules also specify that the sample group of the particular type includes information about the tracks required for each operating point.

9. The method according to claim 3, wherein: The format rule also specifies that information on tracks required for each operation point is omitted from another specific type of sample group including layer information on the number of layers in the bitstream.

10. The method according to claim 1, wherein The visual media data is processed by Versatile Video Codec (VVC), and the plurality of tracks are VVC tracks.

11. The method according to any one of claims 1 to 10, wherein The converting includes generating the visual media file according to the format rules and storing the bitstream in the visual media file.

12. The method according to any one of claims 1 to 10, wherein The converting includes parsing the visual media file according to the format rules to reconstruct the bitstream.

13. An apparatus for processing visual media, comprising a processor configured to implement the method according to any one of claims 1-12.

14. A non-transitory computer-readable storage medium having stored thereon processor-executable code, which, when executed, causes a processor to implement the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Alignment of operation point sample group in multi-layer bitstreams file format

    CN108141617A

  • System and method for signaling scalable video in a media application format - Patents.com

    JP2020515169A