Operation point entity group signaling in video encoding and decoding
By introducing operation point information sample groups and layer information sample groups, and combining the RPR and OLS designs in the VVC standard, the problems of low bandwidth utilization efficiency and high decoding complexity of multi-layer video files in the ISO basic media file format are solved, and efficient storage and processing of multi-layer video files are achieved.
Patent Information
- Application Number
- CN202111092628.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-07
- Filing Date
- 2021-09-17
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2041-09-17
AI Technical Summary
Existing video encoding and decoding technologies, especially in the ISO basic media file format, struggle to efficiently support operation point information and layer-related signaling in the storage and processing of multi-layer video files, resulting in low bandwidth utilization efficiency and increased decoding complexity.
By introducing the concepts of operation point information sample groups and layer information sample groups into the video file format, the correlation between operation points and layers is clarified. By utilizing the RPR and OLS design in the VVC standard, the decoding process of multi-layer bitstreams is simplified, ensuring that the decoder can handle multi-layer video files effectively.
It achieves efficient bandwidth utilization and decoder-friendly operation point information processing in multi-layer video files, reduces decoding complexity, and improves the storage and processing efficiency of video files.
Smart Images

Figure CN114205604B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is filed in accordance with applicable patent law and / or the rules of the Paris Convention to promptly claim priority and benefits to U.S. Provisional Patent Application No. 63 / 079,946, filed September 17, 2020, and U.S. Provisional Patent Application No. 63 / 088,786, filed October 7, 2020. For all purposes under the law, the entire disclosure of the foregoing applications is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of networked user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by visual media processing devices for writing to or parsing visual media files according to file formats.
[0006] In one example aspect, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in multiple tracks according to format rules, the format rules specifying file-level information including syntax elements, the syntax elements identifying one or more tracks from the multiple tracks that contain a specific type of sample point group including operation point information.
[0007] In another example, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to format rules. The visual media file stores one or more tracks comprising one or more video layers. Whether a first set of syntax elements indicating layer relevance information is stored in the visual media file depends on whether a second syntax element indicating that all layers in the visual media file are unrelated has a value of 1.
[0008] In another example, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data in multiple tracks according to format rules specifying how redundant access unit delimiter network access layer (AUD NAL) units stored in the multiple tracks are processed during implicit reconstruction of the bitstream according to the multiple tracks.
[0009] In another example, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to format rules. The visual media file stores one or more tracks comprising one or more video layers. The visual media file includes information about operation points (OPs). Format rules specify whether and how syntax elements are included in sample group entries and the OP's box based on whether the OP contains a single video layer. The syntax elements are configured to indicate an index of the OP's output layer set.
[0010] In another example, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to format rules; wherein the visual media file stores multiple tracks belonging to entity groups of a specific type, and wherein the format rules specify that: in response to multiple tracks having track references to a specific type of group identifier, the multiple tracks (A) omit carrying a sample group of a specific type or (B) carry a sample group of a specific type, such that the information in the sample group of the specific type is consistent with the information in the entity group of the specific type.
[0011] In another example, a visual media processing method is disclosed. The method includes: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data. The visual media file includes multiple tracks, and the visual media file stores an entity group carrying information about operation points in the visual media file and a track carrying each operation point. Formatting rules specify attributes of the visual media file in response to the entity group or sample group storing information about each operation point in the visual media file.
[0012] In yet another example, a visual media writing device is disclosed. This device includes a processor configured to implement the methods described above.
[0013] In yet another example, a visual media parsing apparatus is disclosed. This apparatus includes a processor configured to implement the methods described above.
[0014] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0015] In yet another example, a computer-readable medium having a visual media file stored thereon is disclosed. The visual media is generated or parsed using the methods described in this document.
[0016] This document describes these and other features throughout. Attached Figure Description
[0017] Figure 1 This is a block diagram of an example video processing system.
[0018] Figure 2 This is a block diagram of a video processing device.
[0019] Figure 3 This is a flowchart of an example method for video processing.
[0020] Figure 4 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0021] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0022] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0023] Figure 7 An example of an encoder block diagram is shown.
[0024] Figure 8 An example of a bitstream with two OLS is shown, where vps_max_tid_il_ref_pics_plus1[1][0] of OLS2 is equal to 0.
[0025] Figures 9A to 9F A flowchart depicts an example method for visual media processing. Detailed Implementation
[0026] For ease of understanding, chapter headings are used in this document, and the applicability of the techniques and embodiments disclosed in each chapter is not limited to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Thus, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edits to text relative to the current draft of the VVC specification or the ISOBMFF file format specification are shown with strikethrough and highlighting; strikethrough indicates deleted text, and highlighting indicates added text (including bold and italic).
[0027] 1. Preliminary Discussion
[0028] This document relates to video file formats. Specifically, it concerns the storage of Scalable Universal Video Codec (VVC) video bitstreams within media files based on the ISO Base Media File Format (ISOBMFF). These ideas can be applied, individually or in various combinations, to video bitstreams encoded and decoded by any codec (e.g., the VVC standard), and to any video file format, such as the VVC video file format under development.
[0029] 2. Abbreviation
[0030] ACT adaptive color transform
[0031] ALF adaptive loop filter
[0032] AMVR adaptive motion vector resolution
[0033] APS adaptation parameter set
[0034] AU access unit
[0035] AUD access unit delimiter (access unit delimiter)
[0036] AVC (Advanced Video Coding) (Rec.ITU-T H.264|ISO / IEC14496-10)
[0037] B bi-predictive bidirectional prediction
[0038] BCW bi-prediction with CU-level weights
[0039] BDOF bidirectional optical flow
[0040] BDPCM (Block-based Delta Pulse Code Modulation)
[0041] BP buffering period
[0042] CABAC context-based adaptive binary arithmetic coding
[0043] CB coding block
[0044] CBR constant bit rate
[0045] CCALF cross-component adaptive loop filter
[0046] CPB coded picture buffer (image buffer after encoding / decoding)
[0047] CRA clean random access
[0048] CRC cyclic redundancy check
[0049] CTB coding tree block
[0050] CTU coding tree unit
[0051] CU coding unit
[0052] CVS coded video sequence (video sequence after encoding and decoding)
[0053] DPB decoded picture buffer
[0054] DCI decoding capability information
[0055] DRAP dependent random access point
[0056] DU decoding unit
[0057] DUI decoding unit information
[0058] EG exponential - Golomb index - Columbus
[0059] EGk k-th order exponential-Golomb k-order exponential-Columbus
[0060] EOB (End of Bitstream)
[0061] EOS end of sequence
[0062] FD filler data
[0063] FIFO (First-in, First-out)
[0064] FL fixed-length
[0065] GBR green, blue, and red
[0066] GCI general constraints information
[0067] GDR gradual decoding refresh
[0068] GPM geometric partitioning mode
[0069] HEVC (High Efficiency Video Coding) (Rec.ITU-T H.265|ISO / IEC 23008-2)
[0070] HRD hypothetical reference decoder
[0071] HSS hypothetical stream scheduler
[0072] I intra-frame
[0073] IBC intra-block copy
[0074] IDR instantaneous decoding refresh
[0075] ILRP inter-layer reference picture
[0076] IRAP (Intra-Random Access Point)
[0077] LFNST low-frequency non-separable transform
[0078] LPS least probable symbol
[0079] LSB least significant bit
[0080] LTRP long-term reference picture
[0081] LMCS luma mapping with chroma scaling
[0082] MIP matrix-based intra prediction
[0083] MPS most probable symbol
[0084] MSB most significant bit
[0085] MTS multiple transform selection
[0086] MVP motion vector prediction
[0087] NAL network abstraction layer
[0088] OLS output layer set
[0089] OP operation point
[0090] OPI operating point information
[0091] P predictive prediction
[0092] PH picture header
[0093] POC picture order count
[0094] PPS picture parameter set
[0095] PROF prediction refinement with optical flow
[0096] PT picture timing
[0097] PU picture unit
[0098] QP quantization parameter
[0099] RADL (Random Access Decodable Leading (Picture))
[0100] RASL random access skipped leading (picture)
[0101] RBSP raw byte sequence payload
[0102] RGB red, green, and blue
[0103] RPL reference picture list
[0104] SAO sample adaptive offset
[0105] SAR sample aspect ratio
[0106] SEI supplemental enhancement information
[0107] SH slice header
[0108] SLI subpicture level information
[0109] SODB string of data bits
[0110] SPS sequence parameter set
[0111] STRP short-term reference picture
[0112] STSA step-wise temporal sublayer access
[0113] TR truncated rice (cut short rice)
[0114] VBR variable bit rate
[0115] VCL (Video Coding Layer)
[0116] VPS video parameter set
[0117] VSEI (Various Supplemental Enhancement Information) (Rec.ITU-T H.274|ISO / IEC 23002-7)
[0118] VUI video usability information
[0119] VVC (Various Video Coding) is a universal video codec (Rec.ITU-T H.266|ISO / IEC23090-3).
[0120] 3. Introduction to Video Encoding and Decoding
[0121] 3.1. Video codec standards
[0122] Video codec standards have primarily evolved through the well-known developments of ITU-T and ISO / IEC. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). When the Universal Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC is a new codec standard aimed at reducing the bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which concluded on July 1, 2020.
[0123] The Universal Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the related Universal Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from stored media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple encoded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360-degree immersive media.
[0124] 3.2. File Format Standards
[0125] Media streaming applications are typically based on IP, TCP, and HTTP transport methods and often rely on file formats such as ISO Basic Media File Format (ISOBMFF) [7]. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats that utilize ISOBMFF and DASH, video format-specific file format specifications, such as AVC and HEVC file formats, will be required to encapsulate the video content in ISOBMFF tracks as well as DASH representations and segments. Important information about the video bitstream (e.g., profile, tier, and level) and much other information will need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, such as initialization at the start of a streaming session and selection of appropriate media segments for both stream adaptation during a streaming session.
[0126] Similarly, in order to use image formats that utilize ISOBMFF, image format-specific file format specifications, such as AVC image file format and HEVC image file format, will be required.
[0127] The VVC video file format, which is a file format based on ISOBMFF to store VVC video content, is currently being developed by MPEG. The latest draft specification is MPEG output document N19454 (“Information technology—Coding of audio-visual objects—Part 15: Carriage of network abstraction layer (NAL) unitstructured video in the ISO base media file format—Amendment 2: Carriage of VVC and EVC in ISOBMFF”, July 2020).
[0128] The VVC image file format, which is a file format based on ISOBMFF to store image content encoded and decoded using VVC, is currently being developed by MPEG. The latest draft specification for the VVC image file format is included in MPEG output document N19460 (“Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 12: Image File Format—Amendment 3: Support for VVC, EVC, slideshows and other improvements”, July 2020).
[0129] 3.3. Temporal Scalability Support in VVC
[0130] Like HEVC, VVC includes similar support for temporal scalability. This support includes signaling notification of the temporal ID in the NAL unit header, the restriction that images of specific temporal sublayers cannot be used as inter-frame prediction references by images of lower temporal sublayers, the sub-bitstream extraction process, and the requirement that the output of each sub-bitstream extracted from appropriate inputs must be a consistent bitstream. The Media-aware network element (MANE) can leverage the temporal ID in the NAL unit header based on temporal scalability for stream adaptation purposes.
[0131] 3.4. Changing image resolution within a sequence in VVC
[0132] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC enables image resolution changes at a location within a sequence without requiring encoding of the IRAP image, which is always intra-frame encoded / decoded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference picture when the reference picture used for inter-frame prediction has a different resolution than the current picture being decoded.
[0133] To allow for the reuse of motion compensation modules in existing implementations, the scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). The horizontal and vertical scaling ratios are derived based on the width and height of the images and the left, right, top, and bottom scaling offsets specified for the reference and current images.
[0134] RPR allows for resolution changes without encoding / decoding the IRAP image, which can cause momentary bitrate spikes in streaming or video conferencing scenarios, for example, in response to changes in network conditions. RPR can also be used in applications requiring zooming in on the entire video area or a region of interest. Negative scaling window offsets are allowed to support a wider range of zoom-based applications. Negative scaling window offsets also enable the extraction of sub-image sequences from multiple bitstreams while maintaining the same scaling window for the extracted sub-bitstream as in the original bitstream.
[0135] Unlike spatial scalability in the scalable extension of HEVC, RPR in VVC is performed as part of the same process at the block level, where image resampling and motion compensation are applied in two different stages, with the derivation of the sampling location and motion vector scaling performed during motion compensation.
[0136] To limit implementation complexity, image resolution changes within the CLVS are not permitted when each image in the CLVS has multiple sub-images. Furthermore, when RPR is used between the current image and a reference image, decoder-side motion vector thinning, bidirectional optical flow, and optical flow prediction thinning are not applied. The juxtaposed images used to derive temporal motion vector candidates are also restricted to having the same image size, scaling window offset, and CTU size as the current image.
[0137] To support RPR, VVC differs from HEVC in several other aspects of its design. First, the image resolution, along with the corresponding consistency and scaling windows, is signaled in the PPS (Prefix Per Second) instead of the SPS (Signal Per Second), where the maximum image resolution and corresponding consistency window are signaled. In applications, the maximum image resolution with the corresponding consistency window offset in the SPS can be used as the expected or desired image output size after cropping. Second, for a single-layer bitstream, each image storage (the time slot in the DPB used to store one decoded image) occupies the buffer size required to store the decoded image with the maximum image resolution.
[0138] 3.5. Multi-tier Scalability Support in VVC
[0139] The ability to perform inter-frame prediction using the RPR (Reference Prediction Process) in the VVC core design, based on a reference image with a different size from the current image, allows VVC to easily support bitstreams containing multiple layers with different resolutions, such as two layers with standard definition and high definition resolution respectively. This functionality can be integrated into the VVC decoder without any additional signal processing level encoding / decoding tools, as the upsampling capabilities required for spatial scalability can be provided by reusing the RPR upsampling filter. However, additional high-level syntax design is needed to implement support for bitstream scalability.
[0140] Scalability is supported in VVC, but only in multi-layer profiles. Unlike scalability support in any earlier video codec standards, including extensions to AVC and HEVC, VVC scalability is designed to be as friendly as possible to single-layer decoder implementations. The decoding capability of a multi-layer bitstream is specified as if the bitstream contained only a single layer. For example, decoding capability can be specified in a way independent of the number of layers in the bitstream to be decoded, such as the DPB size. Essentially, a decoder designed for a single-layer bitstream can decode multi-layer bitstreams without significant modifications.
[0141] Compared to the multi-layered extensions of AVC and HEVC, HLS is significantly simplified at the cost of some flexibility. For example, 1) the IRAP AU needs to include a picture of every layer present in CVS, which avoids the need to specify a layer-by-layer decoding process, and 2) a simpler design for POC signaling is included in VVC, instead of a complex POC reset mechanism, to ensure that the derived POC values are the same for all pictures in the AU.
[0142] Similar to HEVC, information about layers and layer dependencies is included in the VPS. OLS information is provided for signaling notifications of which layers are included in the OLS, which layers are output, and other information such as the PTL and HRD parameters associated with each OLS. Similar to HEVC, there are three operating modes for outputting all layers, outputting only the highest layer, or outputting specific indicator layers in a custom output mode.
[0143] There are some differences between the OLS design in VVC and HEVC. First, in HEVC, signaling informs the set of layers, then signaling informs the OLS based on the set of layers, and for each OLS, signaling informs the output layer. The design in HEVC allows a layer belonging to an OLS to be neither an output layer nor a layer required to decode an output layer. In VVC, this design requires that any layer in the OLS is either an output layer or a layer required to decode an output layer. Therefore, in VVC, the OLS is signaled by indicating its output layer, and other layers belonging to the OLS are then simply deduced from the layer dependencies indicated in the VPS. Furthermore, VVC requires that each layer be included in at least one OLS.
[0144] Another difference in the VVC OLS design is that, unlike HEVC, where the OLS consists of all NAL units belonging to the set of identified layers mapped to the OLS, VVC can exclude some NAL units belonging to non-output layers mapped to the OLS. More specifically, the VVC OLS consists of the set of layers mapped to the OLS, where non-output layers only include IRAP or GDR images where ph_recovery_poc_cnt equals 0, or images from sublayers used for inter-layer prediction. This allows only all “necessary” images from all sublayers within the layers forming the OLS to indicate the optimal level value for multi-layer bitstreams, where “necessary” means required for output or decoding. Figure 8 An example of a two-layer bitstream where vps_max_tid_il_ref_pics_plus1[1][0] equals 0 is shown, i.e., when extracting OLS2, only the sub-bitstream of the IRAP picture from layer L0 is retained.
[0145] Considering that allowing different RAP periods at different layers is beneficial in some scenarios, similar to ABC and HEVC, AUs are allowed to have layers with unaligned RAPs. To more quickly identify RAPs in multi-layer bitstreams, i.e., AUs with RAPs at all layers, the Access Element Delimiter (AUD) is extended with a flag indicating whether the AU is an IRAP AU or a GDR AU, compared to HEVC. Furthermore, when the VPS indicates multiple layers, the AUD mandate is present in such IRAP or GDR AUs. However, for single-layer bitstreams indicated by the VPS or bitstreams not involving the VPS, the AUD is entirely optional, as in HEVC, because in this case, the RAP can be easily detected from the first stripe of the NAL cell type in the AU and the corresponding parameter set.
[0146] To enable multiple layers to share SPS, PPS, and APS, and to ensure that the bitstream extraction process does not discard the parameter set required for the decoding process, the VCL NAL unit of the first layer can refer to the SPS, PPS, or APS with the same or lower layer ID value, as long as all OLS of the first layer also include the layer identified by the lower layer ID value.
[0147] 3.6. Some details about the VVC video file format
[0148] 3.6.1. Overview of VVC storage with multiple tiers
[0149] Support for VVC bitstreams with multiple layers includes several tools, and various "models" exist on how they can be used. VVC streams with multiple layers can be placed in tracks in several ways, including the following:
[0150] 1. All layers are in one track, where all layers correspond to operation points;
[0151] 2. All layers are in one track, where there is no operation point that contains all layers;
[0152] 3. One or more layers or sublayers are in separate tracks, containing a bitstream of all samples from the indicated one or more tracks corresponding to the operation point;
[0153] 4. One or more layers or sublayers are in separate orbitals, where there are no operation points for all NAL cells in a set containing one or more orbitals.
[0154] The VVC file format allows one or more layers to be stored in a track. Multiple layers can be stored per track. For example, tracks can be created when a content provider wants to provide a multi-layered bitstream that is not intended for subsetting, or when several predefined sets of bitstreams have already been created for the output layers (where each layer corresponds to a view, such as a stereo pair).
[0155] When a VVC bitstream is represented by multiple tracks and the player uses operation points whose layers are stored in multiple tracks, the player must reconstruct these units before passing the VVC access units to the VVC decoder.
[0156] VVC operation points can be explicitly represented by orbitals, i.e., each sample point in the orbital locally contains access cells or contains access cells by resolving a "subp" orbital reference (if present) and by resolving a "vvcN" orbital reference (if present). Access cells contain NAL cells from all layers and sublayers that are part of the operation point.
[0157] The storage of VVC bitstreams is supported by structures such as the following:
[0158] a) Sample point entries
[0159] b) Operation point information (“vopi”) sample group,
[0160] c) Layer information ('linf') sample group,
[0161] d) Operation point entity group (“opeg”).
[0162] The structure within a sample entry provides information for the decoding or use of the sample, in this case, the codec video and non-VCL data information associated with that sample entry.
[0163] The operation point information sample set records information about operation points, such as the layers and sublayers that make up the operation point, their correlations (if any), the operation point's grade, layer and level parameters, and other such operation point-related information.
[0164] The layer information sample group lists all the layers and sublayers carried in the samples of the orbit.
[0165] The operation point entity group records information about the operation point, such as the layers and sublayers that make up the operation point, their correlations (if any), the operation point's grade, layer and level parameters, and other such operation point-related information, as well as the track identifier carrying each operation point.
[0166] The information in these sample groups, combined with track references, is sufficient to locate track or operation point entities. This information is sufficient for the reader to select operation points according to its capabilities, identify tracks containing the relevant layers and sub-layers required to decode the selected operation point, and extract them effectively.
[0167] 3.6.2. Data Sharing and Reconstruction of VVC Bitstream
[0168] 3.6.2.1. Overview
[0169] In order to reconstruct the access unit based on samples from multiple tracks carrying multi-layer VVC bitstreams, the operation point needs to be determined first.
[0170] Note: When a VVC bitstream is represented by multiple VVC tracks, the file parser can identify the track required for the selected operation point, as follows:
[0171] Find all tracks with VVC sample entries.
[0172] If a track contains a "oref" track reference with the same ID, then that ID is resolved to either a VVC track or an "opeg" entity group.
[0173] Choose an operation point from either the "opeg" entity group or the "vopi" sample group that suits your decoding capabilities and application purpose.
[0174] When the “opeg” entity group exists, it indicates that the set of orbits precisely represents the selected operation point.
[0175] Therefore, the VVC bitstream can be reconstructed and decoded based on the track set.
[0176] When the “opeg” entity group does not exist (i.e. when the “vopi” sample group exists), determine which track set is needed to decode the selected operation point from the “vopi” and “linf” sample groups.
[0177] To reconstruct the bitstream from multiple VVC tracks carrying the VVC bitstream, it may be necessary to first determine the target highest value, TemporalId.
[0178] If several tracks contain data for accessing units, the alignment of the corresponding samples in the tracks is performed based on the sample decoding time, i.e., using a time-sample table without considering the edit list.
[0179] When a VVC bitstream is represented by multiple VVC tracks, the decoding time of the samples should be such that if the tracks are combined into a single stream ordered by increasing the decoding time, the access unit order will be correct, as specified in ISO / IEC 23090-3.
[0180] The sequence of access units is reconstructed from the corresponding samples in the desired orbit according to the implicit reconstruction process described in 3.6.2.2.
[0181] 3.6.2.2. Implicit Reconstruction of VVC Bitstream
[0182] When the operation point information sample group exists, the required trajectory is selected based on the layer it carries and its reference layer, as shown in the operation point information and layer information sample group.
[0183] When an operation point entity group exists, the required track is selected based on the information in the OperatingPointGroupBox.
[0184] When reconstructing a bitstream containing a sub-layer with a TemporalId greater than 0 for a VCL NAL unit, all lower sub-layers within the same layer (i.e., those layers with smaller TemporalIds for VCL NAL units) are also included in the resulting bitstream, and the desired tracks are selected accordingly.
[0185] When reconstructing the access unit, image units from samples with the same decoding time (as specified in ISO / IEC 23090-3) are placed into the access unit in ascending order of nuh_layer_id value.
[0186] When reconstructing access cells with relevant layers and max_tid_il_ref_pics_plus1 is greater than 0, sublayers of reference layers whose TemporalId of VCL NAL cells in the same layer is less than or equal to max_tid_il_ref_pics_plus1-1 (as shown in the operation point information sample group) are also included in the resulting bitstream, and the desired tracks are selected accordingly.
[0187] When reconstructing access units with relevant layers and max_tid_il_ref_pics_plus1 equals 0, only IRAP picture units of the reference layer are included in the resulting bitstream, and the desired tracks are selected accordingly.
[0188] If the VVC track contains a "subp" track reference, then each picture unit is reconstructed according to Clause 11.7.3, with the following additional constraints for the EOS and EOB NAL units: Repeat the process in Clause 11.7.3 for each layer of the target operation point in ascending order of nuh_layer_id. Otherwise, reconstruct each picture unit as described below.
[0189] The reconstructed access units are placed into the VVC bitstream in ascending order of decoding time, and copies of the End of Bitstream (EOB) and End of Sequence (EOS) NAL units are removed from the VVC bitstream, as further described below.
[0190] For access units belonging to different sublayers stored in multiple tracks within the same codec video sequence of a VVC bitstream, there may be more than one track containing an EOS NAL unit with a specific nuh_layer_id value in the corresponding sample. In this case, only one EOS NAL unit should be kept in the last access unit (the one with the longest decoding time) of these access units in the final reconstructed bitstream, placed after all NAL units of the last access unit of these access units except for the EOB NAL unit (if present), and the other EOS NAL units are discarded. Similarly, there may be more than one such track containing an EOB NAL unit in the corresponding sample. In this case, only one EOB NAL unit should be kept in the final reconstructed bitstream, placed at the end of the last access unit of these access units, and the other EOB NAL units are discarded.
[0191] Since a particular layer or sublayer can be represented by more than one orbital, when calculating the orbital required for the operation point, it may be necessary to select from the entire set of orbitals that carry the particular layer or sublayer.
[0192] When the operand entity group is not present, after selecting from the tracks carrying the same layer or sublayer, the final required tracks can still jointly carry some layers or sublayers that do not belong to the target operand. The reconstructed bitstream of the target operand should not contain layers or sublayers carried in the final required tracks that do not belong to the target operand.
[0193] Note: The VVC decoder implementation takes as input the bitstream corresponding to the target output layer set index and the highest TemporalId value of the target operation point. These two values correspond to the TargetOlsIdx and HighestTid variables in Clause 8 of ISO / IEC 23090-3, respectively. Before sending the reconstructed bitstream to the VVC decoder, the file parser needs to ensure that the reconstructed bitstream does not contain any layers or sublayers other than those included in the target operation point.
[0194] 3.6.3. Operation Point Information Sample Group
[0195] 3.6.3.1. Definition
[0196] By using a set of operation point information samples (“vopi”), the application is informed of the different operation points offered by a given VVC bitstream and their composition. Each operation point is associated with the output layer set, the maximum TemporalId, and the grade, level, and layer signaling. All of this information is captured by the “vopi” sample set. In addition to this information, the sample set also provides information on the correlations between layers.
[0197] Both of the following apply when there is more than one VVC track for a VVC bitstream and there is no operation point entity group for the VVC bitstream:
[0198] - In the VVC tracks of the VVC bitstream, there should be one and only one track carrying the "vopi" sample group.
[0199] All other VVC tracks in a VVC bitstream should have a track reference of type "oref" for the track carrying the "vopi" sample group.
[0200] For any specific sample in a given track, a time-domain juxtaposed sample in another track is defined as a sample whose decoding time is the same as that of the specific sample. For a track T carrying a "vopi" sample group... k The orbit T, which is the reference orbit for the "oref" orbit. N Each sample point S in N The following applies:
[0201] -If in orbit T k There are time-domain juxtaposed samples S k Then sample point S N and sample points S k The "vopi" sample group entries are associated with the same "vopi" sample group entries.
[0202] -Otherwise, sample point S N With and orbit T kAccording to the decoding time at sample point S N The "vopi" sample group entry of the previous last sample point is associated with the same "vopi" sample group entry.
[0203] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the sample group description box with grouping_type "vopi". For the more common scenario of a single VPS, it is recommended to use the default sample grouping mechanism defined in ISO / IEC 14496-12 and include the operation point information sample group in the sample table box instead of including it in each track fragment.
[0204] The grouping_type_parameter is not defined for SampleToGroupBox which has a grouping type "vopi".
[0205] 3.6.3.2. Syntax
[0206]
[0207]
[0208]
[0209] 3.6.3.3. Semantics
[0210] The value num_profile_tier_level_minus1 plus 1 gives the number of the following tiers, levels, and combinations of levels, along with the relevant fields.
[0211] ptl_max_temporal_id[i]: Gives the maximum TemporalId of the NAL cell of the relevant bitstream for the specified i-th grade, layer, and level structure.
[0212] Note: The semantics of ptl_max_temporal_id[i] and max_temporal_id given below are different, even though they may carry the same value.
[0213] ptl[i] specifies the i-th grade, level, and hierarchy structure.
[0214] ISO / IEC 23090-3 defines all_independent_layers_flag, each_layer_is_an_ols_flag, ols_mode_idc, and max_tid_il_ref_pics_plus1.
[0215] num_operating_points: Gives the number of operation points that follow the information.
[0216] `output_layer_set_idx` is the index of the set of output layers that defines the operation point. The mapping between `output_layer_set_idx` and the `layer_id` value should be the same as the mapping specified by the VPS for the set of output layers with index `output_layer_set_idx`.
[0217] ptl_idx: A zero-based index of the listed tiers, levels, and layer structures of the output layer set whose signaling notification index is output_layer_set_idx.
[0218] max_temporal_id: Gives the maximum TemporalId of the NAL cell at this operation point.
[0219] Note: The maximum TemporalId value indicated in the layer information sample group has a different semantics than the maximum TemporalId indicated here. However, they may carry the same literal value.
[0220] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.
[0221] layer_id: Provides the nuh_layer_id value of the layer for the operation point.
[0222] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.
[0223] A frame_rate_info_flag value of 0 indicates that frame rate information is not present at the operation point. A value of 1 indicates that frame rate information is present at the operation point.
[0224] A value of 0 for `bit_rate_info_flag` indicates that there is no bit rate information for the operation point. A value of 1 indicates that there is bit rate information for the operation point.
[0225] avgFrameRate gives the average frame rate at the operation point, in frames per (256 seconds). A value of 0 indicates that no average frame rate is specified.
[0226] A constantFrameRate value of 1 indicates that the stream at the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream at the operation point has a constant frame rate. A value of 0 indicates that the stream at the operation point may or may not have a constant frame rate.
[0227] maxBitRate gives the maximum bit rate of the stream at any point in a one-second window, expressed in bits per second.
[0228] avgBitRate gives the average bit rate of the stream at the operation point, in bits per second.
[0229] max_layer_count: The count of all unique layers across all operating points associated with this base orbit.
[0230] layerID: nuh_layer_id of the layer for which all its direct reference layers are given in the following direct_ref_layerID loop.
[0231] num_direct_ref_layers: The number of direct reference layers for the layer whose nuh_layer_id is equal to layerID.
[0232] direct_ref_layerID: nuh_layer_id of the direct reference layer.
[0233] 3.6.4. Layer Information Sample Group
[0234] The list of layers and sublayers carried by the track is signaled in the layer information sample group. When there is more than one VVC track for the same VVC bitstream, each of these VVC tracks should carry the "linf" sample group.
[0235] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the sample group description box with grouping_type "linf". For the more common scenario of a single VPS, it is recommended to use the default sample grouping mechanism defined in ISO / IEC 14496-12 and include the layer information sample group in the sample table box instead of including it in each track fragment.
[0236] The grouping_type_parameter is not defined for SampleToGroupBox which has the grouping type "linf".
[0237] The syntax and semantics of the “linf” sample group are specified in clauses 9.6.3.2 and 9.6.3.3, respectively.
[0238] 3.6.5. Operation Point Entity Group
[0239] 3.6.5.1. Overview
[0240] The operator point entity group is defined as providing the mapping from track to operator point and the level information of the operator point.
[0241] When aggregated samples of tracks mapped to the operation points described in the entity group, the implicit reconstruction process does not require the removal of any further NAL units to produce a consistent VVC bitstream. Tracks belonging to the operation point entity group should have a track reference of type "oref" to the group_id indicated in the operation point entity group.
[0242] All entity_id values included in the operation point entity group should belong to the same VVC bitstream. When present, the OperatingPointGroupBox should be included in the GroupsListBox of the movie-level MetaBox, and not in the file-level or track-level MetaBox.
[0243] 3.6.5.2. Syntax
[0244]
[0245]
[0246] 3.6.5.3. Semantics
[0247] The value num_profile_tier_level_minus1 plus 1 gives the number of the following tiers, levels, and combinations of levels, along with the relevant fields.
[0248] opeg_ptl[i] specifies the i-th grade, level, and hierarchy structure.
[0249] num_operating_points: gives the number of operation points that follow the information.
[0250] `output_layer_set_idx` is the index of the set of output layers that defines the operation point. The mapping between `output_layer_set_idx` and the `layer_id` value should be the same as the mapping specified by the VPS for the set of output layers with index `output_layer_set_idx`.
[0251] ptl_idx: A zero-based index of the listed tiers, levels, and layer structures of the output layer set whose signaling notification index is output_layer_set_idx.
[0252] max_temporal_id: Gives the maximum TemporalId of the NAL cell at this operation point.
[0253] Note: The maximum TemporalId value indicated in the layer information sample group has a different semantics than the maximum TemporalId indicated here. However, they may carry the same literal value.
[0254] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.
[0255] layer_id: Provides the nuh_layer_id value of the layer for the operation point.
[0256] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.
[0257] A frame_rate_info_flag value of 0 indicates that frame rate information is not present at the operation point. A value of 1 indicates that frame rate information is present at the operation point.
[0258] A value of 0 for `bit_rate_info_flag` indicates that there is no bit rate information for the operation point. A value of 1 indicates that there is bit rate information for the operation point.
[0259] avgFrameRate gives the average frame rate at the operation point, in frames per (256 seconds). A value of 0 indicates that no average frame rate is specified.
[0260] A constantFrameRate value of 1 indicates that the stream at the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream at the operation point has a constant frame rate. A value of 0 indicates that the stream at the operation point may or may not have a constant frame rate.
[0261] maxBitRate gives the maximum bit rate of the stream at any point in a one-second window, expressed in bits per second.
[0262] avgBitRate gives the average bit rate of the stream at the operation point, in bits per second.
[0263] entity_count specifies the number of tracks that exist at the operation point.
[0264] entity_idx specifies the index of the list of entity_ids in the entity group belonging to the operation point.
[0265] 4. Examples of technical problems solved by publicly available solutions
[0266] The latest design of the VVC video file format for storing scalable VVC bitstreams has the following problems:
[0267] 1) When a VVC bitstream is represented by multiple VVC tracks, the file parser can identify the track required for the selected operation point by first searching all tracks with VVC sample entries, then searching all tracks containing the "vopi" sample group, and so on, to find information about all operation points provided in the file. However, finding all these tracks can be quite complex.
[0268] 2) The “linf” sample group signaling notifies which layers and / or sublayers are included in the orbit. When using the “vopi” sample group to select the OP, it is necessary to find all orbits containing the “linf” sample group, and the information carried in the “linf” sample group entries of these orbits is used in conjunction with the information in the “vopi” sample group entries on the OP's layers and / or sublayers to identify the desired orbit. This can also be quite complex.
[0269] 3) In the “vopi” sample group entry, layer dependency information is signaled even when the value of all_independent_layers_flag is equal to 1. However, when all_independent_layers_flag is equal to 1, the layer dependency information is already known, so the bits used for signaling are wasted in this case.
[0270] 4) The removal of redundant EOS and EOB NAL cells is specified during the implicit reconstruction of the VVC bitstream based on multiple tracks. However, the removal and / or rewriting of AUD NAL cells may be required during this process, but the corresponding procedure is lacking.
[0271] 5) Specifies that when the “opeg” entity group exists, the bitstream is reconstructed from multiple tracks by including all NAL units of the desired track without removing any NAL units. However, this will, for example, prevent NAL units (such as AUD, EOS, and EOB NAL units for a specific AU) from being included in more than one track carrying the VVC bitstream.
[0272] 6) The container for the “opeg” entity group box is specified as a movie-level MetaBox. However, the entity_id value of the entity group can only refer to the track ID when it is contained in a file-level MetaBox.
[0273] 7) In the “opeg” entity grouping bin, the output_layer_set_idx field is always used for each OP signaling notification. However, if the OP contains only one layer, it is usually not necessary to know the value of the OLS index, and even if knowing the OLS index is useful, it can be easily deduced to be the OLS index of only the OLS of that layer.
[0274] 8) Allowing one of the tracks representing a VVC bitstream to have a "vopi" sample group when an "opeg" entity group exists for the VVC bitstream is permitted. However, allowing both is unnecessary and would only unnecessarily increase the file size and confuse the file parser about which one to use.
[0275] 5. List of solutions
[0276] To address the aforementioned issues, the following outlines a methodology. These items should be considered as examples for interpreting general concepts, rather than interpreted in a narrow sense. Furthermore, these items can be applied individually or in combination in any way.
[0277] 1) To solve problem 1, one or more of the following projects were proposed:
[0278] a. Add signaling for all file-level information about all OPs provided in the file, including the tracks required for each OP, while the tracks required for an OP may carry layers or sublayers not included in that OP.
[0279] b. Add signaling about which tracks contain the "vopi" sample group at the file level.
[0280] i. In one example, a new bin is specified, for example, named the Operation Point Information Track Bin, using a file-level MetaBox as the container, for signaling notifications of tracks carrying "vopi" sample groups.
[0281] A file-level message can be signaled in a file-level bin or a video-level bin, or in a track-level bin, but the location of the track-level bin is identified within the file-level bin or the video-level bin.
[0282] 2) To address problem 2, one or more of the following projects were proposed:
[0283] a. Add information about the required tracks for each OP to the “vopi” sample group entry.
[0284] b. The use of "linf" sample groups is not recommended.
[0285] 3) To solve problem 3, when all_independent_layers_flag equals 1, skip the signaling notification of layer correlation information in the “vopi” sample group entry.
[0286] 4) To address problem 4, an operation for removing redundant AUD NAL units is added during the implicit reconstruction of the VVC bitstream based on multiple tracks.
[0287] a. Alternatively, when necessary, further operations for rewriting the AUD NAL cell may be added.
[0288] i. In one example, it is specified that when reconstructing an access unit based on multiple picture units from different tracks, the value of aud_irap_or_gdr_flag of the AUD NAL unit is set to 0 if the aud_irap_or_gdr_flag of the AUD NAL unit held in the reconstructed access unit is equal to 1 and the reconstructed access unit is not an IRAP or GDR access unit.
[0289] ii. The aud_irap_or_gdr_flag of the AUD NAL cell in the first PU may be equal to 1, while another PU in the same access cell but on a separate track has a picture that is not an IRAP or GDR picture. In this case, the value of aud_irap_or_gdr_flag of the AUD NAL cell in the reconstructed access cell is changed from 1 to 0.
[0290] b. In one example, it is additionally or alternatively specified that when at least one of the multiple picture units from different tracks of the access unit has an AUD NAL unit, the first picture unit (i.e., the picture unit with the smallest nuh_layer_id value) should have an AUD NAL unit.
[0291] c. In one example, it is specified that when multiple picture units from different orbits of the access unit have AUD NAL units present, only the AUD NAL units in the first picture unit are retained in the reconstructed access unit.
[0292] 5) To address problem 5, one or more of the following projects were proposed:
[0293] a. It is specified that when the “opeg” entity group exists and is used, the required track provides the exact set of VCLNAL units required for each OP, but some non-VCL NAL units may become redundant in the reconstructed bitstream and may therefore need to be removed.
[0294] i. Alternatively, it is specified that when the “opeg” entity group exists and is used, the required track provides the exact set of layers and sublayers required for each OP, but some non-VCL NAL units may become redundant in the reconstructed bitstream and may therefore need to be removed.
[0295] b. During the implicit reconstruction of the VVC bitstream based on multiple tracks, operations for removing redundant EOB and EOS NAL units are applied even when the “opeg” entity group exists and is used.
[0296] c. During the implicit reconstruction of the VVC bitstream based on multiple tracks, operations for removing redundant AUD units and operations for rewriting redundant AUD units are applied even when the “opeg” entity group exists and is used.
[0297] 6) To resolve issue 6, specify the container for the "opeg" entity group box as a GroupsListBox within a file-level MetaBox, as follows: When present, the OperatingPointGroupBox should be contained within a GroupsListBox within a file-level MetaBox, and not within a MetaBox at other levels.
[0298] 7) To address issue 7, when an OP contains only one layer, for that OP, skip the signaling of the VvcOperatingPointsRecord in the “vopi” sample group entry and the output_layer_set_idx field in the “opeg” entity group box OperatingPointGroupBox.
[0299] a. In one example, the output_layer_set_idx field in VvcOperatingPointsRecord and / or OperatingPointGroupBox is moved after the loop of signaling notification layer_id, and is conditional on "if(layer_count>1)".
[0300] b. In addition, in one example, it is specified that when output_layer_set_idx does not exist for an OP, its value is inferred to be equal to the OLS index of the OLS containing only the layers in that OP.
[0301] 8) To address problem 8, it was specified that when an “opeg” entity group exists for a VVC bitstream, it means that none of the tracks in the VVC bitstream should have a “vopi” sample group.
[0302] a. Alternatively, both can be allowed, but when both are present, they must be consistent such that choosing either one makes no difference.
[0303] b. In one example, it is specified that tracks belonging to the “opeg” entity group (which all have track references of type “oref” to the group_id indicated in the entity group) should not carry the “vopi” sample group.
[0304] 9) It can be specified that when a VVC bitstream is represented in only one track, it is not allowed to have "opeg" entity groups or "vopi" sample groups.
[0305] 6. Example of an implementation plan
[0306] Below are some example embodiments of the inventions outlined in Section 5 of the previous article, which can be applied to the standard specifications of the VVC video file format. The revised text is based on the latest draft of the VVC specification. Most of the relevant added or modified sections are... Bold underline Highlighted, and some deleted parts are marked with [[ Bold Italic [Highlighted.] There may be other editable changes that are not highlighted.
[0307] 6.1. First Embodiment
[0308] This example applies to projects 2a, 3, 7, 7a, and 7b.
[0309] 6.1.1. Operation Point Information Sample Group
[0310] 6.1.1.1. Definition
[0311] By using a set of operation point information samples (“vopi”), the application is informed of the different operation points offered by a given VVC bitstream and their composition. Each operation point is associated with the output layer set, the maximum TemporalId value, and the grade, level, and layer signaling. All of this information is captured by the “vopi” sample set. In addition to this information, the sample set also provides information on the correlations between layers.
[0312] Both of the following apply when there is more than one VVC track for a VVC bitstream and there is no operation point entity group for the VVC bitstream:
[0313] - In the VVC tracks of the VVC bitstream, there should be one and only one track carrying the "vopi" sample group.
[0314] All other VVC tracks in a VVC bitstream should have a track reference of type "oref" for the track carrying the "vopi" sample group.
[0315] For any specific sample in a given track, a time-domain juxtaposed sample in another track is defined as a sample whose decoding time is the same as that of the specific sample. For a track T carrying a "vopi" sample group... k The orbit T, which is the reference orbit for the "oref" orbit. N Each sample point S in N The following applies:
[0316] -If in orbit T k There are time-domain juxtaposed samples S k Then sample point S N and sample points S k The "vopi" sample group entries are associated with the same "vopi" sample group entries.
[0317] -Otherwise, sample point S N With and orbit T k According to the decoding time at sample point S N The "vopi" sample group entry of the previous last sample point is associated with the same "vopi" sample group entry.
[0318] When several VPSs are referenced by a VVC bitstream, it may be necessary to include several entries in the sample group description box with grouping_type "vopi". For the more common scenario where a single VPS exists, it is recommended to use the default sample group mechanism defined in ISO / IEC 14496-12 and include the operation point information sample group in the sample table box instead of including it in each track fragment.
[0319] The grouping_type_parameter is not defined for SampleToGroupBox which has a grouping type "vopi".
[0320] 6.1.1.2. Syntax
[0321]
[0322]
[0323]
[0324] 6.1.1.3. Semantics
[0325] ...
[0326] num_operating_points: gives the number of operation points that follow the information.
[0327] [[output_layer_set_idx is the index of the set of output layers that defines the operation point. The mapping between output_layer_set_idx and layer_id values should be the same as the mapping specified by the VPS for the set of output layers with index output_layer_set_idx.]]
[0328] ptl_idx: A zero-based index of the listed tiers, levels, and layer structures of the output layer set whose signaling notification index is output_layer_set_idx.
[0329] max_temporal_id: Gives the maximum TemporalId of the NAL cell at this operation point.
[0330] Note: The maximum TemporalId value indicated in the layer information sample group has a different semantics than the maximum TemporalId indicated here. However, they may carry the same literal value.
[0331] layer_count: This field indicates the number of necessary layers for this operation point, as defined in ISO / IEC 23090-3.
[0332] layer_id: Provides the nuh_layer_id value of the layer for the operation point.
[0333] is_outputlayer: A flag indicating whether the layer is an output layer. One indicates an output layer.
[0334] `output_layer_set_idx` is the index of the set of output layers that defines the operation point. The mapping between idx and layer_id values should correspond to the output layer set indexed by output_layer_set_idx on the VPS. The mapping is the same. When output_layer_set_idx does not exist for an OP, its value is inferred to be equal to that containing only that OP. The OLS index of the layer in the OLS.
[0335] op_track_count specifies the number of tracks carrying VCLNAL cells at this operation point.
[0336] op_track_id[j] specifies the j-th track among the tracks carrying VCL NAL units at this operation point. track_ID value.
[0337] A frame_rate_info_flag value of 0 indicates that frame rate information is not present at the operation point. A value of 1 indicates that frame rate information is present at the operation point.
[0338] A value of 0 for `bit_rate_info_flag` indicates that there is no bit rate information for the operation point. A value of 1 indicates that there is bit rate information for the operation point.
[0339] avgFrameRate gives the average frame rate at the operation point, in frames per (256 seconds). A value of 0 indicates that no average frame rate is specified.
[0340] A constantFrameRate value of 1 indicates that the stream at the operation point has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream at the operation point has a constant frame rate. A value of 0 indicates that the stream at the operation point may or may not have a constant frame rate.
[0341] maxBitRate gives the maximum bit rate of the stream at any point in a one-second window, expressed in bits per second.
[0342] avgBitRate gives the average bit rate of the stream at the operation point, in bits per second.
[0343] max_layer_count: The count of all unique layers across all operating points associated with this base orbit.
[0344] layerID: nuh_layer_id of the layer for which all its direct reference layers are given in the following direct_ref_layerID loop.
[0345] num_direct_ref_layers: The number of direct reference layers for the layer whose nuh_layer_id is equal to layerID.
[0346] direct_ref_layerID: nuh_layer_id of the direct reference layer.
[0347] 6.2. Second Embodiment
[0348] This embodiment applies to items 1.bi, 4, 4a, 4.ai, 4b, 4c, 5a, 6, 8, and 8b.
[0349] Implicit reconstruction of VVC bitstream:
[0350] When the operation point information sample group exists, the required trajectory is selected based on the layer it carries and its reference layer, as shown in the operation point information and layer information sample group.
[0351] When an operation point entity group exists, the required track is selected based on the information in the OperatingPointGroupBox.
[0352] When reconstructing a bitstream containing a sub-layer with a TemporalId greater than 0 for VCL NAL units, all lower sub-layers within the same layer (i.e., those layers with smaller TemporalIds for VCL NAL units) are also included in the resulting bitstream, and the required tracks are selected accordingly.
[0353] When reconstructing the access unit, image units from samples with the same decoding time (as specified in ISO / IEC 23090-3) are placed into the access unit in ascending order of nuh_layer_id value. When multiple access units When at least one picture unit in a picture unit has an AUD NAL unit, the first picture unit (i.e., having the smallest nuh_) The image unit with the layer_id value should have AUD NAL units, and only the first image unit should have AUD NAL units. The accessed data is retained within the reconstructed access unit, while other AUD NAL units (if present) are discarded. In this reconstructed access... In the access unit, when the aud_irap_or_gdr_flag of the AUD NAL unit is equal to 1 and the reconstructed access unit is not IRAP When accessing a cell via GDR, the value of aud_irap_or_gdr_flag in the AUD NAL cell is set to 0.
[0354] Note 1: The aud_irap_or_gdr_flag of the AUD NAL unit in the first PU may be equal to 1, while the same access unit... The original image, but another PU in a separate track, has an image that is neither an IRAP nor a GDR image. In this case, the reconstructed image... The value of aud_irap_or_gdr_flag in the AUD NAL cell of the access cell is changed from 1 to 0.
[0355] ...
[0356] Entity Group Other file-level information :
[0357] Sub-image entity group:
[0358] ...
[0359] Operation point entity group:
[0360] Overview:
[0361] The operator point entity group is defined as providing the mapping from track to operator point and the level information of the operator point.
[0362] When the aggregate maps samples of the orbits to the operation points described in the entity group, the implicit reconstruction process does not require the removal of any further samples. VCL NAL units are used to generate a consistent VVC bitstream. Tracks belonging to an operation point entity group should have a track reference of type "oref" to the group_id indicated in the operation point entity group. And should not Carrying "vopi" sample group .
[0363] All entity_id values included in the operating point entity group should belong to the same VVC bitstream. When present, the OperatingPointGroupBox should be included within the [[Movie]]. document It should be placed within a GroupsListBox of a MetaBox, and not included in it. Other levels In the [[file level or track level]] MetaBox.
[0364] Operation Point Information Track Box
[0365] definition:
[0366] Box type: "topi"
[0367] Container: File-level MetaBox
[0368] Mandatory: No
[0369] Quantity: 0 or 1
[0370] The operating point information track box contains the track ID of the track set carrying the "vopi" sample group. The box does not exist. The file does not contain a track with the "vopi" sample group.
[0371] grammar
[0372]
[0373] Semantics
[0374] num_tracks_with_vopi specifies the number of tracks in the file that carry the "vopi" sample group.
[0375] track_ID[i] specifies the track ID of the i-th track carrying the "vopi" sample group.
[0376] Figure 1 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input terminal 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or it may be in a compressed or encoded format. Input terminal 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi interfaces or cellular interfaces).
[0377] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video codec techniques. The output of codec component 1904 may be stored or transmitted via connected communication, as represented by component 1906. Component 1908 may use the bitstream (or codec) representation of the stored or transmitted video received at input 1902 to generate pixel values or displayable video to be sent to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the result of codec are performed by the decoder.
[0378] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0379] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing hardware 3606 may be at least partially included in processor 3602 (e.g., a graphics coprocessor).
[0380] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0381] like Figure 4 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.
[0382] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0383] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded images and associated data. The coded images are coded representations of the images. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via I / O interface 116 through network 130a. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0384] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0385] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or may be external to destination device 120 configured to interface with an external display device.
[0386] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.
[0387] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 4 The video encoder 114 in the system 100 shown.
[0388] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0389] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0390] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0391] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 5 The example is represented separately.
[0392] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0393] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame inter-frame prediction (CIIP) mode, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0394] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0395] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0396] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0397] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0398] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.
[0399] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block to another video block by referencing the motion information of that other video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0400] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0401] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0402] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0403] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0404] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0405] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform the subtraction operation.
[0406] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0407] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0408] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video blocks, respectively, to reconstruct the residual video blocks based on the transform coefficient video blocks. The reconstruction unit 212 can add the reconstructed residual video blocks to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213.
[0409] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0410] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0411] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 4 The video decoder 114 in the system 100 shown.
[0412] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0413] exist Figure 6 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The encoding pass (pass) is the opposite of the decoding pass.
[0414] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.
[0415] The motion compensation unit 302 can generate motion compensation blocks that may perform interpolation based on an interpolation filter. The syntax elements may include an identifier for the interpolation filter to be used at sub-pixel precision.
[0416] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation of the sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.
[0417] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode one or more frames and / or one or more stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information used to decode the encoded video sequence.
[0418] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks based on spatially adjacent blocks. Dequantization unit 303 dequantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0419] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for display on the display device.
[0420] The following is a list of preferred embodiments.
[0421] The following schemes illustrate example embodiments of the techniques discussed in the previous chapter (e.g., items 1 and 2).
[0422] 1. A method for processing visual media (e.g., Figure 3 The method 700 described in the text includes: performing (702) a conversion between visual media data and a file representing a bitstream of the visual media data stored according to format rules; wherein the file includes file-level information for all operation points included in the file, wherein the file-level information includes information about the tracks required for each operation point.
[0423] 2. The method according to Scheme 1, wherein the format rules allow the track to include layers and sub-layers that are not needed for the corresponding operation point.
[0424] 3. The method according to any one of schemes 1-2, wherein the information of the track required for each operating point is included in the vopi sample group entry.
[0425] The following scheme illustrates an example embodiment of the technology discussed in the previous chapter (e.g., Project 3).
[0426] 4. A visual media processing method, comprising: performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the format rules specify skipping layer correlation information from vopi sample group entries when all layers are irrelevant.
[0427] The following schemes illustrate example embodiments of the techniques discussed in the previous chapter (e.g., items 5 and 6).
[0428] 5. A visual media processing method, comprising: performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the format rules define rules associated with the processing of operation point entity groups (opegs) in the bitstream representation.
[0429] 6. The method according to Scheme 5, wherein the format rules specify that, in the presence of opeg, each required track in the file provides an exact set of Video Codec Layer Network Abstraction Layers (VCL NALs) corresponding to each operation point in opeg.
[0430] 7. The method according to Scheme 6, wherein the format rules allow non-VCL units to be included in the track.
[0431] The following scheme illustrates example embodiments of the techniques discussed in the previous chapter (e.g., Project 4).
[0432] 8. A visual media processing method, comprising: performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the conversion includes performing an implicit reconstruction of the bitstream representation according to multiple tracks, wherein redundant access unit delimiter network access units (AUD NAL) are processed according to rules.
[0433] 9. The method according to scheme 8, wherein the rule specifies the removal of the AUD NAL unit.
[0434] 10. The method according to scheme 8, wherein the rule specifies rewriting the AUD NAL unit.
[0435] 11. The method according to any one of schemes 8-10, wherein the rule specifies that, in the case that at least one of the multiple picture units from different tracks of the access unit has an AUD NAL unit, the first picture unit has another AUD NAL unit.
[0436] 12. The method according to any one of schemes 8-10, wherein the rule specifies that, in the case where multiple picture units from different tracks of the access unit have AUD NAL units present, during decoding, only the AUD NAL units in the first picture unit are retained in the reconstructed access unit.
[0437] 13. The method according to any one of schemes 1-12, wherein the conversion includes generating a bitstream representation of the visual media data and storing the bitstream representation in a file according to format rules.
[0438] 14. The method according to any one of schemes 1-12, wherein the conversion includes parsing the file according to format rules to recover visual media data.
[0439] 15. A video decoding apparatus including a processor configured to implement the method according to one or more of claims 1 to 14.
[0440] 16. A video encoding apparatus including a processor configured to implement the method according to one or more of claims 1 to 14.
[0441] 17. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 14.
[0442] 18. A computer-readable medium on which a bitstream representation conforms to a file format generated according to any one of schemes 1 to 14.
[0443] 19. A method, apparatus or system described in this document.
[0444] Some preferred embodiments of the solutions listed above may include the following (e.g., items 1 and 2).
[0445] In some embodiments, a method for processing visual media (e.g., Figure 9A The method described in 910) includes: performing (912) a conversion between visual media data and a visual media file, wherein the visual media file stores a bitstream of visual media data in multiple tracks according to format rules, the format rules specifying file-level information including syntax elements, the syntax elements identifying one or more tracks from the multiple tracks that contain a specific type of sample group including operation point information.
[0446] In the above embodiments, the format rules specify that the visual media file includes file-level information for all operation points provided in the visual media file. The format rules also specify that for each operation point, the file-level information includes information about the corresponding track in the visual media file.
[0447] In some embodiments, format rules allow the required track for a particular operation point to include layers and sublayers that are not required for that particular operation point.
[0448] In some embodiments, the syntax element includes a bin containing a file-level container.
[0449] In some embodiments, the formatting rules specify that file-level information is included in the file-level bin.
[0450] In some embodiments, the format rules specify that file-level information is included in the movie-level bin.
[0451] In some embodiments, the formatting rules specify that file-level information is included in a track-level bin, which is identified in another track-level bin or another file-level bin.
[0452] In some embodiments, the formatting rules also specify that a particular type of sample set includes information about the track required for each operating point.
[0453] In some embodiments, the formatting rules also specify omitting information about the required track for each operating point from another specific type of sample set that includes layer information about the number of layers in the bitstream.
[0454] In some embodiments, visual media data is processed by a generic video codec (VVC), and multiple tracks are VVC tracks.
[0455] Some preferred embodiments may include the following (e.g., item 3).
[0456] In some embodiments, a method for processing visual media (e.g., Figure 9B The method described in (920) includes: performing (922) a conversion between visual media data and a visual media file, wherein the visual media file stores a bitstream of the visual media data according to format rules. The visual media file stores one or more tracks comprising one or more video layers. Whether a first set of syntax elements indicating layer relevance information is stored in the visual media file depends on whether a second syntax element indicating that all layers in the visual media file are unrelated has a value of 1.
[0457] In some embodiments, a first set of syntax elements is stored in a sample group that indicates information about one or more operation points stored in a visual media file.
[0458] In some embodiments, the formatting rule specifies that, in response to a second syntax element having a value of 1, a first set of syntax elements is omitted from the visual media file.
[0459] Some preferred embodiments of the solutions listed above may include the following aspects (e.g., item 4).
[0460] In some embodiments, a method for processing visual media data (e.g., Figure 9C The method described in 930) includes: performing (932) a conversion between visual media data and a visual media file, the visual media file storing a bitstream of visual media data in multiple tracks according to format rules, the format rules specifying how redundant access unit delimiter network access layer (AUD NAL) units stored in the multiple tracks are processed during implicit reconstruction of the bitstream according to the multiple tracks.
[0461] In some embodiments, the formatting rules specify the removal of redundant AUD NAL units during implicit refactoring.
[0462] In some embodiments, the formatting rules specify that redundant AUD NAL cells are rewritten during implicit refactoring.
[0463] In some embodiments, the format rules specify that, in response to implicit reconstruction, generating specific access units with a specific type different from transient random access point type or gradual decode refresh type based on multiple images of multiple tracks, the syntax field of a specific redundant AUD NAL included in the specific access unit is rewritten to a value of 0, indicating that the specific redundant AUD NAL does not represent a transient random access point or gradual decode refresh type.
[0464] In some embodiments, the format rules also specify that the value of the syntax element in the AUD NAL unit in the first picture unit (PU) is rewritten to 0, indicating that in the case that the second PU from a different track includes a picture that is not an intra-frame random access point picture or a progressively decoded refresh picture, the particular AUD NAL does not represent an instantaneous random access point or progressively decoded refresh type.
[0465] In some embodiments, the format rule specifies that at least one of a plurality of picture units from different tracks in response to the access unit has a first AUD NAL unit, and the first picture unit of the access unit generated according to implicit reconstruction includes a second AUD NAL unit.
[0466] In some embodiments, the format rules specify that multiple picture units from different tracks in response to the access unit include AUD NAL units, and a single AUD NAL unit corresponding to the AUD NAL unit of the first picture unit is included in the access unit generated according to the implicit reconstruction.
[0467] Some preferred embodiments of the solutions listed above may include the following aspects (e.g., item 7).
[0468] In some embodiments, a visual media processing method (e.g., Figure 9D The method described in 940 includes: performing (942) a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to format rules. The visual media file stores one or more tracks including one or more video layers. The visual media file includes information about operation points (OPs); wherein, format rules specify whether or how syntax elements are included in sample group entries and bins of OPs based on whether an OP contains a single video layer; wherein, syntax elements are configured to indicate an index of the set of output layers of the OP.
[0469] In some embodiments, the formatting rules specify that, in response to an OP containing a single video layer, syntax elements are omitted from the sample group entries and bins.
[0470] In some embodiments, the formatting rules specify that, in response to an OP containing more than one video layer, a syntax element is included after indicating that more than one video layer is identified.
[0471] In some embodiments, in response to omitting syntax elements from sample group entries and bins, the index of the output layer set of the OP is inferred to be equal to the index of the output layer set including a single video layer.
[0472] Some preferred embodiments of the solutions listed above can be combined with the following aspects (e.g., item 8).
[0473] In some embodiments, a visual media processing method (e.g., Figure 9E The method described in 950) includes: performing (952) a conversion between visual media data and a visual media file, the visual media file storing a bitstream of visual media data according to format rules; wherein the visual media file stores multiple tracks belonging to a specific type of entity group, and wherein the format rules specify that: in response to multiple tracks having a specific type of track reference to a group identifier, multiple tracks (A) omit carrying a specific type of sample group or (B) carry a specific type of sample group, such that the information in the specific type of sample group is consistent with the information in the specific type of entity group.
[0474] In some embodiments, multiple tracks represent bit streams.
[0475] In some embodiments, a particular type of entity group indicates that multiple tracks correspond precisely to operating points.
[0476] In some embodiments, a particular type of sample set includes information about which of a plurality of tracks correspond to the operating point.
[0477] Some preferred embodiments of the solutions listed above can be combined with the following aspects (e.g., items 5, 6, 9).
[0478] In some embodiments, a visual media processing method (e.g., Figure 9F The method described in 960) includes: performing (962) a conversion between visual media data and a visual media file, wherein the visual media file stores a bitstream of visual media data, wherein the visual media file includes multiple tracks; wherein the visual media file stores an entity group carrying information about operation points in the visual media file and a track carrying each operation point; and wherein a format rule specifies attributes of the visual media file in response to the entity group or sample group carrying information about each operation point stored in the visual media file.
[0479] In some embodiments, the format rules specify that the entity group provides a set of tracks for each operation point, carrying an exact set of Video Codec Layer (VCL) Network Abstraction Layer (NAL) units.
[0480] In some embodiments, the format rules also specify that redundant non-VCL NAL cells in the track set are removed during the reconstruction of the bitstream.
[0481] In some embodiments, the format rules specify that the entity group provides a set of tracks for each operation point, carrying an exact set of one or more layers and one or more sublayers.
[0482] In some embodiments, the format rules also specify that redundant non-video codec layer (VCL) network abstraction layer (NAL) units in the track set are removed during the reconstruction of the bitstream.
[0483] In some embodiments, the format rules specify that redundant Bitstream End (EOB) or Stream End (EOS) Network Abstraction Layer (NAL) units are removed during the process of reconstructing the bitstream based on multiple tracks.
[0484] In some embodiments, the format rules specify that access delimiter units (AUDs) are removed or rewritten during the process of reconstructing a bitstream based on multiple tracks.
[0485] In some embodiments, the formatting rules specify that, apart from pre-specified file-level bins, containers of entity group bins associated with entity groups are not allowed to be stored at any level in the visual media file.
[0486] In some embodiments, the pre-specified file-level bin is a group list bin included in the file-level metadata bin.
[0487] In some embodiments, the format rules specify that, in response to the bitstream being stored in a single track within a visual media file, neither the entity group nor the sample group is permitted to store the bitstream.
[0488] In the above embodiments, the conversion includes storing the bitstream into a visual media file according to format rules.
[0489] In the above embodiments, the conversion includes parsing the visual media file according to format rules to reconstruct the bitstream.
[0490] In some embodiments, the visual media file parsing apparatus may include a processor configured to implement the methods disclosed in the above embodiments.
[0491] In some embodiments, the visual media file writing device includes a processor configured to implement the methods disclosed in the above embodiments.
[0492] Some embodiments may include a computer program product on which computer code is stored. When the code is executed by a processor, it causes the processor to perform any of the methods described in the above embodiments.
[0493] Some embodiments may include a computer-readable medium having a visual media file stored thereon, the visual media file conforming to a file format generated according to any of the methods described above.
[0494] In the scheme described herein, the encoder can conform to the format rules by generating a codec representation based on those rules. In the scheme described herein, the decoder can parse the syntax elements in the codec representation using the format rules, knowing whether or not the syntax elements exist, to generate the decoded video.
[0495] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined by the syntax, the bitstream representation of the current video block can, for example, correspond to bits located at different positions within the bitstream or extended at different positions within the bitstream. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the above scheme. Similarly, the encoder can determine whether certain syntax fields are included and generate the codec representation accordingly by including or excluding syntax fields from the codec representation. The term "visual media" can refer to video or image, and the term "visual media processing" can refer to video processing or image processing.
[0496] The disclosed and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a synthetic material that implements a machine-readable propagating signal, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0497] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form (including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on a single computer or on multiple computers located at one site or distributed across multiple sites and interconnected via a communication network.
[0498] The processes and logic flows described in this document can be executed by one or more programmable processors, which execute one or more computer programs to perform functions by manipulating input data and generating output. These processes and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0499] For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The fundamental components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, to receive data from, send data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM discs. The processor and memory may be complemented or integrated therein by dedicated logic circuitry.
[0500] Although this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent document within the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, in some cases, one or more features from a claimed combination may be removed from that combination, and the claimed combination may refer to a sub-combination or a variation of a sub-combination.
[0501] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all of the shown operations to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0502] Only a few implementations and examples have been described, and other implementations, enhancements and variations may be made based on what is described and shown in this patent document.
Claims
1. A method of visual media processing, comprising: performing a conversion between visual media data and a visual media file storing a bitstream of the visual media data according to a format rule; wherein the visual media file stores a plurality of tracks belonging to a group of operation point entities, wherein the format rule specifies that, in response to the plurality of tracks having a particular type of track reference to a group identifier, the plurality of tracks (A) omit a group of operation point information samples or (B) carry a group of operation point information samples such that information in the group of operation point information samples is consistent with information in the group of operation point entities, wherein the group of operation point entities indicates that the plurality of tracks exactly correspond to an operation point, and wherein the group of operation point information samples includes information about which of the plurality of tracks correspond to the operation point.
2. The method of claim 1, wherein, the plurality of tracks represent the bitstream.
3. The method of claim 1, wherein the visual media file stores the bitstream of the visual media data; wherein the visual media file stores a group of operation point entities carrying information about operation points in the visual media file and a track carrying each operation point; and wherein the format rule specifies properties of the visual media file in response to the visual media file storing a group of operation point entities carrying information of each operation point or a group of operation point information samples, wherein the format rule specifies a property that, in addition to a pre-designated file level box, a container of a group box associated with the group of operation point entities is not allowed to be stored in the visual media file at any level.
4. The method of claim 3, wherein, the format rule specifies that the group of operation point entities provides, for each operation point, a set of tracks carrying an exact set of video coding layer (VCL) network abstraction layer (NAL) units.
5. The method of claim 4, wherein, the format rule further specifies that redundant non-VCL NAL units in the set of tracks are removed during reconstruction of the bitstream.
6. The method of claim 3, wherein, the format rule specifies that the group of operation point entities provides, for each operation point, a set of tracks carrying an exact set of one or more layers and one or more sub-layers.
7. The method of claim 6, wherein, the format rule further specifies that redundant non-video coding layer (VCL) network abstraction layer (NAL) units in the set of tracks are removed during reconstruction of the bitstream.
8. The method of claim 3, wherein, the format rule specifies that, in a process of reconstructing a bitstream from the plurality of tracks, redundant bitstream end of bitstream (EOB) or end of stream (EOS) network abstraction layer (NAL) units are removed.
9. The method of claim 3, wherein, the format rule specifies that, in a process of reconstructing a bitstream from the plurality of tracks, access delimiter units (AUDs) are removed or overwritten.
10. The method of claim 3, wherein, the pre-designated file level box is a group list box included in a file level metadata box.
11. The method of claim 3, wherein, the format rule specifies that, in response to the bitstream being stored in a single track in the visual media file, none of the group of operation point entities and / or the group of operation point information samples are allowed to be stored for the bitstream.
12. The method of any one of claims 1-11, wherein, the conversion includes generating the visual media file and storing the bitstream into the visual media file according to the format rule.
13. The method of any one of claims 1-11, wherein, The conversion includes parsing the visual media file according to the format rule to reconstruct the bitstream.
14. An apparatus for processing visual media, comprising a processor configured to implement a method according to any of claims 1-13.
15. A non-transitory computer-readable storage medium having stored thereon processor-executable code, which, when executed, causes a processor to implement a method according to any of claims 1-13.
16. A method for storing a visual media file, comprising: generating a visual media file, wherein the visual media file stores a bitstream of visual media data according to a format rule; storing the visual media file in a non-transitory computer-readable recording medium; wherein the visual media file stores a plurality of tracks belonging to an operation point entity group, and wherein the format rule specifies that, in response to the plurality of tracks having a particular type of track reference to a group identifier, the plurality of tracks (A) omit a group of operation point information samples or (B) carry a group of operation point information samples such that information in the group of operation point information samples is consistent with information in the operation point entity group, wherein the operation point entity group indicates that the plurality of tracks exactly correspond to an operation point, and wherein the group of operation point information samples includes information about which of the plurality of tracks correspond to the operation point.
17. An apparatus for processing visual media, comprising a processor and a non-transitory memory storing instructions, wherein, the instructions, when executed by the processor, cause the processor to perform: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule; wherein the visual media file stores a plurality of tracks belonging to an operation point entity group, wherein the format rule specifies that, in response to the plurality of tracks having a particular type of track reference to a group identifier, the plurality of tracks (A) omit a group of operation point information samples or (B) carry a group of operation point information samples such that information in the group of operation point information samples is consistent with information in the operation point entity group, wherein the operation point entity group indicates that the plurality of tracks exactly correspond to an operation point, and wherein the group of operation point information samples includes information about which of the plurality of tracks correspond to the operation point.
18. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform: performing a conversion between visual media data and a visual media file, the visual media file storing a bitstream of the visual media data according to a format rule; wherein the visual media file storing a plurality of tracks belonging to an operation point entity group, wherein the format rule specifies that, in response to the plurality of tracks having a particular type of track reference to a group identifier, the plurality of tracks (A) omit a group of operation point information samples or (B) carry a group of operation point information samples such that information in the group of operation point information samples is consistent with information in the operation point entity group, wherein the operation point entity group indicates that the plurality of tracks exactly correspond to an operation point, and wherein the group of operation point information samples includes information about which of the plurality of tracks correspond to the operation point. The operation point information sample group includes information about which of the plurality of tracks correspond to the operation point.
Citation Information
Patent Citations
Alignment of operation point sample group in multi-layer bitstreams file format
CN108141617A
Method, device and computer program for dynamically setting operational origin descriptors for retrieving media data and metadata from encapsulated bitstreams
JP2018524877A