Decoder configuration information in vvc video codec
By adjusting the parameter order and the grouping type of the 'rolling' sample group in the VVC decoder configuration record, the problem of incomplete decoder configuration information in the VVC video file format was solved, achieving a more accurate and efficient decoding process.
Patent Information
- Application Number
- CN202180073375.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-26
- Filing Date
- 2021-10-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-10-26
AI Technical Summary
The existing VVC video file format fails to effectively signal necessary parameters, such as the size of the decoded image buffer and the maximum image output reordering, in the decoder configuration information and the signaling notification of the 'rolling' sample group. Furthermore, the signaling notification syntax structure of the PTL information is incomplete, resulting in an unclear decoding process.
Modify the parameter order in the VVC decoder configuration record so that the number of sublayers precedes the PTL record, explicitly specify parameters such as the signaling notification decoder image buffer size, and perform byte alignment before or after the PTL information. Adjust the grouping type parameter of the 'rolling' sample group to clearly describe the relationship between access points and layers.
It implements more complete and clear decoder configuration information signaling, ensuring the accuracy and consistency of parameters during the decoding process, and improving the decoding efficiency and reliability of video file formats.
Smart Images

Figure CN116508322B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This patent application claims the benefit of International Application No. PCT / CN2020 / 123540 by Ye-Kui Wang et al., filed October 26, 2020, and entitled “Signalling of Decoder Configuration Information and the ‘Roll’ Sample Group in VVC Video Files,” which is incorporated by reference herein. TECHNICAL FIELD
[0003] This patent document relates to the generation, storage, and consumption of digital audio-visual media information in file formats. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage can continue to grow. SUMMARY
[0005] A first aspect relates to a method for processing video data, comprising: performing a conversion between visual media data and a visual media data file, the visual media data file comprising a Versatile Video Coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein the VVC decoder configuration record comprises a number of the one or more sub-layers and one or more VVC Profile Tier Level (PTL) records of the one or more sub-layers based on the number of the one or more sub-layers.
[0006] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the conversion comprises: receiving a media file comprising a Versatile Video Coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sub-layers; parsing the VVC decoder configuration record to obtain a number of the one or more sub-layers and one or more VVC PTL records of the one or more sub-layers based on the number of the one or more sub-layers; and decoding the one or more sub-layers based on the VVC PTL records.
[0007] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the converting includes: encoding the plurality of pictures into one or more sub-layers in a visual media file; determining a number of the one or more sub-layers; encoding a VVC decoder configuration record into the media file, the VVC decoder configuration including the number of the one or more sub-layers and one or more VVC PTL records for the one or more sub-layers; and storing the visual media file in a memory.
[0008] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the number of the one or more sub-layers is signaled in the VVC decoder configuration record prior to the VVC PTL record.
[0009] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the VVC decoder configuration record includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL record is located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
[0010] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the number of the one or more sub-layers is located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
[0011] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and wherein the VVC PTL record is located after the reserved bit.
[0012] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the VVC decoder configuration record further includes a reserved bit located after the VVC PTL record.
[0013] Optionally, in any of the preceding aspects, another implementation of this aspect provides that the VVC decoder configuration record further includes a maximum required size of a decoded picture buffer, a maximum picture output reordering, a maximum latency, a gradual decoding refresh (GDR) picture enable flag, a clean random access (CRA) picture enable flag, a reference picture resampling enable flag, a spatial resolution change with coded layer video sequence (CLVS) enable flag, a subpicture partitioning enable flag, a maximum number of subpictures in each picture, a wavefront parallel processing (WPP) enable flag, a tile partitioning enable flag, a maximum number of tiles per picture, a slice partitioning enable flag, a rectangular slice enable flag, a raster-scan slice enable flag, a maximum number of slices per picture, or a combination thereof.
[0014] Optionally, in any of the preceding aspects, another implementation of the aspect provides that when the VVC decoder configuration record includes the VVC PTL record, the VVC decoder configuration record further includes one or more of: maximum required size of a decoded picture buffer, maximum picture output reordering, maximum latency, GDR picture enable flag, CRA picture enable flag, reference picture resampling enable flag, spatial resolution change with CLVS enable flag, subpicture partitioning enable flag, maximum number of subpictures in each picture, WPP enable flag, tile partitioning enable flag, maximum number of tiles per picture, slice partitioning enable flag, rectangular slice enable flag, raster-scan slice enable flag, maximum number of slices per picture.
[0015] Optionally, in any of the preceding aspects, another implementation of the aspect provides that in the VVC decoder configuration record, all syntax elements following the VVC PTL record are byte-aligned.
[0016] A second aspect relates to an apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between visual media data and a visual media data file, the visual media data file comprising a Versatile Video Coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein the VVC decoder configuration record comprises a number of the one or more sub-layers and one or more VVC profile tier level (PTL) records for the one or more sub-layers based on the number of the one or more sub-layers.
[0017] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the conversion comprises: receiving a media file comprising the VVC decoder configuration record and the plurality of pictures coded into the one or more sub-layers; parsing the VVC decoder configuration record to obtain the number of the one or more sub-layers and the one or more VVC PTL records for the one or more sub-layers based on the number of the one or more sub-layers; and decoding the one or more sub-layers based on the VVC PTL records.
[0018] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the conversion comprises: encoding the plurality of pictures into the one or more sub-layers in a visual media file; determining a number of the one or more sub-layers; encoding a VVC decoder configuration record into the media file, the VVC decoder configuration comprising the number of the one or more sub-layers and one or more VVC PTL records for the one or more sub-layers; and storing the visual media file in a memory.
[0019] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the number of one or more sub-layers is signaled in the VVC decoder configuration record prior to the VVC PTL record.
[0020] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the VVC decoder configuration record further includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL record is located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
[0021] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the number of one or more sub-layers is located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
[0022] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and wherein the VVC PTL record is located after the reserved bit.
[0023] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the VVC decoder configuration record further includes a reserved bit located after the VVC PTL record.
[0024] Optionally, in any of the preceding aspects, another implementation of the aspect provides that the VVC decoder configuration record further includes a maximum required size of a decoded picture buffer, a maximum picture output reordering, a maximum latency, a GDR picture enable flag, a CRA picture enable flag, a reference picture resampling enable flag, a spatial resolution change with CLVS enable flag, a subpicture partitioning enable flag, a maximum number of subpictures in each picture, a WPP enable flag, a tile partitioning enable flag, a maximum number of tiles per picture, a slice partitioning enable flag, a rectangular slice enable flag, a raster-scan slice enable flag, a maximum number of slices per picture, or a combination thereof.
[0025] Optionally, in any of the preceding aspects, another implementation of this aspect provides that when the VVC decoder configuration record includes the VVC PTL record, the VVC decoder configuration record further includes one or more of: maximum required size of the decoded picture buffer, maximum picture output reordering, maximum latency, GDR picture enable flag, CRA picture enable flag, reference picture resampling enable flag, spatial resolution change with CLVS enable flag, subpicture partitioning enable flag, maximum number of subpictures in each picture, WPP enable flag, tile partitioning enable flag, maximum number of tiles per picture, slice partitioning enable flag, rectangular slice enable flag, raster-scan slice enable flag, maximum number of slices per picture.
[0026] Optionally, in any of the preceding aspects, another implementation of this aspect provides that in the VVC decoder configuration record, all syntax elements following the VVC PTL record are byte-aligned.
[0027] A third aspect relates to a non-transitory computer readable medium containing a computer program product for use by a video coding device, the computer program product containing computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to perform the method of any of the preceding aspects.
[0028] For the sake of clarity, any one of the preceding embodiments can be combined with any one or more other preceding embodiments to form new embodiments within the scope of the present disclosure.
[0029] These and other features can be more fully understood with reference to the following detailed description in connection with the attached drawings and BRIEF DESCRIPTION OF DRAWINGS
[0030] For a more complete understanding of the present disclosure, reference is now made to the following brief description of the drawings and detailed description in connection with the drawings and the appended claims.
[0031] Figure 1 is a schematic diagram of an example media file containing a versatile video coding (VVC) bitstream of video data.
[0032] Figure 2 is a flowchart of an example method of encoding a VVC decoder configuration record.
[0033] Figure 3 is a flowchart of an example method of decoding a VVC decoder configuration record.
[0034] Figure 4 is a block diagram showing an example video processing system.
[0035] Figure 5is a block diagram of an example video processing device.
[0036] Figure 6 is a flowchart of an example method of video processing.
[0037] Figure 7 is a block diagram illustrating an example video coding system.
[0038] Figure 8 is a block diagram illustrating an example encoder.
[0039] Figure 9 is a block diagram illustrating an example decoder.
[0040] Figure 10 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0041] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations presented herein, but can be modified in any manner within the scope of the appended claims along with their full scope of equivalents.
[0042] VVC, also known as H.266, uses the term in certain descriptions merely for ease of understanding and not for limiting the scope of the disclosed technology. Thus, the technology described herein also applies to other video codec protocols and designs. In this document, editorial changes to the text are indicated by struck-out text that is cancelled and italicized text that is added.
[0043] Example implementations of the aspects described above are described below.
[0044] This document relates to video file formats. In particular, this document relates to signaling of decoder configuration information and “rolling” sample groups in media files carrying Versatile Video Coding (VVC) video bitstreams based on the ISO Base Media File Format (ISOBMFF). These ideas can be applied individually or in various combinations for video bitstreams coded by any codec (e.g., the VVC standard), and for any video file format (e.g., the VVC video file format being developed).
[0045] Adaptive color transform (ACT), adaptive loop filter (ALF), adaptive motion vector resolution (AMVR), adaptive parameter set (APS), access unit (AU), access unit delimiter (AUD), advanced video coding (Rec. ITU-T H.264 | ISO / IEC 14496-10) (AVC), bi-prediction (B), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), block-based delta pulse code modulation (BDPCM), buffering period (BP), context-based adaptive binary arithmetic coding (CABAC), coded block (CB), constant bit rate (CBR), cross-component adaptive loop filter (CCALF), coded layer video stream (CLVS), coded picture buffer (CPB), clean random access (CRA), cyclic redundancy check (CRC), coding tree block (CTB), coding tree unit (CTU), decoding capability information (DCI), dependent random access point (DRAP), decoding unit (DU), decoding unit information (DUI), exponential Golomb (EG), exponential Golomb of order k (EGk), end of bitstream (EOB), end of sequence (EOS), filler data (FD), first-in first-out (FIFO), fixed length (FL), green, blue, and red (GBR), general constraint information (GCI), gradual decoding refresh (GDR), geometric partition mode (GPM), high efficiency video coding (also referred to as Rec. ITU-T H.265|ISO / IEC 23008-2)(HEVC), hypothetical reference decoder (HRD), hypothetical stream scheduler (HSS), intra (I), intra block copy (IBC), instantaneous decoding refresh (IDR), inter-layer reference picture (ILRP), intra random access point (IRAP), low-frequency non-separable transform (LFNST), least probable symbol (LPS), least significant bit (LSB), long-term reference picture (LTRP), luma mapping with chroma scaling (LMCS), matrix-based intra prediction (MIP), most probable symbol (MPS), most significant bit (MSB), multiple transform selection (MTS), motion vector prediction (MVP), network abstraction layer (NAL), output layer set (OLS), operation point (OP), operation point information (OPI), prediction (P), picture header (PH), picture order count (POC), picture parameter set (PPS), prediction refinement with optical flow (PROF), picture timing (PT), picture unit (PU), quantization parameter (QP), random access decodable leading picture (RADL), random access skipped leading picture (RASL), raw byte sequence payload (RBSP), red, green, and blue (RGB), reference picture list (RPL), sample adaptive offset (SAO), sample aspect ratio (SAR), supplemental enhancement information (SEI), slice header (SH), subpicture level information (SLI), string of data bits (SODB), sequence parameter set (SPS), short-term reference picture (STRP), stepping through time sub-layer access (STSA), truncated rice (TR), variable bit rate (VBR), video coding layer (VCL), video parameter set (VPS), versatile supplemental enhancement information (also referred to as Rec. ITU-T H. 274 | ISO / IEC 23002-7) (VSEI), video usability information (VUI), versatile video coding (also referred to as Rec. ITU-T H. 266 | ISO / IEC 23090-3) (VVC), and wavefront parallel processing (WPP).
[0046] Video coding standards have evolved mainly through the development of ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263 standards, ISO / IEC produced MPEG-1 and MPEG-4 Visual standards, and both organizations jointly produced the H.262 / MPEG-2 Video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction is combined with transform coding. To explore further video coding techniques beyond HEVC, the Joint Video Exploration Team (JVET) was founded by the Video Coding Experts Group (VCEG) and MPEG. The JVET took many approaches and incorporated them into a reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as the Joint Video Team (JVT). VVC is a coding standard, targeting at 50% bitrate reduction compared to HEVC. VVC has been finalized by the JVET.
[0047] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for a wide range of applications, including uses such as television broadcasting, video conferencing or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.
[0048] Media streaming applications are typically based on Internet Protocol (IP), Transmission Control Protocol (TCP) and Hyper Text Transfer Protocol (HTTP) transport methods, and often rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats with ISOBMFF and DASH, file format specifications specific to the video format, such as the AVC file format and the HEVC file format, will be used to encapsulate video content in ISOBMFF tracks and in DASH representations and segments. Information about the video bitstream, such as profiles, tiers and levels and many other information, will be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g., for selecting appropriate media segments, both for initialization at the beginning of a streaming session and for stream adaptation during a streaming session.
[0049] Similarly, to use image formats with ISOBMFF, a file format specification specific to the image format will be adopted, such as AVC Image File Format and HEVC Image File Format. MPEG is developing VVC Video File Format, which is an ISOBMFF-based file format for storing VVC video content. MPEG is developing VVC Image File Format, which is an ISOBMFF-based file format for storing image content coded using VVC.
[0050] The following is the design of some VVC file format features based on VVC Image File Format and MPEG. This subclause specifies the decoder configuration information for ISO / IEC 23090-3 video content. This record contains the size of the length field used in each sample entry to indicate the length of NAL units it contains as well as the length of parameter sets, DCI, OPI, and SEI NAL units if stored in the sample entry. This record is an outer framework (whose size is provided by the structure containing it). This record contains a version field. The specification of this version defines version 1 of this record. Changes in the version number indicate non-compatible changes to the record. If the version number is not recognized, the reader / writer shall not attempt to decode this record or the stream to which it applies. Compatible extensions to this record extend it and do not change the configuration version code. Readers / writers shall be prepared to ignore unrecognized data definitions beyond their understanding.
[0051] When a track itself contains a VVC bitstream or by resolving the "subp" track reference, a VVC profile tier level record (VvcPTLRecord) shall be present in the decoder configuration record and, in this case, the specific output layer set of the VVC bitstream is indicated by the field output_layer_set_idx. If ptl_present_flag is equal to zero in the decoder configuration record of a track, the track shall have an "oref" track reference to an ID that can refer to a VVC track or an "opeg" entity group. The values of the syntax elements VvcPTLRecord, chroma_format_idc, and bit_depth_minus8 are valid for all parameter sets (referred to as "all parameter sets" in the following sentences of this paragraph) that are referenced when decoding the stream described by this record. In particular, the following restrictions can apply:
[0052] The tier indication general_profile_idc shall indicate the tier to which the output layer set identified by output_layer_set_idx in this configuration record conforms. If different tiers are marked for different CVSs of the output layer set identified by output_layer_set_idx in this configuration record, the stream can need to be examined to determine which tier, if any, the entire stream conforms to. If the entire stream is not examined, or if the examination shows that there is no tier to which the entire stream conforms, the entire stream shall be split into two or more sub-streams with separate configuration records that can satisfy these rules. The level indication general_tier_flag shall indicate the level equal to or greater than the highest level indicated in all profile_tier_level( ) syntax structures (in all parameter sets) to which the output layer set identified by output_layer_set_idx in this configuration record conforms.
[0053] Each bit in general_constraint_info shall be set only if the bit is set in all general_constraints_info( ) syntax structures of all profile_tier_level( ) syntax structures (in all parameter sets) to which the output layer set identified by output_layer_set_idx in this configuration record conforms. The level indication general_level_idc shall indicate the level equal to or greater than the highest level in all profile_tier_level( ) syntax structures (in all parameter sets) to which the output layer set identified by output_layer_set_idx in this configuration record conforms.
[0054] The following constraint applies to chroma format identifier (chroma format idc). If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps chroma format idc defined in ISO / IEC 23090-3 shall be the same in all SPS referred to by VCL NAL units in the sample described by the current sample entry, and the value of chroma format idc shall be equal to the value of sps chroma format idc. Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), the value of vps ols dpb chroma format[MultiLayerOlsIdx[output layer set idx]] shall be the same for all CVS to which the current sample entry applies, and the value of chroma format idc shall be equal to the value of vps ols dpb chroma format[MultiLayerOlsIdx[output layer set idx]].
[0055] The following constraint applies to bit depth minus 8. If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps bitdepth minus 8 shall be the same in all SPS referred to by VCL NAL units in the sample described by the current sample entry, and the value of bit depth minus 8 shall be equal to the value of sps bitdepth minus 8. Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), the value of vps ols dpb bitdepth minus 8[MultiLayerOlsIdx[output layer set idx]] shall be the same for all CVS to which the current sample entry applies, and the value of bit depth minus 8 shall be equal to the value of vps ols dpb bitdepth minus 8[MultiLayerOlsIdx[output layer set idx]].
[0056] The following constraint applies to picture width. If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_pic_width_max_in_luma_samples as defined in ISO / IEC 23090-3 shall be the same in all SPSs that are referred to by VCL NAL units in the sample to which the current sample entry describes applies, and the value of picture width shall be equal to the value of sps_pic_width_max_in_luma_samples. Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), the value of vps_ols_dpb_pic_width[MultiLayerOlsIdx[output_layer_set_idx]] shall be the same for all CVSs to which the current sample entry describes applies, and the value of pic width shall be equal to the value of vps_ols_dpb_pic_width[MultiLayerOlsIdx[output_layer_set_idx]].
[0057] The following constraint applies to picture height. If the VVC stream to which the configuration record applies is a single-layer bitstream, the value of sps_pic_height_max_in_luma_samples as defined in ISO / IEC 23090-3 shall be the same in all SPSs that are referred to by VCL NAL units in the sample to which the current sample entry describes applies, and the value of picture height shall be equal to the value of sps_pic_height_max_in_luma_samples. Otherwise (the VVC stream to which the configuration record applies is a multi-layer bitstream), the value of vps_ols_dpb_pic_height[MultiLayerOlsIdx[output_layer_set_idx]] shall be the same for all CVSs to which the current sample entry describes applies, and the value of pic height shall be equal to the value of vps_ols_dpb_pic_height[MultiLayerOlsIdx[output_layer_set_idx]].
[0058] Explicit indication about the chroma format and bit depth used by the VVC video elementary stream and other format information is provided in the VVC decoder configuration record. If two sequences indicate different color space or bit depth in their VUI information, then also two different VVC sample entries are taken.
[0059] There are sets of arrays to carry initialization non-VCL NAL units. The NAL unit types are limited to represent DCI, OPI, VPS, SPS, PPS, prefix APS and prefix SEI NAL units. The reserved NAL unit types can be further defined, and the reader should ignore the arrays with the reserved or disallowed NAL unit type values. This lenient behavior is designed not to raise errors, allowing the possibility of backward-compatible extension of these arrays in further specifications. The NAL units carried in the sample entries follow the AUD and OPI NAL units (if any) or are otherwise contained in the beginning of the access unit reconstructed from the first sample of the reference sample entry.
[0060] It is recommended that the arrays are ordered in the sequence of DCI, OPI, VPS, SPS, PPS, prefix APS, prefix SEI.
[0061] The example syntax of VVCPTLRecord and VvcDecoderConfigurationRecord is as follows:
[0062]
[0063]
[0064] The semantics of the above syntax elements are as follows.
[0065] num_bytes_constraint_info specifies the length of the general_constraint_info field. The length of the general_constraint_info field is num_bytes_constraint_info * 8 - 2 bits. The value shall be greater than 0. The value equal to 1 indicates that gci_present_flag in the general_constraint_info() syntax structure represented by the general_constraint_info field is equal to 0.
[0066] general_profile_idc, general_tier_flag, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, general_constraint_info, sublayer_level_present[j], sublayer_level_idc[i], num_sub_profiles, and general_sub_profile_idc[j] contain the matching values of the fields or syntax structures general_profile_idc, general_tier_flag, general_level_idc, ptl_frame_only_constraint_flag, ptl_multilayer_enabled_flag, general_constraint_info(), ptl_sublayer_level_present[i], sublayer_level_idc[i], ptl_num_sub_profiles, and general_sub_profile_idc[j] indicate the stream to which this configuration record applies.
[0067] lengthSizeMinusOne plus 1 specifies the byte length of the NALUnitLength field in VVC video stream samples in the stream to which this configuration record applies. For example, a size of one byte is indicated by the value 0. The value of this field shall be one of 0, 1, or 3, corresponding to lengths coded with 1, 2, or 4 bytes, respectively.
[0068] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to the operation point specified by output_layer_set_idx and numTemporalLayers, and all NAL units in the track belong to this operation point. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular operation point, but can contain VVC bitstreams corresponding to multiple output layer sets, or can contain one or more individual layers that do not form an output layer set or individual sub-layers other than sub-layers with Temporalld equal to 0.
[0069] track_ptl specifies the profile, tier, and level of the output layer set represented by the VVC bitstream contained in the track.
[0070] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided to the VVC decoder by an external device or an OPI NAL unit for decoding the bitstream contained in the track.
[0071] avgFrameRate gives the average frame rate in units of frames / (256 seconds) for the stream to which this configuration record applies. A value of 0 indicates that the average frame rate is not specified. When the track contains multiple layers and is a reconstruction sample for an operation point specified by output_layer_set_idx and numTemporalLayers, this gives the average access unit rate of the bitstream for the operation point.
[0072] constantFrameRate equal to 1 indicates that the stream to which this configuration record applies has a constant frame rate. A value of 2 indicates that the representation of each temporal layer in the stream has a constant frame rate. A value of 0 indicates that the stream can or can not have a constant frame rate. When the track contains multiple layers and is a reconstruction sample for an operation point specified by output_layer_set_idx and numTemporalLayers, this gives an indication of whether the bitstream for the operation point has a constant access unit rate.
[0073] numTemporalLayers greater than 1 indicates that the track to which this configuration record applies is temporally scalable and the number of temporal layers (also referred to as temporal sublayers or sublayers) contained is equal to numTemporalLayers. A value of 1 indicates that the track to which this configuration record applies is not temporally scalable. A value of 0 indicates that it is unknown whether the track to which this configuration record applies is temporally scalable.
[0074] chroma_format_id indicates the chroma format applied to the track.
[0075] picture_width specifies the maximum picture width in luma samples applied to the track.
[0076] picture_height specifies the maximum picture height in luma samples applied to the track.
[0077] bit_depth_minus8 specifies the bit depth applied to the track.
[0078] numArrays specifies the number of arrays of NAL units of the indicated type(s).
[0079] array_completeness equal to 1 indicates that all NAL units of the given type are in the next array and none are in the stream; when equal to 0, it indicates that additional NAL units of the indicated type can be in the stream; the allowed values are constrained by the sample entry name.
[0080] NAL_unit_type indicates the type of NAL units in the following array (all shall be of this type); it is constrained to take one of the values indicating a DCI, OPI, VPS, SPS, PPS, prefix APS, or prefix SEI NAL unit.
[0081] numNalus indicates the number of NAL units of the indicated type included in the configuration record for the stream to which this configuration record applies. SEI arrays shall only contain declarative SEI messages, i.e., those providing information about the entire stream. One example of such SEI can be the user data SEI.
[0082] nalUnitLength specifies the byte length of the NAL unit.
[0083] nalUnit contains a DCI, OPI, VPS, SPS, PPS, APS, or declarative SEI NAL unit.
[0084] Random access recovery point sample groups, also referred to as "roll" sample groups, are used to provide information for recovery points for gradual decoding refresh. When "roll" sample groups are used with a VVC track, the syntax and semantics of grouping_type_parameter are specified to be the same as for "sap" sample groups.
[0085] layer_id_method_idc equal to 0 and 1 are used when the pictures of the target layer of the samples mapped to the "roll" sample group are GDR pictures. When layer_id_method_idc is equal to 0, the "roll" sample group specifies the behavior for all layers in the track.
[0086] The semantics of layer_id_method_idc equal to 1 are specified here.
[0087] When all pictures of the target layers that are not samples mapped to a "rolling" sample group are GDR pictures, layer_id_method_idc equal to 2 and 3 are used, and for pictures of the target layers that are not GDR pictures, the following applies: the PPS referred to has pps_mixed_nalu_types_in_pic_flag equal to 1, and for each subpicture index i in the range of 0 to sps_num_subpics_minus1, both of the following are true: sps_subpic_treated_as_pic_flag[ i ] is equal to 1, and there is at least one IRAP subpicture in the current sample or after in the same CLVS with the same subpicture index i. When layer_id_method_idc is equal to 2, the "rolling" sample group specifies the behavior of all layers present in the track. The semantics of layer_id_method_idc equal to 3 are specified here. When the reader starts decoding using samples marked with layer_id_methoc_idc equal to 2 or 3, the reader-writer needs to further modify the SPS, PPS and PH NAL units of the bitstream so that when the bitstream started with samples marked as belonging to this sample group and layer_id_method_idc equal to 2 and 3 is a consistent bitstream, any SPS referred to by such samples has sps_gdr_enabled_flag equal to 1, any PPS referred to by such samples has pps_mixed_nalu_types_in_pic_flag equal to 0, all VCL NAL units of the AU have nal_unit_type equal to GDR_NUT, and any picture header of the AU has ph_gdr_pic_flag equal to 1 and the value of ph_recovery_poc_cnt corresponds to the roll_distance of the sample group the AU belongs to. When a "rolling" sample group involves a dependent layer instead of its reference layer(s), the sample group indicates features that apply when all reference layers of the dependent layer are available and decoded. A sample group can be used to enable the decoding of a prediction layer.
[0088] When layer_id_method_idc is equal to 1, each bit of the target_layers field represents a layer carried in the track. Since this field is only 28 bits long, the indication of the SAP in the track is limited to a maximum of 28 layers. Each bit of this field starting with the least significant bit (LSB) will be mapped in ascending order of layer_id values to the list of layer_id values signaled in the layer information sample group ("linf") associated with this sample.
[0089] The following are example technical problems solved by the disclosed technical solutions. The latest design of VVC video file format with signaling of decoder configuration information and ‘rolling’ sample groups has the following problems. First, in VvcDecoderConfigurationRecord, when signaling profile, tier, and level information (PTL), picture format parameters, including color format, bit depth, picture width, and picture height, are signaled. These information can be used for content selection purposes. However, there are other parameters that can be used for content selection purposes, such as required decoded picture buffer size, maximum picture output reordering, maximum latency, GDR picture enabled flag, CRA picture enabled flag, reference picture resampling enabled flag, spatial resolution change with CLVS enabled flag, subpicture partitioning enabled flag, maximum number of subpictures in each picture, WPP enabled flag, tile partitioning enabled flag, maximum number of tiles per picture, slice partitioning enabled flag, rectangular slice enabled flag, raster-scan slice enabled flag, maximum number of slices per picture, etc., but can not be signaled in the decoder configuration record.
[0090] Second, in VvcDecoderConfigurationRecord, when signaling PTL information, the numTemporalLayers field is also signaled after the signaling of PTL information. However, the syntax structure of the signaling of PTL information depends on the numTemporalLayers field.
[0091] Third, in the description of random access recovery point sample groups, i.e., ‘rolling’ sample groups, the semantics of the field layer_id_method_idc equal to 1 or 3 is not correctly specified. In particular, when layer_id_method_idc is equal to 1, the signaling of applicable layers can be specified, but when layer_id_method_idc is equal to 3, it can not be specified.
[0092] Disclosed herein are mechanisms to address one or more of the above-listed problems. In one example, the VVC decoder configuration record is modified to locate the number of sub-layers before the PTL record. In this way, the decoder can first obtain the number of sub-layers and use that number to obtain the PTL record for each sub-layer. In another example, the grouping type parameter of the rolling sample group is modified to more clearly describe the correlation between the access points in the rolling sample group and the layers to which these access points apply. For example, the target layer can indicate the layer to which the access point relates. In addition, a layer identifier method identification code can be set to indicate whether the access point applies to all layers or only the layer in the target layer parameter. Further, the layer identifier method identification code can be set to indicate whether the access point consists of only GDR pictures or includes a combination of GDR pictures and mixed NAL unit pictures.
[0093] To address the above and other problems, the following methods are disclosed as summarized below. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow way. In addition, these items can be applied individually or in any combination.
[0094] Embodiment 1
[0095] To address the first problem, one or more of the following parameters can be signaled in the VvcDecoderConfigurationRecord: maximum required size of the decoded picture buffer, maximum picture output reordering (e.g., the maximum allowed number of pictures that can precede any picture in decoding order and follow that picture in output order), maximum latency (e.g., the maximum number of pictures that can precede any picture in output order and follow that picture in decoding order), GDR picture enable flag, CRA picture enable flag, reference picture resampling enable flag, spatial resolution change with CLVS enable flag, subpicture partitioning enable flag, maximum number of subpictures in each picture, WPP enable flag, tile partitioning enable flag, maximum number of tiles per picture, slice partitioning enable flag, rectangular slice enable flag, and raster-scan slice enable flag, and maximum number of slices per picture.
[0096] (a) In one example, one or more of the above parameters are signaled only when PTL information is signaled in the VvcDecoderConfigurationRecord.
[0097] (b) In one example, one or more of the parameters can exist before the signaling of the PTL information. In addition, for all parameters signaled before the PTL information, byte alignment can be required. In one example, a reserved bit can be further signaled.
[0098] (c) In one example, one or more parameters can exist before the signaling of the PTL information. In addition, byte alignment can be required for all parameters signaled before the PTL information. In one example, a reserved bit can be further signaled.
[0099] (d) In one example, a subset of one or more parameters can exist before the signaling of the PTL information, while the remaining can exist after the signaling. In addition, byte alignment can be required for all parameters signaled before the PTL information. In one example, a reserved bit can be further signaled.
[0100] (e) In addition, byte alignment can be required for all parameters signaled after the PTL information. In one example, a reserved bit can be further signaled.
[0101] Embodiment 2
[0102] To address the second issue, the VvcDecoderConfigurationRecord is modified such that when the PTL information is signaled, the numTemporalLayers field is also signaled before the PTL information.
[0103] (a) In one example, when the PTL information is signaled in the VvcDecoderConfigurationRecord, it is signaled after the fields chroma_format_idc, bit_depth_minus8, numTemporalLayers, and constantFrameRate. In one example, the PTL information is directly signaled after all the above fields and some reserved bits.
[0104] (b) In one example, when the PTL information is signaled in the VvcDecoderConfigurationRecord, it is signaled after the fields numTemporalLayers and constantFrameRate. In one example, the PTL information is directly signaled after all the above fields and some reserved bits. In addition, additional reserved bits are further signaled after the PTL information.
[0105] (c) In another example, when the PTL information is signaled in the VvcDecoderConfigurationRecord, it is signaled as the last field among all the fields conditioned with “if (ptl_present_flag)”.
[0106] (d) In one example, a reserved bit is signaled prior to signaling the PTL information.
[0107] Embodiment 3
[0108] To address the third problem 3, one or more of the following modifications are made: The following sentence in clause 9.5.7:
[0109] (a) “The semantics of layer_id_method_idc equal to 1 are specified in clause 9.5.7.” is modified as follows: “When layer_id_method_idc is equal to 1, the layers whose behavior is specified by the ‘rolling’ sample group are specified in clause 9.5.7.” As used herein, clause 9.5.7 refers to the corresponding numbered clause in the document ISO / IEC 14496-15:2021(E) entitled “Information technology - Coding of audio-visual objects - Part 15: Network Abstraction Layer (NAL) unit structured video transport” of ISO / IEC JTC1.
[0110] (b) “The semantics of layer_id_method_idc equal to 3 are specified in clause 9.5.7.” is modified as follows: “When layer_id_method_idc is equal to 3, the layers whose behavior is specified by the ‘rolling’ sample group are specified in the same way as when layer_id_method_idc is equal to 1 as specified in clause 9.5.7.”
[0111] Embodiment 4
[0112] To address problem 3, optionally, one or more of the following changes are made:
[0113] (a) The following sentence in clause 9.5.7: “When layer_id_method_idc is equal to 1, each bit in the field target_layers represents a layer carried in the track.” is changed as follows: “When layer_id_method_idc is equal to 1 or 3, each bit in the field target_layers represents a layer carried in the track.”
[0114] (b) The following sentence: “The semantics of layer_id_method_idc equal to 1 are specified in clause 9.5.7.” is modified as follows: “When layer_id_method_idc is equal to 1, the layers whose behavior is specified by the ‘rolling’ sample group are specified in clause 9.5.7.”
[0115] (c) The following sentence: "Clause 9.5.7 specifies semantics for layer_id_method_idc equal to 3." is modified as follows: "When layer_id_method_idc is equal to 3, the behavior of layers specified by'scrolling' sample groups is specified in Clause 9.5.7."
[0116] The following are some example embodiments of the aspects summarized above, which can be applied to the standard specification of the VVC video file format. The modified text is based on the latest draft specification of the relevant features described above. The relevant parts that are added or modified are shown in bold underlined text, and the parts that are deleted are shown in bold italic text.
[0117] In one example, the syntax of VvcDecoderConfigurationRecord is modified as follows:
[0118]
[0119]
[0120] In one example, the semantics of VvcDecoderConfigurationRecord is modified as follows:
[0121] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to the operation point specified by output_layer_set_idx and numTemporalLayers, and all NAL units in the track belong to that operation point. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular operation point, but can contain VVC bitstreams corresponding to multiple output layer sets, or can contain one or more individual layers that do not form an output layer set or individual sub-layers other than sub-layers with Temporalld equal to 0.
[0122]
[0123] track_ptl specifies the profile, tier, and level of the output layer set represented by the VVC bitstream contained in the track.
[0124] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided to the VVC decoder by an external device or OPINAL unit, as specified in ISO / IEC 23090-3, for decoding the bitstream contained in the track.
[0125]
[0126]
[0127] picture_width specifies the maximum picture width in luma samples that applies to the track. picture_height specifies the maximum picture height in luma samples that applies to the track.
[0128]
[0129] numArrays specifies the number of arrays of NAL units of the indicated type(s).
[0130] In one example, the description of random access recovery point sample group is modified as follows: The random access recovery point sample group "rolling" is used to provide information about recovery points for gradual decoding refresh. When the "rolling" sample group is used with a VVC track, the syntax and semantics of grouping_type_parameter are specified to be the same as the syntax and semantics of the "sap" sample group. When the pictures of the target layer of the samples mapped to the "rolling" sample group are GDR pictures, layer_id_method_idc equal to 0 and 1 are used. When layer_id_method_idc is equal to 0, the "rolling" sample group specifies the behavior for all layers in the track.
[0131]
[0132] Is specified in clause 9.5.7.
[0133] When not all pictures of the target layer of the samples mapped to the "rolling" sample group are GDR pictures, layer_id_method_idc equal to 2 and 3 are used, and for the pictures of the target layer that are not GDR pictures, the following applies: The referenced PPS has pps_mixed_nalu_types_in_pic_flag equal to 1, and for each subpicture index i in the range of 0 to sps_num_subpics_minus1, inclusive, both of the following are true: sps_subpic_treated_as_pic_flag[ i ] is equal to 1, and there is at least one IRAP subpicture in the same CLVS that has the same subpicture index i in the current sample or after. When layer_id_method_idc is equal to 2, the "rolling" sample group specifies the behavior for all layers present in the track.
[0134] When layer_id_method_idc is equal to 3, the layer whose behavior is specified by the 'rolling' sample group. The semantics of layer_id_method_idc equal to 3 are specified in the same way as specified in clause 9.5.7 for layer_id_method_idc equal to 1.
[0135] When a reader starts decoding using samples marked with layer_id_method_idc equal to 2 or 3, the reader needs to further modify the SPS, PPS and PH NAL units of the bitstream reconstructed according to clause 11.6 (of the ISO / IEC 14496-15:2021(E) document) such that the bitstream starts with samples marked as belonging to the sample group and layer_id_method_idc equal to 2 and 3 is a conforming bitstream when sps_gdr_enabled_flag of any SPS such samples refer to is equal to 1, pps_mixed_nalu_types_in_pic_flag of any PPS such samples refer to is equal to 0, nal_unit_type of all VCL NAL units of the AU is equal to GDR_NUT, and ph_gdr_pic_flag of any picture header of the AU is equal to 1, and the value of ph_recovery_poc_cnt corresponds to the roll_distance of the sample group the AU belongs to.
[0136] When a 'rolling' sample group refers to dependent layers instead of its reference layer(s), the sample group indicates features that apply when all the reference layers of the dependent layer are available and decoded. A sample group can be used to initiate the decoding of a prediction layer.
[0137] In one example, the syntax of VvcDecoderConfigurationRecord is modified as follows:
[0138]
[0139]
[0140] In one example, the semantics of VvcDecoderConfigurationRecord are modified as follows:
[0141] ptl_present_flag equal to 1 specifies that the track contains a VVC bitstream corresponding to the operation point specified by output_layer_set_idx and numTemporalLayers, and all NAL units in the track belong to that operation point. ptl_present_flag equal to 0 specifies that the track can not contain a VVC bitstream corresponding to a particular operation point, but can contain VVC bitstreams corresponding to multiple output layer sets, or can contain one or more individual layers that do not form an output layer set or individual sub-layers other than sub-layers with Temporalld equal to 0.
[0142]
[0143] output_layer_set_idx specifies the output layer set index of the output layer set represented by the VVC bitstream contained in the track. The value of output_layer_set_idx can be used as the value of the TargetOlsIdx variable provided to the VVC decoder by an external device or an OPI NAL unit, as specified in ISO / IEC 23090-3, for decoding the bitstream contained in the track.
[0144] avgFrameRate gives the average frame rate in units of frames / (256 seconds) for the stream to which this configuration record applies. The value 0 indicates that the average frame rate is not specified. When the track contains multiple layers and samples are reconstructed for the operation point specified by output_layer_set_idx and numTemporalLayers, this gives the average access unit rate of the bitstream for the operation point.
[0145] constantFrameRate equal to 1 indicates that the stream to which this configuration record applies has a constant frame rate. The value 2 indicates that the representation of each temporal layer in the stream has a constant frame rate. The value 0 indicates that the stream can or can not have a constant frame rate. When the track contains multiple layers and samples are reconstructed for the operation point specified by output_layer_set_idx and numTemporalLayers, this gives an indication of whether the bitstream for the operation point has a constant access unit rate.
[0146] numTemporalLayers greater than 1 indicates that the track to which this configuration record applies is temporally scalable and the number of temporal layers (also referred to as temporal sub-layers or sub-layers in ISO / IEC 23090-3) contained is equal to numTemporalLayers. The value 1 indicates that the track to which this configuration record applies is not temporally scalable. The value 0 indicates that it is unknown whether the track to which this configuration record applies is temporally scalable.
[0147] chroma_format_id indicates the chroma format applied to the track.
[0148]
[0149] picture_width specifies the maximum picture width, in luma samples, applied to the track.
[0150] picture_height specifies the maximum picture height, in luma samples, applied to the track.
[0151]
[0152] numArrays specifies the number of arrays of NAL units of the indicated type(s).
[0153] Figure 1 diagram of an example media file 100 for a VVC bitstream 127 containing video data. The media file includes pictures 125 that can be displayed to create a video sequence. The pictures 125 are compressed in the VVC bitstream 127. The bitstream 127 also includes various parameter sets 123 that indicate to a decoder parameters used to compress the pictures 125. The parameter sets 123 can include a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), and an adaptation parameter set (APS), which include parameters for an entire video, a video sequence, one or more pictures, and a region of one or more pictures, respectively.
[0154] The compression can include intra prediction and inter prediction. In intra prediction, the pictures 125 are partitioned into blocks and each block is coded relative to other blocks in the same picture 125. In inter prediction, the pictures 125 are partitioned into blocks and each block is coded relative to other blocks in other pictures 125. Pictures 125 coded according to inter prediction or intra prediction can be referred to as inter coded pictures or intra coded pictures, respectively. One benefit of inter coded pictures is that such pictures 125 are substantially more compressed than intra coded pictures. However, because inter coded pictures are coded relative to other pictures 125, a video decoder cannot start decoding a video sequence at an inter coded picture. Instead, a video decoder can start decoding a video at any intra coded picture. Intra coded pictures can also be referred to as IRAP pictures. This is because any intra coded picture can serve as an access point 135 for a video stream. An access point 135 is any location in a video stream at which a decoder can start decoding the video stream, typically without encountering decoding errors due to missing information, except for GDR pictures as described below.
[0155] In some cases, picture 125 can be partitioned into sub-pictures. A sub-picture is a rectangular region in picture 125. Sub-pictures are beneficial in that they can be processed separately during decoding and display. For example, in a picture-in-picture application, a virtual reality application, and the like, sub-pictures can be displayed instead of displaying the entire picture 125. Also, in a video call application, for example, sub-pictures can be rearranged and stitched together in different configurations. In some cases, the set of access points 135 can be different for different sub-pictures in the same picture 125. For example, a sub-picture with less important video can have fewer access points 135 to increase compression. When this occurs, picture 125 can include intra-coded sub-pictures and inter-coded sub-pictures, also referred to as IRAP sub-pictures and non-IRAP sub-pictures. Bitstream 127 is a set of network abstraction layer (NAL) units, which are elements of video data sized to fit communication network packets. Thus, parameter sets 123 and picture 125 are carried in bitstream 127 in NAL units. Thus, a picture 125 with IRAP sub-pictures and non-IRAP sub-pictures can be referred to as a mixed NAL unit picture.
[0156] Another access point 135 scheme involves the use of GDR pictures. GDR pictures include an intra coded portion and one or more inter coded portions. GDR pictures are used in groups to create access points 135. Specifically, a first GDR picture contains an intra coded region on the leftmost portion of a picture 125, with the remaining portion of the picture coded according to inter coding. A second GDR picture contains an intra coded region that is shifted to the right to abut but not overlap the intra coded region of the first GDR picture. The remaining portion of the second GDR picture is inter coded. In this way, the intra coded region sweeps from left to right across multiple pictures. One constraint of GDR pictures is that the inter coded region to the left of the intra coded region can only refer back to a previous GDR picture in the current GDR picture group. A decoder can start decoding from the first GDR picture in the group. In this case, the decoder is able to decode the intra coded region, but not the inter coded region. The decoder can then proceed to the second GDR picture, in which case both the intra coded region and the inter coded region to the left of the intra coded region can be decoded. Once the decoder reaches the last GDR picture, all regions can be decoded and the video can be displayed. GDR pictures create errors when used as access points 135, but this error does not persist beyond the last GDR picture in the group. Therefore, GDR pictures are not typically displayed when the group is used as an access point 135. The benefit of GDR pictures is that each GDR picture is smaller than an entire IRAP picture, which reduces the data burst associated with each access point 135. When a decoder does not use a GDR picture as an access point 135, the video before the GDR picture group is available, so the decoder can decode all of the GDR pictures in the group without errors in the inter coded region. It should be noted that GDR pictures are typically prohibited from being used with hybrid NAL unit pictures.
[0157] Pictures 125 and parameter sets 123 can be organized into layers 120 and / or sub-layers. Layers 120 are groupings of pictures 125 and parameter sets 123 that can be decoded and output as part of an output layer set. For example, different layers 120 can be coded at different resolutions. In another example, an output layer set can include a base layer and enhancement layers. This allows a decoder to decode the base layer and obtain a video at a first resolution, and then decode a desired number of enhancement layers to increase the resolution based on device and network capabilities. Sub-layers 121 are a type of layer 120 that allows for temporal scaling. For example, pictures 125 can be assigned to different sub-layers 121 based on a temporal identifier (Id). In this way, each sub-layer 121 contains a subset of pictures 125. This allows a decoder to decode and display a selected sub-layer 121 to achieve a desired frame rate.
[0158] Layers 120 and / or sub-layers 121 of bitstream 127 can be arranged in tracks 110. Tracks 110 contain sequences of a particular type of timed sample that can be decoded and displayed by a decoder. In this context, a sample is a unit of media data. For example, tracks 110 can include a set of timed compressed video samples (e.g., pictures 125 over time), compressed audio samples, hint data samples, parameter samples, etc. It should be noted that the term sample can also refer to a color value of a pixel, but this is not the intended definition in this context. Tracks 110 can contain any number of layers 120 and / or any number of sub-layers 121 containing such samples.
[0159] As can be appreciated from the foregoing description, data in media file 100 can be arranged in a variety of ways. Accordingly, media file 100 also contains a sample table box 130 that contains parameters describing the samples (e.g., media data) contained in tracks 110. For example, a decoder can read sample table box 130 to determine how to begin processing data contained in individual tracks 110. Among many other parameters, sample table box 130 can contain a rolling sample group 131 and a VVC decoder configuration record 141.
[0160] Rolling sample group 131 is also referred to as a random access recovery sample group. Rolling sample group 131 is a unit of data in layer 120 of VVC bitstream 127 used to signal access points 135 to a VVC decoder and is primarily used to signal access points 135 occurring at GDR pictures. It should be noted that random access point (RAP) sample groups can be used to signal access points occurring at other IRAP pictures, such as IDR, CRA, broken link access (BLA), etc. Accordingly, rolling sample group 131 contains a list of pointers to access points 135 of GDR pictures contained in VVC bitstream 127. Access points 135 are considered samples of rolling sample group 131. In some example implementations, the operation of rolling sample group 131 is unclear. The present disclosure addresses these issues by providing parameters that clearly describe the relationship between access points 135 in rolling sample group 131 and layers 120.
[0161] The set of rolling sample points 131 includes a group type parameter 137, which can also be denoted as group type parameter. The group type parameter 137 is a parameter that specifies the dependency / correspondence between the access point 135 and the layers 120. It should be noted that when the access point 135 applies to the layers 120, the layers 120 can be referred to as dependent layers. Thus, the layers 120 include a set of dependent layers, which can be the same as the set of all layers 120 or a subset of the layers 120. The group type parameter 137 also includes a target layers parameter 136 and a layer identifier method identification code 138, which can be denoted as target layers and layer id method idc, respectively. In an example implementation, the target layers parameter 136 includes a plurality of bits, each bit specifying one of the dependent layers. In one example, the target layers parameter 136 can be 24 bits long, thus being able to specify up to 24 dependent layers.
[0162] The layer identifier method identification code 138 specifies the nature of the access point 135 and clarifies the dependency between the access point 135 and the layers. In an example, the layer identifier method identification code 138 can include a four-bit value. In a particular implementation, the layer identifier method identification code 138 can be set to zero or two to indicate that the access point 135 applies to all layers 120. In this case, all layers are dependent layers, and the target layers parameter 136 can be omitted from the media file 100 and / or ignored by the decoder. Further, the layer identifier method identification code 138 can be set to one or four to indicate that the access point 135 applies only to the dependent layers specified by the target layers parameter 136. Further, the layer identifier method identification code 138 can indicate the nature of the pictures 125 present at the access point 135. For example, the layer identifier method identification code 138 can be set to zero or one to indicate that the access point 135 are all GDR pictures. Further, the layer identifier method identification code 138 can be set to two or three to indicate that the access point 135 can be GDR pictures or mixed NAL unit pictures having both IRAP sub-pictures and non-IRAP sub-pictures.
[0163] In a specific implementation, layer id method idc can be set to zero when all access points in the relevant layer are GDR pictures and the access point applies to all layers. Further, layer id method idc is set to 1 when all access points in the relevant layer are GDR pictures and the access point applies to only the relevant layer. Further, layer id method idc is set to 2 when access points in the relevant layer are GDR pictures, mixed NAL unit pictures, or a combination thereof, and the access point applies to all layers. Finally, layer id method idc is set to 3 when access points in the relevant layer are GDR pictures, mixed NAL unit pictures, or a combination thereof, and the access point applies to only the relevant layer. In this way, a decoder can parse access point 135, grouping type parameter 137, target layer 136, and layer identifier method identification code 138 to determine the relevance between access point 135 in the roll sample group 131 and layer 120. The decoder can then use access point 135 to begin decoding pictures 125 in the relevant layer.
[0164] Further, sample table box 130 can include a VVC decoder configuration record 141, which can be represented as VVCDecoderConfigurationRecord. VVC decoder configuration record 141 contains data that a decoder can use to select content. For example, VVC decoder configuration record 141 can contain data describing output layer sets in track 110 and corresponding layers 120. A decoder can then use such data to select tracks 110 that should be decoded and displayed. For example, VVC decoder configuration record 141 can contain data describing a VVC profile tier level (PTL) record 143, output layer set indexes, frame rates, a number of sub-layers 121, bit depths, chroma formats, picture sizes, and the like.
[0165] The VVC PTL record 143 indicates tier, profile, and level information for the layer 120 and / or sub-layer 121. The tier, profile, and level specify restrictions on the bitstream and thus limit the capabilities needed to decode the bitstream. The tier, profile, and level can also be used to indicate interoperability points between various decoder implementations. A tier is a defined set of coding tools used to create compatible or conforming bitstreams. Each tier specifies a subset of algorithmic features and restrictions that all decoders conforming to that tier should support. A level is a set of constraints on the bitstream (e.g., maximum luma sample rate, maximum bitrate for resolution, etc.). For example, a level can refer to a set of constraints (e.g., hardware constraints) that indicate the decoder performance needed to play back a bitstream of a specified tier. These levels are divided into two tiers: a main tier and a high tier. The main tier is lower than the high tier. These tiers are used to handle applications that differ in maximum bitrate. The main tier is designed for most applications, while the high tier is designed for very demanding applications. For any given tier, the level of the tier typically corresponds to a particular decoder processing load and memory capability. Thus, a decoder should select a layer 120 and / or sub-layer 121 for playback by determining a layer 120 and / or sub-layer 121 with PTL information that matches the decoder’s capabilities.
[0166] In some example implementations, the VVC decoder configuration record 141 is unclear because the number of sub-layers 145 is signaled in the VVC decoder configuration record 141 after the VVC PTL record 143. This is a problem because the number of sub-layers 145 is needed by the decoder before the decoder is able to interpret the VVC PTL record 143. In the present disclosure, the number of sub-layers 145 is signaled in the VVC decoder configuration record 141 before the VVC PTL record 143. The decoder can then parse the VVC decoder configuration record to obtain the number of sub-layers 145 and use the number of sub-layers 145 to determine the number of VVC PTL records for the sub-layers 121. In one example, the VVC decoder configuration record 141 includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element. The VVC PTL record 143 can be located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record 141. Further, the number of sub-layers 145 can be located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record 141.
[0167] In a specific implementation, the VVC decoder configuration record 141 can be implemented as follows to position the number of sub-layers 145 before the VVC PTL record 143 for determining the PTL information for the track 110, layer 120, and / or sub-layer 121.
[0168]
[0169] In another example, various additional information can be included in the VVC decoder configuration record 141 to support selection of tracks 110, layers 120, and / or sub-layers 121 in the decoder. Such information can include maximum required size of the decoded picture buffer, maximum picture output reordering, maximum latency, GDR picture enable flag, CRA picture enable flag, reference picture resampling enable flag, spatial resolution change with CLVS enable flag, subpicture partitioning enable flag, maximum number of subpictures in each picture, WPP enable flag, tile partitioning enable flag, maximum number of tiles per picture, slice partitioning enable flag, rectangular slices enable flag, raster-scan slices enable flag, maximum number of slices per picture, or a combination thereof. In some examples, such information can be included only when the VVC decoder configuration record 141 includes a VVC PTL record 143.
[0170] By including such information and / or by rearranging the order of data, the VVC decoder configuration record 141 is improved to allow for additional features and / or more efficient selection of tracks 110, layers 120, and / or sub-layers 121 by the decoder.
[0171] Figure 2 is a flowchart of an example method 200 of encoding a VVC decoder configuration record, such as by encoding the VVC decoder configuration record into a media file 100. At step 201, the codec encodes a plurality of pictures into one or more sub-layers in a media file. At step 203, the codec determines a number of sub-layers. At step 205, the encoder encodes a VVC decoder configuration record into the media file, the VVC decoder configuration including the number of sub-layers and one or more VVC PTL records for the sub-layers. In one example, the number of sub-layers can be signaled in the VVC decoder configuration record prior to the VVC PTL records. In one example, the VVC decoder configuration record further includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element. The VVC PTL records can be located in the VVC decoder configuration record after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element. Further, the number of sub-layers can be located in the VVC decoder configuration record prior to the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element. In an example, the VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and the VVC PTL records are located after the reserved bit. In one example, the VVC decoder configuration record further includes a reserved bit located after the VVC PTL records.
[0172] In one example, the VVC decoder configuration record further includes a maximum required size of a decoded picture buffer, a maximum picture output reordering, a maximum latency, a GDR picture enabled flag, a CRA picture enabled flag, a reference picture resampling enabled flag, a spatial resolution change with CLVS enabled flag, a subpicture partitioning enabled flag, a maximum number of subpictures in each picture, a WPP enabled flag, a tile partitioning enabled flag, a maximum number of tiles per picture, a slice partitioning enabled flag, a rectangular slice enabled flag, a raster-scan slice enabled flag, a maximum number of slices per picture, or a combination thereof. In one example, when the VVC decoder configuration record includes the VVC PTL record, the VVC decoder record only discloses one or more or the aforementioned parameters.
[0173] In one example, in the VVC decoder configuration record, all syntax elements following the VVC PTL record are required to be byte-aligned.
[0174] At step 207, the encoder stores the media file in a memory, for example, for later transmission to a decoder. In one embodiment, the media file is transmitted to a decoder.
[0175] Figure 3 is a flowchart of an example method 300 of decoding a VVC decoder configuration record, for example, by receiving a media file 100 as a result of method 200. At step 301, the decoder receives a media file including a VVC decoder configuration record and a plurality of pictures coded into one or more sub-layers. At step 303, the decoder parses the VVC decoder configuration record to obtain a number of sub-layers and one or more VVC PTL records of the sub-layers based on the number of sub-layers. In an example, the number of sub-layers can be signaled in the VVC decoder configuration record before the VVC PTL record. In an example, the VVC decoder configuration record further includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element. The VVC PTL record can be located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record. Further, the number of sub-layers can be located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record. In an example, the VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and the VVC PTL record is located after the reserved bit. In an example, the VVC decoder configuration record further includes a reserved bit located after the VVC PTL record.
[0176] In an example, the VVC decoder configuration record includes a maximum required size of a decoded picture buffer, a maximum picture output reordering, a maximum latency, a GDR picture enable flag, a CRA picture enable flag, a reference picture resampling enable flag, a spatial resolution change with CLVS enable flag, a subpicture partitioning enable flag, a maximum number of subpictures in each picture, a WPP enable flag, a tile partitioning enable flag, a maximum number of tiles per picture, a slice partitioning enable flag, a rectangular slice enable flag, a raster-scan slice enable flag, a maximum number of slices per picture, or a combination thereof. In one example, when the VVC decoder configuration record includes a VVC PTL record, the VVC decoder record discloses only one or more or the aforementioned parameters.
[0177] In an example, in the VVC decoder configuration record, all syntax elements following the VVC PTL record are required to be byte-aligned.
[0178] At step 305, the decoder decodes one or more sub-layers based on the VVC PTL record. The decoder can then forward the decoded media file or portions thereof (e.g., particular layers and / or sub-layers) to a display for viewing by a user.
[0179] Figure 4 is a block diagram of an example video processing system 400 that can implement the various techniques disclosed herein. Various implementations can include some or all of the components in the system 400. The system 400 can include an input 402 to receive video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values) or can be received in a compressed or encoded format. The input 402 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0180] The system 400 can include a codec component 404 that can implement various encoding or coding methods described in this document. The codec component 404 can reduce the average bitrate of a video from the input 402 to the output of the codec component 404 to produce a coded representation of the video. Thus, the codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 404 can be stored or transmitted via a connected communication, as represented by component 406. The bitstream (or coded) representation of the video received at the input 402 can be used by component 408 to generate pixel values or displayable video that is sent to a display interface 410. The process of generating user- visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0181] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Displayport, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other digital data processing and / or video display enabled devices.
[0182] Figure 5 is a block diagram of an example video processing apparatus 500. The apparatus 500 can be used to implement one or more of the methods described herein. The apparatus 500 can be implemented in a smart phone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 500 can include one or more processors 502, one or more memories 504, and video processing hardware 506. The processor(s) 502 can be configured to implement one or more methods described in this document. The memory(ies) 504 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 506 can be used to implement, in hardware circuitry, some of the techniques described in this document. In some embodiments, the video processing hardware 506 can be at least partially included in the processor 502, for example, a graphics co-processor.
[0183] Figure 6is a flowchart of an example method 600 of video processing. The method 600 performs a conversion between visual media data and a file storing information corresponding to the visual media data according to a video file format. In the context of an encoder, this conversion can be performed by encoding visual media data into a visual media data file of the video file format. In the context of a decoder, this conversion can be performed by decoding a visual media data file of the video file format to obtain visual media data for display.
[0184] Figure 7 is a block diagram illustrating an example video coding system 700 that can utilize the techniques of this disclosure. As shown, video coding system 700 can include a source device 710 and a destination device 720. Source device 710 generates encoded video data, which can be referred to as a video encoding device. Destination device 720 can decode the encoded video data generated by source device 710, and the destination device 720 can be referred to as a video decoding device. Figure 7
[0185] Source device 710 can include a video source 712, a video encoder 714, and an input / output (I / O) interface 716.
[0186] Video source 712 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 714 encodes video data from video source 712 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax elements. I / O interface 716 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 720 from I / O interface 716 via network 730. The encoded video data can also be stored onto a storage medium / server 740 for access by destination device 720.
[0187] Destination device 720 can include an I / O interface 726, a video decoder 724, and a display device 722. I / O interface 726 can include a receiver and / or a modem. I / O interface 726 can acquire encoded video data from source device 710 or storage medium / server 740. Video decoder 724 can decode the encoded video data. Display device 722 can display the decoded video data to a user. Display device 722 can be integrated with destination device 720, or can be external to destination device 720, configured to interface with destination device 720.
[0188] The video encoder 714 and the video decoder 724 can operate according to video compression standards such as High Efficiency Video Codec (HEVC), Multi-Functional Video Codec (VVC), and other current and / or other standards.
[0189] Figure 8 This is a block diagram illustrating an example of a video encoder 800, which can be... Figure 7 The system 700 shown contains a video encoder 714. The video encoder 800 can be configured to perform any or all of the techniques disclosed herein. Figure 8 In the example, the video encoder 800 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 800. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0190] The functional components of the video encoder 800 may include a segmentation unit 801, a prediction unit 802 (which may include a mode selection unit 803, a motion estimation unit 804, a motion compensation unit 805, and an intra-frame prediction unit 806), a residual generation unit 807, a transform processing unit 808, a quantization unit 809, an inverse quantization unit 810, an inverse transform unit 811, a reconstruction unit 812, a buffer 813, and an entropy coding unit 814.
[0191] In other examples, the video encoder 800 may include more, fewer, or different functional components. In one example, the prediction unit 802 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0192] Furthermore, some components, such as the motion estimation unit 804 and the motion compensation unit 805, can be highly integrated, but for interpretive purposes... Figure 8 The examples are shown separately.
[0193] The segmentation unit 801 can segment an image into one or more video blocks. The video encoder 800 and video decoder 900 can support various video block sizes.
[0194] The mode selection unit 803 can select one of intra or inter coding modes, e.g., based on the error results, and provide the resulting intra or inter coded block to the residual generation unit 807 to generate residual block data and to the reconstruction unit 812 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 803 can select a combined intra and inter prediction (CIIP) mode in which the prediction is based on both an inter prediction signal and an intra prediction signal. The mode selection unit 803 can also select a resolution of motion vectors (e.g., sub-pixel or integer-pixel precision) for a block in the case of inter prediction.
[0195] To perform inter prediction for a current video block, the motion estimation unit 804 can generate motion information for the current video block by comparing one or more reference frames from the buffer 813 to the current video block. The motion compensation unit 805 can determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 813 (other than the picture associated with the current video block).
[0196] The motion estimation unit 804 and the motion compensation unit 805 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0197] In some examples, the motion estimation unit 804 can perform single prediction for a current video block, and the motion estimation unit 804 can search for a reference video block for the current video block in a reference picture of list 0 or list 1. The motion estimation unit 804 can then generate a reference index indicating the reference picture of list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 804 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. The motion compensation unit 805 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0198] In other examples, the motion estimation unit 804 can perform bi-prediction for a current video block, and the motion estimation unit 804 can search for a reference video block for the current video block in a reference picture of list 0 and also search for another reference video block for the current video block in a reference picture of list 1. The motion estimation unit 804 can then generate a reference index indicating the reference picture of list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. The motion estimation unit 804 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 805 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0199] In some examples, the motion estimation unit 804 can output the entire set of motion information for use in decoding processing by the decoder. In some examples, the motion estimation unit 804 can not output the entire set of motion information for the current video block. Instead, the motion estimation unit 804 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 804 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0200] In one example, the motion estimation unit 804 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 900 that the current video block has the same motion information as another video block.
[0201] In another example, the motion estimation unit 804 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector that indicates the video block. The video decoder 900 can use the motion vector that indicates the video block and the motion vector difference to determine the motion vector of the current video block.
[0202] As discussed above, the video encoder 800 can predictively signal motion vectors. Two examples of predictively signaling techniques that can be implemented by the video encoder 800 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0203] The intra prediction unit 806 can perform intra prediction on the current video block. When the intra prediction unit 806 performs intra prediction on the current video block, the intra prediction unit 806 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0204] The residual generation unit 807 can generate residual data for the current video block by subtracting (e.g., represented by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks that correspond to different sample components of samples in the current video block.
[0205] In other examples, such as in skip mode, there can be no residual data for the current video block, and the residual generation unit 807 can not perform the subtraction operation.
[0206] The transform processing unit 808 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0207] After transform processing unit 808 generates a transform coefficient video block associated with the current video block, quantization unit 809 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0208] Inverse quantization unit 810 and inverse transform unit 811 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 812 can add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 802 to produce a reconstructed video block associated with the current block for storage in buffer 813.
[0209] After reconstruction unit 812 reconstructs the video block, loop filtering operations can be performed to reduce video blockiness artifacts in the video block.
[0210] Entropy encoding unit 814 can receive data from other functional components of video encoder 800. When entropy encoding unit 814 receives data, entropy encoding unit 814 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0211] Figure 9 is a block diagram illustrating an example of a video decoder 900 that can be Figure 7 the video decoder 724 in the system 700 illustrated in FIG. 13.
[0212] Video decoder 900 can be configured to perform any or all of the techniques of this disclosure. In Figure 9 examples, video decoder 900 includes a plurality of functional components. The techniques described in this disclosure can be shared amongst the various components of video decoder 900. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0213] In Figure 9 examples, video decoder 900 includes an entropy decoding unit 901, a motion compensation unit 902, an intra-prediction unit 909, an inverse quantization unit 904, an inverse transform unit 905, and a reconstruction unit 906 and a buffer 907. In some examples, video decoder 900 can perform a decoding process generally inverse to the encoding process described with respect to video encoder 800 Figure 8 ).
[0214] The entropy decoding unit 901 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., coded blocks of video data). The entropy decoding unit 901 can decode the entropy encoded video and, from the entropy decoded video data, the motion compensation unit 902 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 902 can determine such information, for example, by performing AMVP and merge mode.
[0215] The motion compensation unit 902 can generate a motion compensated block, possibly interpolated based on an interpolation filter. An identifier of the interpolation filter to be used at sub-pixel precision can be included in a syntax element.
[0216] The motion compensation unit 902 can use the interpolation filter used by the video encoder 800 during encoding of the video block to calculate values of the interpolation of sub-integer pixel values of the reference block. The motion compensation unit 902 can determine the interpolation filter used by the video encoder 800 from the received syntax information and use the interpolation filter to generate the prediction block.
[0217] The motion compensation unit 902 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information to decode the encoded video sequence.
[0218] The intra prediction unit 903 can form a prediction block from spatial neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 904 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 901. The inverse transform unit 905 applies an inverse transform.
[0219] The reconstruction unit 906 can sum a residual block with a corresponding prediction block generated by the motion compensation unit 902 or the intra prediction unit 903 to form a decoded block. As desired, a deblocking filter can also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in the buffer 907, which provides reference blocks for subsequent motion compensation / intra prediction, and also produces decoded video for presentation on a display device.
[0220] Figure 10is a schematic diagram of an example encoder 1000. The encoder 1000 is suitable for implementing VVC techniques. The encoder 1000 comprises three in-loop filters, namely a deblocking filter (DF) 1002, a sample adaptive offset (SAO) 1004 and an adaptive loop filter (ALF) 1006. Unlike the DF 1002 which uses a pre-defined filter, the SAO 1004 and the ALF 1006 exploit the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter respectively, with the offset and filter coefficients being signaled with the aid of coded side information. The ALF 1006 is located at the last processing stage of each picture and can be regarded as a tool that attempts to capture and fix artifacts produced by previous stages.
[0221] The encoder 1000 further comprises an intra prediction component 1008 and a motion estimation / compensation (ME / MC) component 1010 configured to receive an input video. The intra prediction component 1008 is configured to perform intra prediction, while the ME / MC component 1010 is configured to perform inter prediction with the aid of reference pictures obtained from a reference picture buffer 1012. Residual blocks from either inter or intra prediction are fed into a transform component 1014 and a quantization component 1016 to generate quantized residual transform coefficients, which are fed into an entropy coding component 1018. The entropy coding component 1018 entropy encodes the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 1016 can be fed into an inverse quantization component 1020, an inverse transform component 1022 and a reconstruction (REC) component 1024. The REC component 1024 is capable of outputting pictures to the DF 1002, the SAO 1004 and the ALF 1006 for filtering before these pictures are stored in the reference picture buffer 1012.
[0222] Next, a list of some embodiments preferred solutions is provided.
[0223] The following solutions show examples of the techniques discussed herein.
[0224] 1. A visual media processing method (e.g., Figure 6The method 600 shown in the middle, comprising: performing (602) a conversion between visual media data and a file storing information corresponding to the visual media data according to a video file format; wherein the video file format comprises a decoder configuration record configured with information for content selection, wherein the decoder configuration record comprises one or more fields: required decoding picture buffer size, maximum picture output reordering, maximum latency, gradual decoding refresh picture enable flag, clean random access picture enable flag, reference picture resampling enable flag, spatial resolution change with coded video layer sequence enable flag, subpicture partitioning enable flag, maximum number of subpictures in each picture, wavefront parallel processing enable flag, tile partitioning enable flag, maximum number of tiles per picture, slice partitioning enable flag, rectangular slice enable flag, raster-scan slice enable flag, maximum number of slices per picture.
[0225] 2. A visual media processing method, comprising: performing a conversion between visual media data and a file storing information corresponding to the visual media data according to a video file format according to a rule; wherein the rule specifies that a field indicating a number of temporal layers is included in a decoder configuration record according to whether a level tier level information of the visual media data is included in the file; wherein the rule further specifies that the field is included before the level tier level information.
[0226] 3. The method according to solution 2, wherein the rule further specifies an order in which the level tier level information appears in the video file format relative to one or more additional information fields.
[0227] 4. The method according to solution 3, wherein the one or more additional information fields comprise a chroma format indication field, a bit depth field, a field indicating a number of temporal layers, or a field indicating whether constant frame rate is used for the visual media data.
[0228] 5. The method according to solution 3, wherein the one or more additional information fields comprise a reserved bits field.
[0229] 6. The method according to any of solutions 2-5, wherein the rule specifies that the level tier level information is included as a last field of the decoder configuration record.
[0230] 7. The method according to any of solutions 1-6, wherein the conversion comprises generating a bitstream representation of the visual media data, and storing the bitstream representation to the file according to the format rule.
[0231] 8. The method according to any of solutions 1-6, wherein the conversion comprises parsing the file to recover the visual media data according to the format rule.
[0232] 9. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 8.
[0233] 10. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 8.
[0234] 11. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any of solutions 1 to 8.
[0235] 12. A computer readable medium having a bitstream representation conforming to a file format generated according to any of solutions 1 to 8.
[0236] 13. The methods, apparatuses or systems described in this document. In the solutions described herein, an encoder can conform to a format rule by producing a coded representation according to the format rule. In the solutions described herein, a decoder can parse syntax elements in a coded representation using a format rule and understand the presence and absence of syntax elements according to the format rule to produce a decoded video.
[0237] In this document, the term "video processing" can refer to video encoding, video decoding, video compression or video decompression. For example, during a conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm can be applied, and vice versa. As defined by the syntax, a bitstream representation of a current video block can correspond to, for example, bits that are co-located or scattered at different locations within the bitstream. For example, a macroblock can be encoded according to transformed and coded error residual values and also using bits in a header and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can be present or absent, as described in the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields and generate a coded representation accordingly by including or excluding the syntax fields from the coded representation.
[0238] The disclosed and other aspects, examples, implementations, modules and functional operations can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0239] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0240] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0241] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0242] Although the patent document contains many details, these should not be construed as limiting the subject matter or the scope of any patent covering implementation, but as a description of features particular to certain embodiments of the technology. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0243] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to implement such operations, and each operation can be performed in a different order or concurrently with other operations disclosed herein. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0244] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
[0245] A first component is directly coupled to a second component when there are no intervening components between the first component and the second component other than a wire, trace, or another medium. A first component is indirectly coupled to a second component when there are intervening components between the first component and the second component other than a wire, trace, or another medium. The term “coupled” and variations thereof include both direct and indirect couplings. Unless otherwise stated, the use of the term “about” means a range of plus or minus 10% of the subsequent value.
[0246] While several embodiments have been provided in the present disclosure, it is understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are therefore to be considered as illustrative and not restrictive, and the intention is not to limit the details given herein to the precise forms disclosed. For example, the various elements or components can be combined or integrated in another system or certain features can be omitted, or not implemented.
[0247] In addition, the various embodiments described and illustrated herein can be combined with other systems, modules, techniques, or methods to produce further embodiments. Other items shown or discussed as separate from other items or collectively can be combined. The examples described herein are intended to be illustrative and not exclusive. Numerous modifications, substitutions, and changes can be made to the systems and methods described and illustrated herein without departing from the spirit and scope of the disclosure.
Claims
1. A method of processing video data, comprising: performing a conversion between visual media data and a visual media data file, the visual media data file comprising a Versatile Video Coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein the VVC decoder configuration record comprises a number of the one or more sub-layers and one or more VVC Profile Tier Level (PTL) records of the one or more sub-layers based on the number of the one or more sub-layers, wherein the VVC decoder configuration record comprises a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL records are located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
2. The method of claim 1, wherein, the conversion comprises: receiving a media file comprising the VVC decoder configuration record and the plurality of pictures coded into one or more sub-layers; parsing the VVC decoder configuration record to obtain the number of the one or more sub-layers and the one or more VVC PTL records of the one or more sub-layers based on the number of the one or more sub-layers; and decoding the one or more sub-layers based on the VVC PTL records.
3. The method of claim 1, wherein, the conversion comprises: encoding the plurality of pictures into the one or more sub-layers in the visual media data file; determining a number of the one or more sub-layers; encoding the VVC decoder configuration record into a media file, the VVC decoder configuration record comprising the number of the one or more sub-layers and one or more VVC PTL records of the one or more sub-layers; and storing the visual media data file in a memory.
4. The method of claim 1, wherein, the number of the one or more sub-layers is signaled in the VVC decoder configuration record before the VVC PTL records.
5. The method of claim 1, wherein, the number of the one or more sub-layers is located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
6. The method of claim 1, wherein, the VVC decoder configuration record further comprises a reserved bit located after the bit depth minus eight syntax element, and wherein the VVC PTL records are located after the reserved bit.
7. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: performing a conversion between visual media data and a visual media data file, the visual media data file including a versatile video coding, VVC, decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein, the VVC decoder configuration record comprises a number of the one or more sub-layers and one or more VVC Profile Tier Level (PTL) records of the one or more sub-layers based on the number of the one or more sub-layers, wherein the VVC decoder configuration record includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL record is located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
8. The apparatus of claim 7, wherein, The conversion includes: receiving a media file including the VVC decoder configuration record and the plurality of pictures coded into one or more sub-layers; parsing the VVC decoder configuration record to obtain a number of the one or more sub-layers and one or more VVC PTL records of the one or more sub-layers based on the number of the one or more sub-layers; and decoding the one or more sub-layers based on the VVC PTL records.
9. The apparatus of claim 7, wherein, The conversion includes: encoding the plurality of pictures into the one or more sub-layers in the visual media data file; determining a number of the one or more sub-layers; encoding the VVC decoder configuration record into a media file, the VVC decoder configuration record including the number of the one or more sub-layers and one or more VVC PTL records of the one or more sub-layers; and storing the visual media data file in a memory.
10. The apparatus of claim 7, wherein, The number of the one or more sub-layers is signaled in the VVC decoder configuration record before the VVC PTL record.
11. The apparatus of claim 7, wherein, The number of the one or more sub-layers is located before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
12. The apparatus of claim 7, wherein, The VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and wherein the VVC PTL record is located after the reserved bit.
13. A non-transitory computer-readable medium containing a computer program product for use by a video coding device, the computer program product containing instructions executable by a processor such that, when executed by the processor, the video coding device is caused to: performing a conversion between visual media data and a visual media data file, the visual media data file including a versatile video coding, VVC, decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein, The VVC decoder configuration record includes the number of the one or more sub-layers and one or more VVC profile tier level PTL records of the one or more sub-layers based on the number of the one or more sub-layers, wherein the VVC decoder configuration record includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL record is located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.
14. The non-transitory computer-readable medium of claim 13, wherein, The conversion includes: receiving a media file including the VVC decoder configuration record and the plurality of pictures coded into one or more sub-layers; parsing the VVC decoder configuration record to obtain a number of the one or more sub-layers and one or more VVC profile tier level (PTL) records of the one or more sub-layers based on the number of the one or more sub-layers; and decoding the one or more sub-layers based on the VVC PTL records.
15. The non-transitory computer-readable medium of claim 13, wherein, The converting includes: encoding the plurality of pictures into the one or more sub-layers in the visual media data file; determining a number of the one or more sub-layers; encoding the VVC decoder configuration record into a media file, the VVC decoder configuration record including the number of the one or more sub-layers and one or more VVC PTL records of the one or more sub-layers; and storing the visual media data file in a memory.
16. The non-transitory computer-readable medium of claim 13, wherein, The number of the one or more sub-layers is signaled in the VVC decoder configuration record before the VVC PTL records.
17. The non-transitory computer-readable medium of claim 13, wherein, The number of the one or more sub-layers is located in the VVC decoder configuration record before the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element, and wherein the VVC decoder configuration record further includes a reserved bit located after the bit depth minus eight syntax element, and wherein the VVC PTL records are located after the reserved bit.
18. A method of storing a bitstream of a video, comprising: generating a visual media data file, storing the visual media data file into a non-transitory computer readable storage medium, wherein the visual media data file includes a Versatile Video Coding (VVC) decoder configuration record and a plurality of pictures coded into one or more sub-layers, wherein the VVC decoder configuration record includes a number of the one or more sub-layers and one or more VVC profile tier level (PTL) records of the one or more sub-layers based on the number of the one or more sub-layers, wherein the VVC decoder configuration record includes a constant frame rate syntax element, a chroma format identifier syntax element, and a bit depth minus eight syntax element, and wherein the VVC PTL records are located after the constant frame rate syntax element, the chroma format identifier syntax element, and the bit depth minus eight syntax element in the VVC decoder configuration record.