Horizontal information in video coding
By optimizing video encoding and decoding methods and utilizing the format rules of intra-frame random access points and inter-layer prediction information, the problems of low bandwidth requirements and low decoding efficiency in multi-layer video coding are solved, achieving more efficient video processing and wider support for application scenarios.
Patent Information
- Application Number
- CN202180025125.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-29
- Filing Date
- 2021-03-23
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-03-23
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively utilize intra-frame random access points and inter-layer prediction information when processing multi-layer video encoding, leading to increased bandwidth requirements and low decoding efficiency.
A video processing method is adopted to transform the relationship between video and codec representation through format rules, including intra-frame random access point images or intra-frame codec images, and to limit the maximum temporal layer ID of signaling notifications. A hypothetical reference decoder and special effects mode access representation are used to optimize the bitstream structure to support multi-layer video coding.
It improves the efficiency and flexibility of video encoding and decoding, reduces bandwidth requirements, adapts to the interoperability requirements of different decoders, and supports a wider range of application scenarios such as television broadcasting, video conferencing, and immersive adaptive 360° media.
Smart Images

Figure CN115428438B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a Chinese national phase application of International Patent Application No. PCT / US2021 / 023595, filed on March 23, 2021, and timely claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 000,941, filed on March 27, 2020, and U.S. Provisional Patent Application No. 63 / 085,107, filed on September 29, 2020. The entire disclosures of the aforementioned applications are incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Art
[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing a codec representation of a video using control information useful for decoding the codec representation.
[0006] In one exemplary aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more output layer sets (OLSs) and a codec representation of the video, wherein the one or more video pictures are coded or decoded as intra random access point pictures or intra codec pictures in the codec representation, wherein the codec representation conforms to a format rule that specifies the location and type of information included in the codec representation for decoding the one or more video pictures.
[0007] In another exemplary aspect, another video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a codec representation, the one or more pictures comprising one or more slices coded into a codec representation of one or more temporal video layers; wherein the codec representation conforms to a format rule that specifies constraints on signaling of one or more inter-layer prediction information syntax elements, does not require signaling a two-dimensional syntax element indicating a maximum temporal layer id of reference pictures used for coding a current video layer, and does not require directly signaling a layer id of the maximum temporal layer id of reference pictures used for coding the current video layer.
[0008] In another exemplary aspect, another video processing method is disclosed, the method comprising performing a conversion between a video comprising one or more pictures and a codec representation of the video according to a rule, wherein each picture is coded as an intra random access picture, and wherein the rule specifies that operation of a hypothetical reference decoder for the codec representation uses a maximum sub-layer value equal to zero.
[0009] In another exemplary aspect, another video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video including one or more output layer sets according to a format rule, wherein at least one of the one or more output layer sets consists of a trick mode access representation including only intra random access point pictures or only intra codec pictures, and wherein the format rule specifies whether or how level information of the trick mode representation is indicated in the bitstream.
[0010] In another exemplary aspect, another video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream includes a trick mode access representation of one or more output layer sets according to a format rule, wherein the format rule specifies that the trick mode access representation includes only intra random access point pictures, and wherein the format rule specifies whether or how a hypothetical reference decoder is to be operated.
[0011] In yet another exemplary aspect, a video encoder apparatus is disclosed, wherein the video encoder includes a processor configured to implement the above method.
[0012] In yet another exemplary aspect, a video decoder apparatus is disclosed, wherein the video decoder includes a processor configured to implement the above method.
[0013] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.
[0014] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram of an exemplary video processing system.
[0016] Figure 2 It is a block diagram of a video processing device.
[0017] Figure 3 is a flow chart of an exemplary method of video processing.
[0018] Figure 4 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0019] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0020] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0021] Figure 7A and 7B is a flow chart of an exemplary method of video processing based on some implementations of the disclosed technology. DETAILED DESCRIPTION
[0022] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to that section alone. Furthermore, the use of H.266 terminology in some descriptions is for ease of understanding only and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0023] 1. Initial Discussion
[0024] This document relates to video coding techniques. Specifically, the video coding techniques are about the signaling of the horizontal information of the intra-only random access point (IRAP) representation and / or intra-only representation of the bitstream, the signaling of sub-layers not used for inter-layer prediction, and the signaling of virtual boundaries. The intra-only random access point representation of the bitstream is a sub-bitstream consisting only of IRAP pictures and associated non-VCL NAL units in the bitstream. The intra-only representation of the bitstream is a sub-bitstream consisting only of intra-coded pictures and associated non-VCL NAL units in the bitstream. These ideas can be applied alone or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding (for example, the Versatile Video Codec (VVC) under development).
[0025] 2. Abbreviation
[0026] ACT Adaptive Color Transformation
[0027] ALF Adaptive Loop Filter
[0028] AMVR Adaptive Motion Vector Resolution
[0029] APS Adaptation Parameter Set
[0030] AU Access Unit
[0031] AUD Access Unit Delimiter
[0032] AVC Advanced Video Codec (Rec.ITU-T H.264 | ISO / IEC 14496-10)
[0033] B. Bidirectional Prediction
[0034] BCW Bidirectional prediction with CU-level weights
[0035] BDOF Bidirectional Optical Flow
[0036] BDPCM Block-based delta pulse coding modulation
[0037] BP buffer period
[0038] CABAC Context-based Adaptive Binary Arithmetic Coding
[0039] CB codec block
[0040] CBR Constant Bit Rate
[0041] CCALF Cross-Component Adaptive Loop Filter
[0042] CLVS codec layer video sequence
[0043] CLVSS codec layer video sequence starts
[0044] CPB codec picture buffer
[0045] CRA Clear Random Access
[0046] CRC Cyclic Redundancy Check
[0047] CRR Cross-RAP Reference
[0048] CTB Codec Tree Block
[0049] CTU Codec Tree Unit
[0050] CU codec unit
[0051] CVS encoded and decoded video sequence
[0052] CVSS encoded video sequence starts
[0053] DPB decoded picture buffer
[0054] DCI decoding capability information
[0055] DRAP Related Random Access Point
[0056] DU decoding unit
[0057] DUI decoding unit information
[0058] EG Index Columbus
[0059] EGk kth order index Columbus
[0060] EOB End of bitstream
[0061] EOS sequence end
[0062] FD fill data
[0063] FIFO First In First Out
[0064] FL fixed length
[0065] GBR Green, Blue and Red
[0066] GCI General Constraint Information
[0067] GDR Progressive Decode Refresh
[0068] GPM geometric partitioning mode
[0069] HEVC High-Efficiency Video Codec (Rec.ITU-T H.265 | ISO / IEC 23008-2)
[0070] HRD Hypothetical Reference Decoder
[0071] HSS Hypothetical Stream Scheduler
[0072] Intra-I frame
[0073] IBC Intra Block Copy
[0074] IDR Instant Decode Refresh
[0075] ILRP inter-layer reference image
[0076] IRAP Intra-frame Random Access Point
[0077] LFNST Low-frequency non-separable transform
[0078] LPS Least Probable Symbol
[0079] LSB Least Significant Bit
[0080] LTRP Long Term Reference Picture
[0081] LMCS chroma scaling and luma mapping
[0082] MIP matrix-based intra prediction
[0083] MPS Most Probable Symbol
[0084] MSB Most Significant Bit
[0085] MTS Multi-Transformation Selection
[0086] MVP Motion Vector Prediction
[0087] NAL Network Abstraction Layer
[0088] OLS output layer set
[0089] OP operating point
[0090] OPI Operating Point Information
[0091] P prediction
[0092] PH Image Header
[0093] POC picture sequence counting
[0094] PPS Picture Parameter Set
[0095] PROF Prediction refinement using optical flow
[0096] PT Picture Timing
[0097] PU picture unit
[0098] QP quantization parameter
[0099] RADL Random Access Decodable Preamble (Image)
[0100] RAP Random Access Point
[0101] RASL Random Access Skip Preamble (Image)
[0102] RBSP Raw Byte Sequence Payload
[0103] RGB red, green, and blue
[0104] RPL Reference Image List
[0105] SAO Sample Adaptive Offset
[0106] SAR sample aspect ratio
[0107] SEI Supplemental Enhancement Information
[0108] SH Strip Header
[0109] SLI sub-picture level information
[0110] SODB data bit string
[0111] SPS sequence parameter set
[0112] STRP Short-Term Reference Picture
[0113] STSA Stepwise Time Sublayer Access
[0114] TR Truncated RICE
[0115] TU Transform Unit
[0116] VBR variable bit rate
[0117] VCL video codec layer
[0118] VPS Video Parameter Set
[0119] VSEI Generic Supplementary Enhancement Information (Rec.ITU-T H.274 | ISO / IEC 23002-7)
[0120] VUI Video Availability Information
[0121] VVC (Versatile Video Codec) (Rec. ITU-T H.266 | ISO / IEC 23090-3)
[0122] 3. Video Codec Introduction
[0123] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Video, and the two organizations jointly produced the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by the Video Coding Experts Group (VCEG) and MPEG. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially began, JVET was renamed the Joint Video Exploration Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC, which was finalized by JVET at its 19th meeting that ended on July 1, 2020.
[0124] The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and immersive adaptive 360° media.
[0125] 3.1. Grade, level and level
[0126] Video codec standards usually specify profiles and levels. Some video codec standards also specify levels, such as HEVC and the currently under-development VVC.
[0127] Profiles, levels, and grades specify constraints on the bitstream, and therefore on the capabilities required to decode the bitstream. Profiles, levels, and grades may also be used to indicate points of interoperability between various decoder implementations.
[0128] Each profile specifies a subset of algorithmic features and restrictions that should be supported by all decoders conforming to that profile. Note that encoders are not required to utilize all codecs or features supported in a profile, but decoders conforming to a profile are required to support all codecs or features.
[0129] Each level of a profile specifies a set of constraints on the values that can be taken by bitstream syntax elements. The same set of level and level definitions is typically used with all profiles, but individual implementations may support different levels, and within a level, different levels for each supported profile. For any given profile, the level of the profile typically corresponds to a specific decoder processing load and memory capabilities.
[0130] Specifies the capabilities of a video decoder conforming to a video codec specification in terms of its ability to decode video streams conforming to the constraints of the profile, level, and grade specified in the video codec specification. When expressing the capabilities of a decoder for a specified profile, the grade and level supported by that profile shall also be expressed.
[0131] 3.2. Random Access and Support in HEVC and VVC
[0132] Random access refers to accessing and decoding the bitstream starting from a picture that is not the first picture in the bitstream in decoding order. In order to support tuning and channel switching in broadcast / multicast and multi-party video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are usually intra-frame codec pictures, but can also be inter-frame codec pictures (for example, in the case of gradual decoding refresh).
[0133] HEVC includes signaling of intra random access point (IRAP) pictures via NAL unit types in the NAL unit header. Three types of IRAP pictures are supported, namely instantaneous decoder refresh (IDR), cleanup random access (CRA) and broken link access (BLA) pictures. IDR pictures restrict the inter prediction structure so as not to reference any pictures before the current group of pictures (GOP) and are commonly referred to as closed GOP random access points. CRA pictures are less restricted by allowing certain pictures to reference pictures before the current GOP, all of which are discarded in the case of random access. CRA pictures are commonly referred to as open GOP random access points. BLA pictures typically result from the concatenation of two bitstreams or parts thereof at a CRA picture, for example during stream switching. To enable better system usage of IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures that can be used to better match stream access point types as defined in the ISO based media file format (ISOBMFF) for random access support in Dynamic Adaptive Streaming over HTTP (DASH).
[0134] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with an associated RADL picture or the other type without an associated RADL picture), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC, primarily for two reasons: i) the basic functionality of BLA pictures can be implemented with CRA pictures plus the end of a sequence NAL unit, the presence of which indicates that the subsequent picture starts a new CVS in a single-layer bitstream; ii) during the development of VVC, it was desirable to specify fewer NAL unit types than in HEVC, as indicated by using five bits instead of six for the NAL unit type field in the NAL unit header.
[0135] Another key difference in random access support between VVC and HEVC is that GDR is supported in VVC in a more standardized way. In GDR, decoding of the bitstream can start from an inter-frame coded picture, and although not the entire picture region can be decoded correctly at the beginning, after a number of pictures, the entire picture region will be correct. AVC and HEVC also support GDR, using the recovery point SEI message for signaling of GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and bitstreams are allowed to start with GDR pictures. This means that the entire bitrate is allowed to contain only inter-frame coded pictures and not a single intra-frame coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded slices or blocks across multiple pictures instead of intra-coding the entire picture, allowing for significant end-to-end latency reduction, which is considered more important today than ever before as ultra-low latency applications such as those based on wireless displays, online gaming, and drones become more popular.
[0136] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., correctly decoded area) and the non-refreshed area at the picture between the GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, in-loop filtering across the boundary will not be applied, so decoding mismatches of some samples at or near the boundary will not occur. This may be useful when the application determines to display correctly decoded areas during the GDR process.
[0137] IRAP pictures and GDR pictures may be collectively referred to as random access point (RAP) pictures.
[0138] 3.3. General Scalable Video Codec (SVC) and SVC in VVC
[0139] Scalable Video Codec (SVC, sometimes also referred to simply as scalability in video codec) refers to a video codec that uses a base layer (BL), sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a base quality level. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously coded layers. For example, the bottom layer can serve as the BL, while the top layer can serve as the EL. Intermediate layers can serve as either the EL or the RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be the EL for layers below the intermediate layer (such as the base layer or any intervening enhancement layers), while simultaneously serving as the RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, multiple views may exist, and information of one view may be utilized to encode (e.g., encode or decode) information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0140] In SVC, parameters used by an encoder or decoder are grouped into parameter sets based on the codec level at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be utilized by one or more codecs of different layers in a bitstream can be included in a video parameter set (VPS), and parameters that can be utilized by one or more pictures in a codec of a video sequence can be included in a sequence parameter set (SPS). Similarly, parameters that can be utilized by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters that are specific to individual slices can be included in a slice header. Similarly, an indication of which parameter set(s) a particular layer is using at a given time domain can be provided at various codec levels.
[0141] Due to the support of reference picture resampling (RPR) in VVC, support for bitstreams containing multiple layers (for example, two layers with SD and HD resolutions in VVC) can be designed without the need for any additional signal processing horizontal codec tools, because the upsampling required for spatial scalability support can use only RPR upsampling filters. However, high-level syntax changes (compared to not supporting scalability) are required for scalability support. Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards, including extensions of AVC and HEVC, the design of VVC scalability has been made as friendly to single-layer decoder design as possible. The decoding capabilities of multi-layer bitstreams are specified as if there is only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not require too many changes to be able to decode multi-layer bitstreams. Compared to the design of multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, the IRAP AU needs to contain pictures for each layer present in the CVS.
[0142] A VVC bitstream can consist of one or more output layer sets (OLS). An OLS is a set of layers that specifies one or more layers as output layers. An output layer is a layer that is output after being decoded.
[0143] Parameter Set
[0144] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0145] The SPS is designed to carry sequence-level header information, while the PPS is designed to carry infrequently changing picture-level header information. Using SPS and PPS eliminates the need to repeat infrequently changing information for every sequence or picture, thus avoiding redundant signaling of this information. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, eliminating the need for redundant transmission and improving error resilience.
[0146] The VPS is introduced to carry sequence level header information that is common to all levels in a multi-level bitstream.
[0147] APS is introduced to carry picture-level or slice-level information that requires relatively many bits to encode or decode, can be shared by multiple pictures, and can have relatively many different variations in a sequence.
[0148] 3.5. Grade, Level, and Horizontal Syntax and Semantics in VVC
[0149] In the latest VVC text (in JVET-Q2001-vE / v15), the PTL syntax and semantics are as follows.
[0150] 7.3.3.1 General Grade, Level and Level Grammar
[0151]
[0152] 7.4.4.1 General grade, level, and level semantics
[0153] The profile_tier_level() syntax structure provides level information and optionally profile, level, sub-profile and general constraint information.
[0154] When the profile_tier_level() syntax structure is included in a VPS, OlsInScope is one or more OLSs specified by the VPS. When the profile_tier_level() syntax structure is included in an SPS, OlsInScope is an OLS including only the layer that is the lowest layer among the layers referencing the SPS, and the lowest layer is an independent layer.
[0155] general_profile_idc indicates compliance with OlsInScope as specified in Annex A. The bitstream shall not contain values of general_profile_idc other than those specified in Annex A. Other values of general_profile_idc are reserved by ITU-T | ISO / IEC for future use.
[0156] general_tier_flag specifies the tier context used to interpret general_level_idc specified in Appendix A.
[0157] general_level_idc indicates compliance with OlsInScope as specified in Annex A. The bitstream shall not contain values of general_level_idc other than those specified in Annex A. Other values of general_level_idc are reserved by ITU-T | ISO / IEC for future use.
[0158] NOTE 1 - A larger value of general_level_idc indicates a higher level.The maximum level signaled in the DCI NAL unit for OlsInScope may be higher but not lower than the level signaled in the SPS for the CLVS contained within OlsInScope.
[0159] NOTE 2 - When OlsInScope conforms to multiple profiles, general_profile_idc shall indicate the profile that provides the preferred decoding result or preferred bitstream identification, as determined by the encoder (in a manner not specified in this specification).
[0160] NOTE 3 - When the CVS of OlsInScope conforms to different profiles, multiple profile_tier_level() syntax structures may be included in the DCI NAL unit so that for each CVS of OlsInScope, there is at least one set of indicated profiles, tiers and levels for a decoder capable of decoding the CVS.
[0161] num_sub_profiles specifies the number of general_sub_profile_idc[i] syntax elements.
[0162] general_sub_profile_idc[i] indicates the i-th registered interoperability metadata as specified in Rec. ITU-T T.35, the content of which is not specified in this specification.
[0163] subtial_level_present_flag[i] equal to 1 specifies that level information is present in the profile_tier_level() syntax structure for the sublayer representation with TemporalId equal to i. subtial_level_present_flag[i] equal to 0 specifies that level information is not present in the profile_tier_level() syntax structure for the sublayer representation with TemporalId equal to i.
[0164] ptl_alignment_zero_bits shall be equal to 0.
[0165] Apart from the provisions for the inference of the absence of a value, the semantics of the syntax element sublayer_level_idc[i] are identical to the syntax element general_level_idc, but apply to the sublayer representation with TemporalId equal to i.
[0166] When not present, the value of subsidiary_level_idc[i] is inferred as follows:
[0167] - sublayer_level_idc[maxNumSubLayersMinus1] is inferred to be equal to general_level_idc of the same profile_tier_level() structure,
[0168] - For i from maxNumSubLayersMinus1-1 to 0 (in descending order of the value of i) (inclusive), sublayer_level_idc[i] is inferred to be equal to sublayer_level_idc[i+1].
[0169] 3.6. HRD for IRAP-only representation of bitstream
[0170] One design for supporting HRD operation to fully specify the conformance of a sub-bitstream consisting only of IRAP AUs in the bitstream is described below.
[0171] 3.6.1. IRAP-only HRD Information SEI Message Syntax
[0172]
[0173] 3.6.2. IRAP-only HRD Information SEI Message Semantics
[0174] The IRAP-only HRD Information (IOH) SEI message contains information about the level of conformance of the sub-bitstream consisting only of IRAP AU sequences in the CVS set of the OLS to which the SEI message applies, denoted as targetCvss, when testing the conformance of the extracted bitstream containing the IRAP AU sequence according to Annex A. The OLS to which the IOH message applies is also referred to as the applicable OLS or associated OLS. The CVS in the rest of this subclause refers to the CVS of the applicable OLS. The IRAP AU sequence consists of all IRAP AUs within the targetCvss.
[0175] When an IOH SEI message exists for any AU of a CVS (in the bitstream or provided by external means not specified in this specification), an IOH SEI message will exist for the first AU of the CVS. IOH SEI messages continue in decoding order from the current AU until the next AU or the end of the bitstream containing an IOH SEI message whose content is different from the current IOH SEI message. All IOH SEI messages applied to the same CVS will have the same content.
[0176] irap_only_level_idc indicates that only the sub-bitstream corresponding to the IRAP AU of targetCvss conforms to the level specified in Annex A. The HRD-only IRAP information SEI message shall not contain the value of irap_only_level_idc unless specified in Annex A. Other values of irap_only_level_idc are reserved by ITU-T | ISO / IEC for future use.
[0177] irap_only_max_speedup_minus100 plus 100 divided by 100 specifies the maximum value of the speedup to be applied to the HRD timing of sub-bitstreams corresponding to IRAP-only AUs that still conform to the targetCvss of the indicated irap_only_level_idc. When not present, the value of irap_only_max_speedup is inferred to be 1.
[0178] irap_only_general_nal_hrd_params_present_flag equal to 1 specifies that NAL HRD parameters (belonging to the Type II bitstream conformance point) are present in the HRD-only IRAP information SEI message. irap_only_general_nal_hrd_params_present_flag equal to 0 specifies that NAL HRD parameters are not present in the HRD-only IRAP information SEI message.
[0179] irap_only_general_vcl_hrd_params_present_flag equal to 1 specifies that VCL HRD parameters (belonging to the Type I bitstream conformance point) are present in the HRD-only IRAP information SEI message. irap_only_general_vcl_hrd_params_present_flag equal to 0 specifies that VCL HRD parameters are not present in the HRD-only IRAP information SEI message.
[0180] irap_only_cpb_cnt_minus1 specifies the number of alternative CPB delivery plans plus 1. The value of irap_only_cpb_cnt_minus1 shall be in the range of 0 to 31 (inclusive).
[0181] The requirement for bitstream conformance is that irap_only_general_nal_hrd_params_present_flag, irap_only_general_vcl_hrd_params_present_flag, and irap_only_cpb_cnt_minus1 are equal to general_nal_hrd_params_present_flag, general_vcl_hrd_params_present_flag, and hrd_cpb_cnt_minus1, respectively.
[0182] irap_only_nal_bit_rate_value_minus1[i] (together with bit_rate_scale) specifies the input bit rate of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_nal_bit_rate_value_minus1[i] shall be between 0 and 2 32 The bit rate is given in bits per second by:
[0183] BitRate[maxSubLayer][i]=(irap_only_nal_bit_rate_value_minus1[i]+1)*2 (6+bit_rate_scale) *speedupFactor
[0184] where speeupFactor is a value in the range of 1 to irap_only_max_speeup_minus100 plus 100 divided by 100.
[0185] irap_only_nal_cpb_size_value_minus1[i] is used together with cpb_size_scale to specify the CPB size of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_nal_cpb_size_value_minus1[i] shall be between 0 and 2 32 -2 (inclusive). The CPB size in bits is given by:
[0186] CpbSize[maxSubLayer][i]=(irap_only_nal_cpb_size_value_minus1[i]+1)*2 (4+cpb_size_scale) .
[0187] irap_only_vcl_bit_rate_value_minus1[i] (together with bit_rate_scale) specifies the input bit rate of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_vcl_bit_rate_value_minus1[i] shall be between 0 and 2 32 The bit rate is given in bits per second by:
[0188] BitRate[maxSubLayer][i]=(irap_only_vcl_bit_rate_value_minus1[i]+1)*2 (6+bit_rate_scale) *speedupFactor
[0189] where speeupFactor is a value in the range of 1 to irap_only_max_speeup_minus100 plus 100 divided by 100.
[0190] irap_only_vcl_cpb_size_value_minus1[i] is used together with cpb_size_scale to specify the CPB size of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_vcl_cpb_size_value_minus1[i] shall be between 0 and 2 32 -2 (inclusive). The CPB size in bits is given by:
[0191] CpbSize[maxSubLayer][i]=(irap_only_vcl_cpb_size_value_minus1[i]+1)*2 (4+cpb_size_scale) .
[0192] 3.6.3. General SEI Payload Semantics
[0193] In the following text, changes are highlighted in bold and italics. Deleted text is marked with double brackets (e.g., [[a]] indicates the deletion of the character "a"). ...
[0195] The list VclAssociatedSeiList is set to include PayloadType values 3, 19, 45, 129, 137, 144, 145, 147 to 150 (inclusive), 153 to 156 (inclusive), 168 and 204.
[0196] The list PicUnitReppConSeiList is set to include PayloadType values 0, 1, 19, 45, 129, 133, 137, 147 to 150 (inclusive), 153 to 156 (inclusive), 168, 203, [[, and ]] 204
[0197] NOTE 4 - VclAssociatedSeiList consists of SEI message payload type values that infer constraints on the NAL unit headers of SEI NAL units based on the NAL unit headers of associated VCL NAL units when scalable nesting is not performed. PicUnitReppConSeiList consists of SEI message payload type values that are subject to the 4 repetitions per PU constraint.
[0198] A requirement for bitstream conformance is that the following restrictions apply to the inclusion of SEI messages in SEI NAL units:
[0199] - When general_same_pic_timing_in_all_ols_flag is equal to 1, there shall be no SEI NAL unit containing a scalable nesting SEI message with payloadType(PT) equal to 1, and when a SEI NAL unit contains a non-scalable nesting SEI message with payloadType(PT) equal to 1, the SEI NAL unit shall not contain any other SEI message with payloadType not equal to 1.
[0200] - When the SEI NAL unit contains a value equal to 0 (BP), 1 (PT), 130 (DUI), [[or]] 203 (SLI) When a non-scalable nested SEI message with a payload type of any other SEI message with a payload type of
[0201] - When a SEINAL unit contains a scalable nesting SEI message with a payload type equal to 0 (BP), 1 (PT), 130 (DUI), or 203 (SLI), the SEINAL unit shall not contain a payload type not equal to 0, 1, 130, 203, [or] 133 (scalable nesting) Any other SEI message.
[0202] - When a SEI NAL unit contains a SEI message with payload type equal to 3 (filler payload), the SEI NAL unit shall not contain any other SEI message with payload type not equal to 3.
[0203] The following applies to the applicable OLS or layer of the non-scalable nested SEI message:
[0204] - For non-scalable nested SEI messages, when payloadType is equal to 0 (BP), 1 (PT), 130 (DUI), [[or]] 203 (SLI) The non-scalable nesting SEI message applies to all OLSs (when present), which consist of all layers in the current CVS in the entireBitstream. When there is no OLS consisting of all layers in the current CVS, the entireBitstream, there shall be no non-scalable nesting SEI message with payload type equal to 0 (BP), 1 (PT), 130 (DUI), [[or]] 203 (SLI), or 205 (IOH).
[0205] - For a non-scalable nesting SEI message, when payloadType is equal to any value in VclAssociatedSeiList, the non-scalable nesting SEI message applies only to layers whose VCL NAL units have nuh_layer_id equal to the nuh_layer_id of the SEI NAL unit containing the SEI message.
[0206] A requirement for bitstream conformance is that the following restrictions apply to the value of nuh_layer_id of SEI NAL units:
[0207] - When a non-scalable nesting SEI message has payloadType equal to any value in VclAssociatedSeiList, the SEI NAL unit containing the non-scalable nesting SEI message shall have nuh_layer_id equal to the value of nuh_layer_id of the VCL NAL unit associated with the SEI NAL unit.
[0208] - The SEI NAL unit containing the scalable nesting SEI message shall have nuh_layer_id equal to the lowest value of nuh_layer_id of all layers to which the scalable nesting SEI message applies (when sn_ols_flag of the scalable nesting SEI message is equal to 0) or the lowest value of nuh_layer_id of all layers in the OLS to which the scalable nesting SEI message applies (when sn_ols_flag of the scalable nesting SEI message is equal to 1).
[0209] NOTE 5 - Same as DCI, OPI, VPS, AUD, and EOB NAL units, containing payload type equal to 0 (BP), 1 (PT), or 130 (DUI), [[or]] 203 (SLI) The value of nuh_layer_id of the SEI NAL unit of a non-scalable nesting SEI message is not restricted.
[0210] 3.6.4. Scalable Nested SEI Message Semantics
[0211] Scalable nesting SEI messages provide a mechanism to associate an SEI message with a specific OLS, a specific layer, or a specific sub-picture set.
[0212] The scalable nesting SEI message contains one or more SEI messages. The SEI message contained in the scalable nesting SEI message is also called a scalable nesting SEI message.
[0213] The requirement for bitstream conformance is that the following restrictions apply to the inclusion of SEI messages within scalable nested SEI messages:
[0214] - SEI messages with payload type equal to 3 (filler payload) or 133 (scalable nesting) shall not be included in a scalable nesting SEI message.
[0215] - When the scalable nesting SEI message contains a BP, PT, DUI, [[or]] SLI or 205 (IOH) SEI message, the scalable nesting SEI message shall not contain a payload type not equal to 0 (BP), 1 (PT), 130 (DUI), [[or]] 203 (SLI) Any other SEI message.
[0216] The requirement for bitstream conformance is that the following restrictions apply to the value of nal_unit_type for SEI NAL units containing scalable nesting SEI messages:
[0217] - When the scalable nesting SEI message contains an SEI message with payloadType (decoded picture hash) not equal to 132, the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT.
[0218] - When the scalable nesting SEI message contains an SEI message with payloadType equal to 132 (decoded picture hash), the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to SUFFIX_SEI_NUT.
[0219] sn_ols_flag = 0 specifies that the scalable nesting SEI message applies to a specific layer.
[0220] The requirement for bitstream conformance is that the following restrictions apply to the value of sn_ols_flag:
[0221] - When the scalable nesting SEI message contains a value equal to 0 (BP), 1 (PT), 130 (DUI), [[, or]] 203 (SLI) When the SEI message contains a payloadType of , the value of sn_ols_flag shall be equal to 1.
[0222] - When the scalable nesting SEI message contains a SEI message with payloadType equal to the value in VclAssociatedSeiList, the value of sn_ols_flag shall be equal to 0.
[0223] 3.6.5. Proposed (highlighted) changes to Appendix C:
[0224] In the following text, changes are highlighted in bold and italics. Deleted text is marked with double brackets (e.g., [[a]] indicates the deletion of the character "a").
[0225] C.1 Overview ...
[0227] For each test, the following sequenced steps are applied in the order listed, followed by the procedures described after those steps in this subclause:
[0228] 1. By selecting the table with OLS index opOlsIdx, the highest TemporalId value opTid, and optionally a list of target sub-picture indices opSubpicIdxList[j] for j ranging from 0 to NumLayersInOls[opOlsIdx]-1 (inclusive) The test operation point denoted as targetOp is selected. The value of opOlsIdx is in the range of 0 to TotalNumOlss-1 (inclusive). The value of opTid is in the range of 0 to vps_max_substins_minus1 (inclusive).
[0229] If opSubpicIdxList[] is not present, targetOp consists of pictures, and each pair of selected values of opOlsIdx and opTid shall be such that the sub-bitstream BitstreamToDecode is output by invoking the sub-bitstream extraction process as specified in subclause C.6, with entireBitstream, opOlsIdx, opTid, and The following conditions are met as input:
[0230] - There is at least one VCL NAL unit in BitstreamToDecode with TemporalId equal to opTid.
[0231] Otherwise (opSubpicIdxList[] is present), targetOp consists of sub-pictures, and for each set of selected values of opOlsIdx, opTid, and opSubpicIdxList[j] for j in the range 0 to NumLayersInOls[opOlsIdx]-1, inclusive, the sub-bitstream bitstream_todecode shall be such that it is output by invoking the sub-picture sub-bitstream extraction process as specified in subclause C.7, for entireBitstream, opOlsIdx, opTid, [[, and ]]opSubpicIdxList[j] for j in the range 0 to NumLayersInOls[opOlsIdx]-1, inclusive, The following conditions are met as input:
[0232] - There is at least one VCL NAL unit in BitstreamToDecode with TemporalId equal to opTid.
[0233] - For each j in the range 0 to NumLayerInOls[opOlsIdx]-1, there is at least one VCL NAL unit with nuh_layer_id equal to SubpicIdVal[opSubpicIdxList[j]] (inclusive).
[0234] NOTE 1 - Regardless of the presence of opSubpicIdxList[], due to the bitstream conformance requirement for each IRAP or GDR AU to be fulfilled, there is at least one VCL NAL unit with nuh_layer_id equal to LayerIdInOls[opOlsIdx][j] in the range from 0 to NumLayersInOls[opOlsIdx]-1 (inclusive) for each j.
[0235] 2. If opSubpicIdxList[] does not exist, the following applies:
[0236] - If the layers in targetOp include all layers in entireBitstream, and opTid is equal to the highest TemporalId value among all NAL units in entireBitstream, then BitstreamToDecode is set to the same as entireBitstream.
[0237] - Otherwise, set BitstreamToDecode as output by invoking the sub-bitstream extraction process as specified in subclause C.6, with entireBitstream, opOlsIdx and opTid as input.
[0238] Otherwise (opSubpicIdxList[] exists), BitstreamToDecode is set as output by calling the sub-picture sub-bitstream extraction process as specified in subclause C.7, with entireBitstream, opOlsIdx, opTid and opSubpicIdxList[j] for j in the range from 0 to NumLayersInOls[opOlsIdx]-1 (inclusive) as input.
[0239] 3. The values of TargetOlsIdx and Htid are set equal to the opOlsIdx and opTid of targetOp respectively.
[0240] 4. Select the general_timing_hrd_parameters() syntax structure, ols_timing_hrd_parameters() syntax structure, and sublayer_hrd_parameters() syntax structure applied to BitstreamToDecode as follows
[0241] - If NumLayersInOls[TargetOlsIdx] is equal to 1, the general_timing_hrd_parameters() syntax structure and the ols_timing_hrd_parameters() syntax structure in the SPS are selected (or provided by external means not specified in this specification). Otherwise, the general_timing_hrd_parameters() syntax structure and the vps_ols_timing_hrd_idx[MultiLayerOlsIdx[TargetOlsIdx]]-th ols_timing_hrd_parameters() syntax structure in the VPS are selected (or provided by external means not specified in this specification).
[0242] - Then, within the selected ols_timing_hrd_parameters() syntax structure, to test the Type I bitstream compliance point, the sublayer_hrd_parameters(Htid) syntax structure following the condition "if(general_vcl_hrd_params_present_flag)" is selected, and the variable NalHrdModeFlag is set to be equal to 0, and to test the Type II bitstream compliance point, the sublayer_hrd_parameters(Htid) syntax structure following the condition "if(general_nal_hrd_params_present_flag)" is selected, and the variable NalHrdModeFlag is set to be equal to 1. When BitstreamToDecode is a type II bitstream and NalHrdModeFlag is equal to 0, all non-VCL NAL units except PH and filler data NAL units, and all leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes, and trailing_zero_8bits syntax elements that form the byte stream from the NAL unit stream (as specified in Annex B) when present, are discarded from BitstreamToDecode, and the remaining bitstream is assigned to BitstreamToDecode.
[0243]
[0244] 5. The AU associated with the BP SEI message (present in BitstreamToDecode or available through external means not specified in this specification) that applies to the target operation is selected as the HRD initialization point and is called AU 0.
[0245] 6. When general_DU_hrd_params_present_flag in the selected general_timing_hrd_parameters() syntax structure is equal to 1, the CPB is scheduled to operate at the AU level (in this case, the variable DecodingUnitHrdFlag is set to 0) or at the DU level (in this case, the variable DecodingUnitHrdFlag is set to 1). Otherwise, DecodingUnitHrdFlag is set to 0 and the CPB is scheduled to operate at the AU level.
[0246] 7. For each AU in BitstreamToDecode starting from AU 0, select the BP SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the AU and applied to TargetOlsIdx, and select the PT SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the AU and applied to TargetOlsIdx, and when DecodingUnitHrdFlag is equal to 1 and bp_du_cpb_params_in_pic_timing_sei_flag is equal to 0, select the DUI SEI message (present in BitstreamToDecode or available through external means not specified in this specification) associated with the DU in the AU and applied to TargetOlsIdx.
[0247] 8. Select the value of ScIdx. The selected ScIdx should be in the range of 0 to hrd_cpb_cnt_minus1 (inclusive).
[0248] 9. When the bp_alt_cpb_params_present_flag of the BP SEI message associated with AU 0 is equal to 0 When BP SEI message associated with AU 0 has bp_alt_cpb_params_present_flag equal to 1, the variable DefaultInitCpbParamsFlag is set to 1. Otherwise, when the BP SEI message associated with AU 0 has bp_alt_cpb_params_present_flag equal to 1, any of the following shall be used to select the initial CPB removal delay and delay offset:
[0249] - If NalHrdModeFlag is equal to 1, the default initial CPB removal delay and delay offset indicated by bp_nal_initial_cpb_removal_delay[Htid][scidx] and bp_nal_initial_cpb_removal_offset[Htid][scidx], respectively, in the selected BP SEI message are selected. Otherwise, the default initial CPB removal delay and delay offset indicated by bp_vcl_initial_cpb_remotion_delay[Htid][ScIdx] and bp_vcl_initial_cpb_remotion_offset[Htid][ScIdx], respectively, in the selected BP SEI message are selected. The variable DefaultInitCpbCaramsFlag is set to 1.
[0250] If NalHrdModeFlag is equal to 1, the alternative initial CPB removal delay and delay offset denoted by bp_nal_initial_cpb_removal_delay[Htid][scidx] and bp_nal_initial_cpb_removal_offset[Htid][scidx], respectively, in the selected BP SEI message, and by pt_nal_cpb_alt_initial_removal_delay_delta[Htid][ScIdx] and pt_nal_cpb_alt_initial_removal_offset_delta[Htid][ScIdx], respectively, in the PT SEI message associated with the AU that follows AU 0 in decoding order are selected. Otherwise, the alternative initial CPB removal delay and delay offset represented by bp_vcl_initial_cpb_remotion_delay[Htid][ScIdx] and bp_vcl_initial_cpb_remotion_offset[Htid][ScIdx], respectively, in the selected BP SEI message and by pt_vcl_cpb_alt_initial_remotion_offset_delta[Htid][ScIdx], respectively, in the PT SEI message associated with the AU that follows AU 0 in decoding order are selected. The variable DefaultInitCpbCaramsFlag is set equal to 0, and one of the following applies:
[0251] - RASL AUs containing RASL pictures with pps_mixed_nalu_type_in_pic_flag equal to 0 and associated with CRA pictures contained in AU 0 are discarded from BitstreamToDecode, and the remaining bitstream is allocated to BitstreamToDecode.
[0252] - All AUs following AU 0 in decoding order up to the AU associated with the DRAP indication SEI message are discarded from BitstreamToDecode, and the remaining bitstream is allocated to BitstreamToDecode.
[0253] Each compliance test consists of a combination of one option selected in each of these steps. When there is more than one option for a step, only one option is selected for any particular compliance test. All possible combinations of all steps form the entire set of compliance tests. For each operating point in the test, the number of bitstream compliance tests to be performed is equal to The values of n0, n1, n2, n3, n4 and n5 are specified as follows:
[0254] -n0 is set equal to 2.
[0255] -n1 is equal to hrd_cpb_cnt_minus1+1.
[0256] -n2 is the number of AUs in each BitstreamToDecode associated with the BP SEI message that applies to TargetOlsIdx and for which all of the following conditions are true:
[0257] -nal_unit_type is equal to CRA_NUT.
[0258] - The associated BP SEI message has bp_alt_cpb_params_present_flag equal to 1.
[0259] - There is at least one RASL picture with pps_mixed_nalu_type_in_pic_flag equal to 0 associated with the AU.
[0260] -n3 is the number of IRAPs or GDR AUs in BitstreamToDecode, each IRAP or GDR AU is associated with a BP SEI message that applies to TargetOlsIdx, and for which at least one of the following conditions is false:
[0261] -nal_unit_type is equal to CRA_NUT.
[0262] - The associated BP SEI message has bp_alt_cpb_params_present_flag equal to 1.
[0263] - There is at least one RASL picture with pps_mixed_nalu_type_in_pic_flag equal to 0 associated with the AU.
[0264] -n4 is the number of AUs in BitstreamToDecode that are each associated with a DRAP indication SEI message that applies to TargetOlsIdx, and for each of which the associated PT SEI message has pt_cpb_alt_timing_info_present_flag equal to 1.
[0265] -n5 is derived as follows:
[0266] - If general_du_hrd_params_present_flag in the selected general_timing_hrd_parameters() syntax structure is equal to 0, then n5 is equal to 1.
[0267] -Otherwise, n5 is equal to 2.
[0268]
[0269] NOTE 2 - n0 corresponds to conformance testing for Type I bitstream conformance and Type II bitstream conformance. n1 corresponds to conformance testing per CPB delivery schedule. n2 corresponds to conformance testing of the bitstream starting from each CRA picture with the associated RASL picture and the presence of an alternative initial CPB removal delay and delay offset. These tests are performed twice: once for bitstream preservation and once for bitstream removal with the RASL picture associated with the CRA. n3 corresponds to conformance testing of the bitstream starting from each IRAP or GDR AU that is not a CRA with an associated RASL picture and the presence of an alternative initial CPB removal delay and delay offset. n4 corresponds to conformance testing of the bitstream starting from each IRAP with an associated DRAP picture with the presence of alternative timing information and the result of removing all AUs between the DRAP picture with the presence of alternative timing information and the preceding IRAP. n5 corresponds to AU-based conformance and DU-based conformance testing when general_du_hrd_params_present_flag is equal to 1.
[0270] When BitstreamToDecode is a type II bitstream, the following applies:
[0271] - and the sublayer_hrd_parameters(htid) syntax structure following the conditional "if (general_vcl_hrd_params_present_flag)" is selected, the test is performed at the Type I conformance point shown in Figure C.1 of Appendix C, and only VCL and filler data NAL units are counted for input bitrate and CPB storage.
[0272] -otherwise, and the sublayer_hrd_parameters(Htid) syntax structure following the condition "if (general_nal_hrd_params_present_flag)" is selected, the Type II conformance point shown in Figure C.1 of Appendix C is tested and all bytes of the Type II bitstream (which can be a NAL unit stream or a byte stream) are counted for the input bitrate and CPB storage.
[0273]
[0274] NOTE 3 - For the variable bit rate (VBR) case (cbr_flag[Htid][ScIdx] equal to 0), for the same values of InitCpbRemovalDelay[ScIdx], BitRate[Htid][ScIdx], and CpbSize[Htid][ScIdx], the NAL HRD parameters established by the ScIdx values for the Type II conformance point shown in C.1 of Annex C are also sufficient to establish VCL HRD conformance for the Type I conformance point shown in Figure C.1 of Annex C. This is because the data stream entering the Type I conformance point is a subset of the data stream entering the Type II conformance point, and because for the VBR case, the CPB is allowed to become empty and remain empty until the time domain at which the next picture starts to arrive is scheduled. ...
[0276] For each bitstream conformance test, the CPB size (in bits) is as given in subclause where the ScIdx and HRD parameters are as specified above in this subclause, and the DPB parameters dpb_max_dec_pic_buffering_minus1[Htid], dpb_max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found in the dpb_parameters() syntax structure applied to the target OLS or are derived as follows:
[0277] - If NumLayersInOls[TargetOlsIdx] is equal to 1, then find the dpb_parameters() syntax structure in the SPS and set the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY,MaxChromaFormat, and MaxBitDepthMinus8 to be equal to sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, sps_chroma_format_idc, and ps_bitdepth_minus8, respectively, found in the SPS.
[0278] Otherwise (NumLayersInOls[TargetOlsIdx] is greater than 1), the dpb_parameters() syntax structure is identified by vps_ols_dpb_params_idx[MultiLayerOlsIdx[TargetOlsIdx]] found in the VPS, and the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, and MaxBitDepthMinus8 are set equal to the vps_ ols_dpb_pic_width[MultiLayerOlsIdx[TargetOlsIdx]], vps_ols_dpb_pic_height[MultiLayerOlsIdx[TargetOlsIdx]], vps_ol s_dpb_chroma_format[MultiLayerOlsIdx[TargetOlsIdx]] and vps_ols_dpb_bitdepth_minus8[MultiLayerOlsIdx[TargetOlsIdx]].
[0279] If DecodingUnitHrdFlag is equal to 0, HRD operates at the AU level and each DU is an AU. Otherwise, HRD operates at the DU level and each DU is a subset of AU.
[0280] NOTE 6 - If the HRD operates at the AU level, each time some bits are removed from the CPB, the DU as the entire AU is removed from the CPB. Otherwise (HRD operates at the DU level), each time some bits are removed from the CPB, the DU as a subset of the AU is removed from the CPB. Regardless of whether the HRD operates at the access UNT level or the DU level, each time a picture is output from the DPB, the entire decoded picture is output from the DPB, but the picture output time domain is derived based on the CPB removal time derived in different ways and the DPB output delay signaled in different ways.
[0281] To express the constraints in this appendix, the following are specified:
[0282] Each AU is referred to as AU n, where the number n identifies the specific AU. AU 0 is selected per step 5 above. The value of n is +1 for each subsequent AU in decoding order.
[0283] - Each DU is referred to as DUm, where the number m identifies the specific DU. The first DU in decoding order in AU 0 is referred to as DU0. The value of m is incremented by 1 for each subsequent DU in decoding order.
[0284] NOTE 7 - The numbering of DUs is relative to the first DU in AU 0.
[0285] -Picture n refers to the coded or decoded picture of AU n.
[0286] HRD operates as follows:
[0287] - Initialize HRD to DU 0, set both CPB and DPB to empty (set DPB fullness equal to 0).
[0288] NOTE 8 - After initialization, HRD is not initialized again through subsequent BP SEI messages.
[0289] - Data associated with DUs arriving into the CPB as scheduled is delivered by the Hypothetical Stream Scheduler (HSS).
[0290] - The data associated with each DU is removed and decoded on-the-fly by the on-the-fly decoding process at the DU's CPB removal time.
[0291] - Place each decoded picture in the DPB.
[0292] - When a decoded picture no longer requires inter-frame prediction reference and no longer needs to be output, the decoded picture is removed from the DPB.
[0293] For each bitstream conformance test, the operation of the CPB is specified in subclause C.2, the transient decoder operation is specified in clauses 2 to 9, the operation of the DPB is specified in subclause C.3, and output clipping is specified in subclauses C.3.3 and C.5.2.2.
[0294] In subclauses 7.3.5.1 and 7.4.6.1 The HSS and HRD information regarding the number of enumerated delivery schedules and their associated bit rates and buffer sizes is specified. The HRD is initialized as specified in the BP SEI message (specified in subclause D.3). The timing of removing DUs from the CPB and outputting decoded pictures from the DPB is specified using information in the PT SEI message (specified in subclause D.4) or the DUI SEI message (specified in subclause D.5). All timing information associated with a particular DU shall arrive before the CPB removal time of the DU. ...
[0296] C.2.2 DU arrival timing ...
[0298] The final arrival time of DU m is derived as follows:
[0299] if (!decodingUnitParamsFlag)
[0300] AuFinalArrivalTime[m]=initArrivalTime[m]+sizeInbits[m]÷BitRate[ Htid][ScIdx] (1584)
[0301] else
[0302] DuFinalArrivalTime[m]=initArrivalTime[m]+sizeInbits[m]÷BitRate[ Htid][ScIdx]
[0303] where sizeInbits[m] is the size of DU m in bits, counting the bits of VCL NAL units, PH NAL units, and filler data NAL units for a Type I conformance point, or all bits of the Type II bitstream for a Type II conformance point, where Type I and Type II conformance points are as described in C.1 of Annex C.
[0304] The values of ScIdx,BitRate[ Htid][ScIdx]and CpbSize[ Htid][ScIdx]areconstrained as follows:
[0305] - If the content of the general_timing_hrd_parameters() syntax structure selected for the AU containing AU m is different from that of the previous AU, the HSS selects the value of ScIdx ScIdx1 from the values of ScIdx provided in the general_timing_hrd_parameters() syntax structure selected for the AU containing AU m, which results in a BitRate[
[0306] Htid][ScIdx1] or CpbSize[ Htid][ScIdx1]. BitRate[ Htid][ScIdx1] or CpbSize[ The value of [Htid][ScIdx1] may be different from the value of ScIdx used for the previous AU, ScIdx0. Htid][ScIdx0] or CpbSize[ The value of [Htid][ScIdx0].
[0307] Otherwise, HSS continues to use the previous ScIdx, BitRate[ Htid][ScIdx] and CpbSize[ Htid][ScIdx] value to operate.
[0308] When HSS selects BitRate different from those of previous AU[ Htid][ScIdx]or CpbSize[ Htid][ScIdx], the following applies:
[0309] -Variable BitRate[ Htid][ScIdx] takes effect at the initial CPB arrival time of the current AU.
[0310] -Variable CpbSize[ Htid][ScIdx] takes effect as follows:
[0311] - If CpbSize[ If the new value of Htid][ScIdx] is larger than the old CPB size, it takes effect at the initial CPB arrival time of the current AU.
[0312] - Otherwise, at the CPB removal time of the current AU, CpbSize[ The new value of [Htid][ScIdx] takes effect.
[0313] C.2.3 Timing of DU removal and DU decoding ...
[0315] The nominal removal time of AU n from CPB is specified as follows:
[0316] If AU n is an AU with n equal to 0 (the AU that initializes HRD), the nominal removal time of the AU from the CPB is defined as:
[0317]
[0318] - Otherwise, the following applies:
[0319] - When AU n is the first AU of the BP that does not initialize HRD, the following applies:
[0320] The nominal removal time of AU n from CPB is expressed by the following formula:
[0321]
[0322] where AuNominalRemovalTime[firstAuInPrevBuffPeriod] is the nominal removal time of the first AU of the previous BP, AuNominalRemovalTime[prevNonDiscardableAu] is the nominal removal time of the previous AU in decoding order with TemporalId equal to 0, which has at least one picture with TemporalId equal to 0 that is not a RASL or RADL picture, AuCpbRemovalDelayVal is the value of CpbRemovalDelayVal[Htid] derived from pt_cpb_removal_delay_minus1[Htid] and pt_cpb_removal_delay_delta_idx[Htid] in the PT SEI message, and the selected value of the AU as specified in subclause C.1. bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[Htid]] in the BP SEI message associated with AU n, and concatenationFlag and auCpbRemovalDelayDeltaMinus1 are the values of the syntax elements bp_concatenation_flag and bp_cpb_removal_delay_delta_minus1, respectively, in the BP SEI message associated with AU n, selected as specified in subclause C.1.
[0323] After deriving the nominal CPB removal time and before deriving the DPB output time domain for access unit n, the variables DpbDelayOffset and CpbDelayOffset are derived as:
[0324] - If one or more of the following conditions are true, DpbDelayOffset is set equal to the value of the PT SEI message syntax element pt_nal_dpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 1) or pt_vcl_dpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 0) of AU n+1, and CpbDelayOffset is set equal to the value of the PT SEI message syntax element pt_nal_cpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 1) or pt_vcl_cpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 0) of AU n+1, where the PT SEI message containing the syntax elements is selected as specified in subclause C.1:
[0325] - UseAltCpbParamsFlag for AU n is equal to 1.
[0326] -DefaultInitCpbParamsFlag is equal to 0.
[0327] Otherwise, both DpbDelayOffset and CpbDelayOffset are set equal to 0.
[0328] - When AU n is not the first AU of BP, the nominal removal time of AU n from CPB is specified as:
[0329]
[0330] where AunominalRemovalTime[firstAudInCurrBuffPeriod] is the nominal removal time of the first AU of the current BP, and AuCpbRemovalDelayVal is the value of CpbRemovery_delay_val[OpTid] derived from pt_cpb_removery_delay_minus1[OpTid] and pt_cpb_removery_delay_delta_idx[OpTid] in the PT SEI message, and bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[OpTid]] in the BP SEI message associated with AU n, selected as specified in subclause C.1. ...
[0332] C.6 IRAP AU sub-bitstream extraction process
[0333] The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, and the target highest TemporalId value tIdTarget
[0334] The output of this process is the sub-bitstream outBitstream.
[0335] The OLS with the OLS index TargetOlsIdx is called the target OLS.
[0336] Any output sub-bitstream that meets all of the following conditions shall be a conforming bitstream, as required for bitstream conformity of the input bitstream:
[0337] - The output sub-bitstream is the output of the process specified in this subclause, with the bitstream, targetOlsIdx equal to the index into the OLS list specified by the VPS, and tIdTarget equal to any value in the range of 0 to vps_ptl_max_tid[vps_ols_ptl_idx[targetOlsIdx]] (inclusive) as input.
[0338] - The output sub-bitstream contains at least one VCL NAL unit where nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0339] - The output sub-bitstream contains at least one VCL NAL unit with TemporalId equal to tIdTarget.
[0340] NOTE - A conforming bitstream contains one or more slice NAL units of a codec with TemporalId equal to 0, but does not necessarily contain a slice NAL unit of a codec with nuh_layer_id equal to 0.
[0341] The output sub-bitstream OutBitstream is derived by applying the following sequence of steps:
[0342] 1. The bitstream outBitstream is set to be the same as the bitstream inBitstream.
[0343] 2. Remove all NAL units whose TemporalId is greater than tIdTarget from outBitStream.
[0344] 3. Remove from outBitStream all NAL units with a nuh_layer_id not included in the list LayerIdInfols[targetOlsIdx] and that are not DCI, OPI, VPS, AUD, or EOB NAL units and are not SEI NAL units containing a non-scalable nesting SEI message with payload PayloadType equal to 0, 1, 130, or 203.
[0345] 4. Remove from outBitStream all APS and VCL NAL units for which all of the following conditions are true and whose associated non-VCL NAL units have NAL_unit_type equal to PH_NUT or FD_NUT, or have NAL_unit_type equal to SUFFIX_SEI_NUT or PREFIX_SEI_NUT and contain SEI messages with PayloadType not equal to any of 0 (BP), 1 (PT), 130 (DUI), and 203 (SLI):
[0346] - nal_unit_type is equal to APS_NUT, TRAIL_NUT, STSA_NUT, RADL_NUT, or RASL_NUT, or nal_unit_type is equal to GDR_NUT and the associated ph_recovery_poc_cnt is greater than 0.
[0347] - TemporalId is greater than or equal to NumSubLayersInLayerInOLS[targetOlsIdx][GeneralLayerIdx[nuh_layer_id]].
[0348]
[0349] 6. When all VCL NAL units of an AU are removed by steps 2, 3, or 4 above and there are AUD or OPI NAL units in the AU, remove the AUD or OPI NAL units from the bitstream.
[0350] 7. For each OPI NAL unit in outBitStream, set opi_htid_info_present_flag equal to 1, set opi_ols_info_present_flag equal to 1, set opi_htid_plus1 equal to tIdTarget+1, and set opi_ols_idx equal to targetOlsIdx.
[0351] 8. When an AUD exists in an AU in the outer bitstream and the AU becomes an IRAP or GDR AU, the aud_irap_or_gdr_flag of the AUD is set equal to 1.
[0352] 9. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages with sn_ols_flag equal to 1 and no i value in the range of 0 to sn_num_olss_minus1 (inclusive) such that NestingOlsIdx[i] is equal to targetOlsIdx.
[0353] 10. Remove from outBitstream all SEI NAL units containing scalable nesting SEI messages with sn_ols_flag equal to 0 and no value in list NestingLayerId equal to the value in list LayerIdOls[targetOlsIdx].
[0354] 11. When LayerIdInols[targetOlsIdx] does not include all values of nuh_layer_id in all VCL NAL units in the bitstream, the following applies in the order listed:
[0355] a. Remove all SEI NAL units from outBitStream that contain bits equal to 0 (BP), 130 (DUI), [[, or]] 203 (SLI) Non-scalable nested SEI message with a payload type of
[0356] b. When general_same_pic_timing_in_all_ols_flag is equal to 0, all SEI NAL units containing non-scalable nesting SEI messages with payload type equal to 1 are removed from outBitStream.
[0357] c. When outBitstream contains SEI NAL unit seiNalUnitA, which contains a scalable nesting SEI message with sn_subpic_flag equal to 1 and sn_subpic_flag equal to 0 that applies to the target OLS, or when NumLayerInols[targetOlsIdx] is equal to 1 and outBitstream contains SEI NAL unit seiNalUnitA, which contains a scalable nesting SEI message with sn_subpic_flag equal to 0 and sn_subpic_flag equal to 0 that applies to the layer in outBitstream, generate a new SEI NAL unit seiNalUnitB, include it in the PU that includes seiNalUnitA immediately following seiNalUnitA, extract the scalable nesting SEI messages from the scalable nesting SEI message and include them directly in seiNalUnitB (as a non-scalable nesting SEI message), and remove seiNalUnitA from outBitstream.
[0358] C.7 Sub-picture sub-bitstream extraction process
[0359] The input to this process is a list consisting of the bitstream, the target OLS index targetOlsIdx, the target highest TemporalId value tIdTarget, [[ and ]]i, where target sub-picture index values subpicIdxTarget[i] are in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive).
[0360] The output of this process is the sub-bitstream outBitstream.
[0361] The OLS with OLS index TargetOlsIdx is called target OLS. Among the layers in the target OLS, those layers whose referenced SPS has sps_num_subpics_minus1 greater than 0 are called multiSubpicLayers.
[0362] Any output sub-bitstream that meets all of the following conditions shall be a conforming bitstream, as required for bitstream conformity of the input bitstream:
[0363] - The output sub-bitstream is the output of the process specified by the bitstream in this sub-range, where targetOlsIdx is equal to the index of the OLS list specified by the VPS, tIdTarget is equal to any value in the range 0 to vps_max_sublayers_minus1 (inclusive), and the list of i subpicIdxTarget[i] from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive) satisfies the following conditions as input:
[0364] - The value of subpicIdxTarget[i] is equal to a value in the range of 0 to sps_num_subpics_minus1 (inclusive) such that sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] is equal to 1, where sps_num_subpics_minus1 and sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] are found in or inferred based on the SPS referenced by the layer with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i].
[0365] NOTE 1 - When sps_num_subpics_minus1 of the layer whose nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] is equal to 0, the value of subpicIdxTarget[i] is equal to 0.
[0366] - For any two different integer values of m and n, when sps_num_subpics_minus1 is greater than 0 for both layers with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][m] and LayerIdInOls[targetOlsIdx][n], respectively, subpicIdxTarget[m] is equal to subpicIdxTarget[n].
[0367] - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to each of the nuh_layer_id values in the list LayerIdInOls[targetOlsIdx].
[0368] - The output sub-bitstream contains at least one VCL NAL unit with TemporalId equal to tIdTarget.
[0369] NOTE 2 - A conforming bitstream contains one or more slice NAL units of a codec with TemporalId equal to 0, but does not necessarily contain a slice NAL unit of a codec with nuh_layer_id equal to 0.
[0370] - The output sub-bitstream contains at least one VCL NAL unit where nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is equal to SubpicIdVal[subpicIdxTarget[i]], where each i is in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive).
[0371] The output sub-bitstream outBitstream is derived by the following sequential steps:
[0372] a. Use inBitstream,targetOlsIdx,[[and]]tIdTarget The sub-bitstream extraction process specified in Appendix C.6 is called as input and the output of the process is assigned to outBitstream.
[0373] 4. Technical problems solved by the disclosed technical solution
[0374] In the intra-only trick mode used by some applications (e.g., for fast forward), the decoder decodes only IRAP-only representations consisting of only IRAP pictures in the bitstream, or only intra-only representations consisting of only intra-coded pictures in the bitstream. However, there is no mechanism to signal the level information of IRAP-only and / or intra-only representations. Knowing the level information is useful for the decoder to know whether it can consume the video bitstream at least in trick mode playback.
[0375] The second problem is as follows. JVET-R0193-v2 recommends signaling the max_tid_il_ref_pics_plus1 value separately for each direct reference layer of the layer, i.e., max_tid_il_ref_pics_plus1[i][j] for each direct reference layer j less than i, instead of the current single max_tid_il_ref_pics_plus1[i]. JVET-R0046 proposes adding the following constraint: the picture referenced by each ILRP in the RefPicList[0] or RefPicList[1] of the slice of the current picture shall be an IRAP picture or have a TemporalId less than max_tid_il_ref_pics_plus1[refPicVpsLayerId], where refPicVpsLayerId is equal to the nuh_layer_id of the referenced picture. However, the above constraint has two problems. First, the index of max_tid_il_ref_pics_plus1 should be a layer index rather than a layer ID. Second, for the JVET-R0193-v2 proposal, the max_tid_il_ref_pics_plus1 syntax element becomes two-dimensional.
[0376] The third issue is as follows. JVET-R0266 proposes signaling to replace the u(13) codec with the ue(v) codec for the virtual boundary position syntax elements sps_virtual_boundaries_pos_x[i], sps_virtual_boundaries_pos_y[i], ph_virtual_boundaries_pos_x[i], and ph_virtual_boundaries_pos_y[i]. However, on the other hand, the minimum allowed value for all of these syntax elements is 1. Therefore, it makes more sense to codec them with "_minus1" in the syntax element name.
[0377] The fourth issue is as follows. In the design of the HRD for IRAP-only representation of the bitstream, the variable maxSubLayer is used in many places but is not specified, so the HRD operation is not clearly defined. For IRAP AU sequences, since the TemporalId of all AUs in the sub-bitstream will be equal to 0, the variable should be replaced with the value 0.
[0378] The fifth issue is as follows. In the design of an HRD for IRAP-only representation of a bitstream, the speedup factor should not be selected for compliance testing, but rather signaled in the IOH SEI message as a property of the content. Once only IRAP AUs are included in the output of the sub-bitstream extraction process, the speedup factor for timing compared to the original bitstream is fixed. This means that irap_only_max_speedup_minus100 should be changed to irap_only_speedup_minus100, and the value of SpeedupFactor should be derived to be equal to (irap_only_speedup_minus100 + 100) / 100.
[0379] 5. Exemplary Embodiments and Solutions
[0380] To solve the above problems and other problems, the following methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these items can be applied alone or in any combination.
[0381] 1) In one example, level information and / or an indication of the presence of level information for the IRAP-only representation of each OLS may be signaled in the bitstream. The IRAP-only representation of an OLS consists only of IRAP pictures and associated non-VCL NAL units in the bitstream of the OLS.
[0382] 2) In one example, level information and / or an indication of the presence of level information for an intra-only representation of each OLS may be signaled in the bitstream. The intra-only representation of an OLS consists only of intra-coded pictures and associated non-VCL NAL units in the bitstream of the OLS.
[0383] 3) In one example, level information for each OLS and / or an indication of the presence of level information for IRAP-only representation and level information for intra-only representation may be signaled in the bitstream.
[0384] 4) In one example, level information and / or an indication of the presence of level information for IRAP-only and / or intra-only representation for OLS may be signaled in one or more of VPS, SPS, and DCI.
[0385] a. Alternatively, the level information and / or indication of the presence of the level information for IRAP-only and / or intra-only representation for OLS may be signaled in a SEI message.
[0386] i. Alternatively, the level information of the OLS is signaled in the SEI message, and a flag is further signaled in the SEI message to indicate whether the level information is for IRAP-only representation or for intra-only representation.
[0387] 5) In one example, level information for IRAP-only and / or intra-only representation of OLS and / or an indication of the presence of level information may be signaled in a profile, level, and level (PTL) syntax structure.
[0388] 6) In the above example, whether level information is signaled may depend on an indication of the presence of level information for IRAP-only and / or intra-only representations.
[0389] a. In one example, two indications (eg, two flags) are signaled to control the presence of level information for both IRAP-only representation and intra-only representation, respectively.
[0390] i. Alternatively, the level information may be signaled separately for IRAP-only presentation and intra-only presentation.
[0391] b. In one example, when the indication of the presence of level information for IRAP-only (and / or intra-only) representation specifies that level information is present, the level information for IRAP-only (and / or intra-only) representation is present in the profile_tier_level() syntax structure. Otherwise, when the indication of the presence of level information specifies that level information is not present for IRAP-only (and / or intra-only) representation, the level information for IRAP-only (and / or intra-only) representation is not present in the profile_tier_level() syntax structure.
[0392] c. In one example, only one indication (eg, one flag) is signaled to control the presence of level information for both IRAP-only representation and intra-only representation.
[0393] i. Furthermore, alternatively, the primary level information may be signaled for both IRAP-only presentation and intra-only presentation.
[0394] 7) In another example, HRD parameters (both sequence level and AU / picture level) are signaled for each of the IRAP-only representation and / or intra-only representation, and HRD conformance tests are additionally specified to ensure conformance of each of the IRAP-only representation bitstream and / or the intra-only representation bitstream.
[0395] 8) To solve the second problem, the following constraints are specified in VVC:
[0396] The reference picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be an IRAP picture or have a TemporalId less than max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx], where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refPicLayerId] respectively, and refPicLayerId is the nuh_layer_id of the reference picture.
[0397] 9) Alternatively, to solve the second problem, the following constraints are adopted in VVC:
[0398] The pictures referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall exist in the DPB and shall have a nuh_layer_id that is smaller than that of the current picture.
[0399] Modify it as follows:
[0400] The picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB and shall have a nuh_layer_idrefPicLayerId less than the nuh_layer_id of the current picture and shall be an IRAP picture or have a TemporalId less than max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx], where CurrLayerIdx and RefLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[RefPicLayerId], respectively.
[0401] 10) Alternatively, to solve the second problem, the following constraints are specified in VVC:
[0402] The reference picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be an IRAP picture or have a TemporalId less than or equal to Max Max(0,max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId] respectively, and refpicLayerId is the nuh_layer_id of the reference picture.
[0403] 11) Alternatively, to solve the second problem, the following constraints are adopted in VVC:
[0404] The pictures referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall exist in the DPB and shall have a nuh_layer_id that is smaller than the current picture.
[0405] Modify it as follows:
[0406] The picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB, shall have a nuh_layer_idrefPicLayerId less than the nuh_layer_id of the current picture, and shall be an IRAP picture or have a TemporalId less than or equal to Max(0, max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where CurrLayerIdx and RefLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[RefPicLayerId], respectively.
[0407] 12) In order to solve the third problem, the following solution is proposed.
[0408] Virtual boundary position syntax element and Renamed to and
[0409] a. In addition, alternatively, at least one or all of them are ue(v) coded or u(v) coded.
[0410] b. and The semantics of are as follows:
[0411] [i] plus 1 specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8. The value of sps_virtual_boundary_pos_x_minus1[i] shall be in the range of 0 to Ceil(pic_width_max_in_luma_samples ÷ 8) - 2, inclusive.
[0412] [i] plus 1 specifies the position of the i-th horizontal virtual boundary in units of luma samples divided by 8. The value of sps_virtual_boundary_pos_y_minus1[i] shall be in the range of 0 to Ceil(pic_height_max_in_luma_samples ÷ 8) - 2, inclusive.
[0413] [i] plus 1 specifies the position of the i-th vertical virtual boundary in units of luma samples divided by 8. The value of ph_virtual_boundary_pos_x_minus1[i] shall be in the range of 0 to Ceil(pic_width_in_luma_samples ÷ 8) - 2 (inclusive).
[0414] The list VirtualBoundariesPosX[i], where i ranges from 0 to NumVerVirtualBoundaries-1 (inclusive), specifies the position of the vertical virtual boundaries in units of luma samples, derived as follows:
[0415] for(i=0;i <NumVerVirtualBoundaries;i++)
[0416] VirtualBoundariesPosX[i]=(sps_virtual_boundaries_present_flag?
[0417] (sps_virtual_boundary_pos_x_minus1[i]+1):(ph_virtual_boundary_pos_x_minus1[i]+1))*8 (84)
[0418] The distance between any two vertical virtual boundaries shall be greater than or equal to CtbSizeY luma samples.
[0419] [i] plus 1 specifies the position of the i-th horizontal virtual boundary in units of luma samples divided by 8. The value of ph_virtual_boundary_pos_y_minus1[i] shall be in the range of 0 to Ceil(pic_height_in_luma_samples ÷ 8) - 2 (inclusive).
[0420] The list VirtualBoundariesPosY[i] in the range from 0 to NumHorVirtualBoundaries-1 (inclusive) for i specifies the position of the horizontal virtual boundaries in units of luma samples, derived as follows:
[0421] for(i=0;i <NumHorVirtualBoundaries;i++)
[0422] VirtualBoundariesPosY[i]=(sps_virtual_boundaries_present_flag?
[0423] (sps_virtual_boundary_pos_y_minus1[i]+1):(ph_virtual_boundary_pos_y_minus1[i]+1))*8 (86)
[0424] The distance between any two horizontal virtual boundaries shall be greater than or equal to CtbSizeY luma samples.
[0425] 13) How to signal the number of horizontal virtual boundaries (e.g., in ) and / or the value range of the number of horizontal virtual boundaries may depend on the number of vertical virtual boundaries (e.g., ).
[0426] a. In one example, when the number of vertical virtual borders is equal to 0, instead of signaling the number of horizontal virtual borders, the number of horizontal virtual borders minus 1 is signaled.
[0427] b. Alternatively, when the number of vertical virtual boundaries is equal to 0, the number of horizontal virtual boundaries may be in the range of [1, Ceil(pic_height_in_luma_samples÷8)–1]. Otherwise, it may be in the range of [0, Ceil(pic_height_in_luma_samples÷8)–1].
[0428] 14) Whether and / or how to signal the disabling of in-loop filtering on virtual boundaries applied in coded pictures in CLVS (e.g. ) and / or the number of horizontal / vertical virtual boundaries in the SPS (e.g., ) may depend on the layout of the sub-picture.
[0429] a. In one example, whether signaling Depends on whether there is more than 1 sub-picture.
[0430] i. In one example, only when there is more than 1 sub-picture, it can be signaled
[0431] ii. In addition, alternatively, when If not present, it is inferred to be 0.
[0432] b. In one example, instead of indicating a control for disabling in-loop filtering across virtual boundaries to be applied to all coded pictures in the CLVS, a control for disabling in-loop filtering across virtual boundaries to be applied to sub-pictures may be indicated in the bitstream.
[0433] i. In one example, assuming there are N sub-pictures, N control flags may be signaled, and each flag is specific to a sub-picture.
[0434] 15) To address the fourth issue, all occurrences of the variable maxSubLayer are replaced with the value 0 in the design of the HRD for IRAP-only representation of the bitstream.
[0435] 16) To address the fifth issue, in the design of HRD for IRAP-only representation of the bitstream, the acceleration factor is not selected for compliance testing but is signaled in the IOH SEI message.
[0436] a. In one example, the irap_only_max_speedup_minusM (eg, M=100) syntax element in the IOH SEI message syntax is changed to irap_only_speeup_minusM, and the value of SpeedupFactor is derived to be equal to (irap_only_speeup_minusM+M) / M.
[0437] b. In one example, (irap_only_speedup_minusM+M) should be no less than 2 K -1 (eg, K=32 or K depends on how the syntax is signaled).
[0438] c. In one example, instead of signaling irap_only_speedup_minusM (e.g., M=100), a syntax element named irap_only_speedup may be directly signaled, which specifies the maximum value of the acceleration of the HRD timing to be applied to the sub-bitstream corresponding to the IRAP AU sequence that still conforms to the targetCvss of the indicated irap_only_level_idc.
[0439] i. In one example, irap_only_speedup is signaled with u(k) (eg, k=32).
[0440] ii. In one example, the signaled value of irap_only_speedup is not less than 100.
[0441] iii. In one example, when irap_only_speedup is not present, it is inferred to be equal to 100.
[0442] iv. Furthermore, alternatively, the value of SpeedupFactor is derived to be equal to (irap_only_speedup_minus+offset) / M.
[0443] 1. In one example, offset is set to 0.
[0444] 2. In one example, offset is set to (M / 2).
[0445] 6. Examples
[0446] The following are some exemplary embodiments of some aspects of the invention outlined in Section 5 above that can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-Q2001-vE / v15. Most relevant sections that have been added or modified are highlighted in bold italics, and some deleted sections are marked with double brackets (e.g., [[a]] represents the deletion of the character "a"). There are some other changes that are editorial in nature and are therefore not highlighted.
[0447] 6.1. First embodiment
[0448] This embodiment applies to items 1, 4, and 5.
[0449] 7.3.3.1 General Grade, Level and Level Grammar
[0450]
[0451] 7.4.4.1 General grade, level, and level semantics ...
[0453] sublayer_level_present_flag[i] equal to 1 specifies the presence of level information in the profile_tier_level() syntax structure for the sublayer presentation with TemporalId equal to i-1, sublayer_level_present_flag[i] equal to 0 specifies that there is no level information in the profile_tier_level() syntax structure of the sublayer representation with TemporalId equal to i-1,
[0454] ptl_alignment_zero_bits shall be equal to 0.
[0455] The semantics of the syntax element subtial_level_idc[i] are identical to those of the syntax element general_level_idc, except for the specification of the inference of a non-existent value. Applied to the sub-layer representation with TemporalId equal to i-1,
[0456] When not present, the value of subsidiary_level_idc[i] is inferred as follows:
[0457] - sublayer_level_idc[maxNumSubLayersMinus1+1] is inferred to be equal to general_level_idc of the same profile_tier_level() structure,
[0458] - For i from maxNumSubLayersMinus1-1 to 0 (in descending order of the value of i) (inclusive), sublayer_level_idc[i] is inferred to be equal to sublayer_level_idc[i+1]. ...
[0460] A.1 Overview of grades, levels and grades ...
[0462] For each operating point identified by TargetOlsIdx and Htid, the profile, tier and level information are indicated by general_profile_idc and general_tier_flag in the VPS and sublayer_level_idc[Htid+1] found in or derived from the VPS.
[0463] When no VPS is available, the profile and tier information are indicated by general_profile_idc and general_tier_flag in the SPS, and the level information is indicated as follows:
[0464] – If Htid is provided by external means indicating the highest TemporalId of any NAL unit in the bitstream, the level information is indicated by sublayer_level_idc[Htid+1] found in or derived from the SPS.
[0465] – Otherwise (Htid is not provided by external means), the level information is indicated in the SPS via general_level_idc. ...
[0467] 6.2. Second embodiment
[0468] This embodiment is used for items 3, 4 and 5.
[0469] 7.3.3.1 General Grade, Level and Level Grammar
[0470]
[0471] 7.4.4.1 General grade, level, and level semantics ...
[0473] When i is greater than 1, subtial_level_present_flag[i] equal to 1 specifies the presence of level information in the profile_tier_level() syntax structure for the sublayer representation with TemporalId equal to i-2; subtial_level_present_flag[i] equal to 0 specifies that there is no level information in the profile_tier_level() syntax structure of the sublayer representation with TemporalId equal to i-2;
[0474] ptl_alignment_zero_bits shall be equal to 0.
[0475] The semantics of the syntax element subtial_level_idc[i] are identical to those of the syntax element general_level_idc, except for the provisions regarding the inference of an absent value.
[0476] When not present, the value of subsidiary_level_idc[i] is inferred as follows:
[0477] - sublayer_level_idc[maxNumSubLayersMinus1+2] is inferred to be equal to general_level_idc of the same profile_tier_level() structure,
[0478] - For i from maxNumSubLayersMinus1-1 to 0 (in descending order of the value of i) (inclusive), sublayer_level_idc[i] is inferred to be equal to sublayer_level_idc[i+1]. ...
[0480] A.1 Overview of grades, levels and grades ...
[0482] For each operating point identified by TargetOlsIdx and Htid, the profile, grade and level information are indicated by general_profile_idc and general_tier_flag in the VPS and sublayer_level_idc[Htid+2] found in or derived from the VPS.
[0483] When no VPS is available, the profile and tier information are indicated by general_profile_idc and general_tier_flag in the SPS, and the level information is indicated as follows:
[0484] – If Htid is provided by external means indicating the highest TemporalId of any NAL unit in the bitstream, the level information is indicated by sublayer_level_idc[Htid+2] found in or derived from the SPS.
[0485] – Otherwise (Htid is not provided by external means), the level information is indicated in the SPS via general_level_idc. ...
[0487] 6.3. Third embodiment
[0488] This embodiment addresses items 15 and 16. The textual changes are related to the HRD design for IRAP-only representation of the bitstream in section 3.7 of this document.
[0489] 6.3.1. IRAP-only HRD Information SEI Message Syntax
[0490]
[0491] 6.3.2. IRAP-only HRD Information SEI Message Semantics
[0492] The IRAP-only HRD Information (IOH) SEI message contains information about the level of conformance of the sub-bitstream consisting only of IRAP AU sequences in the CVS set of the OLS to which the SEI message applies, denoted as targetCvss, when testing the conformance of the extracted bitstream containing the IRAP AU sequence according to Annex A. The OLS to which the IOH message applies is also referred to as the applicable OLS or associated OLS. The CVS in the rest of this subclause refers to the CVS of the applicable OLS. The IRAP AU sequence consists of all IRAP AUs within the targetCvss.
[0493] When an IOH SEI message exists for any AU of a CVS (in the bitstream or provided by external means not specified in this specification), an IOH SEI message will exist for the first AU of the CVS. IOH SEI messages continue in decoding order from the current AU until the next AU or the end of the bitstream containing an IOH SEI message whose content is different from the current IOH SEI message. All IOH SEI messages applied to the same CVS will have the same content.
[0494] Indicates the level to which the sub-bitstream of the IRAP AU sequence corresponding to the target Cvss conforms as specified in Annex A. The IOH SEI message shall not contain the value of irap_only_level_idc except the value specified in Annex A. Other values of irap_only_level_idc are reserved by ITU-T | ISO / IEC for future use.
[0495] Specifies the [[maximum]] speedup value to be applied to the HRD timing of the sub-bitstream corresponding to the IRAP AU sequence that conforms to the targetCvss of the indicated irap_only_level_idc plus 100 divided by 100. When not present, the value of irap_only_[[max]]_speedup is inferred to be 1.
[0496]
[0497] equal to 1 specifies that NAL HRD parameters (regarding Type II bitstream conformance point) are present in the IOH SEI message. irap_only_general_nal_hrd_params_present_flag equal to 0 specifies that NAL HRD parameters are not present in the IOH SEI message.
[0498] =1 specifies that the VCL HRD parameters (regarding the Type I bitstream conformance point) are present in the IOH SEI message. =irap_only_general_vcl_hrd_params_present_flag = 0 specifies that the VCL HRD parameters are not present in the IOH SEI message.
[0499] The number of alternative CPB delivery plans is specified by adding 1. The value of irap_only_cpb_cnt_minus1 shall be in the range of 0 to 31 (inclusive).
[0500] The requirement for bitstream conformance is that irap_only_general_nal_hrd_params_present_flag, irap_only_general_vcl_hrd_params_present_flag, and irap_only_cpb_cnt_minus1 are equal to general_nal_hrd_params_present_flag, general_vcl_hrd_params_present_flag, and hrd_cpb_cnt_minus1, respectively.
[0501] (Together with bit_rate_scale) Specifies the input bit rate of the i-th CPB of the sub-bitstream of the IRAP AU sequence corresponding to targetCvss when the CPB operates at the AU level. irap_only_nal_bit_rate_value_minus1[i] shall be between 0 and 2 32 The bit rate is given in bits per second by:
[0502] [[BitRate[maxSubLayer][i]=(irap_only_nal_bit_rate_value_minus1[i]+1)*2 (6+bit_rate_scale) *speedupFactor
[0503] where speeupFactor is a value in the range of 1 to irap_only_max_speeup_minus100 plus 100 divided by 100. ]]
[0504]
[0505] [i] (together with cpb_size_scale) specifies the CPB size of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_nal_cpb_size_value_minus1[i] shall be between 0 and 2. 32 -2 (inclusive). The CPB size in bits is given by:
[0506] [[CpbSize[maxSubLayer][i]=(irap_only_nal_cpb_size_value_minus1[i]+1)*2 (4+cpb_size_scale) .]]
[0507]
[0508] (Together with bit_rate_scale) Specifies the input bit rate of the i-th CPB of the sub-bitstream of the IRAP AU sequence corresponding to targetCvss when the CPB operates at the AU level. irap_only_vcl_bit_rate_value_minus1[i] shall be between 0 and 2 32 The bit rate is given in bits per second by:
[0509] [[BitRate[maxSubLayer][i]=(irap_only_vcl_bit_rate_value_minus1[i]+1)*2 (6+bit_rate_scale) *speedupFactor
[0510] where speeupFactor is a value in the range of 1 to irap_only_max_speeup_minus100 plus 100 divided by 100.
[0511]
[0512] [i] (together with cpb_size_scale) specifies the CPB size of the i-th CPB of the sub-bitstream corresponding only to the IRAP AU of targetCvss when CPB operates at AU level. irap_only_vcl_cpb_size_value_minus1[i] shall be between 0 and 2. 32 -2 (inclusive). The CPB size in bits is given by:
[0513] [[CpbSize[maxSubLayer][i]=(irap_only_vcl_cpb_size_value_minus1[i]+1)*2 (4+cpb_size_scale) .]]
[0514]
[0515] 6.3.3. Proposed (highlighted) changes to Appendix C:
[0516] C.1 Overview ...
[0518] For each test, the following sequenced steps are applied in the order listed, followed by the procedures described after those steps in this subclause:
[0519] 1. Select the test operation point denoted as targetOp by selecting the list of target sub-picture index values opSubpicIdxList[j] for j from 0 to NumLayersInOls[opOlsIdx]-1, inclusive [[and an optional playback speedup value trickPlaySpeedup]] with OLS index opOlsIdx, the highest TemporalId value opTid, onlyIrapAusFlag, and optionally, a playback speedup value trickPlaySpeedup. The value of opOlsIdx is in the range of 0 to TotalNumOlss-1, inclusive. The value of opTid is in the range of 0 to vps_max_substins_minus1, inclusive. [[The value of speedupFactor is in the range of 1 to irap_only_max_speedup_minus100 plus 100 divided by 100. ]]
[0520] If opSubpicIdxList[] is not present, targetOp consists of pictures, and each pair of selected values of opOlsIdx and opTid shall be such that the sub-bitstream BitstreamToDecode is output by invoking the sub-bitstream extraction process as specified in subclause C.6 with entireBitstream, opOlsIdx, opTid and onlyIrapAusFlag as inputs satisfying the following conditions:
[0521] - There is at least one VCL NAL unit in BitstreamToDecode with TemporalId equal to opTid.
[0522] Otherwise (opSubpicIdxList[] is present), targetOp consists of sub-pictures, and for each set of selected values of opOlsIdx, opTid, and opSubpicIdxList[j] for j in the range 0 to NumLayersInOls[opOlsIdx]-1, inclusive, the sub-bitstream bitstream_todecode shall be output by invoking the sub-picture sub-bitstream extraction process as specified in subclause C.7, with entireBitstream, opOlsIdx, opTid, opSubpicIdxList[j] in the range 0 to NumLayersInOls[opOlsIdx]-1, inclusive, and onlyIrapAusFlag for j satisfying the following conditions as input:
[0523] - There is at least one VCL NAL unit in BitstreamToDecode with TemporalId equal to opTid.
[0524] - For each j in the range 0 to NumLayerInOls[opOlsIdx]-1, there is at least one VCL NAL unit with nuh_layer_id equal to SubpicIdVal[opSubpicIdxList[j]] (inclusive).
[0525] NOTE 1 - Regardless of the presence of opSubpicIdxList[], due to the bitstream conformance requirement for each IRAP or GDR AU to be fulfilled, there is at least one VCL NAL unit with nuh_layer_id equal to LayerIdInOls[opOlsIdx][j] in the range from 0 to NumLayersInOls[opOlsIdx]-1 (inclusive) for each j. ...
[0527] C.2.2 DU arrival timing ...
[0529] The final arrival time of DU m is derived as follows:
[0530]
[0531] where sizeInbits[m] is the size of DU m in bits, counting the bits of VCL NAL units, PH NAL units, and filler data NAL units for a Type I conformance point, or all bits of the Type II bitstream for a Type II conformance point, where Type I and Type II conformance points are as described in C.1 of Annex C.
[0532] ScIdx, BitRate[onlyIrapAusFlag?0:Htid][ScIdx] and CpbSize[onlyIrapAusFlag?0:Htid][ScIdx] are constrained as follows:
[0533] - If the content of the general_timing_hrd_parameters() syntax structure selected for the AU containing AU m is different from that of the previous AU, the HSS selects the value of ScIdx ScIdx1 from the value of ScIdx provided in the general_timing_hrd_parameters() syntax structure selected for the AU containing AU m, which results in BitRate[onlyIrapAusFlag?0:Htid][ScIdx1] or CpbSize[onlyIrapAusFlag?0:Htid][ScIdx1] for the AU containing AU m. The value of BitRate[onlyIrapAusFlag?0:Htid][ScIdx1] or CpbSize[onlyIrapAusFlag?0:Htid][ScIdx1] may be different from the value of ScIdx ScIdx1 for the previous AU. 0:Htid][ScIdx0] or the value of CpbSize[onlyIrapAusFlag? 0:Htid][ScIdx0].
[0534] Otherwise, the HSS continues to operate with the previous values of ScIdx, BitRate[onlyIrapAusFlag?0:Htid][ScIdx] and CpbSize[onlyIrapAusFlag?0:Htid][ScIdx].
[0535] When the HSS selects values of BitRate[onlyIrapAusFlag?0:Htid][ScIdx] or CpbSize[onlyIrapAusFlag?0:Htid][ScIdx] that are different from those of the previous AU, the following applies:
[0536] - The variable BitRate[onlyIrapAusFlag?0:Htid][ScIdx] takes effect at the initial CPB arrival time of the current AU.
[0537] - The variable CpbSize[onlyIrapAusFlag?0:Htid][ScIdx] is implemented as follows:
[0538] - Is there a new CpbSize[onlyIrapAusFlag]? 0:Htid][ScIdx] value that is larger than the old CPB size, which takes effect at the initial CPB arrival time of the current AU.
[0539] Otherwise, the new value of CpbSize[onlyIrapAusFlag?0:Htid][ScIdx] takes effect at the CPB removal time of the current AU.
[0540] C.2.3 DU removal and DU decoding time ...
[0542] The nominal removal time of AU n from CPB is specified as follows:
[0543] If AU n is an AU with n equal to 0 (the AU that initializes HRD), the nominal removal time of the AU from the CPB is defined as:
[0544] AuNominalRemovalTime[0]=InitCpbRemovalDelay[ScIdx]÷SpeedupFactor÷90000 (1585)
[0545] - Otherwise, the following applies:
[0546] - When AU n is the first AU of the BP that does not initialize HRD, the following applies:
[0547] The nominal removal time of AU n from CPB is expressed by the following formula:
[0548]
[0549]
[0550] where AuNominalRemovalTime[firstAuInPrevBuffPeriod] is the nominal removal time of the first AU of the previous BP, AuNominalRemovalTime[prevNonDiscardableAu] is the nominal removal time of the previous AU in decoding order with TemporalId equal to 0, which has at least one picture with TemporalId equal to 0 that is not a RASL or RADL picture, AuCpbRemovalDelayVal is the value of CpbRemovalDelayVal[Htid] derived from pt_cpb_removal_delay_minus1[Htid] and pt_cpb_removal_delay_delta_idx[Htid] in the PT SEI message, and the selected value of the AU as specified in subclause C.1. bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[Htid]] in the BP SEI message associated with AU n, and concatenationFlag and auCpbRemovalDelayDeltaMinus1 are the values of the syntax elements bp_concatenation_flag and bp_cpb_removal_delay_delta_minus1, respectively, in the BP SEI message associated with AU n, selected as specified in subclause C.1.
[0551] After deriving the nominal CPB removal time and before deriving the DPB output time domain for access unit n, the variables DpbDelayOffset and CpbDelayOffset are derived as:
[0552] - If one or more of the following conditions are true, DpbDelayOffset is set equal to the value of the PT SEI message syntax element pt_nal_dpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 1) or pt_vcl_dpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 0) of AU n+1, and CpbDelayOffset is set equal to the value of the PT SEI message syntax element pt_nal_cpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 1) or pt_vcl_cpb_delay_offset[Htid] (when NalHrdModeFlag is equal to 0) of AU n+1, where the PT SEI message containing the syntax elements is selected as specified in subclause C.1:
[0553] - UseAltCpbParamsFlag for AU n is equal to 1.
[0554] -DefaultInitCpbParamsFlag is equal to 0.
[0555] Otherwise, both DpbDelayOffset and CpbDelayOffset are set equal to 0.
[0556] - When AU n is not the first AU of BP, the nominal removal time of AU n from CPB is specified as:
[0557] AuNominalRemovalTime[n]=AuNominalRemovalTime[firstAuInCurrBuffPeriod]+ClockTick*(AuCpbRemovalDelayVal-CpbDelayOffset)÷SpeedupFactor (1587)
[0558] where AuNominalRemovalTime[firstAuInCurrBuffPeriod] is the nominal removal time of the first AU of the current BP, and AuCpbRemovalDelayVal is a value derived from pt_cpb_removal_delay_minus1[OpTid] and pt_cpb_removal_delay_delta_idx[OpTid]] in the PT SEI message, and bp_cpb_removal_delay_delta_val[pt_cpb_removal_delay_delta_idx[OpTid]] in the BP SEI message associated with AU n selected as specified in subclause C.1. ...
[0560] C.6 IRAP AU sub-bitstream extraction process ...
[0562] 5. When onlyIrapAusFlag is equal to 1, remove all NAL units applicable to the following AUs from outBitstream:
[0563] -NumLayersInOls[targetOlsIdx] is equal to 1 and Not equal to IDR_W_RADL, IDR_NLP, or CRA_NUT.
[0564] -NumLayersInOls[targetOlsIdx] is greater than 1 and There are no AUD NAL units with aud_irap_or_gdr_flag equal to 1.
[0565] Figure 1 1 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or a cellular interface.
[0566] System 1900 may include a codec component 1904 that implements various codecs or encoding methods described herein. Codec component 1904 may reduce the average bit rate of the video from input 1902 to output of codec component 1904 to generate a coded representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via connected communications, as represented by component 1906. The stored or transmitted bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or displayable video that is sent to display interface 1910. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it will be understood that codec tools or operations are used at the encoder, and corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0567] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0568] Figure 2 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more of the methods described herein. Device 3600 can be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more of the methods described herein. Memory 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.
[0569] Figure 4 is a block diagram illustrating an exemplary video coding system 100 that may utilize the techniques of this disclosure.
[0570] like Figure 4 As shown in , a video codec system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.
[0571] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0572] Video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. A codec picture is a codec representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to destination device 120 via network 130a via I / O interface 116. The encoded video data may also be stored on storage media / server 130b for access by destination device 120.
[0573] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0574] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage media / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, configured to interface with an external display device.
[0575] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or other standards.
[0576] Figure 5 is a block diagram illustrating an example of a video encoder 200, which may be Figure 4 The video encoder 114 in the system 100 is shown in FIG.
[0577] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 5In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0578] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0579] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.
[0580] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated but are not shown for purposes of explanation. Figure 5 In the example, they are represented separately.
[0581] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0582] The mode selection unit 203 may, for example, select one of the intra or inter coding modes based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select a resolution (e.g., sub-pixel or integer pixel precision) for the motion vector of the block in the case of inter prediction.
[0583] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.
[0584] Motion estimation unit 204 and motion compensation unit 205 may perform different operations for the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0585] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures of list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0586] In other examples, motion estimation unit 204 may perform bidirectional prediction on the current video block. Motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. Motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0587] In some examples, motion estimation unit 204 may output the full set of motion information for use in the decoding process of the decoder.
[0588] In some examples, motion estimation unit 204 may not output a full set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information for the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.
[0589] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0590] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0591] As discussed above, video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0592] Intra-prediction unit 206 may perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, intra-prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0593] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0594] In other examples, such as in skip mode, there may be no residual data for the current video block, and residual generation unit 207 may not perform a subtraction operation.
[0595] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0596] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0597] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0598] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0599] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0600] Figure 6 is a block diagram illustrating an example of a video decoder 300, which may be Figure 4 The video decoder 114 in the system 100 is shown in FIG.
[0601] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 6 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0602] exist Figure 6 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform operations generally similar to those described with respect to the video encoder 200 ( Figure 5 ) is a decoding round that is the inverse of the encoding described.
[0603] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 may decode the entropy-encoded video data, and the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information from the entropy-decoded video data. For example, the motion compensation unit 302 may determine such information by performing AMVP and merge mode.
[0604] Motion compensation unit 302 may generate motion compensated blocks, possibly performing interpolation based on interpolation filters.Identifiers of interpolation filters to be used with sub-pixel precision may be included in the syntax elements.
[0605] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used during encoding of the video block by video encoder 200. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 from received syntax information and use the interpolation filters to produce a predictive block.
[0606] The motion compensation unit 302 may use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.
[0607] The intra-frame prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra-frame prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0608] The reconstruction unit 306 may sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0609] A list of some preferred examples of embodiments is provided below.
[0610] The first set of clauses illustrates exemplary embodiments of the techniques discussed in the previous section.The following clauses illustrate exemplary embodiments of the techniques discussed in the previous section (eg, item 1).
[0611] 1. A video processing method (e.g., Figure 3 ), comprising: performing (3002) a conversion between a video comprising one or more output layer sets (OLS) and a codec representation of the video, the OLS comprising one or more video pictures, wherein the one or more video pictures are coded as intra random access point pictures or intra codec pictures in the codec representation, wherein the codec representation conforms to a format rule that specifies the location and type of information included in the codec representation for decoding the one or more video pictures.
[0612] 2. A method according to clause 1, wherein the format rule provides that: in the case where all pictures of one or more pictures are included as intra-frame random access pictures in the codec representation, level information corresponding to each OLS and / or an indication of the presence of level information is included in the first set of syntax elements of the codec representation.
[0613] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 2).
[0614] 3. A method according to any of clauses 1-2, wherein the format rule provides that: when all of one or more pictures are included in the codec representation using only intra-frame coding, level information corresponding to each OLS and / or an indication of the presence of level information is included in the second set of syntax elements of the codec representation.
[0615] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 3).
[0616] 4. The method according to any of clauses 2-3, wherein the format rule specifies that the first set of syntax elements and the second set of syntax elements are included in a codec representation.
[0617] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 4).
[0618] 5. The method of clause 4, wherein the first set of syntax elements and / or the second set of syntax elements are included in a video parameter set or a sequence parameter set.
[0619] 6. The method of clause 5, wherein the first set of syntax elements and / or the second set of syntax elements are included in a decoding capability information (DCI) syntax structure.
[0620] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 5).
[0621] 7. The method of clause 5, wherein the first set of syntax elements and / or the second set of syntax elements are included in a profile, level and level (PTL) syntax structure.
[0622] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 6).
[0623] 8. A method according to any of clauses 2-7, wherein the format rule further specifies whether an indication of the presence of the level information is included in the codec representation controls whether the level information is included.
[0624] 9. A method according to clause 1, wherein the format rule specifies that the indication of the presence of level information in the first set of syntax elements and the second set of syntax elements corresponds to two different fields.
[0625] The following items illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 7).
[0626] 10. The method according to any of clauses 2-9, wherein the format rule further provides for including hypothetical reference decoder parameters having the first set of syntax elements and / or the second set of syntax elements.
[0627] 11. A video processing method, comprising: performing conversion between a video comprising one or more pictures and a codec representation, the one or more pictures comprising one or more slices encoded and decoded into one or more time domain video layers in the codec representation; wherein the codec representation complies with a format rule that specifies constraints on the signaling notification of one or more inter-layer prediction information syntax elements, without having to signal a two-dimensional syntax element indicating a maximum time domain layer id of a reference picture used to encode and decode a current video layer, and without having to directly signal the layer id of the maximum time domain layer id of the reference picture used to encode and decode the current video layer.
[0628] 12. A method according to clause 11, wherein the format rule specifies that the reference picture referenced by each intra-layer residual prediction in RefPicList[0] or RefPicList[1] of the slice of the current picture is an intra-frame random access point picture or has a TaemporalId less than max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx], where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId] respectively, and refpicLayerId is the nuh_layer_id of the reference picture.
[0629] 13. A method according to clause 11, wherein the format rule specifies that: the picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB, shall have a nuh_layer_id refPicLayerId less than the nuh_layer_id of the current picture, and shall be an IRAP picture or have a TemporalId less than max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx], where CurrLayerIdx and RefLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[RefPicLayerId], respectively.
[0630] 14. A method according to clause 11, wherein the format rule specifies that the reference picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be an IRAP picture or have a TemporalId less than or equal to Max Max(0,max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where currLayerIdx and refLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId] respectively, and refpicLayerId is the nuh_layer_id of the reference picture.
[0631] 15. A method according to clause 11, wherein the format rule specifies that a picture referenced by each ILRP in RefPicList[0] or RefPicList[1] of the slice of the current picture shall be present in the DPB, shall have a nuh_layer_id refPicLayerId less than the nuh_layer_id of the current picture, and shall be an IRAP picture or have a TemporalId less than or equal to Max(0, max_tid_il_ref_pics_plus1[currLayerIdx][refLayerIdx]-1), where CurrLayerIdx and RefLayerIdx are equal to GeneralLayerIdx[nuh_layer_id] and GeneralLayerIdx[refpicLayerId], respectively.
[0632] 16. A video processing method, the method comprising: performing a conversion between a video comprising one or more pictures and a codec representation of the video, wherein each picture is coded as an intra-frame random access picture according to a rule; wherein the rule stipulates that the operation of an imaginary reference decoder for the codec representation uses a maximum sub-layer value equal to zero.
[0633] 17. A method according to clause 16, wherein the acceleration factor is signalled in a Supplemental Enhancement Information (SEI) message.
[0634] 18. A method according to any of clauses 1 to 17, wherein the converting comprises encoding the video into the codec representation.
[0635] 19. A method according to any of clauses 1 to 17, wherein the converting comprises decoding the codec representation to generate pixel values of the video.
[0636] 20. A video decoding apparatus comprising a processor configured to implement the method of one or more of clauses 1 to 19.
[0637] 21. A video encoding apparatus comprising a processor configured to implement the method of one or more of clauses 1 to 19.
[0638] 22. A computer program product having computer code stored thereon which, when executed by a processor, causes the processor to carry out the method according to any one of clauses 1 to 19.
[0639] 23. A method, apparatus or system as described in this document.
[0640] The second set of items illustrates exemplary embodiments of the techniques discussed in the previous section (eg, items 1-7, 15, and 16).
[0641] 1. A method for video processing (e.g., Figure 7A The method 710 shown comprises: performing (712) conversion between a video and a bitstream of the video comprising one or more output layer sets according to format rules, and wherein at least one of the one or more output layer sets consists of a trick mode access representation comprising only intra random access point pictures or only intra codec pictures, and wherein the format rules specify whether or how horizontal information of the trick mode representation is indicated in the bitstream.
[0642] 2. A method according to clause 1, wherein the format rule specifies that in the case where at least one of the one or more output layer sets has the track mode access representation comprising only intra random access point pictures, the level information of the track mode access representation and / or an indication of the presence of level information is indicated by a first group of syntax elements in the syntax structure.
[0643] 3. A method according to clause 1, wherein the format rule specifies that in the case where at least one of the one or more output layer sets has the track mode access representation comprising only intra-frame coded pictures, the level information of the track mode access representation and / or an indication of the presence of level information is indicated by a second set of syntax elements in the syntax structure.
[0644] 4. A method according to clause 2 or 3, wherein the syntax structure is included in the bitstream.
[0645] 5. A method according to clause 2 or 3, wherein each output layer set comprises one or more non-video codec layer (VCL) network abstraction layer (NAL) units associated with the one or more pictures.
[0646] 6. A method according to any of clauses 2-5, wherein the format rule specifies including the first set of syntax elements and the second set of syntax elements in the syntax structure.
[0647] 7. A method according to any of clauses 2-6, wherein the syntax structure is a video parameter set or a sequence parameter set.
[0648] 8. A method according to any of clauses 2-6, wherein the syntax structure is a decoding capability information (DCI) syntax structure.
[0649] 9. A method according to any of clauses 2-6, wherein the syntax structure is a Supplemental Enhancement Information (SEI) message.
[0650] 10. A method according to any of clauses 2-6, wherein the grammatical structure is a profile, level and level (PTL) grammatical structure.
[0651] 11. A method according to any of clauses 2-6, wherein the format rules further specify that whether the horizontal information is included in the syntax structure depends on an indication of the presence of horizontal information of one or more track mode access representations.
[0652] 12. A method according to clause 11, wherein the format rule specifies that the indication of the presence of level information of the one or more track mode access indications in the first set of syntax elements and the second set of syntax elements corresponds to two different fields.
[0653] 13. A method according to clause 11 or 12, wherein horizontal information for a track mode access representation including only intra-coded pictures and horizontal information for a track mode access representation including only intra random access point pictures are signaled separately.
[0654] 14. A method according to clause 11, wherein the horizontal information of the one or more track mode access representations is present in a profile, level and level (PTL) syntax structure, in the case where the indication of the presence of the horizontal information of the one or more track mode access representations specifies the presence of the horizontal information.
[0655] 15. A method according to clause 11, wherein the horizontal information of the one or more track mode access representations is not present in a profile, level and level (PTL) syntax structure in a case where the indication of the presence of the horizontal information of the one or more track mode access representations specifies that the horizontal information is not present.
[0656] 16. A method according to clause 11, wherein the format rule specifies that the indication of the presence of level information in the first set of syntax elements and the second set of syntax elements corresponds to one field.
[0657] 17. A method according to clause 16, wherein the level information is signaled once for a track mode access representation comprising only intra random access point pictures and a track mode access representation comprising only intra codec pictures.
[0658] 18. A method according to any of clauses 2-5, wherein the format rule further specifies including hypothetical reference decoder parameters with the first set of syntax elements and / or the second set of syntax elements.
[0659] 19. A method for video processing (e.g., Figure 7B The method 720 shown comprises: performing a conversion between a video comprising one or more pictures and a bitstream of the video, and wherein, according to a format rule, the bitstream comprises a trick mode access representation of the one or more output layer sets, wherein the format rule specifies that the trick mode access representation only includes intra random access point pictures, and wherein the format rule specifies whether or how to operate a hypothetical reference decoder.
[0660] 20. A method according to clause 19, wherein the format rule specifies that operation of the hypothetical reference decoder for the bitstream uses a value of a maximum sub-layer equal to zero.
[0661] 21. A method according to clause 19, wherein the format rules provide for signalling information about the acceleration factor in a Supplemental Enhancement Information (SEI) message.
[0662] 22. The method of clause 21, wherein the SEI message includes syntax elements for deriving the value of the acceleration factor.
[0663] 23. A method according to clause 22, wherein the syntax element corresponds to irap_only_speedup_minusM, and the value of the acceleration factor is derived as (irap_only_speedup_minusM+M) / M, where M is an integer.
[0664] 24. The method of clause 23, wherein (irap_only_speedup_minusM+M) is not less than 2 K -1, where K is an integer.
[0665] 25. A method according to clause 23, wherein the SEI message includes a syntax element that specifies a maximum value of the acceleration factor to be applied to a hypothetical reference decoder timing of a sub-bitstream, the sub-bitstream corresponding to an intra-frame random access point access unit sequence in a set of video sequences encoded or decoded for one or more output layer sets of the bitstream to which the SEI message is applied.
[0666] 26. A method according to clause 25, wherein the syntax element is signaled using u(k), where K is an integer.
[0667] 27. A method according to clause 25, wherein the value of the syntax element is not less than 100.
[0668] 28. A method according to clause 21, wherein a syntax element for specifying a maximum value of an acceleration factor to be applied to a hypothetical reference decoder timing of a sub-bitstream corresponding to a sequence of intra random access point access units in a set of coded video sequences of one or more output layer sets of the bitstream to which the SEI message is applied is not present in the bitstream, and the value of the syntax element is inferred to be equal to 100.
[0669] 29. The method of clause 21, wherein the value of the acceleration factor is derived to be equal to (irap_only_speedup_minusM+offset) / M, where offset and M are integers and irap_only_speedup_minusM is a syntax element having a value for deriving the acceleration factor.
[0670] 30. The method of clause 29, wherein offset is set to 0 or M / 2.
[0671] 31. A method according to any one of clauses 1 to 30, wherein the converting comprises encoding the video into the bitstream.
[0672] 32. A method according to any one of clauses 1 to 30, wherein the converting comprises decoding the video from the bitstream.
[0673] 33. The method of clauses 1 to 30, wherein the converting comprises generating the bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0674] 34. A video processing apparatus comprising a processor configured to implement the method of any one or more of clauses 1 to 33.
[0675] 35. A method of storing a bitstream of video, comprising the method of any one of clauses 1 to 34, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0676] 36. A computer-readable medium storing program code which, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 34.
[0677] 37. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0678] 38. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 34.
[0679] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are co-located or interspersed at different locations within the bitstream as defined by the syntax. For example, a macroblock may be encoded based on a transformed and encoded error residual value, and may also be encoded using bits in a header and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on the knowledge that some fields may or may not be present, as described in the solution above. Similarly, the encoder may determine whether certain syntax fields will be included, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0680] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of materials that implement a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0681] A computer program (also referred to as a program, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that stores other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer, or to execute on multiple computers that are located in one location or distributed in multiple locations and interconnected by a communication network.
[0682] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0683] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices for storing data to receive data from or transfer data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0684] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or the scope that may be claimed, but rather should be interpreted as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be removed from the combination, and a claimed combination may be directed to a subcombination or variations of a subcombination.
[0685] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0686] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: performing conversion between a video and a bitstream of said video comprising one or more sets of output layers according to format rules, and wherein at least one of the one or more output layer sets comprises a trick mode representation, the trick mode representation comprising an IRAP (Intra Random Access Point)-only representation or an intra-only representation, wherein the IRAP-only representation comprises only an IRAP picture and one or more non-VCL (Video Codec Layer) NAL (Network Abstraction Layer) units associated with the IRAP picture, and wherein the intra-only representation comprises only intra-codec pictures and one or more non-VCL NAL units associated with the intra-codec pictures, and The format rules specify whether or how to indicate level information of the trick mode representation in a bitstream.
2. The method according to claim 1, wherein The format rules specify that level information of the trick mode representation and / or an indication of the presence of level information is indicated by one or more syntax elements in a syntax structure included in the bitstream.
3. The method according to claim 2, wherein: The syntax structure is a video parameter set, a sequence parameter set, a decoding capability information (DCI) syntax structure, a supplemental enhancement information (SEI) message, or a profile, level and level (PTL) syntax structure.
4. The method according to claim 2, wherein: The format rule further provides for including hypothetical reference decoder parameters with the one or more syntax elements.
5. The method according to claim 2, wherein: The format rules further specify that whether the level information is included in the syntax structure depends on an indication of the presence of level information for the trick mode representation.
6. The method according to claim 2, wherein: The one or more syntax elements include sublayer_level_present_flag[i], Where sublayer_level_present_flag[i] is equal to 1 specifies: When i is greater than 0, the level information is present in the profile_tier_level() syntax structure of the sub-layer representation with TemporalId equal to i-1, and When i is equal to 0, the level information is present in the profile_tier_level() syntax structure of the IRAP-only representation.
7. The method according to claim 2, wherein: The one or more syntax elements include sublayer_level_present_flag[i], Where sublayer_level_present_flag[i] is equal to 0, it specifies: When i is greater than 0, the level information is not present in the profile_tier_level() syntax structure of the sub-layer representation with TemporalId equal to i-1, and When i is equal to 0, level information is not present in the profile_tier_level() syntax structure of the IRAP-only representation.
8. The method according to claim 2, wherein: The one or more output layer sets include a trick mode representation comprising an IRAP-only representation and a trick mode representation comprising an intra-only representation, and The format rule specifies that for a trick mode representation including an IRAP-only representation and for a trick mode representation including an intra-only representation, the presence indication of the level information corresponds to one field.
9. The method according to claim 8, wherein The one or more syntax elements include sublayer_level_present_flag[i], Where sublayer_level_present_flag[i] is equal to 1 specifies: When i is greater than 1, the level information is present in the profile_tier_level() syntax structure of the sub-layer representation with TemporalId equal to i-2. When i is equal to 1, the level information is present in the profile_tier_level() syntax structure of the intra-only representation, and When i is equal to 0, the level information is present in the profile_tier_level() syntax structure of the IRAP-only representation.
10. The method of claim 8, wherein: The one or more syntax elements include sublayer_level_present_flag[i], Where sublayer_level_present_flag[i] is equal to 0, it specifies: When i is greater than 1, the level information is not present in the profile_tier_level() syntax structure of the sub-layer representation with TemporalId equal to i-2. When i is equal to 1, the level information is not present in the profile_tier_level() syntax structure of the intra-only representation, and When i is equal to 0, level information is not present in the profile_tier_level() syntax structure of the IRAP-only representation.
11. The method of claim 2, wherein the one or more syntax elements include sublayer_level_idc[i], Wherein when i is greater than 0, sublayer_level_idc[i] applies to the sublayer representation with TemporalId equal to i-1, and when i is equal to 0, sublayer_level_idc[i] applies to the IRAP-only representation.
12. The method of claim 1, wherein the one or more output layer sets include a trick mode representation comprising an IRAP-only representation and a trick mode representation comprising an intra-only representation, and Wherein, for the trick mode representation including the IRAP-only representation and the trick mode representation including the intra-only representation, the level information of the two is included separately.
13. The method of claim 1 , wherein the one or more output layer sets include a trick mode representation comprising an IRAP-only representation and a trick mode representation comprising an intra-only representation, and For the trick mode representation including IRAP-only representation and the trick mode representation including intra-only representation, the level information of both is included in the bitstream at once.
14. The method according to claim 1, wherein The converting includes encoding the video into the bitstream.
15. The method according to claim 1, wherein The converting includes decoding the video from the bitstream.
16. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: performing conversion between a video and a bitstream of said video comprising one or more sets of output layers according to format rules, and wherein at least one of the one or more output layer sets comprises a trick mode representation, the trick mode representation comprising an IRAP (Intra Random Access Point)-only representation or an intra-only representation, wherein the IRAP-only representation comprises only an IRAP picture and one or more non-VCL (Video Codec Layer) NAL (Network Abstraction Layer) units associated with the IRAP picture, and wherein the intra-only representation comprises only intra-codec pictures and one or more non-VCL NAL units associated with the intra-codec pictures, and in, The format rules specify whether or how the level information of the trick mode representation is indicated in the bitstream.
17. The device according to claim 16, wherein The format rules specify that level information of the trick mode representation and / or an indication of the presence of level information is indicated by one or more syntax elements in a syntax structure included in the bitstream, and The syntax structure is a video parameter set, a sequence parameter set, a decoding capability information (DCI) syntax structure, a supplemental enhancement information (SEI) message, or a profile, level and level (PTL) syntax structure.
18. A non-transitory computer-readable storage medium storing instructions that cause a processor to: performing conversion between a video and a bitstream of said video comprising one or more sets of output layers according to format rules, and wherein at least one of the one or more output layer sets comprises a trick mode representation, the trick mode representation comprising an IRAP (Intra Random Access Point)-only representation or an intra-only representation, wherein the IRAP-only representation comprises only an IRAP picture and one or more non-VCL (Video Codec Layer) NAL (Network Abstraction Layer) units associated with the IRAP picture, and wherein the intra-only representation comprises only intra-codec pictures and one or more non-VCL NAL units associated with the intra-codec pictures, and in, The format rules specify whether or how the level information of the trick mode representation is indicated in the bitstream.
19. The non-transitory computer-readable storage medium of claim 18, wherein: The format rules specify that level information of the trick mode representation and / or an indication of the presence of level information is indicated by one or more syntax elements in a syntax structure included in the bitstream, and The syntax structure is a video parameter set, a sequence parameter set, a decoding capability information (DCI) syntax structure, a supplemental enhancement information (SEI) message, or a profile, level and level (PTL) syntax structure.
20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: generating the bitstream for the video according to a format rule, The bitstream includes one or more output layer sets, wherein at least one of the one or more output layer sets comprises a trick mode representation, the trick mode representation comprising an IRAP (Intra Random Access Point)-only representation or an intra-only representation, wherein the IRAP-only representation comprises only an IRAP picture and one or more non-VCL (Video Codec Layer) NAL (Network Abstraction Layer) units associated with the IRAP picture, and wherein the intra-only representation comprises only intra-codec pictures and one or more non-VCL NAL units associated with the intra-codec pictures, and The format rules specify whether or how to indicate level information of the trick mode representation in a bitstream.
21. A method for storing a video bitstream, comprising: generating a bitstream for the video according to a format rule, the bitstream comprising one or more sets of output layers; as well as storing the bitstream in a non-transitory computer-readable storage medium, wherein at least one of the one or more output layer sets comprises a trick mode representation, the trick mode representation comprising an IRAP (Intra Random Access Point)-only representation or an intra-only representation, wherein the IRAP-only representation comprises only an IRAP picture and one or more non-VCL (Video Codec Layer) NAL (Network Abstraction Layer) units associated with the IRAP picture, and wherein the intra-only representation comprises only intra-codec pictures and one or more non-VCL NAL units associated with the intra-codec pictures, and The format rules specify whether or how to indicate level information of the trick mode representation in a bitstream.
22. The method according to claim 1, wherein said bitstream comprising a trick mode access representation of said one or more output layer sets, wherein the format rule further stipulates that the trick mode access representation only includes the IRAP picture, and The format rules specify whether or how to operate the hypothetical reference decoder.
23. The method according to claim 22, wherein The format rules further specify that the operation of the hypothetical reference decoder for the bitstream uses a value of the largest sub-layer equal to zero.
24. The method according to claim 22, wherein The format rules also stipulate that information about the acceleration factor is signaled in the supplemental enhancement information SEI message.
25. The method according to claim 24, wherein The SEI message includes syntax elements for deriving a value of the acceleration factor.
26. The method according to claim 25, wherein The syntax element corresponds to irap_only_speedup_minusM, and the value of the acceleration factor is derived as (irap_only_speedup_minusM+M) / M, where M is an integer.
27. The method according to claim 26, wherein (irap_only_speedup_minusM+M) is not less than 2 K -1, where K is an integer.
28. The method according to claim 26, wherein The SEI message includes a syntax element that specifies the maximum value of the acceleration factor to be applied to the hypothetical reference decoder timing of the sub-bitstream, which corresponds to the intra-frame random access point access unit sequence in the video sequence set of the codec of one or more output layer sets of the bitstream to which the SEI message is applied.
29. The method according to claim 28, wherein The syntax element is signaled using u(k), where k is an integer.
30. The method of claim 28, wherein The value of the syntax element is not less than 100.
31. The method of claim 24, wherein: The syntax element for specifying the maximum value of the acceleration factor to be applied to the hypothetical reference decoder timing of the sub-bitstream corresponding to the intra random access point access unit sequence in the set of coded video sequences of one or more output layer sets of the bitstream to which the SEI message is applied is not present in the bitstream, and the value of the syntax element is inferred to be equal to 100.
32. The method of claim 24, wherein: The value of the acceleration factor is derived to be equal to (irap_only_speedup_minusM+offset) / M, where offset and M are integers, and irap_only_speedup_minusM is a syntax element having a value for deriving the acceleration factor.
33. The method according to claim 32, wherein The offset is set to 0 or M / 2.
34. The method according to any one of claims 22 to 33, wherein The converting includes encoding the video into the bitstream.
35. The method according to any one of claims 22 to 33, wherein The converting includes decoding the video from the bitstream.
36. The method according to any one of claims 22 to 33, wherein The converting includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
37. A video processing device comprising a processor configured to implement the method of any one of claims 22 to 36.
38. A method of storing a bitstream of a video, comprising the method according to any one of claims 22 to 36, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
39. A computer readable medium storing program code which, when executed, causes a processor to implement the method according to any one of claims 22 to 36.
40. A computer-readable medium storing a bitstream generated by the method according to any one of claims 22 to 36 executed by a video processing device.
Citation Information
Patent Citations
Signaling video samples for trick mode video representations
CN103081488A
Signaling information for coding
CN105556975A