Decoding Picture Buffer Memory Allocation and Picture Output in Scalable Video Coding
By introducing format rules to manage the decoder buffer in video encoding and decoding, the problem of unreasonable management of decoding image buffers in multi-layer videos is solved, the decoding efficiency is improved and the delay is reduced, and it is suitable for multi-layer video encoding and decoding standards such as VVC.
Patent Information
- Application Number
- CN202180024092.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-17
- Filing Date
- 2021-03-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-03-16
AI Technical Summary
When existing video encoding and decoding technologies are processed with multi-layer video, it is difficult to effectively manage the decoding image buffer, resulting in unreasonable resource allocation and affecting decoding efficiency and delay.
By introducing format rules, the image output in the decoder buffer includes maximum image width and height, bit depth, chroma format, decoding order and buffer management, ensuring that the image is removed from the buffer when appropriate, and supporting multi-layer video encoding and decoding standards such as VVC.
Improves video decoding efficiency, reduces end-to-end latency, adapts to the needs of low-latency applications such as wireless displays and online gaming, and simplifies the decoder design of multi-layer bitstreams.
Smart Images

Figure CN115315941B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is being filed to claim the priority and benefit of U.S. Provisional Patent Application No. 62 / 990,749, filed on March 17, 2020, under the patent laws and / or rules applicable under the Paris Convention. For all purposes under the law, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This patent document relates to image and video encoding and decoding. Background art
[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. With the increasing number of connected user devices capable of receiving and displaying video, the bandwidth demand for digital video use is expected to continue to grow. Summary of the invention
[0005] This document discloses techniques that can be used by video encoders and decoders for processing the encoded - decoded representation of video using control information useful for decoding the encoded - decoded representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and an encoded - decoded representation of the video; wherein the encoded - decoded representation includes a video parameter set that indicates a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels representing the video.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and an encoded - decoded representation of the video, wherein the encoded - decoded representation complies with format rules that specify a maximum picture width and / or a maximum picture height of video pictures of all video layers and a value of a variable that controls whether a picture in a decoder buffer is output before being removed from the decoder buffer.
[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and an encoded - decoded representation of the video, wherein the encoded - decoded representation complies with format rules that specify a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels representing the video and a value of a variable that controls whether a picture in a decoder buffer is output before being removed from the decoder buffer.
[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to format rules that specify that the value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is independent of whether separate color planes are used for encoding the video.
[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to format rules that specify that the value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is included in the coded representation at an access unit (AU) level.
[0011] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to format rules that specify that the picture output flag of a video picture in an access unit is determined based on the pic_output_flag variable of another video picture in the access unit.
[0012] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to format rules that specify the value of the picture output flag for video pictures that do not belong to an output layer.
[0013] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to format rules, where the bitstream includes one or more output layer sets (OLSs), each OLS including one or more coded layer video sequences, and where the format rules specify that a video parameter set indicates, for each of the one or more OLSs, a maximum allowed value of a chroma format indicator for pixels representing the video and / or a maximum allowed value of a bit depth.
[0014] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and where the format rules specify that the maximum picture width and / or maximum picture height of video pictures of all video layers controls the value of a variable indicating whether a picture in a decoded picture buffer that is decoded before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer.
[0015] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify a maximum allowable value of a chroma format indicator for representing pixels of the video and / or a maximum allowable value of a bit depth, and a value of a variable that controls whether a picture in a decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer.
[0016] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to rules, and wherein the rules specify that a value of a variable that controls whether a picture in a decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is independent of whether separate color planes are used to encode the video.
[0017] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer when decoding a certain type of access unit is included in the bitstream.
[0018] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a first flag that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer when decoding a specific type of access unit is not indicated in a picture header.
[0019] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag associated with an access unit that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer depends on a value of a flag of each picture of the access unit.
[0020] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a variable that indicates whether to output a picture in an access unit is determined based on a flag that indicates whether to output another picture in the access unit.
[0021] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when a picture in an access unit does not belong to an output layer, the value of a variable indicating whether to output the picture is set to be equal to a certain value.
[0022] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when the video includes only one output layer, access units that do not include the output layer are encoded and decoded by setting a variable indicating whether to output a picture in the access unit to a first value for a picture having the highest layer ID (identification) value and a second value for all other pictures.
[0023] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0024] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0025] In yet another example aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0026] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a block diagram of an example video processing system.
[0028] Figure 2 is a block diagram of a video processing device.
[0029] Figure 3 is a flowchart of an example method of video processing.
[0030] Figure 4 is a block diagram illustrating a video codec system according to some embodiments of the present disclosure.
[0031] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0032] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0033] Figures 7A to 7D shows a flowchart of an example method of video processing based on some embodiments of the disclosed technology.
[0034] Figures 8A to 8C A flowchart showing an example method of video processing based on some embodiments of the disclosed technology.
[0035] Figures 9A to 9C A flowchart showing an example method of video processing based on some embodiments of the disclosed technology. Detailed Description
[0036] The section headings used in this document are for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each section to that section only. Additionally, the use of H.266 terms in some descriptions is for ease of understanding only and does not limit the scope of the disclosed technology. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.
[0037] 1. Preliminary Discussion
[0038] This patent document relates to video codec technology. Specifically, it is about the signaling of decoding picture buffer (DPB) parameters for DPB memory allocation and the specification of the output of decoded pictures in scalable video coding, where the video bitstream can contain more than one layer. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Coding (VVC) being developed.
[0039] 2. Abbreviations
[0040] APS Adaptive Parameter Set
[0041] AU Access Unit
[0042] AUD Access Unit Delimiter
[0043] AVC Advanced Video Coding
[0044] CLVS Coded Layer Video Sequence
[0045] CPB Coded Picture Buffer
[0046] CRA Completely Random Access
[0047] CTU Coding Tree Unit
[0048] CVS Coded Video Sequence
[0049] DCI Decoding Capability Information
[0050] DPB Decoding Picture Buffer
[0051] EOB End of Bitstream
[0052] EOS Sequence End
[0053] GDR Gradual Decoding Refresh
[0054] HEVC High Efficiency Video Coding
[0055] HRD Hypothetical Reference Decoder
[0056] IDR Instantaneous Decoding Refresh
[0057] JEM Joint Exploration Model
[0058] MCTS Motion Constrained Tile Set
[0059] NAL Network Abstraction Layer
[0060] OLS Output Layer Set
[0061] PH Picture Header
[0062] PPS Picture Parameter Set
[0063] PTL Profile, Tier and Level
[0064] PU Picture Unit
[0065] RAP Random Access Point
[0066] RBSP Raw Byte Sequence Payload
[0067] SEI Supplemental Enhancement Information
[0068] SPS Sequence Parameter Set
[0069] SVC Scalable Video Coding
[0070] VCL Video Coding Layer
[0071] VPS Video Parameter Set
[0072] VTM VVC Test Model
[0073] VUI Video Usability Information
[0074] VVC Versatile Video Coding
[0075] 3. Introduction to Video Coding
[0076] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, which employs temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts in VVC standardization, new coding technologies have been incorporated into the VVC standard at each JVET meeting. The working draft and test model VTM of VVC are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.
[0077] 3.1. Overview and Scalable Video Coding (SVC) in VVC
[0078] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding that uses a Base Layer (BL) (sometimes referred to as a Reference Layer (RL)) and one or more Scalable Enhancement Layers (ELs). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can act as the BL, while the top layer can act as the EL. Intermediate layers can act as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below it (such as the base layer or any intervening enhancement layer) and at the same time act as an RL for one or more enhancement layers above it. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to code (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).
[0079] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be utilized by one or more coded video sequences of different layers in a bitstream can be included in a Video Parameter Set (VPS), and the parameters that can be utilized by one or more pictures in a coded video sequence can be included in a Sequence Parameter Set (SPS). Similarly, the parameters utilized by one or more slices in a picture can be included in a Picture Parameter Set (PPS), and other parameters specific to an individual slice can be included in the slice header. Similarly, an indication of which (which) parameter set a particular layer uses at a given time can be provided at various coding levels.
[0080] Due to the support for Reference Picture Resampling (RPR) in VVC, it is possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without any additional signal processing level codec tools, because the upsampling required for spatial scalability support can be achieved using only the RPR upsampling filter. However, for scalability support, high-level syntax changes are required (compared to non-scalability support). Scalability support was specified in VVC version 1. Different from the scalability support in any earlier video codec standards (including extensions of AVC and HEVC), the design of VVC scalability has been made as friendly as possible to single-layer decoder design. The decoding capabilities of multi-layer bitstreams are specified in a way that is as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified independently of the number of layers in the bitstream to be decoded. Basically, a decoder designed for single-layer bitstreams does not need much modification to be able to decode multi-layer bitstreams. Compared to the design of multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, it is required that an IRAP AU contains pictures of each layer present in the CVS.
[0081] 3.2. Random Access and Its Support in HEVC and VVC
[0082] Random access means accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, search in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include randomly accessible points that are close together, which are usually intra-coded pictures but can also be inter-coded pictures (e.g., in the case of progressive decoding refresh).
[0083] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the NAL unit header by NAL unit type. Three types of IRAP pictures are supported, namely Instantaneous Decoder Refresh (IDR), Complete Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any pictures before the current Group of Pictures (GOP), which is traditionally referred to as a closed GOP random access point. By allowing specific pictures to reference pictures before the current GOP, CRA pictures are less restrictive, where in the case of random access, all pictures are discarded. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures typically result from the concatenation of two bitstreams or parts thereof at a CRA picture, for example during stream switching. To better enable the system to use IRAP pictures, a total of six different NAL units are defined to signal the attributes of IRAP pictures, which can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF), which is used for random access support in HTTP-based Dynamic Adaptive Streaming over HTTP (DASH).
[0084] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or the other type without associated RADL pictures), and one type of CRA picture. These are basically the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic function of BLA pictures can be achieved by CRA pictures plus the end of the sequence NAL unit, the presence of which indicates that the subsequent pictures start a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, as indicated by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.
[0085] Another key difference between VVC and HEVC in terms of random access support is that GDR is supported in a more standardized way in VVC. In GDR, the decoding of the bitstream can start from inter-coded pictures. Although not the entire picture area can be correctly decoded at the beginning, after several pictures, the entire picture area will be correct. AVC and HEVC also support GDR by using recovery point SEI messages to signal GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. It is allowed for the CVS and the bitstream to start with GDR pictures. This means that it is allowed for the entire bitstream to contain only inter-coded pictures without a single intra-coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior of GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded stripes or blocks over multiple pictures, as opposed to intra-coding the entire picture, thus allowing a significant reduction in end-to-end latency, which is considered more important today as ultra-low latency applications such as wireless display, online gaming, and drone-based applications become more popular.
[0086] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., the correctly decoded area) and the non-refreshed area at the pictures between a GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so there will be no decoding mismatch for some samples at or near the boundary. This can be useful when the application determines to display the correctly decoded area during the GDR process.
[0087] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.
[0088] 3.3. Parameter Sets
[0089] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced starting from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.
[0090] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry picture-level header information that does not change frequently. With the SPS and PPS, the information that does not change frequently does not need to be repeated for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of the SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.
[0091] The VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0092] The APS is introduced to carry picture-level or slice-level information that requires a significant number of bits to encode and decode, can be shared by multiple pictures, and can have a significant number of different variations in the sequence.
[0093] 3.4. Related Definitions in VVC
[0094] The related definitions in the latest VVC text (in JVET-Q2001-vE / v15) are as follows.
[0095] Associated IRAP picture (of a specific picture): The previous IRAP picture in decoding order (when present) has the same nuh_layer_id value as the specific picture.
[0096] Coded Video Sequence (CVS): A sequence of AUs consisting of CVSS AUs in decoding order, followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU.
[0097] Coded Video Sequence Start (CVSS) AU: An AU that has PUs for each layer in the CVS and the coded pictures in each PU are CLVSS pictures.
[0098] Gradual Decoding Refresh (GDR) AU: An AU in which the coded pictures in each current PU are GDR pictures.
[0099] Gradual Decoding Refresh (GDR) PU: A PU in which the coded pictures are GDR pictures.
[0100] Gradual Decoding Refresh (GDR) picture: A picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT.
[0101] Intra Random Access Point (IRAP) AU: An AU that has PUs for each layer in the CVS and the coded pictures in each PU are IRAP pictures.
[0102] Intra Random Access Point (IRAP) picture: A coded picture in which all VCL NAL units have the same nal_unit_type value in the range from IDR_W_RADL to CRA_NUT, inclusive of IDR_W_RADL and CRA_NUT.
[0103] Previous picture: A picture in the same layer as the associated IRAP picture and preceding the associated IRAP picture in output order.
[0104] Trailing picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not a STSA picture.
[0105] Note – Trailing pictures associated with an IRAP picture are also after the IRAP picture in decoding order. A picture that follows the associated IRAP picture in output order and precedes the associated IRAP picture in decoding order is not allowed.
[0106] 3.5. VPS Syntax and Semantics in VVC
[0107] VVC supports scalability, also known as Scalable Video Coding, where multiple layers can be encoded in a single coded video bitstream.
[0108] In the latest VVC text (in JVET-Q2001-vE / v15), scalability information is signaled in the VPS, and its syntax and semantics are as follows.
[0109] 7.3.2.2 Video Parameter Set Syntax
[0110]
[0111]
[0112]
[0113]
[0114] 7.4.3.2 Video Parameter Set RBSP Semantics
[0115] The VPS RBSP shall be available for the decoding process before being referenced, including in at least one AU with a TemporalId equal to 0, or provided externally.
[0116] All VPS NAL units with a specific vps_video_parameter_set_id value in a CVS shall have the same content.
[0117] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.
[0118] vps_max_layers_minus1 plus 1 specifies the maximum number of allowed layers in each CVS that refers to the VPS.
[0119] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in the layers of each CVS that refers to the VPS. The value of vps_max_sublayers_minus1 shall be in the range of 0 to 6 (including the end values).
[0120] vps_all_layers_same_num_sublayers_flag being equal to 1 specifies that the number of temporal sublayers is the same for all layers in each CVS that refers to the VPS. vps_all_layers_same_num_sub layers_flag being equal to 0 specifies that the layers in each CVS that refers to the VPS may or may not have the same number of temporal sublayers. When not present, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.
[0121] vps_all_independent_layers_flag being equal to 1 specifies that all layers in the CVS are independently encoded and decoded without using inter-layer prediction. vps_all_independent_layers_flag being equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. When not present, the value of vps_all_independent_layers_flag is inferred to be equal to 1.
[0122] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] shall be less than the value of vps_layer_id[n].
[0123] vps_independent_layer_flag[i] being equal to 1 specifies that the layer with index i does not use inter-layer prediction. vps_independent_layer_flag[i] being equal to 0 specifies that the layer with index i can use inter-layer prediction and there exists a syntax element vps_direct_ref_layer_flag[i][j] in vps with j in the range from 0 to i–1 (including the end values). When it does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1.
[0124] vps_direct_ref_layer_flag[i][j] being equal to 0 specifies that the layer with index j is not a direct reference layer of the layer with index i. vps_direct_ref_layer_flag[i][j] being equal to 1 specifies that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[i][j] does not exist with i and j in the range from 0 to vps_max_layers_minus1 (including the end values), it is inferred to be equal to 0. When vps_independent_layer_flag[i] is equal to 0, there should be at least one j value in the range from 0 to i - 1 (including the end values) such that the value of vps_direct_ref_layer_flag[i][j] is equal to 1.
[0125] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived as follows:
[0126]
[0127]
[0128] The variable GeneralLayerIdx[i] specifies the layer index of the layer where nuh_layer_id is equal to vps_layer_id[i], and is derived as follows:
[0129] for(i = 0; i <= vps_max_layers_minus1; i++) (38)
[0130] GeneralLayerIdx[vps_layer_id[i]] = i
[0131] For any two different values of i and j (both within the range from 0 to vps_max_layers_minus1, inclusive), when dependencyFlag[i][j] equals 1, the requirement for bitstream conformance is that the values of chroma_format_idc and bit_depth_minus8 applicable to layer i shall be equal to the values of chroma_format_idc and bit_depth_minus8 applicable to layer j, respectively.
[0132] max_tid_ref_present_flag[i] being equal to 1 specifies the presence of the syntax element max_tid_il_ref_pics_plus1[i]. max_tid_ref_present_flag[i] being equal to 0 specifies the absence of the syntax element max_tid_il_ref_pics_plus1[i].
[0133] max_tid_il_ref_pics_plus1[i] being equal to 0 specifies that inter-layer prediction is not used for non-IRAP pictures of layer i. max_tid_il_ref_pics_plus1[i] being greater than 0 specifies that, for decoding pictures of layer i, no pictures with a TemporalId greater than max_tid_il_ref_pics_plus1[i] - 1 are used as ILRP. When not present, the value of max_tid_il_ref_pics_plus1[i] is inferred to be equal to 7.
[0134] each_layer_is_an_ols_flag being equal to 1 specifies that each OLS contains only one layer, and each layer in the CVS of the reference VPS itself is an OLS, where the single included layer is the only output layer. each_layer_is_an_ols_flag being equal to 0 specifies that an OLS may contain multiple layers. If vps_max_layers_minus1 equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.
[0135] ols_mode_idc being equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the highest layer in the OLS is output.
[0136] ols_mode_idc being equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i (including the end values), and for each OLS, all the layers in the output OLS are output.
[0137] ols_mode_idc being equal to 2 specifies that the total number of OLSs specified by the VPS is signaled explicitly, and for each OLS, the output layers are signaled explicitly, and the other layers are the layers that are direct or indirect reference layers of the output layers of the OLSs.
[0138] The value of ols_mode_idc shall be in the range from 0 to 2 (including the end values). The value 3 of ols_mode_idc is reserved by ITU-T|ISO / IEC for future use.
[0139] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.
[0140] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.
[0141] The variable TotalNumOlss specifies the total number of OLSs specified by the VPS, which is derived as follows:
[0142]
[0143] ols_output_layer_flag[i][j] being equal to 1 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS when ols_mode_idc is equal to 2. ols_output_layer_flag[i][j] being equal to 0 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS when ols_mode_idc is equal to 2.
[0144] The variable NumOutputLayersInOls[i] specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] specifies the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] specifies whether the k-th layer is used as an output layer in at least one OLS, as derived below:
[0145]
[0146]
[0147]
[0148] For each i value in the range from 0 to vps_max_layers_minus1 (including the end values), the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] should not both be equal to 0. In other words, there should not be a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.
[0149] For each OLS, there should be at least one layer as an output layer. In other words, for any i value in the range from 0 to TotalNumOlss–1 (including the end values), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.
[0150] The variable NumLayersInOls[i] specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] specifies the nuh_layer_id value of the j-th layer in the i-th OLS, as derived below:
[0151]
[0152]
[0153] Note 1 – The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, only the contained layer is output.
[0154] The variable OlsLayerIdx[i][j] specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j], as derived below:
[0155]
[0156] The lowest layer in each OLS shall be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss–1 (including the end values), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] shall be equal to 1.
[0157] Each layer shall be included in at least one OLS specified in the VPS. In other words, for each layer where the specific value of nuh_layer_id nuhLayerId is equal to one of vps_layer_id[k] in the range from 0 to vps_max_layers_minus1 (including the end values), there should be at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 (including the end values), and j is in the range from 0 to NumLayersInOls[i] - 1 (including the end values), such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.
[0158] vps_num_ptls_minus1 plus 1 specifies the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 shall be less than TotalNumOlss.
[0159] pt_present_flag[i] being equal to 1 specifies the tier, layer, and general constraint information present in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] being equal to 0 specifies that the tier, layer, and general constraint information is not present in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (i - 1)-th profile_tier_level() syntax structure in the VPS.
[0160] ptl_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation, the level information of which is present in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[i] shall be in the range (including the end values) from 0 to vps_max_sublayers_minus1. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.
[0161] vps_ptl_alignment_zero_bit shall be equal to 0.
[0162] ols_ptl_idx[i] specifies the index (to the list of profile_tier_level() syntax structures in the VPS) of the profile_tier_level() syntax structure applied to the i-th OLS. When present, the value of ols_ptl_idx[i] shall be in the range (including the end values) from 0 to vps_num_ptls_minus1. When vps_num_ptls_minus1 is equal to 0, the value of ols_ptl_idx[i] is inferred to be equal to 0.
[0163] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS is also present in the SPS referenced by the layers in the i-th OLS. The requirement for bitstream consistency is that when NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structures signaled for the i-th OLS in the VPS and the SPS should be the same.
[0164] vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params shall be in the range (including the end values) from 0 to 16. When absent, the value of vps_num_dpb_params is inferred to be equal to 0.
[0165] vps_sublayer_dpb_params_present_flag is used to control the presence of max_dec_pic_buffering_minus1[], max_num_reorder_pics[], max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax structure in the VPS. When not present, vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.
[0166] dpb_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation in the i-th dpb_parameters() syntax structure for which the DPB parameter may be present in the VPS. The value of dpb_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.
[0167] ols_dpb_pic_width[i] specifies the width of each picture storage buffer of the i-th OLS in units of luma samples.
[0168] ols_dpb_pic_height[i] specifies the height of each picture storage buffer of the i-th OLS in units of luma samples.
[0169] ols_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure that applies to the i-th OLS when NumLayersInOls[i] is greater than 1 (index into the list of dpb_parameters() syntax structures in the VPS). When present, the value of ols_dpb_params_idx[i] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[i] is not present, the value of ols_dpb_params_idx[i] is inferred to be equal to 0.
[0170] When NumLayersInOls[i] is equal to 1, the dpb_parameters() syntax structure applied to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.
[0171] A vps_general_hrd_params_present_flag equal to 1 specifies that the VPS contains the general_hrd_parameters() syntax structure and other HRD parameters. A vps_general_hrd_params_present_flag equal to 0 specifies that the VPS does not contain the general_hrd_parameters() syntax structure or other HRD parameters. When not present, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.
[0172] When NumLayersInOls[i] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS exist in the SPS referenced by the layer in the i-th OLS.
[0173] A vps_sublayer_cpb_params_present_flag equal to 1 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters for sublayer representation, where TemporalId ranges from 0 to hrd_max_tid[i] (including the end values). A vps_sublayer_cpb_params_present_flag equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters for sublayer representation, where TemporalId is only equal to hrd_max_tid[i]. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.
[0174] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented by the sublayer where the TemporalId is in the range from 0 to hrd_max_tid[i] – 1 (including the end values) are inferred to be the same as the HRD parameters represented by the sublayer where the TemporalId is equal to hrd_max_tid[i]. These include the HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element until the sublayer_hrd_parameters(i) syntax structure immediately following under the condition "if(general_vcl_hrd_params_present_flag)" in the ols_hrd_parameters syntax structure.
[0175] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the VPS when vps_general_hrd_params_present_fla is equal to 1. The value of num_ols_hrd_params_minus1 shall be in the range from 0 to TotalNumOlss - 1 (including the end values).
[0176] hrd_max_tid[i] specifies the highest sublayer representation TemporalId for which the HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range from 0 to vps_max_sublayers_minus1 (including the end values). When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[i] is inferred to be equal to vps_max_sublayers_minus1.
[0177] ols_hrd_idx[i] specifies the index (index into the list of ols_hrd_parameters() syntax structures in the VPS) of the ols_hrd_parameters() syntax structure applied to the i-th OLS when NumLayersInOls[i] is greater than 1. The value of ols_hrd_idx[[i] shall be in the range from 0 to num_ols_hrd_params_minus1 (including the end values).
[0178] When NumLayersInOls[i] is equal to 1, the ols_hrd_parameters() syntax structure applied to the i-th OLS exists in the SPS of the layer reference in the i-th OLS.
[0179] If the value of num_ols_hrd_param_minus1 + 1 is equal to TotalNumOlss, the value of ols_hrd_idx[i] is inferred to be equal to i. Otherwise, when NumLayersInOls[i] is greater than 1 and num_ols_hrd_params_minus1 is equal to 0, the value of ols_hrd_idx[[i] is inferred to be equal to 0.
[0180] vps_extension_flag being equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. vps_extension_flag being equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure.
[0181] vps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of this specification. Decoders compliant with this version of this specification shall ignore all vps_extension_data_flag syntax elements.
[0182] 3.6. SPS Syntax and Semantics in VVC
[0183] In the latest VVC text (in JVET-Q2001-vE / v15), the SPS syntax and semantics most relevant to the invention of this document are as follows.
[0184] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0185]
[0186] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0188] gdr_enabled_flag being equal to 1 specifies that GDR pictures may exist in the CLVS of the reference SPS. gdr_enabled_flag being equal to 0 specifies that GDR pictures do not exist in the CLVS of the reference SPS.
[0189] chroma_format_idc specifies the chroma sampling related to the luma sampling specified in Clause 6.2. ...
[0191] bit_depth_minus8 specifies the bit depth BitDepth of the samples in the luma and chroma arrays and the value of the luma and chroma quantization parameter range offset QpBdOffset as follows:
[0192] BitDepth = 8 + bit_depth_minus8 (45)
[0193] QpBdOffset = 6 * bit_depth_minus8 (46)
[0194] bit_depth_minus8 shall be in the range from 0 to 8 (including the end values). ...
[0196] 3.7. Picture Header Structure Syntax and Semantics in VVC
[0197] In the latest VVC text (in JVET-Q2001-vE / v15), the picture header structure syntax and semantics most relevant to the invention of this document are as follows.
[0198] 7.3.2.7 Picture Header Structure Syntax
[0199]
[0200]
[0201] 7.4.3.7 Picture Header Structure Semantics
[0202] The PH syntax structure contains information common to all slices of the coded picture related to the PH syntax structure.
[0203] gdr_or_irap_pic_flag being equal to 1 specifies that the current picture is a GDR or IRAP picture. gdr_or_irap_pic_flag being equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture.
[0204] gdr_pic_flag being equal to 1 specifies that the picture associated with the PH is a GDR picture. gdr_pic_flag being equal to 0 specifies that the picture associated with the PH is not a GDR picture. When absent, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0.
[0205] Note 1 – When gdr_or_irap_pic_flag equals 1 and gdr_pic_flag equals 0, the picture related to PH is an IRAP picture. ...
[0207] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb - 1 (including the end values).
[0208] no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a CLVSS picture that is not the first picture in the bitstream specified in Annex C.
[0209] recovery_poc_cnt specifies the recovery point of the decoded pictures in output order. If the current picture is a GDR picture associated with PH and there is a picture picA in CLVS that follows the current GDR picture in decoding order and whose PicOrderCntVal is equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then picture picA is called a recovery point picture. Otherwise, the first picture in output order whose PicOrderCntVal is greater than the PicOrderCntVal of the current picture plus the value of recovery_poc_cnt is called a recovery point picture. The recovery point picture shall not be before the current GDR picture in decoding order. The value of recovery_poc_cnt shall be in the range of 0 to MaxPicOrderCntLsb - 1 (including the end values).
[0210] When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:
[0211] RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt(81)
[0212] Note 2--When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the relevant GDR picture, the current and subsequent decoded pictures in the output order exactly match the corresponding pictures generated by decoding the process starting from the previous IRAP picture, and when present, are located before the relevant GDR picture in the decoding order. ...
[0214] 3.8. Set PictureOutputFlag
[0215] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable PictureOutputFlag is as follows (as part of Clause 8.1.2 Decoding process for coded pictures).
[0216] 8.1.2 Decoding process for coded pictures
[0217] The decoding process specified in this clause applies to each coded picture in BitstreamToDecode (referred to as the current picture and represented by the variable CurrPic).
[0218] According to the value of chroma_format_idc, the number of sample arrays of the current picture is as follows:
[0219] --If chroma_format_idc is equal to 0, the current picture consists of 1 sample array S L and.
[0220] --Otherwise (chroma_format_idc is not equal to 0), the current picture consists of 3 sample arrays S L , S Cb , S Cr and.
[0221] The decoding process of the current picture takes the syntax elements and uppercase variables from Clause 7 as input. When interpreting the semantics of each syntax element in each NAL unit and in the remainder of Clause 8, the term "bitstream" (or a part thereof, such as the CVS of the bitstream) refers to BitstreamToDecode (or a part thereof).
[0222] According to the value of separate_colour_plane_flag, the structure of the decoding process is as follows:
[0223] --If separate_colour_plane_flag is equal to 0, the decoding process is called once, with the current picture as the output.
[0224] -- Otherwise (separate_colour_plane_flag equals 1), the decoding process is called three times. The input to the decoding process is all the NAL units of the coded picture having the same colour_plane_id value. The decoding process for the NAL units having a particular colour_plane_id value is specified as if only the CVS of the monochrome colour format having that particular colour_plane_id value would be present in the bitstream. The output of each of the three decoding processes is assigned to one of the 3 sample arrays of the current picture, where the NAL units with colour_plane_id equal to 0, 1, and 2 are assigned to S L , S Cb and S Cr .
[0225] Note – When separate_colour_plane_flag equals 1 and chroma_format_idc equals 3, the variable ChromaArrayType is derived to be equal to 0. During the decoding process, the value of this variable is evaluated, resulting in the same operations as for a monochrome picture (when chroma_format_idc equals 0).
[0226] For the current picture CurrPic, the decoding process operates as follows:
[0227] 1. The decoding of the NAL units is specified in Clause 8.2.
[0228] 2. The following decoding processes are specified by the processes in Clause 8.3 using the syntax elements in the slice header layer and above:
[0229] -- The variables and functions related to picture order count are derived as specified in Clause 8.3.1. This is only required to be called for the first slice of the picture.
[0230] -- At the start of the decoding process for each slice of a non-IDR picture, the decoding process for reference picture list construction specified in Clause 8.3.2 is called to derive the reference picture list 0 (RefPicList[0]) and the reference picture list 1 (RefPicList[1]).
[0231] -- The decoding process for reference picture marking in Clause 8.3.3 is called, where a reference picture can be marked as "not used for reference" or "used for long-term reference". This is only required to be called for the first slice of the picture.
[0232] --When the current picture is a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, call the decoding process specified in Subclause 8.3.4 for generating unavailable reference pictures, and this process only needs to be called for the first slice of the picture.
[0233] --PictureOutputFlag is set as follows:
[0234] --If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:
[0235] --The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.
[0236] --gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0237] --gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.
[0238] --sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:
[0239] --The PictureOutputFlag of PicA is equal to 1.
[0240] --The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.
[0241] --PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0242] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and the ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0243] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag.
[0244] 3. The processing in Clauses 8.4, 8.5, 8.6, 8.7, and 8.8 specifies the decoding process using syntax elements in all syntax structure layers. The requirement for bitstream conformance is that the coded and decoded slices of a picture will contain the slice data for each CTU of the picture, such that the partitioning of the picture into slices and the partitioning of a slice into CTUs each form a partitioning of the picture.
[0245] 4. After all slices of the current picture have been decoded, the current decoded picture is marked as "for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "for short-term reference".
[0246] 3.9. Setting DPB Parameters for HRD Operations
[0247] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting DPB parameters for HRD operations is as follows (as part of Clause C.1).
[0248] C.1 General ...
[0250] For each bitstream conformance test, the CPB size (in bits) is CpbSize[Htid][ScIdx], as specified in Clause 7.4.6.3, where ScIdx and the HRD parameters are specified in that clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found or derived from the dpb_parameters() syntax structure applied to the target OLS as follows:
[0251] -- If the target OLS contains only one layer, the dpb_parameters() syntax structure is found in the SPS and is referred to as the layer in the target OLS.
[0252] -- Otherwise (if the target OLS contains more than one layer), dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsIdx] found in the VPS. ...
[0254] 3.10. Setting of NoOutputOfPriorPicsFlag
[0255] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable NoOutputOfPriorPicsFlag is as follows (as part of the specification for removing pictures from the DPB).
[0256] C.3.2 Removing Pictures from the DPB before Decoding the Current Picture
[0257] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (including the current picture) and is as follows:
[0258] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.
[0259] -- When the current AU is a CVSS AU other than AU 0, apply the following ordered steps:
[0260] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:
[0261] -- If the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for any picture in the current AU are different from the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous picture in the same CLVS, respectively, then NoOutputOfPriorPicsFlag may (but shall not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag.
[0262] Note -- Although it is preferable that NoOutputOfPriorPicsFlag be set to be equal to no_output_of_prior_pics_flag under these conditions, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1 in this case.
[0263] -- Otherwise, NoOutputOfPriorPicsFlag is set to be equal to no_output_of_prior_pics_flag.
[0264] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set to be equal to 0.
[0265] -- All such pictures k in the DPB will be removed from the DPB when the following two conditions are true for any picture k in the DPB:
[0266] -- Picture k is marked as "not used for reference".
[0267] -- The PictureOutputFlag of picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU of the current picture n (denoted as DU m); that is, DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].
[0268] -- For each picture removed from the DPB, the DPB fullness is decremented by 1.
[0269] C.5.2.2 Output and removal of pictures from the DPB
[0270] Output and removal of pictures from the DPB occur instantaneously when the first DU of the AU containing the current picture is removed from the CPB, before decoding the current picture (but after parsing the slice header of the first slice of the current picture) and proceed as follows:
[0271] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.
[0272] -- If the current picture is a CLVSS picture other than picture 0, apply the following ordered steps:
[0273] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:
[0274] -- If the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for any picture of the current AU are different from the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous picture in the same CLVS, respectively, then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag.
[0275] Note--Although it is preferable that NoOutputOfPriorPicsFlag be set equal to no_output_of_prior_pics_flag under these conditions, in this case, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1.
[0276] -- Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0277] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD as follows:
[0278] -- If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set equal to 0.
[0279] -- Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as "not to be output" and "not used for reference" are cleared (without output) by repeatedly invoking the "collision" procedure specified in Clause C.5.2.4, and all non-empty picture storage buffers in the DPB are cleared, and the DPB fullness is set equal to 0.
[0280] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not to be output" and "not used for reference" are cleared (without output). For each picture storage buffer cleared, the DPB fullness is decremented by 1. When one or more of the following conditions are true, the "collision" procedure specified in Clause C.5.2.4 is repeatedly invoked, and for each additional picture storage buffer cleared, the DPB fullness is further decremented by 1 until none of the following conditions are true:
[0281] -- The number of pictures marked as "to be output" in the DPB is greater than max_num_reorder_pics[Htid].
[0282] -- max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "to be output" and its associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].
[0283] -- The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid] + 1.
[0284] 4. Technical problems solved by the disclosed technical solution
[0285] The existing scalability design in the latest VVC text (in JVET - Q2001 - vE / v15) has the following problems:
[0286] 1) Currently, the maximum values of the picture width and height of all pictures in all layers are signaled in the VPS so that the decoder can correctly allocate memory for the DPB. Similar to the picture width and height, the chroma format and bit depth, currently specified by the SPS syntax elements chroma_format_idc and bit_depth_minus8 respectively, also affect the size of the picture storage buffer in the DPB. However, the maximum values of chroma_format_idc and bit_depth_minus8 for all pictures in all layers are not signaled.
[0287] 2) Currently, the setting of the value of the variable NoOutputOfPriorPicsFlag involves a change in the value of pic_width_max_in_luma_samples or pic_height_max_in_luma_samples. However, it should be changed to use the maximum values of the picture width and height of all pictures in all layers.
[0288] 3) Currently, the setting of NoOutputOfPriorPicsFlag involves a change in the value of chroma_format_idc or bit_depth_minus8. However, it should be changed to use the maximum values of the chroma format and bit depth of all pictures in all layers.
[0289] 4) Currently, the setting of NoOutputOfPriorPicsFlag involves a change in the value of separate_colour_plane_flag. However, separate_colour_plane_flag only exists and is used when chroma_format_idc is equal to 3, which specifies the 4:4:4 chroma format, and for the 4:4:4 chroma format, the value of separate_colour_plane_flag being 0 or 1 does not affect the buffer size required to store decoded pictures. Therefore, the setting of NoOutputOfPriorPicsFlag should not involve a change in the value of separate_colour_plane_flag.
[0290] 5) Currently, in the PH of IRAP and GDR pictures, the no_output_of_prior_pics_flag is signaled, and both the semantics of the flag and the procedure for setting the NoOutputOfPriorPicsFlag are specified in a way that the no_output_of_prior_pics_flag is layer-specific or PU-specific. However, since the DPB operation is OLS-specific or AU-specific, both the semantics of the no_output_of_prior_pics_flag and the use of the flag in the setting of NoOutputOfPriorPicsFlag should be specified in an AU-specific way.
[0291] 6) The current text for setting the value of the variable PictureOutputFlag of the current picture is related to the use of the PictureOutputFlag of pictures in the same AU as the current picture and in higher layers than the current picture. However, for the picA picture with the nuh_layer_id greater than that of the current picture, the PictureOutputFlag of picA has not been derived when deriving the PictureOutputFlag of the current picture.
[0292] 7) The current text for setting the value of the variable PictureOutputFlag has the following problem. There are two layers in the bitstream of the OLS, and only the higher layer is the output layer, and in a specific AU auA, the pic_output_flag of the picture in the higher layer is equal to 0. On the decoder side, the picture in the higher layer of auA does not exist (due to, for example, loss or layer down-switching), while the picture in the lower layer of auA exists and its pic_output_flag is equal to 1. Then the value of the PictureOutputFlag of the picture in the lower layer of auA will be set to be equal to 1. However, when the OLS has only one output layer and the pic_output_flag of the picture in the output layer is equal to 0, it should be interpreted that the encoder (or content provider) does not want to output the picture for the AU containing this picture.
[0293] 8) The current text for setting the value of the variable PictureOutputFlag has the following problem. In the bitstream of OLS, there are three or more layers, and only the top layer is the output layer. On the decoder side, if the top layer picture of the current AU does not exist (due to, for example, loss or layer down-switching), while two or more pictures of the lower layers of the current AU exist and the pic_output_flag of these pictures is equal to 1, then for this AU, more than one picture will be output. However, this is problematic because for OLS there is only one output layer, so the codec or content provider expects only one picture to be output.
[0294] 9) The current text for setting the value of the variable PictureOutputFlag has the following problem. OLS mode 2 (when ols_mode_idc is equal to 2) can also specify only one output layer like mode 0, but the behavior of outputting the lower layer pictures of the AU when the output layer picture (which is also the top layer picture) does not exist is only specified for mode 0.
[0295] 10) For OLS that contains only one output layer, when the picture of the output layer (which is also the top layer picture) is not available to the decoder (due to, for example, loss or layer down-switching), the decoder will not be able to know whether the pic_output_flag of this picture is equal to 1 or 0. If it is equal to 1, then it makes sense to output the lower layer pictures, but if it is equal to 0, then from the perspective of user experience, outputting the lower layer pictures may be worse because the encoder (content provider) has set this value to 0 for some reason, for example, for this specific OLS, this AU should not have picture output.
[0296] 5. List of Embodiments and Techniques
[0297] To solve the above problems and other problems, methods outlined below are disclosed. The items listed should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these items can be applied individually or combined in any way.
[0298] Solutions to Problems 1 to 5
[0299] 1) To solve problem 1, one or both of the maximum values of chroma_format_idc and bit_depth_minus8 of all pictures of all layers can be signaled in the VPS.
[0300] 2) To solve problem 2, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be at least based on one or both of the maximum picture width and height of all pictures of all layers that can be signaled in the VPS.
[0301] 3) To solve Problem 3, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be at least based on one or both of the maximum values of chroma_format_idc and bit_depth_minus8 of all pictures of all layers that can be signaled in the VPS.
[0302] 4) To solve Problem 4, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be independent of the value of separate_colour_plane_flag.
[0303] 5) To solve Problem 5, both the semantics of no_output_of_prior_pics_flag and its use in the setting of NoOutputOfPriorPicsFlag can be specified in an AU-specific manner.
[0304] a. In one example, it can be required that when present, for all pictures in an AU, the value of no_output_of_prior_pics_flag should be the same, and the value of no_output_of_prior_pics_flag of the AU is considered to be the value of no_output_of_prior_pics_flag of the pictures of the AU.
[0305] b. Alternatively, in one example, when irap_or_gdr_au_flag is equal to 1, no_output_of_prior_pics_flag can be removed from the PH syntax and can be signaled in the AUD syntax.
[0306] i. For a single-layer bitstream, since AUD is optional, when AUD is absent for an IRAP or GDR AU, the value of no_output_of_prior_pics_flag can be inferred to be equal to 1 (which means that if the encoder wants to signal the value 0 of no_output_of_prior_pics_flag for an IRAP or GDR AU in a single-layer bitstream, it must signal AUD for that AU in the bitstream).
[0307] c. Alternatively, in one example, the value of the no_output_of_prior_pics_flag for an AU can be considered equal to 0 if and only if the no_output_of_prior_pics_flag for each picture in the AU is equal to 0; otherwise, the value of the no_output_of_prior_pics_flag for the AU can be considered equal to 1.
[0308] i. The disadvantage of this method is that the setting of the NoOutputOfPriorPicsFlag and the output of pictures in the CVSS AU need to wait for all pictures in the AU to arrive.
[0309] Solutions to Problems 6 to 10
[0310] 6) To solve Problem 6, the setting of the PictureOutputFlag for the current picture can be specified to be at least based on the pic_output_flag (instead of the PictureOutputFlag) of pictures that are in the same AU as the current picture and in a layer higher than the current picture.
[0311] 7) To solve Problems 7 to 9, whenever the current picture does not belong to the output layer, the value of the PictureOutputFlag for the current picture is set to be equal to 0.
[0312] a. Alternatively, to solve Problems 7 and 8, when there is only one output layer and there is no output layer for the AU (which must be the top layer when there is only one output layer), then for the picture among all pictures in the AU available to the decoder that has the highest value of nuh_layer_id and pic_output_flag equal to 1, the PictureOutputFlag is set to be equal to 1, and for all other pictures in the AU available to the decoder, the PictureOutputFlag is set to be equal to 0.
[0313] 8) To solve Problem 10, the value of the pic_output_flag for the output layer pictures of the AU can be signaled in the AUD or SEI message in the AU, or in the PH of one or more other pictures in the AU.
[0314] 6. Embodiment
[0315] The following are some example embodiments of some aspects of the invention summarized in Section 5 above, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-Q2001-vE / v15. Most of the relevant parts that have been added or modified are highlighted in bold italics, and some deleted parts are marked with double brackets (e.g., [[a]] means deleting the character "a"). There are also some other changes that are editorial in nature and are therefore not highlighted.
[0316] 6.1. First embodiment
[0317] This embodiment is used for items 1, 2, 3, 4, 5 and 5a.
[0318] 7.3.2.2 Video parameter set syntax
[0319] ...
[0321] 7.4.3.2 Video Parameter Set RBSP Semantics ...
[0323] ols_dpb_pic_width[i] specifies the width of each picture storage buffer of the i-th OLS in units of luma samples.
[0324] ols_dpb_pic_height[i] specifies the height of each picture storage buffer of the i-th OLS in units of luma samples.
[0325]
[0326] ols_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure that applies to the i-th OLS when NumLayersInOls[i] is greater than 1 (index into the list of dpb_parameters() syntax structures in the VPS). When present, the value of ols_dpb_params_idx[i] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[i] is not present, the value of ols_dpb_params_idx[i] is inferred to be equal to 0.
[0327] When NumLayersInOls[i] is equal to 1, the dpb_parameters() syntax structure that applies to the i-th OLS is present in the SPS referenced by the layers in the i-th OLS. ...
[0329] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...
[0331] The gdr_enabled_flag being equal to 1 specifies that GDR pictures may be present in the CLVS of the reference SPS. The gdr_enabled_flag being equal to 0 specifies that GDR pictures are not present in the CLVS of the reference SPS.
[0332] The chroma_format_idc specifies the chroma sampling related to the luma sampling specified in Clause 6.2.
[0333] ...
[0335] The bit_depth_minus8 specifies the bit depth BitDepth of the samples of the luma and chroma arrays and the value of the luma and chroma quantization parameter range offset QpBdOffset as follows:
[0336] BitDepth = 8 + bit_depth_minus8 (45)
[0337] QpBdOffset = 6 * bit_depth_minus8 (46)
[0338] The bit_depth_minus8 shall be in the range of 0 to 8 (inclusive of the end values).
[0339] ...
[0341] 7.4.3.7 Picture Header Structure Semantics ...
[0343] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after the bitstream specified in Appendix C. in the bitstream specified in Appendix C after that.
[0344] ...
[0346] C.1 General ...
[0348] For each bitstream conformance test, the CPB size (number of bits) is CpbSize[Htid][ScIdx] as specified in Clause 7.4.6.3, where ScIdx and the HRD parameters are specified in that clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found or derived from the dpb_parameters() syntax structure applied to the target OLS as follows:
[0349] -- If then the dpb_parameters() syntax structure is found in the SPS, which is referred to as the layer in the target OLS,
[0350] -- Otherwise (the target OLS contains more than one layer), dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsIdx] found in the VPS, ...
[0352] C.3.2 Removing pictures from the DPB before decoding the current picture
[0353] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (containing the current picture) and is done as follows:
[0354] -- Call the decoding process for reference picture list construction specified in Clause 8.3.2 and call the decoding process for reference picture marking specified in Clause 8.3.3.
[0355] -- When the current AU is a CVSS AU other than AU 0, apply the following ordered steps:
[0356] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:
[0357] -- If derived or the value of max_dec_pic_buffering_minus1[Htid] respectively for the previous Derived or if the value of max_dec_pic_buffering_minus1[Htid] is different, the NoOutputOfPriorPicsFlag may (but shall not) be set to 1 by the DUT decoder, regardless of the value of no_output_of_prior_pics_flag.
[0358] Note -- Although it is preferred that the NoOutputOfPriorPicsFlag be set equal to no_output_of_prior_pics_flag under these conditions, in this case the DUT decoder is allowed to set the NoOutputOfPriorPicsFlag to 1.
[0359] -- Otherwise, the NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0360] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT decoder is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set equal to 0.
[0361] -- For any picture k in the DPB, all such pictures k in the DPB will be removed from the DPB when the following two conditions are true:
[0362] -- Picture k is marked as "not used for reference".
[0363] -- The PictureOutputFlag of picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU of the current picture n (denoted as DU m); i.e., DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].
[0364] -- For each picture removed from the DPB, the DPB fullness is decremented by 1.
[0365] C.5.2.2 Outputting and Removing Pictures from the DPB
[0366] Output and remove the picture from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture). This occurs instantaneously when removing the first DU of the AU containing the current picture from the CPB and is as follows:
[0367] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.
[0368] -- If the current is not CLVSS as then apply the following ordered steps:
[0369] 1. The variable NoOutputOfPriorPicsFlag is derived for the DUT as follows:
[0370] -- If the derived or the value of max_dec_pic_buffering_minus1[Htid] is different from the value for the previous or max_dec_pic_buffering_minus1[Htid], respectively, then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the DUT, regardless of the value of no_output_of_prior_pics_ .
[0371] Note -- Although it is preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag in these conditions, the DUT is allowed to set NoOutputOfPriorPicsFlag to 1 in this case.
[0372] -- Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.
[0373] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT is applied to the HRD as follows:
[0374] -- If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set equal to 0.
[0375] -- Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output) by repeatedly invoking the "collision" process specified in Clause C.5.2.4, and all non-empty picture storage buffers in the DPB are cleared, and the DPB fullness is set equal to 0.
[0376] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output). For each picture storage buffer that is cleared, the DPB fullness is decremented by 1. When one or more of the following conditions are true, the "collision" process specified in Clause C.5.2.4 is repeatedly invoked, and for each additional picture storage buffer that is cleared, the DPB fullness is further decremented by 1 until none of the following conditions are true:
[0377] -- The number of pictures marked as "required for output" in the DPB is greater than max_num_reorder_pics[Htid].
[0378] -- max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "required for output" and its associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].
[0379] -- The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1.
[0380] 6.2. Second Embodiment
[0381] This embodiment is for Items 1, 2, 3, 4, 5, and 5c, where the text varies with respect to the text of the first embodiment.
[0382] 7.4.3.7 Picture Header Structure Semantics ...
[0384] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after pictures in the CVSS AU that decode the first AU in a bitstream not specified in Annex C.
[0385] [[The requirement for bitstream conformance is that when present, the value of the no_output_of_prior_pics_flag shall be the same for all pictures of an AU.
[0386] When the no_output_of_prior_pics_flag is present in the PH of a picture of an AU, the value of the no_output_of_prior_pics_flag of the AU is the value of the no_output_of_prior_pics_flag of the picture of the AU.]] ...
[0388] C.3.2 Removing pictures from the DPB before decoding the current picture
[0389] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (including the current picture) and is as follows:
[0390] -- Call the decoding process for reference picture list construction specified in Clause 8.3.2 and call the decoding process for reference picture marking specified in Clause 8.3.3.
[0391] -- When the current AU is a CVSS AU other than AU 0, apply the following ordered steps:
[0392] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:
[0393] -- If the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous AU in decoding order, then NoOutputOfPriorPicsFlag may (but shall not) be set to 1 by the DUT, regardless of [[the value of]] no_output_of_prior_pics_flag [[for the current AU]]
[0394] Note-- [[Although]] under these conditions, it is preferable for NoOutputOfPriorPicsFlag to be set equal to [[equal to the no_output_of_prior_pics_flag of the current AU]], but in this case, the DUT is allowed to set NoOutputOfPriorPicsFlag to 1.
[0395] -- Otherwise, [[NoOutputOfPriorPicsFlag is set equal to the no_output_of_prior_pics_flag of the current AU]]
[0396] --
[0397] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set equal to 0.
[0398] -- All such pictures k in the DPB will be removed from the DPB when the following two conditions are true for any picture k in the DPB:
[0399] -- Picture k is marked as "not used for reference".
[0400] -- The PictureOutputFlag of Picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current Picture n; that is, DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].
[0401] -- For each picture removed from the DPB, the DPB fullness is decremented by 1.
[0402] C.5.2.2 Output and Removal of Pictures from the DPB
[0403] Output and removal of pictures from the DPB occur instantaneously when removing the first DU of the AU containing the current picture from the CPB, before decoding the current picture (but after parsing the slice header of the first slice of the current picture) and proceed as follows:
[0404] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.
[0405] -- If the current AU is a CLVSS AU that is not AU 0, apply the following ordered steps:
[0406] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:
[0407] -- If the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous AU in decoding order, then NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of the value of [[the current AU's]] no_output_of_prior_pics_flag
[0408] Note-- [[Although]] under these conditions, it is preferable that the NoOutputOfPriorPicsFlag be set equal to [[the no_output_of_prior_pics_flag of the current AU]], but in this case, the DUT decoder is allowed to set the NoOutputOfPriorPicsFlag to 1.
[0409] -- Otherwise, then the NoOutputOfPriorPicsFlag is [[set equal to the no_output_of_prior_pics_flag of the current AU]].
[0410] --
[0411] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT decoder is applied to the HRD as follows:
[0412] -- If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set equal to 0.
[0413] -- Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output) by repeatedly calling the "collision" procedure specified in Clause C.5.2.4, and all non-empty picture storage buffers in the DPB are cleared, and the DPB fullness is set equal to 0.
[0414] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output). For each picture storage buffer cleared, the DPB fullness is decremented by 1. When one or more of the following conditions are true, the "collision" procedure specified in Clause C.5.2.4 is repeatedly called, and for each additional picture storage buffer cleared, the DPB fullness is further decremented by 1 until none of the following conditions are true:
[0415] -- The number of pictures marked as "required for output" in the DPB is greater than max_num_reorder_pics[Htid].
[0416] --max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "to be output", and its related variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].
[0417] --The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1.
[0418] 6.3. Third Embodiment
[0419] This embodiment is used for Item 6, Item 7 (the changed text does not include the notes added in Clause 8.1.2), and Item 7a (the notes added in Clause 8.1.2).
[0420] 7.4.3.7 Picture Header Structure Semantics ...
[0422] recovery_poc_cnt specifies the recovery point of the decoded pictures in the output order.
[0423]
[0424]
[0425] If the current picture is a [[PH-associated]] GDR picture, and there is a picture picA in the CLVS that follows the current GDR picture in the decoding order, and its PicOrderCntVal is equal to [[the value of PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt]], then the picture picA is called a recovery point picture. Otherwise, the first picture in the output order whose PicOrderCntVal is greater than [[the value of PicOrderCntVal of the current picture plus the value of recovery_poc_cnt]] is called a recovery point picture. The recovery point picture should not be before the current GDR picture in the decoding order. The value of recovery_poc_cnt should be in the range from 0 to MaxPicOrderCntLsb - 1 (including the end values).
[0426] [[When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:
[0427] RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81)]
[0428] Note 2--When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to that of the relevant GDR picture [[RpPicOrderCntVal]], the current and subsequent decoded pictures in the output order exactly match the corresponding pictures generated by decoding the process starting from the previous IRAP picture, and when present, are located before the relevant GDR picture in the decoding order. ...
[0430] 8.1.2 Decoding Process of Encoded and Decoded Pictures ...
[0432] --
[0433]
[0434] --PictureOutputFlag is set as follows:
[0435] --If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:
[0436] --The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.
[0437] --gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0438] --gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.
[0439] --sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:
[0440] --The PictureOutputFlag of PicA is equal to 1.
[0441] -- The nuh_layer_id of PicA, nuhLid, is greater than the nuh_layer_id of the current picture.
[0442] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0443] -- sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0444] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag. ...
[0446] 6.4. Fourth Embodiment
[0447] This embodiment is for Project 6 and Project 7a.
[0448] 8.1.2 Decoding Process of Encoded / Decoded Pictures ...
[0450] -- PictureOutputFlag is set as follows:
[0451] -- If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:
[0452] -- The current picture is a RASL picture, and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.
[0453] -- gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0454] -- gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.
[0455]
[0456] -- When the sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:
[0457] -- The PictureOutputFlag of PicA is equal to 1.
[0458] -- The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.
[0459] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0460] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0461] -- Otherwise, the PictureOutputFlag is set to be equal to pic_output_flag. ...
[0463] 6.5. Fifth Embodiment
[0464] This embodiment is only used for Project 6.
[0465] 8.1.2 Decoding Process of Encoded and Decoded Pictures ...
[0467] -- The PictureOutputFlag is set as follows:
[0468] -- If one of the following conditions is true, the PictureOutputFlag is set to be equal to 0:
[0469] -- The current picture is a RASL picture, and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.
[0470] -- The gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.
[0471] -- The gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.
[0472] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:
[0473] -- For PicA [[PictureOutputFlag]] is equal to 1.
[0474] -- The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.
[0475] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).
[0476] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.
[0477] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag.
[0478] Figure 1 FIG. 24 is a block diagram showing an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0479] System 1900 may include a codec component 1904 that may implement various codec or encoding methods described in this document. The codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Codec techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 may be stored or transmitted via a communication connection as represented by component 1906. The stored or communicatively transmitted bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or a displayable video for transmission to a display interface 1910. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, while certain video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0480] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0481] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be embodied in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 may be configured to implement one or more methods described in this document. The memory (multiple memories) 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry.
[0482] Figure 4 is a block diagram showing an example video codec system 100 that may utilize the techniques of the present disclosure.
[0483] As Figure 4As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, where the source device 110 may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, where the destination device 120 may be referred to as a video decoding device.
[0484] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0485] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0486] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0487] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.
[0488] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or additional standards.
[0489] Figure 5 is a block diagram showing an example of a video encoder 200, which may be the video encoder 114 in the system 100 shown Figure 4 in the system 100 shown.
[0490] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 the example of, video encoder 200 includes multiple functional components. The techniques described in the present disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0491] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0492] In other examples, video encoder 200 may include more, fewer, or different functional components. In an example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0493] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in the Figure 5 example for purposes of explanation.
[0494] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.
[0495] Mode selection unit 203 may select one of the coding / decoding modes (e.g., intra or inter) based on error results, and provide the resulting intra-coded / decoded block or inter-coded / decoded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block to be used as a reference picture. In some examples, mode selection unit 203 may select a combination of intra and inter prediction modes (CIIP), where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 may also select the resolution of the motion vector of the block (e.g., sub-pixel or integer-pixel accuracy).
[0496] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0497] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.
[0498] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference picture in list 0 or list 1 of reference video blocks for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1, the reference index including the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0499] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block, the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in list 1. The motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 including the reference video blocks and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0500] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder.
[0501] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0502] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0503] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0504] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0505] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0506] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0507] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0508] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0509] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0510] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0511] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the block effect in the video block.
[0512] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0513] Figure 6 is a block diagram showing an example of a video decoder 300, and the video decoder 300 can be Figure 4 the video decoder 114 in the system 100 shown.
[0514] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 6 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0515] In Figure 6 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described for the video encoder 200 ( Figure 5 ).
[0516] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 can determine motion information including a motion vector, a motion vector precision, a reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes.
[0517] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0518] The motion compensation unit 302 may use an interpolation filter such as that used by the video encoder 200 during encoding of a video block to compute an interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information and may use the interpolation filter to generate a prediction block.
[0519] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode (a) frame(s) and / or (a) strip(s) of an encoded video sequence, partitioning information describing how each macroblock of a picture describing the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.
[0520] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in a bitstream. The inverse quantization unit 303 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 303 applies an inverse transform.
[0521] The reconstruction unit 306 may add a residual block to a corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block in order to remove blocking artifacts. The decoded video block is then stored in the buffer 307, providing a reference block for subsequent motion compensation / intra prediction and also producing decoded video for presentation on a display device.
[0522] Next, a list of preferred examples of some embodiments is provided.
[0523] A first set of clauses illustrates example embodiments of the techniques discussed in the previous chapter. The following clauses illustrate example embodiments of the techniques discussed in the previous chapter (e.g., item 1).
[0524] 1. A video processing method (e.g., Figure 3 method 3000 as illustrated), comprising: performing a conversion (3002) between a video and a coded representation of the video having one or more video layers including one or more video pictures; wherein the coded representation includes a video parameter set that indicates a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels used to represent the video.
[0525] The following clauses illustrate example embodiments of the technology discussed in the previous chapter (e.g., item 2).
[0526] 2. A video processing method, comprising: performing a conversion between a video having one or more video layers and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies a maximum picture width and / or a maximum picture height of video pictures of all video layers to control a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer.
[0527] 3. The method according to clause 2, wherein the variable is signaled in a video parameter set.
[0528] The following clauses illustrate example embodiments of the technology discussed in the previous chapter (e.g., item 3).
[0529] 4. A video processing method, comprising: performing a conversion between a video having one or more video layers and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies a maximum value of a chroma format indicator and / or a maximum bit depth of pixels representing the video to control a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer.
[0530] 5. The method according to clause 4, wherein the variable is signaled in a video parameter set.
[0531] The following clauses illustrate example embodiments of the technology discussed in the previous chapter (e.g., item 4).
[0532] 6. A video processing method, comprising: performing a conversion between a video having one or more video layers and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies that a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is independent of whether separate color planes are used to encode the video.
[0533] The following clauses illustrate example embodiments of the technology discussed in the previous chapter (e.g., item 5).
[0534] 7. A video processing method, comprising: performing a conversion between a video having one or more video layers and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies that a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is included in the coded representation at an access unit (AU) level.
[0535] 8. The method according to clause 7, wherein the format rule specifies that the value is the same for all AUs in the coded representation.
[0536] 9. The method according to any one of clauses 7 - 8, wherein the variable is indicated in the picture header.
[0537] 10. The method according to any one of clauses 7 - 8, wherein the variable is indicated in the access unit delimiter.
[0538] The following clauses illustrate example embodiments of the technology discussed in the previous section (e.g., item 6).
[0539] 11. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies that a picture output flag of a video picture in an access unit is determined based on a pic_output_flag variable of another video picture in the access unit.
[0540] The following clauses illustrate example embodiments of the technology discussed in the previous section (e.g., item 7).
[0541] 12. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies a value of a picture output flag for a video picture that does not belong to an output layer.
[0542] 13. The method according to clause 12, wherein the format rule specifies that the value of the picture output flag of the video picture is set to zero.
[0543] 14. The method according to clause 12, wherein the video includes only one output layer, and wherein an access unit that does not include an output layer is encoded / decoded by setting the picture output flag value to logical 1 for a picture having the highest layer id value and setting the picture output flag value to logical 0 for all other pictures.
[0544] The following clauses illustrate example embodiments of the technology discussed in the previous section (e.g., item 8).
[0545] 15. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the access unit delimiter.
[0546] 16. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the auxiliary enhancement information field.
[0547] 17. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the picture header of one or more pictures.
[0548] 18. The method according to any one of clauses 1 to 17, wherein the conversion includes encoding the video into a codec representation.
[0549] 19. The method according to any one of clauses 1 to 17, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0550] 20. A video decoding apparatus, comprising a processor configured to implement the method according to one or more of clauses 1 to 19.
[0551] 21. A video encoding apparatus, comprising a processor configured to implement the method according to one or more of clauses 1 to 19.
[0552] 22. A computer program product storing computer code, which when executed by a processor causes the processor to implement the method according to any one of clauses 1 to 19.
[0553] 23. A method, apparatus or system described in this document.
[0554] The second set of clauses shows example embodiments of the techniques discussed in the previous chapter (e.g., items 1 - 4).
[0555] 1. A method for video processing (e.g., the method 710 as Figure 7A shown), comprising: performing a conversion 712 between a video and a bitstream of the video according to format rules, wherein the bitstream includes one or more output layer sets (OLSs), each OLS including one or more codec layer video sequences, and wherein the format rules specify that a video parameter set indicates, for each of the one or more OLSs, a maximum allowable value of a chroma format indicator for representing pixels of the video and / or a maximum allowable value of a bit depth.
[0556] 2. The method according to clause 1, wherein the maximum allowable value of the chroma format indicator of the OLS applies to all sequence parameter sets referenced by one or more codec layer video sequences in the OLS.
[0557] 3. The method according to clause 1 or 2, wherein the maximum allowable value of the bit depth of the OLS applies to all sequence parameter sets referenced by one or more codec layer video sequences in the OLS.
[0558] 4. The method according to any one of clauses 1 to 3, wherein, in order to perform a conversion on an OLS that includes a video sequence with more than one coding / decoding layer and has an OLS index i, the rule specifies allocating memory for the decoded picture buffer according to at least one value among the syntax elements including ols_dpb_pic_width[i] indicating the width of each picture storage buffer of the i-th OLS, ols_dpb_pic_height[i] indicating the height of each picture storage buffer of the i-th OLS, a syntax element indicating the maximum allowable value of the chroma format indicator of the i-th OLS, and a syntax element indicating the maximum allowable value of the bit depth of the i-th OLS.
[0559] 5. The method according to any one of clauses 1 to 4, wherein the video parameter set is included in the bitstream.
[0560] 6. The method according to any one of clauses 1 to 4, wherein the video parameter set is indicated separately from the bitstream.
[0561] 7. A method for video processing (e.g., the method 720 as shown in Figure 7B ), including: performing a conversion 722 between a video with one or more video layers and the bitstream of the video according to format rules, and wherein the format rules specify a variable value that controls whether pictures in the decoded picture buffer before the current picture in decoding order in the bitstream are output before the pictures are removed from the decoded picture buffer for the maximum picture width and / or maximum picture height of the video pictures of all video layers.
[0562] 8. The method according to clause 7, wherein the variable is derived based on at least one or more syntax elements included in the video parameter set.
[0563] 9. The method according to clause 7, wherein, in a case where the value of the maximum width of each picture, the maximum height of each picture, the maximum allowable value of the chroma format indicator, or the maximum allowable value of the bit depth derived for the current access unit is different from the value of the maximum width of each picture, the maximum height of each picture, the maximum allowable value of the chroma format indicator, or the maximum allowable value of the bit depth derived for the previous access unit in decoding order, the value of the variable is set to 1.
[0564] 10. The method according to clause 9, wherein the value of the variable being equal to 1 indicates that pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures are removed from the decoded picture buffer.
[0565] 11. The method according to any one of clauses 7 to 10, wherein the value of the variable is further based on the maximum allowable value of the chrominance format indicator and / or the maximum allowable value of the bit depth for the pixels representing the video.
[0566] 12. The method according to any one of clauses 7 to 11, wherein the video parameter set is included in the bitstream.
[0567] 13. The method according to any one of clauses 7 to 11, wherein the video parameter set is indicated separately from the bitstream.
[0568] 14. A method for video processing (e.g., as shown in 730), comprising: performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules 732, and wherein the format rules specify the maximum allowable value of the chrominance format indicator and / or the maximum allowable value of the bit depth for the pixels representing the video to control the value of a variable indicating whether a picture in the decoded picture buffer before the current picture in decoding order is output before the picture is removed from the decoded picture buffer. Figure 7C shown in 730), including: performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules 732, and wherein the format rules specify the maximum allowable value of the chrominance format indicator and / or the maximum allowable value of the bit depth for the pixels representing the video to control the value of a variable indicating whether a picture in the decoded picture buffer before the current picture in decoding order is output before the picture is removed from the decoded picture buffer.
[0569] 15. The method according to clause 14, wherein the variable is derived based on at least one or more syntax elements signaled in the video parameter set.
[0570] 16. The method according to clause 14, wherein, in a case where the value of the maximum width per picture, the maximum height per picture, the maximum allowable value of the chrominance format indicator, or the maximum allowable value of the bit depth for each picture derived for the current access unit is different from the value of the maximum width per picture, the maximum height per picture, the maximum allowable value of the chrominance format indicator, or the maximum allowable value of the bit depth for each picture derived for the previous access unit in decoding order, the value of the variable is set to 1.
[0571] 17. The method according to clause 16, wherein the value of the variable being equal to 1 indicates that a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is removed from the decoded picture buffer.
[0572] 18. The method according to any one of clauses 14 to 17, wherein the value of the variable is further based on the maximum picture width and / or the maximum picture height of the video pictures of all video layers.
[0573] 19. The method according to any one of clauses 14 to 18, wherein the video parameter set is included in the bitstream.
[0574] 20. The method according to any one of clauses 14 to 18, wherein the video parameter set is indicated separately from the bitstream.
[0575] 21. A method for video processing (e.g., the method 740 as Figure 7D shown), comprising: performing a conversion between a video having one or more video layers and a bitstream of the video according to a rule 742, and wherein the rule specifies that the value of a variable indicating whether a picture in a decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is independent of whether separate color planes are used for encoding the video.
[0576] 22. The method according to clause 21, wherein, in the case where separate color planes are not used for encoding the video, the rule specifies that the video picture is decoded only once, or in the case where separate color planes are used for encoding the video, the rule specifies that the picture decoding is called three times.
[0577] 23. The method according to any one of clauses 1 to 22, wherein the conversion includes encoding the video into a bitstream.
[0578] 24. The method according to any one of clauses 1 to 22, wherein the conversion includes decoding the video from the bitstream.
[0579] 25. The method according to clauses 1 to 22, wherein the conversion includes generating a bitstream from the video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.
[0580] 26. A video processing apparatus, comprising a processor configured to implement the method according to any one or more of clauses 1 to 25.
[0581] 27. A method for storing a bitstream of a video, comprising the method according to any one of clauses 1 to 25, and further comprising storing the bitstream into a non-transitory computer-readable recording medium.
[0582] 28. A computer-readable medium storing program code that, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 25.
[0583] 29. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0584] 30. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 25.
[0585] The third set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., item 5).
[0586] 1. A method for video processing (e.g., the method asFigure 8A The method (810) shown includes: performing a conversion (812) between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag indicating whether to remove a previously decoded and stored picture in a decoded picture buffer when decoding a certain type of access unit is included in the bitstream.
[0587] 2. The method according to clause 1, wherein the format rules specify that the value is the same for all pictures in the access unit.
[0588] 3. The method according to clause 1 or 2, wherein the format rules specify that a value of a variable indicating whether a picture in the decoded picture buffer that is in decoding order before the current picture in the bitstream is output before the picture is removed from the decoded picture buffer is based on the value of the flag.
[0589] 4. The method according to any one of clauses 1 to 3, wherein the flag is indicated in a picture header.
[0590] 5. The method according to any one of clauses 1 to 3, wherein the flag is indicated in a slice header.
[0591] 6. A method for video processing (e.g., the method (820) as Figure 8B shown), includes: performing a conversion (822) between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a first flag indicating whether to remove a previously decoded and stored picture in a decoded picture buffer when decoding a specific type of access unit is not indicated in a picture header.
[0592] 7. The method according to clause 6, wherein the first flag is indicated in an access unit delimiter.
[0593] 8. The method according to clause 6, wherein a second flag indicating an IRAP (Intra Random Access Point picture) or GDR (Gradual Decoding Refresh) access unit has a certain value.
[0594] 9. The method according to clause 6, wherein in a case where there is no access unit delimiter for an IRAP (Intra Random Access Point picture) or GDR (Gradual Decoding Refresh) access unit, the value of the first flag is inferred to be equal to 1.
[0595] 10. A method for video processing (e.g., as Figure 8CThe method shown (830) includes: performing a conversion (832) between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that the value of a flag associated with an access unit indicating whether to remove a previously decoded and stored picture from a decoded picture buffer depends on the value of the flag of each picture of the access unit.
[0596] 11. The method according to clause 10, wherein the formatting rules specify that, in the case where the flag of each picture of the access unit is equal to 0, the value of the flag of the access unit is considered equal to 0; otherwise, the value of the flag of the access unit is considered equal to 1.
[0597] 12. The method according to any one of clauses 1 to 11, wherein the conversion includes encoding the video into a bitstream.
[0598] 13. The method according to any one of clauses 1 to 11, wherein the conversion includes decoding the video from the bitstream.
[0599] 14. The method according to any one of clauses 1 to 11, wherein the conversion includes generating a bitstream from the video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.
[0600] 15. A video processing apparatus, including a processor configured to implement the method according to any one or more of clauses 1 to 14.
[0601] 16. A method of storing a bitstream of a video, including the method according to any one of clauses 1 to 14, and further including storing the bitstream into a non-transitory computer-readable recording medium.
[0602] 17. A computer-readable medium storing program code that, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 14.
[0603] 18. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0604] 19. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 14.
[0605] The fourth set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., items 6-8).
[0606] 1. A method of video processing (e.g., as Figure 9AThe method shown (910) includes: performing a conversion (912) between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that a value of a variable indicating whether to output a picture in an access unit is determined based on a flag indicating whether to output another picture in the access unit.
[0607] 2. The method according to clause 1, wherein the other picture is in a higher layer than the picture.
[0608] 3. The method according to clause 1 or 2, wherein the flag controls a decoded picture output and removal process.
[0609] 4. The method according to clause 1 or 2, wherein the flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a picture group header.
[0610] 5. The method according to any one of clauses 1 to 4, wherein the value of the variable is further based on at least one of the following: i) a flag of a value of an identifier specifying a video parameter set (VPS), ii) whether the current video layer is an output layer, ii) whether the current picture is a random access skipped previous picture, a progressive decoded refresh picture, a recovery picture of a progressive decoded refresh picture, or iii) whether a picture in a decoded picture buffer before the current picture in decoding order is output before the picture is recovered.
[0611] 6. The method according to clause 5, wherein the value of the variable is set to be equal to 0 when i) the flag of the value of the identifier specifying the VPS is greater than 0 and the current layer is not an output layer, or ii) one of the following conditions is true:
[0612] The current picture is a random access skipped previous picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the intra random access point picture is recovered; or
[0613] The current picture is a progressive decoded refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is recovered, or the current picture is a recovery picture of a progressive decoded refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is recovered.
[0614] 7. The method according to clause 6, wherein the value of the variable is set to be equal to the value of the flag when neither i) nor ii) is satisfied.
[0615] 8. The method according to clause 6, wherein a value of 0 of the variable indicates not to output the picture in the access unit.
[0616] 9. The method according to clause 1, wherein the variable is PictureOutputFlag, and the flag is pic_output_flag.
[0617] 10. The method according to any one of clauses 1 to 9, wherein the flag is included in the access unit delimiter.
[0618] 11. The method according to any one of clauses 1 to 9, wherein the flag is included in the supplementary enhancement information field.
[0619] 12. The method according to any one of clauses 1 to 9, wherein the flag is included in the picture header of one or more pictures.
[0620] 13. A method for video processing (e.g., the method 920 as shown in Figure 9B ) includes: performing conversion 922 between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when the picture in the access unit does not belong to the output layer, the value of the variable indicating whether to output the picture is set to be equal to a certain value.
[0621] 14. The method according to clause 13, wherein the certain value is 0.
[0622] 15. A method for video processing (e.g., the method 930 as shown in Figure 9C ) includes: performing conversion 932 between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when the video includes only one output layer, the access unit not including the output layer is encoded and decoded by setting the variable indicating whether to output the picture in the access unit to a first value for the picture having the highest layer ID (identification) value and a second value for all other pictures.
[0623] 16. The method according to clause 15, wherein the first value is 1 and the second value is 0.
[0624] 17. The method according to any one of clauses 1 to 16, wherein the conversion includes encoding the video into a bitstream.
[0625] 18. The method according to any one of clauses 1 to 16, wherein the conversion includes decoding the video from the bitstream.
[0626] 19. The method according to clauses 1 to 16, wherein the conversion includes generating a bitstream from a video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.
[0627] 20. A video processing apparatus, including a processor configured to implement the method according to any one or more of clauses 1 to 19.
[0628] 21. A method for storing a bitstream of a video, including the method according to any one of clauses 1 to 19, and further including storing the bitstream in a non-transitory computer-readable recording medium.
[0629] 22. A computer-readable medium storing program code that, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 19.
[0630] 23. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0631] 24. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 19.
[0632] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from the pixel representation of a video to the corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits that are co-located or scattered in different places within the bitstream. For example, a macroblock may be encoded according to the transform and coding / decoding error residual values and also using bits in the headers and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on this determination, knowing whether some fields may be present or absent, as described in the above solution. Similarly, the encoder may determine whether to include or exclude a particular syntax field and generate the coded / decoded representation accordingly by including the syntax field or excluding the syntax field from the coded / decoded representation.
[0633] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal that is generated to encode information for transmission to a suitable receiver apparatus, e.g., a machine-generated electrical, optical, or electromagnetic signal.
[0634] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers distributed across one site or multiple sites and interconnected by a communication network.
[0635] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0636] Processors suitable for executing computer programs include, for example, any one or more processors of general and special purpose microprocessors, as well as any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to receive data from or transfer data to or receive data from and transfer data to the one or more mass storage devices. However, a computer does not necessarily require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0637] Although this patent document contains many details, these details should not be construed as limitations on any subject or the scope that may be claimed, but rather as descriptions of features of particular embodiments specific to a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variations of a sub-combination.
[0638] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as required in all embodiments.
[0639] Only some embodiments and examples have been described, and other embodiments, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: Performing conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag indicating whether to remove a previously decoded and stored picture in a decoded picture buffer when decoding a certain type of access unit is included in the bitstream; wherein the format rules further specify that when a first condition related to a maximum allowable value of a chroma format indicator or a maximum allowable value of a bit depth for representing pixels of the video is satisfied, a value of a variable indicating whether a picture in the decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is set to a preset value; wherein when the first condition is not satisfied, the value of the variable is set to the value of the flag; wherein when the chroma format of a picture or the bit depth of a picture changes, but the maximum allowable value of the chroma format indicator or the maximum allowable value of the bit depth does not change, the first condition is not satisfied.
2. The method according to claim 1, wherein The format rules specify that the value of the flag for the access unit is the same for all pictures in the access unit.
3. The method according to claim 1, wherein, The format rules specify that the value of the flag for the access unit is the same for all slices in the access unit.
4. The method according to claim 1, wherein The format rules specify that the value of the flag is considered as the value of the flag of the picture / slice of the access unit.
5. The method according to claim 1, wherein The flag is indicated in a picture header.
6. The method according to claim 1, wherein, The flag is indicated in a slice header.
7. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
8. The method according to claim 1, wherein, The conversion includes decoding the video from the bitstream.
9. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform conversion between a video having one or more video layers and a bitstream of the video according to format rules, and Among them, The format rules specify that a value of a flag indicating whether to remove a previously decoded and stored picture in a decoded picture buffer when decoding a certain type of access unit is included in the bitstream; wherein the format rules further specify that when a first condition related to a maximum allowable value of a chroma format indicator or a maximum allowable value of a bit depth for representing pixels of the video is satisfied, a value of a variable indicating whether a picture in the decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is set to a preset value; wherein when the first condition is not satisfied, the value of the variable is set to the value of the flag; wherein when the chroma format of a picture or the bit depth of a picture changes, but the maximum allowable value of the chroma format indicator or the maximum allowable value of the bit depth does not change, the first condition is not satisfied.
10. The apparatus according to claim 9, wherein, The format rule specifies that the value of the flag of the access unit is the same for all pictures in the access unit.
11. The device according to claim 9, wherein, The format rule specifies that the value of the flag of the access unit is the same for all slices in the access unit.
12. The device according to claim 9, wherein, The format rule specifies that the value of the flag is considered as the value of the flag of the picture / slice of the access unit.
13. The device according to claim 9, wherein The flag is indicated in the picture header.
14. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform conversion between a video having one or more video layers and a bitstream of the video according to a format rule, and Among them, the format rule specifies that the value of a flag indicating whether to remove a previously decoded picture stored in the decoded picture buffer from the decoded picture buffer when decoding a certain type of access unit is included in the bitstream; wherein the format rule further specifies that when a first condition related to a maximum allowable value of a chroma format indicator or a maximum allowable value of a bit depth for pixels representing the video is satisfied, the value of a variable indicating whether a picture in the decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is set to a preset value; wherein when the first condition is not satisfied, the value of the variable is set to the value of the flag; wherein when the chroma format of a picture or the bit depth of a picture changes, but the maximum allowable value of the chroma format indicator or the maximum allowable value of the bit depth does not change, the first condition is not satisfied.
15. The medium according to claim 14, wherein, The format rule specifies that the value of the flag of the access unit is the same for all slices in the access unit.
16. The medium according to claim 14, wherein, The format rule specifies that the value of the flag is considered as the value of the flag of the picture / slice of the access unit.
17. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: generating the bitstream of the video having one or more video layers according to a format rule, and wherein the format rule specifies that the value of a flag indicating whether to remove a previously decoded picture stored in the decoded picture buffer from the decoded picture buffer when decoding a certain type of access unit is included in the bitstream; wherein the format rule further specifies that when a first condition related to a maximum allowable value of a chroma format indicator or a maximum allowable value of a bit depth for pixels representing the video is satisfied, the value of a variable indicating whether a picture in the decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is set to a preset value; wherein when the first condition is not satisfied, the value of the variable is set to the value of the flag; wherein when the chroma format of a picture or the bit depth of a picture changes, but the maximum allowable value of the chroma format indicator or the maximum allowable value of the bit depth does not change, the first condition is not satisfied.
18. A method for storing a bitstream of a video, comprising: generating the bitstream of the video having one or more video layers according to format rules; and storing the bitstream in a non-transitory computer-readable recording medium, and wherein the format rules specify that a value of a flag indicating whether to remove a previously decoded and stored picture in a decoded picture buffer when decoding a certain type of access unit is included in the bitstream; wherein the format rules further specify that when a first condition related to a maximum allowable value of a chrominance format indicator or a maximum allowable value of a bit depth for representing pixels of the video is satisfied, a value of a variable indicating whether a picture in the decoded picture buffer before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is set to a preset value; wherein when the first condition is not satisfied, the value of the variable is set to the value of the flag; wherein when the chrominance format of a picture or the bit depth of a picture changes, but the maximum allowable value of the chrominance format indicator or the maximum allowable value of the bit depth does not change, the first condition is not satisfied.
Citation Information
Patent Citations
Device and method for scalable coding of video information
CN105637862A
Methods and Systems for Signaling Multi-Layer Bitstream Data
US20080007438A1