Picture Output Flag Indication in Video Coding and Decoding

By introducing format rules to control the image output of the decoder buffer in video encoding and decoding, the resource efficiency problem in multi-layer video management is solved, and low-latency and high-efficiency video processing is achieved.

CN115315942BActive Publication Date: 2025-07-25DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180022454.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-17
Filing Date
2021-03-16
Publication Date
2025-07-25
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies process multi-layer video, it is difficult to effectively control the image output in the decoder buffer, resulting in inefficient resource management and unable to meet the needs of modern video applications for low latency and high efficiency.

Method used

By introducing format rules, the output of the picture in the decoder buffer is controlled, including parameters such as maximum image width and height, bit depth, chroma format indicator, etc., and combined with the picture output flag in the access unit, determine whether the picture is output, and realize efficient management of multi-layer video.

Benefits of technology

Improve the resource management efficiency of decoder buffers, meet the needs of modern video applications for low latency and high efficiency, and is suitable for multi-layer video encoding and decoding standards such as VVC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115315942B_ABST
    Figure CN115315942B_ABST
Patent Text Reader

Abstract

Methods and apparatuses for video processing including encoding and decoding are described. An example video processing method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that a value of a variable indicating whether to output a picture in an access unit is determined based on a flag indicating whether to output another picture in the access unit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is being filed to claim the priority and benefit of U.S. Provisional Patent Application No. 62 / 990,749, filed on March 17, 2020, under the patent laws and / or rules applicable under the Paris Convention. For all purposes under the law, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders for processing encoded - decoded representations of video using control information useful for decoding the encoded - decoded representations.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers including one or more video pictures and an encoded - decoded representation of the video; wherein the encoded - decoded representation includes a video parameter set that indicates a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels representing the video.

[0007] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and an encoded - decoded representation of the video, wherein the encoded - decoded representation conforms to a format rule that specifies a maximum picture width and / or a maximum picture height of video pictures of all video layers and a value of a variable that controls whether a picture in a decoder buffer is output before being removed from the decoder buffer.

[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and an encoded - decoded representation of the video, wherein the encoded - decoded representation conforms to a format rule that specifies a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels representing the video and a value of a variable that controls whether a picture in a decoder buffer is output before being removed from the decoder buffer.

[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to a format rule that specifies that the value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is independent of whether separate color planes are used for encoding the video.

[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to a format rule that specifies that the value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is included in the coded representation at an access unit (AU) level.

[0011] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to a format rule that specifies that a picture output flag of a video picture in an access unit is determined based on a pic_output_flag variable of another video picture in the access unit.

[0012] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a coded representation of the video, where the coded representation conforms to a format rule that specifies a value of a picture output flag for a video picture that does not belong to an output layer.

[0013] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to a format rule, where the bitstream includes one or more output layer sets (OLSs), each OLS including one or more coded layer video sequences, and where the format rule specifies that a video parameter set indicates, for each of the one or more OLSs, a maximum allowed value of a chroma format indicator for pixels representing the video and / or a maximum allowed value of a bit depth.

[0014] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to a format rule, and where the format rule specifies that the maximum picture width and / or maximum picture height of video pictures of all video layers controls the value of a variable indicating whether a picture in a decoded picture buffer that is decoded before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer.

[0015] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify a maximum allowable value of a chroma format indicator for representing pixels of the video and / or a maximum allowable value of a bit depth, and a value of a variable that controls whether a picture in a decoded picture buffer that is prior to a current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer.

[0016] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to rules, and wherein the rules specify that a value of a variable that controls whether a picture in a decoded picture buffer that is prior to a current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is independent of whether separate color planes are used to encode the video.

[0017] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer when decoding a certain type of access unit is included in the bitstream.

[0018] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a first flag that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer when decoding a particular type of access unit is not indicated in a picture header.

[0019] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag associated with an access unit that indicates whether a previously decoded picture stored in the decoded picture buffer is removed from the decoded picture buffer depends on a value of a flag of each picture of the access unit.

[0020] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a variable that indicates whether a picture in an access unit is output is determined based on a flag that indicates whether another picture in the access unit is output.

[0021] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that when a picture in an access unit does not belong to an output layer, the value of a variable indicating whether to output the picture is set to be equal to a certain value.

[0022] In another exemplary aspect, another video processing method is disclosed. The method includes performing a conversion between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that when the video includes only one output layer, access units that do not include the output layer are encoded and decoded by setting a variable indicating whether to output a picture in the access unit to a first value for a picture having the highest layer ID (identification) value and a second value for all other pictures.

[0023] In yet another exemplary aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.

[0024] In yet another exemplary aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.

[0025] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0026] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a block diagram of an exemplary video processing system.

[0028] Figure 2 is a block diagram of a video processing apparatus.

[0029] Figure 3 is a flowchart of an exemplary method of video processing.

[0030] Figure 4 is a block diagram illustrating a video codec system according to some embodiments of the present disclosure.

[0031] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0032] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0033] Figures 7A to 7D shows a flowchart of an exemplary method of video processing based on some implementations of the disclosed technology.

[0034] Figures 8A to 8C A flowchart showing an example method of video processing based on some embodiments of the disclosed technology.

[0035] Figures 9A to 9C A flowchart showing an example method of video processing based on some embodiments of the disclosed technology. Detailed Description

[0036] The section headings used in this document are for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each section to that section only. Additionally, the use of H.266 terms in some descriptions is for ease of understanding only and does not limit the scope of the disclosed technology. Thus, the technologies described herein are also applicable to other video codec protocols and designs.

[0037] 1. Preliminary Discussion

[0038] This patent document relates to video codec technology. Specifically, it is about signaling of decoding picture buffer (DPB) parameters for DPB memory allocation and specifying the output of decoded pictures in scalable video coding, where the video bitstream may contain more than one layer. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Coding (VVC) being developed.

[0039] 2. Abbreviations

[0040] APS Adaptive Parameter Set

[0041] AU Access Unit

[0042] AUD Access Unit Delimiter

[0043] AVC Advanced Video Coding

[0044] CLVS Coded Layer Video Sequence

[0045] CPB Coded Picture Buffer

[0046] CRA Completely Random Access

[0047] CTU Coding Tree Unit

[0048] CVS Coded Video Sequence

[0049] DCI Decoding Capability Information

[0050] DPB Decoding Picture Buffer

[0051] EOB End of Bitstream

[0052] End of EOS sequence

[0053] GDR Gradual Decoding Refresh

[0054] HEVC High Efficiency Video Coding

[0055] HRD Hypothetical Reference Decoder

[0056] IDR Instantaneous Decoding Refresh

[0057] JEM Joint Exploration Model

[0058] MCTS Motion Constrained Tile Set

[0059] NAL Network Abstraction Layer

[0060] OLS Output Layer Set

[0061] PH Picture Header

[0062] PPS Picture Parameter Set

[0063] PTL Profile, Tier and Level

[0064] PU Picture Unit

[0065] RAP Random Access Point

[0066] RBSP Raw Byte Sequence Payload

[0067] SEI Supplemental Enhancement Information

[0068] SPS Sequence Parameter Set

[0069] SVC Scalable Video Coding

[0070] VCL Video Coding Layer

[0071] VPS Video Parameter Set

[0072] VTM VVC Test Model

[0073] VUI Video Usability Information

[0074] VVC Versatile Video Coding

[0075] 3. Introduction to Video Coding

[0076] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, in which temporal prediction plus transform coding is employed. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly simultaneously. The goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts on VVC standardization, new coding technologies have been adopted into the VVC standard at each JVET meeting. The working draft and test model VTM of VVC are updated after each meeting. The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.

[0077] 3.1. Overview and Scalable Video Coding (SVC) in VVC

[0078] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding that uses a Base Layer (BL) (sometimes referred to as a Reference Layer (RL)) and one or more Scalable Enhancement Layers (ELs). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can act as the BL, while the top layer can act as the EL. Intermediate layers can act as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below it (such as the base layer or any intermediate enhancement layer) and at the same time act as an RL for one or more enhancement layers above it. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and the information of one view can be used to code (e.g., encode or decode) the information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).

[0079] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be utilized by one or more coded video sequences of different layers in a bitstream can be included in a Video Parameter Set (VPS), and the parameters that can be utilized by one or more pictures in a coded video sequence can be included in a Sequence Parameter Set (SPS). Similarly, the parameters utilized by one or more slices in a picture can be included in a Picture Parameter Set (PPS), and other parameters specific to a single slice can be included in the slice header. Similarly, an indication of which (which) parameter set a given layer uses at a given time can be provided at various coding levels.

[0080] Due to the support for reference picture resampling (RPR) in VVC, it is possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without any additional signal processing level codec tools, because the upsampling required for spatial scalability support can be achieved using only the RPR upsampling filter. However, for scalability support, high-level syntax changes are required (compared to non-scalability support). Scalability support is specified in VVC version 1. Different from the scalability support in any earlier video codec standards (including extensions of AVC and HEVC), the design of VVC scalability has been made as friendly as possible to single-layer decoder designs. The decoding capabilities of multi-layer bitstreams are specified in a way that is as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified independently of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not need much modification to be able to decode a multi-layer bitstream. Compared to the designs of multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, it is required that an IRAP AU contains pictures of each layer present in the CVS.

[0081] 3.2. Random Access and Its Support in HEVC and VVC

[0082] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast / multicast and multi-party video conferencing, search in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include randomly accessible points that are close together, which are usually intra-coded pictures, but can also be inter-coded pictures (e.g., in the case of progressive decoding refresh).

[0083] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the NAL unit header by NAL unit type. Three types of IRAP pictures are supported, namely Instantaneous Decoder Refresh (IDR), Complete Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference any pictures prior to the current Group of Pictures (GOP), which is traditionally referred to as a closed GOP random access point. CRA pictures are less restrictive by allowing a particular picture to reference pictures prior to the current GOP, where in the case of random access, all pictures are discarded. CRA pictures are traditionally referred to as open GOP random access points. BLA pictures typically result from the concatenation of two bitstreams or parts thereof at a CRA picture, e.g., during stream switching. To better enable the system to use IRAP pictures, a total of six different NAL units are defined to signal the attributes of IRAP pictures, which can be used to better match the stream access point types defined in the ISO Base Media File Format (ISOBMFF), which is used for random access support in HTTP-based Dynamic Adaptive Streaming over HTTP (DASH).

[0084] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with or the other type without associated RADL pictures), and one type of CRA picture. These are basically the same as in HEVC. The BLA picture type in HEVC is not included in VVC, mainly for two reasons: i) The basic functionality of BLA pictures can be achieved by a CRA picture plus the sequence NAL unit end, the presence of which indicates that the subsequent pictures start a new CVS in a single-layer bitstream. ii) During the development of VVC, it was desired to specify fewer NAL unit types than in HEVC, as indicated by using 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0085] Another key difference in random access support between VVC and HEVC is that GDR is supported in a more standardized way in VVC. In GDR, the decoding of the bitstream can start from an inter-coded picture. Although not the entire picture area can be correctly decoded at the beginning, after several pictures, the entire picture area will be correct. AVC and HEVC also support GDR by using recovery point SEI messages to signal GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. It is allowed for the CVS and the bitstream to start with GDR pictures. This means that it is allowed for the entire bitstream to contain only inter-coded pictures without a single intra-coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior of GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded stripes or blocks over multiple pictures, as opposed to intra-coding the entire picture, thus allowing a significant reduction in end-to-end latency, which is considered more important today as ultra-low latency applications such as wireless displays, online games, and drone-based applications become more popular.

[0086] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., the correctly decoded area) and the non-refreshed area at the picture between a GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, loop filtering across the boundary will not be applied, so there will be no decoding mismatch for some samples at or near the boundary. This can be useful when the application determines to display the correctly decoded area during the GDR process.

[0087] IRAP pictures and GDR pictures can be collectively referred to as random access point (RAP) pictures.

[0088] 3.3. Parameter Sets

[0089] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced starting from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.

[0090] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry picture-level header information that changes infrequently. With the SPS and PPS, the information that changes infrequently does not need to be repeated for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of the SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.

[0091] The VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0092] The APS is introduced to carry such picture-level or slice-level information that requires a significant number of bits to encode and decode, can be shared by multiple pictures, and can have a significant number of different variations in the sequence.

[0093] 3.4. Related Definitions in VVC

[0094] The related definitions in the latest VVC text (in JVET-Q2001-vE / v15) are as follows.

[0095] Associated IRAP picture (of a specific picture): The previous IRAP picture in decoding order (when present) has the same nuh_layer_id value as the specific picture.

[0096] Coded Video Sequence (CVS): A sequence of AUs consisting of CVSS AUs in decoding order, followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU.

[0097] Coded Video Sequence Start (CVSS) AU: An AU that has PUs for each layer in the CVS and the coded pictures in each PU are CLVSS pictures.

[0098] Gradual Decoding Refresh (GDR) AU: An AU in which the coded pictures in each current PU are GDR pictures.

[0099] Gradual Decoding Refresh (GDR) PU: A PU in which the coded pictures are GDR pictures.

[0100] Gradual Decoding Refresh (GDR) picture: A picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT.

[0101] Intra Random Access Point (IRAP) AU: An AU that has PUs for each layer in the CVS and the coded pictures in each PU are IRAP pictures.

[0102] Intra Random Access Point (IRAP) picture: A coded picture in which all VCL NAL units have the same nal_unit_type value in the range from IDR_W_RADL to CRA_NUT, inclusive of IDR_W_RADL and CRA_NUT.

[0103] Previous picture: A picture in the same layer as the associated IRAP picture and preceding the associated IRAP picture in output order.

[0104] Trailing picture: A non-IRAP picture that follows the associated IRAP picture in output order and is not a STSA picture.

[0105] Note – Trailing pictures associated with an IRAP picture also follow the IRAP picture in decoding order. A picture that follows the associated IRAP picture in output order and precedes the associated IRAP picture in decoding order is not allowed.

[0106] 3.5. VPS Syntax and Semantics in VVC

[0107] VVC supports scalability, also known as scalable video coding, in which multiple layers can be encoded in a single coded video bitstream.

[0108] In the latest VVC text (in JVET-Q2001-vE / v15), scalability information is signaled in the VPS, and its syntax and semantics are as follows.

[0109] 7.3.2.2 Video Parameter Set Syntax

[0110]

[0111]

[0112]

[0113]

[0114] 7.4.3.2 Video Parameter Set RBSP Semantics

[0115] The VPS RBSP shall be available for the decoding process before being referenced, including in at least one AU with a TemporalId equal to 0, or provided externally.

[0116] All VPS NAL units with a specific vps_video_parameter_set_id value in the CVS shall have the same content.

[0117] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.

[0118] vps_max_layers_minus1 plus 1 specifies the maximum number of allowed layers in each CVS that refers to the VPS.

[0119] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in the layers of each CVS that refers to the VPS. The value of vps_max_sublayers_minus1 shall be in the range from 0 to 6 (including the end values).

[0120] vps_all_layers_same_num_sublayers_flag being equal to 1 specifies that the number of temporal sublayers is the same for all layers in each CVS that refers to the VPS. vps_all_layers_same_num_sub layers_flag being equal to 0 specifies that the layers in each CVS that refers to the VPS may or may not have the same number of temporal sublayers. When not present, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.

[0121] vps_all_independent_layers_flag being equal to 1 specifies that all layers in the CVS are independently encoded and decoded without using inter-layer prediction. vps_all_independent_layers_flag being equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. When not present, the value of vps_all_independent_layers_flag is inferred to be equal to 1.

[0122] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] shall be less than the value of vps_layer_id[n].

[0123] vps_independent_layer_flag[i] being equal to 1 specifies that the layer at index i does not use inter-layer prediction. vps_independent_layer_flag[i] being equal to 0 specifies that the layer at index i can use inter-layer prediction and that there is a syntax element vps_direct_ref_layer_flag[i][j] in the vps with j in the range from 0 to i–1 (inclusive of the end values). When it does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1.

[0124] vps_direct_ref_layer_flag[i][j] being equal to 0 specifies that the layer at index j is not a direct reference layer of the layer at index i. vps_direct_ref_layer_flag[i][j] being equal to 1 specifies that the layer at index j is a direct reference layer of the layer at index i. When vps_direct_ref_layer_flag[i][j] does not exist with i and j in the range from 0 to vps_max_layers_minus1 (inclusive of the end values), it is inferred to be equal to 0. When vps_independent_layer_flag[i] is equal to 0, there should be at least one j value in the range from 0 to i - 1 (inclusive of the end values) such that the value of vps_direct_ref_layer_flag[i][j] is equal to 1.

[0125] The variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] are derived as follows:

[0126]

[0127]

[0128] The variable GeneralLayerIdx[i] specifies the layer index of the layer with nuh_layer_id equal to vps_layer_id[i] and is derived as follows:

[0129] for(i = 0; i <= vps_max_layers_minus1; i++) (38)

[0130] GeneralLayerIdx[vps_layer_id[i]] = i

[0131] For any two different values of i and j (both within the range from 0 to vps_max_layers_minus1, inclusive), when dependencyFlag[i][j] equals 1, the requirement for bitstream consistency is that the values of chroma_format_idc and bit_depth_minus8 applicable to layer i should be equal to the values of chroma_format_idc and bit_depth_minus8 applicable to layer j, respectively.

[0132] max_tid_ref_present_flag[i] being equal to 1 specifies the presence of the syntax element max_tid_il_ref_pics_plus1[i]. max_tid_ref_present_flag[i] being equal to 0 specifies the absence of the syntax element max_tid_il_ref_pics_plus1[i].

[0133] max_tid_il_ref_pics_plus1[i] being equal to 0 specifies that inter-layer prediction is not used for non-IRAP pictures of layer i. max_tid_il_ref_pics_plus1[i] being greater than 0 specifies that, for decoding pictures of layer i, no pictures with a TemporalId greater than max_tid_il_ref_pics_plus1[i] - 1 are used as ILRP. When not present, the value of max_tid_il_ref_pics_plus1[i] is inferred to be equal to 7.

[0134] each_layer_is_an_ols_flag being equal to 1 specifies that each OLS contains only one layer, and each layer in the CVS of the reference VPS itself is an OLS, where the single included layer is the only output layer. each_layer_is_an_ols_flag being equal to 0 specifies that an OLS may contain multiple layers. If vps_max_layers_minus1 equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.

[0135] ols_mode_idc being equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes layers with layer indices from 0 to i (inclusive), and for each OLS, only the highest layer in the OLS is output.

[0136] ols_mode_idc being equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1. The i-th OLS includes the layers with layer indices from 0 to i (including the end values), and for each OLS, all the layers in the output OLS are output.

[0137] ols_mode_idc being equal to 2 specifies that the total number of OLSs specified by the VPS is signaled explicitly, and for each OLS, the output layers are signaled explicitly, and the other layers are the layers that are direct or indirect reference layers of the output layers of the OLSs.

[0138] The value of ols_mode_idc shall be in the range of 0 to 2 (including the end values). The value 3 of ols_mode_idc is reserved by ITU-T|ISO / IEC for future use.

[0139] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.

[0140] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.

[0141] The variable TotalNumOlss specifies the total number of OLSs specified by the VPS and is derived as follows:

[0142]

[0143] ols_output_layer_flag[i][j] being equal to 1 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS when ols_mode_idc is equal to 2. ols_output_layer_flag[i][j] being equal to 0 specifies that the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS when ols_mode_idc is equal to 2.

[0144] The variable NumOutputLayersInOls[i] specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerInOLS[i][j] specifies the number of sub-layers in the j-th layer in the i-th OLS, the variable OutputLayerIdInOls[i][j] specifies the nuh_layer_id value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k] specifies whether the k-th layer is used as an output layer in at least one OLS, as derived below:

[0145]

[0146]

[0147]

[0148] For each i value in the range from 0 to vps_max_layers_minus1 (including the end values), the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] should not both be equal to 0. In other words, there should not be a layer that is neither an output layer of at least one OLS nor a direct reference layer of any other layer.

[0149] For each OLS, there should be at least one layer as an output layer. In other words, for any i value in the range from 0 to TotalNumOlss–1 (including the end values), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.

[0150] The variable NumLayersInOls[i] specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] specifies the nuh_layer_id value of the j-th layer in the i-th OLS, as derived below:

[0151]

[0152]

[0153] Note 1 – The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, only the contained layer is output.

[0154] The variable OlsLayerIdx[i][j] specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j], as derived below:

[0155] for(i = 0; i < TotalNumOlss; i++)

[0156] for j = 0; j < NumLayersInOls[i]; j++) (42)

[0157] OlsLayerIdx[i][LayerIdInOls[i][j]] = j

[0158] The lowest layer in each OLS should be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss–1 (including the end values), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] should be equal to 1.

[0159] Each layer should be included in at least one OLS specified by the VPS. In other words, for each layer where the specific value of nuh_layer_id nuhLayerId is equal to one of vps_layer_id[k] in the range from 0 to vps_max_layers_minus1 (including the end values), there should be at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 (including the end values), and j is in the range from 0 to NumLayersInOls[i] - 1 (including the end values), such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.

[0160] vps_num_ptls_minus1 plus 1 specifies the number of profile_tier_level() syntax structures in the VPS. The value of vps_num_ptls_minus1 should be less than TotalNumOlss.

[0161] pt_present_flag[i] being equal to 1 specifies the tier, layer, and general constraint information present in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] being equal to 0 specifies that the tier, layer, and general constraint information is not present in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (i - 1)-th profile_tier_level() syntax structure in the VPS.

[0162] ptl_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation, the level information of which is present in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[i] shall be in the range (including the end values) from 0 to vps_max_sublayers_minus1. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0163] vps_ptl_alignment_zero_bit shall be equal to 0.

[0164] ols_ptl_idx[i] specifies the index (index into the list of profile_tier_level() syntax structures in the VPS) of the profile_tier_level() syntax structure applied to the i-th OLS. When present, the value of ols_ptl_idx[i] shall be in the range (including the end values) from 0 to vps_num_ptls_minus1. When vps_num_ptls_minus1 is equal to 0, the value of ols_ptl_idx[i] is inferred to be equal to 0.

[0165] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS also exists in the SPS referred to by the layer in the i-th OLS. The requirement for bitstream consistency is that when NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure signaled for the i-th OLS in the VPS and SPS should be the same.

[0166] vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params shall be in the range of 0 to 16 (inclusive). When not present, the value of vps_num_dpb_params is inferred to be equal to 0.

[0167] vps_sublayer_dpb_params_present_flag is used to control the presence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure in the VPS. When not present, vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.

[0168] dpb_max_temporal_id[i] specifies the highest sublayer-represented TemporalId for which DPB parameters may be present in the i-th dpb_parameters() syntax structure in the VPS. The value of dpb_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1 (inclusive). When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0169] ols_dpb_pic_width[i] specifies the width in luma samples of each picture storage buffer for the i-th OLS.

[0170] ols_dpb_pic_height[i] specifies the height of each picture storage buffer of the i-th OLS in units of luma samples.

[0171] ols_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure that applies to the i-th OLS when NumLayersInOls[i] is greater than 1 (index into the list of dpb_parameters() syntax structures in the VPS). When present, the value of ols_dpb_params_idx[i] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[i] is not present, the value of ols_dpb_params_idx[i] is inferred to be equal to 0.

[0172] When NumLayersInOls[i] is equal to 1, the dpb_parameters() syntax structure that applies to the i-th OLS is present in the SPS referenced by the layers in the i-th OLS.

[0173] vps_general_hrd_params_present_flag equal to 1 specifies that the VPS contains a general_hrd_parameters() syntax structure and other HRD parameters. vps_general_hrd_params_present_flag equal to 0 specifies that the VPS does not contain a general_hrd_parameters() syntax structure or other HRD parameters. When not present, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.

[0174] When NumLayersInOls[i] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS are present in the SPS referenced by the layer in the i-th OLS.

[0175] The vps_sublayer_cpb_params_present_flag being equal to 1 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented by sublayers, where the TemporalId ranges from 0 to hrd_max_tid[i] (including the end values). The vps_sublayer_cpb_params_present_flag being equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented by sublayers, where the TemporalId is only equal to hrd_max_tid[i]. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.

[0176] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented by sublayers with TemporalId in the range from 0 to hrd_max_tid[i] – 1 (including the end values) are inferred to be the same as the HRD parameters represented by the sublayer with TemporalId equal to hrd_max_tid[i]. These include the HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element until the sublayer_hrd_parameters(i) syntax structure immediately following under the condition "if(general_vcl_hrd_params_present_flag)" in the ols_hrd_parameters syntax structure.

[0177] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the VPS when vps_general_hrd_params_present_fla is equal to 1. The value of num_ols_hrd_params_minus1 shall be in the range from 0 to TotalNumOlss - 1 (including the end values).

[0178] hrd_max_tid[i] specifies the TemporalId represented by the highest sublayer in which the HRD parameter is included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range (including the end values) from 0 to vps_max_sublayers_minus1. When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[i] is inferred to be equal to vps_max_sublayers_minus1.

[0179] ols_hrd_idx[i] specifies the index (index into the list of ols_hrd_parameters() syntax structures in the VPS) of the ols_hrd_parameters() syntax structure applied to the i-th OLS when NumLayersInOls[i] is greater than 1. The value of ols_hrd_idx[[i] shall be in the range (including the end values) from 0 to num_ols_hrd_params_minus1.

[0180] When NumLayersInOls[i] is equal to 1, the ols_hrd_parameters() syntax structure applied to the i-th OLS is present in the SPS referenced by the layer in the i-th OLS.

[0181] If the value of num_ols_hrd_param_minus1 + 1 is equal to TotalNumOlss, the value of ols_hrd_idx[i] is inferred to be equal to i. Otherwise, when NumLayersInOls[i] is greater than 1 and num_ols_hrd_params_minus1 is equal to 0, the value of ols_hrd_idx[[i] is inferred to be equal to 0.

[0182] vps_extension_flag being equal to 0 specifies that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. vps_extension_flag being equal to 1 specifies that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure.

[0183] The vps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of this specification. A decoder compliant with this version of this specification shall ignore all vps_extension_data_flag syntax elements.

[0184] 3.6. SPS Syntax and Semantics in VVC

[0185] In the latest VVC text (in JVET-Q2001-vE / v15), the SPS syntax and semantics most relevant to the invention herein are as follows.

[0186] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0187]

[0188] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...

[0190] When gdr_enabled_flag is equal to 1, it specifies that GDR pictures may exist in the CLVS of the reference SPS. When gdr_enabled_flag is equal to 0, it specifies that GDR pictures do not exist in the CLVS of the reference SPS.

[0191] chroma_format_idc specifies the chroma sampling related to the luma sampling specified in Clause 6.2. ...

[0193] bit_depth_minus8 specifies the bit depth BitDepth of the samples in the luma and chroma arrays and the value of the luma and chroma quantization parameter range offset QpBdOffset as follows:

[0194] BitDepth = 8 + bit_depth_minus8 (45)

[0195] QpBdOffset = 6 * bit_depth_minus8 (46)

[0196] bit_depth_minus8 shall be in the range from 0 to 8 (including the end values). ...

[0198] 3.7. Picture Header Structure Syntax and Semantics in VVC

[0199] In the latest VVC text (in JVET-Q2001-vE / v15), the picture header structure syntax and semantics most relevant to the invention herein are as follows.

[0200] 7.3.2.7 Picture Header Structure Syntax

[0201]

[0202]

[0203] 7.4.3.7 Picture Header Structure Semantics

[0204] The PH syntax structure contains information common to all slices of the coded picture associated with the PH syntax structure.

[0205] A gdr_or_irap_pic_flag equal to 1 specifies that the current picture is a GDR or IRAP picture. A gdr_or_irap_pic_flag equal to 0 specifies that the current picture may or may not be a GDR or IRAP picture.

[0206] A gdr_pic_flag equal to 1 specifies that the picture associated with the PH is a GDR picture. A gdr_pic_flag equal to 0 specifies that the picture associated with the PH is not a GDR picture. When not present, the value of gdr_pic_flag is inferred to be equal to 0. When gdr_enabled_flag is equal to 0, the value of gdr_pic_flag shall be equal to 0.

[0207] Note 1 – When gdr_or_irap_pic_flag is equal to 1 and gdr_pic_flag is equal to 0, the picture associated with the PH is an IRAP picture. ...

[0209] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the ph_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of ph_pic_order_cnt_lsb shall be in the range (including the end points) of 0 to MaxPicOrderCntLsb-1.

[0210] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a CLVSS picture that is not the first picture in the bitstream specified in Annex C.

[0211] The recovery_poc_cnt specifies the recovery points of decoded pictures in the output order. If the current picture is a GDR picture associated with PH and there is a picture picA in the CLVS that follows the current GDR picture in the decoding order and whose PicOrderCntVal is equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then the picture picA is called a recovery point picture. Otherwise, the first picture in the output order whose PicOrderCntVal is greater than the PicOrderCntVal of the current picture plus the value of recovery_poc_cnt is called a recovery point picture. The recovery point picture shall not be before the current GDR picture in the decoding order. The value of recovery_poc_cnt shall be in the range of 0 to MaxPicOrderCntLsb-1 (including the end values).

[0212] When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0213] RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81)

[0214] Note 2--When the gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to the RpPicOrderCntVal of the relevant GDR picture, the current and subsequent decoded pictures in the output order exactly match the corresponding pictures generated by starting the decoding process from the previous IRAP picture, and when present, are before the relevant GDR picture in the decoding order. ...

[0216] 3.8. Setting PictureOutputFlag

[0217] In the latest VVC text (in JVET-Q2001-vE / v15), the specification for setting the value of the variable PictureOutputFlag is as follows (as part of Clause 8.1.2 Decoding Process for coded pictures).

[0218] 8.1.2 Decoding Process for Coded Pictures

[0219] The decoding process specified in this clause applies to each coded picture in BitstreamToDecode (referred to as the current picture and represented by the variable CurrPic).

[0220] According to the value of chroma_format_idc, the number of sample arrays of the current picture is as follows:

[0221] -- If chroma_format_idc is equal to 0, the current picture consists of 1 sample array S L and is composed of.

[0222] -- Otherwise (chroma_format_idc is not equal to 0), the current picture consists of 3 sample arrays S L , S Cb , S Cr and is composed of.

[0223] The decoding process of the current picture takes the syntax elements and uppercase variables from Clause 7 as input. When interpreting the semantics of each syntax element in each NAL unit and in the remainder of Clause 8, the term "bitstream" (or a part thereof, such as the CVS of the bitstream) refers to BitstreamToDecode (or a part thereof).

[0224] According to the value of separate_colour_plane_flag, the structure of the decoding process is as follows:

[0225] -- If separate_colour_plane_flag is equal to 0, the decoding process is called once, with the current picture as the output.

[0226] -- Otherwise (separate_colour_plane_flag is equal to 1), the decoding process is called three times. The input to the decoding process is all the NAL units of the coded picture with the same colour_plane_id value. The decoding process for the NAL units with a specific colour_plane_id value is specified as if only the CVS of the monochrome colour format with that specific colour_plane_id value would be present in the bitstream. The output of each of the three decoding processes is assigned to one of the 3 sample arrays of the current picture, where the NAL units with colour_plane_id equal to 0, 1, and 2 are assigned to S L , S Cb , and S Cr respectively.

[0227] Note – When separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3, the variable ChromaArrayType is deduced to be equal to 0. During the decoding process, the value of this variable is evaluated, resulting in the same operations as for a monochrome picture (when chroma_format_idc is equal to 0).

[0228] For the current picture CurrPic, the decoding process operates as follows:

[0229] 1. The decoding of NAL units is specified in Clause 8.2.

[0230] 2. The process in Clause 8.3 uses the syntax elements in the slice header layer and above to specify the following decoding processes:

[0231] -- The variables and functions related to picture order counting are derived as specified in Clause 8.3.1. This only needs to be called for the first slice of the picture.

[0232] -- At the start of the decoding process for each slice of a non-IDR picture, the decoding process for reference picture list construction specified in Clause 8.3.2 is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0233] -- The decoding process for reference picture marking in Clause 8.3.3 is called, where a reference picture can be marked as "not used for reference" or "used for long-term reference". This only needs to be called for the first slice of the picture.

[0234] -- When the current picture is a CRA picture with NoOutputBeforeRecoveryFlag equal to 1 or a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, the decoding process for generating unavailable reference pictures specified in Subclause 8.3.4 is called, and this only needs to be called for the first slice of the picture.

[0235] -- PictureOutputFlag is set as follows:

[0236] -- If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:

[0237] -- The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0238] -- gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0239] -- The gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture where NoOutputBeforeRecoveryFlag is equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0240] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0241] -- The PictureOutputFlag of PicA is equal to 1.

[0242] -- The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.

[0243] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).

[0244] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0245] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag.

[0246] 3. The processing in Clauses 8.4, 8.5, 8.6, 8.7, and 8.8 specifies the decoding process using syntax elements in all syntax structure layers. The requirement for bitstream conformance is that the coded and decoded slices of a picture will contain the slice data for each CTU of the picture, such that the partitioning of the picture into slices and the partitioning of a slice into CTUs each form a partitioning of the picture.

[0247] 4. After all slices of the current picture have been decoded, the current decoded picture is marked as "for short-term reference", and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as "for short-term reference".

[0248] 3.9. Set DPB parameters for HRD operations

[0249] In the latest VVC text (in JVET-Q2001-vE / v15), the specification of the setting of DPB parameters for HRD operation is as follows (as part of Clause C.1).

[0250] C.1 General ...

[0252] For each bitstream conformance test, the CPB size (number of bits) is CpbSize[Htid][ScIdx] as specified in Clause 7.4.6.3, where ScIdx and HRD parameters are specified in that clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found or derived from the dpb_parameters() syntax structure applied to the target OLS as follows:

[0253] -- If the target OLS contains only one layer, the dpb_parameters() syntax structure is found in the SPS, which is referred to as the layer in the target OLS.

[0254] -- Otherwise (the target OLS contains more than one layer), the dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsIdx] found in the VPS. ...

[0256] 3.10. Setting of NoOutputOfPriorPicsFlag

[0257] In the latest VVC text (in JVET-Q2001-vE / v15), the specification of the setting of the value of the variable NoOutputOfPriorPicsFlag is as follows (as part of the specification of removing pictures from the DPB).

[0258] C.3.2 Removing pictures from the DPB before decoding the current picture

[0259] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (containing the current picture) and is done as follows:

[0260] -- Call the decoding process of reference picture list construction specified in Clause 8.3.2 and call the decoding process of reference picture marking specified in Clause 8.3.3.

[0261] --When the current AU is a CVSS AU other than AU 0, the following ordered steps are applied:

[0262] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:

[0263] --If the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for any picture in the current AU are different from the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous picture in the same CLVS, respectively, then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag.

[0264] Note--Although it is preferable to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag under these conditions, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1 in this case.

[0265] --Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag.

[0266] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set equal to 0.

[0267] --When the following two conditions are true for any picture k in the DPB, all such pictures k in the DPB are removed from the DPB:

[0268] --Picture k is marked as "not used for reference".

[0269] --The PictureOutputFlag of picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current picture n; that is, DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].

[0270] --For each picture removed from the DPB, the DPB fullness is decremented by 1.

[0271] C.5.2.2 Output and removal of pictures from the DPB

[0272] Output and removal of pictures from the DPB occur instantaneously when the first DU of the AU containing the current picture is removed from the CPB (but after parsing the slice header of the first slice of the current picture) and are as follows:

[0273] --Call the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.

[0274] --If the current picture is a CLVSS picture other than picture 0, the following ordered steps are applied:

[0275] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:

[0276] -- If the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for any picture of the current AU are different from the values of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous picture in the same CLVS, respectively, then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag.

[0277] Note -- Although it is preferred that NoOutputOfPriorPicsFlag be set to be equal to no_output_of_prior_pics_flag under these conditions, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1 in this case.

[0278] -- Otherwise, NoOutputOfPriorPicsFlag is set to be equal to no_output_of_prior_pics_flag.

[0279] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD as follows:

[0280] -- If NoOutputOfPriorPicsFlag is equal to 1, then all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set equal to 0.

[0281] -- Otherwise (NoOutputOfPriorPicsFlag equals 0), empty (no output) all picture storage buffers containing pictures marked as "not needed for output" and "not used for reference" and empty all non-empty picture storage buffers in the DPB by repeatedly calling the "collision" process specified in Clause C.5.2.4, and set the DPB fullness to be equal to 0.

[0282] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not needed for output" and "not used for reference" are emptied (no output). For each picture storage buffer that is emptied, the DPB fullness is decremented by 1. When one or more of the following conditions are true, repeatedly call the "collision" process specified in Clause C.5.2.4, and for each additional picture storage buffer that is emptied, further decrement the DPB fullness by 1 until none of the following conditions are true:

[0283] -- The number of pictures marked as "needed for output" in the DPB is greater than max_num_reorder_pics[Htid].

[0284] -- max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "needed for output" and its associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].

[0285] -- The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1.

[0286] 4. Technical problems solved by the disclosed technical solution

[0287] The existing scalability designs in the latest VVC text (in JVET-Q2001-vE / v15) have the following problems:

[0288] 1) Currently, the maximum values of the picture width and height of all pictures of all layers are signaled in the VPS so that the decoder can correctly allocate memory for the DPB. Like the picture width and height, the chroma format and bit depth, currently specified by the SPS syntax elements chroma_format_idc and bit_depth_minus8 respectively, also affect the size of the picture storage buffers in the DPB. However, the maximum values of chroma_format_idc and bit_depth_minus8 of all pictures of all layers are not signaled.

[0289] 2) Currently, the setting of the value of the variable NoOutputOfPriorPicsFlag involves a change in the value of pic_width_max_in_luma_samples or pic_height_max_in_luma_samples. However, it should be changed to use the maximum values of the picture width and height of all pictures of all layers.

[0290] 3) Currently, the setting of NoOutputOfPriorPicsFlag involves a change in the value of chroma_format_idc or bit_depth_minus8. However, it should be changed to use the maximum values of the chroma format and bit depth of all pictures of all layers.

[0291] 4) Currently, the setting of NoOutputOfPriorPicsFlag involves a change in the value of separate_colour_plane_flag. However, separate_colour_plane_flag only exists and is used when chroma_format_idc is equal to 3, which specifies the 4:4:4 chroma format, and for the 4:4:4 chroma format, the value of separate_colour_plane_flag being equal to 0 or 1 does not affect the buffer size required to store the decoded pictures. Therefore, the setting of NoOutputOfPriorPicsFlag should not involve a change in the value of separate_colour_plane_flag.

[0292] 5) Currently, no_output_of_prior_pics_flag is signaled in the PH of IRAP and GDR pictures, and both the semantics of the flag and the process for setting NoOutputOfPriorPicsFlag are specified in a way that no_output_of_prior_pics_flag is layer-specific or PU-specific. However, since the DPB operation is AU-specific or OLS-specific, both the semantics of no_output_of_prior_pics_flag and the use of the flag in the setting of NoOutputOfPriorPicsFlag should be specified in an AU-specific way.

[0293] 6) The current text for setting the value of the variable PictureOutputFlag for the current picture concerns the PictureOutputFlag for using pictures in the same AU as the current picture and in higher layers than the current picture. However, for picA pictures where nuh_layer_id is greater than that of the current picture, the PictureOutputFlag of picA has not been derived yet when deriving the PictureOutputFlag of the current picture.

[0294] 7) The current text for setting the value of the variable PictureOutputFlag has the following problem. There are two layers in the OLS bitstream, and only the higher layer is the output layer. And in a specific AU auA, the pic_output_flag of the picture in the higher layer is equal to 0. On the decoder side, the picture in the higher layer of auA does not exist (due to, for example, loss or layer down-switching), while the picture in the lower layer of auA exists and its pic_output_flag is equal to 1. Then the value of the PictureOutputFlag of the picture in the lower layer of auA will be set to be equal to 1. However, when OLS has only one output layer and the pic_output_flag of the picture in the output layer is equal to 0, it should be interpreted that the encoder (or content provider) does not wish to output a picture for the AU containing that picture.

[0295] 8) The current text for setting the value of the variable PictureOutputFlag has the following problem. There are three or more layers in the OLS bitstream, and only the top layer is the output layer. On the decoder side, the picture in the top layer of the current AU does not exist (due to, for example, loss or layer down-switching), while two or more pictures in the lower layers of the current AU exist and their pic_output_flag is equal to 1. Then, for that AU, more than one picture will be output. However, this is problematic because for OLS there is only one output layer, so the codec or content provider expects to output only one picture.

[0296] 9) The current text for setting the value of the variable PictureOutputFlag has the following problem. OLS mode 2 (when ols_mode_idc is equal to 2) can also specify only one output layer like mode 0, but the behavior of outputting the lower layer picture of the AU when the picture in the output layer (which is also the top layer picture) does not exist is specified only for mode 0.

[0297] 10) For an OLS that contains only one output layer, when the picture of the output layer (which is also the picture of the highest layer) is not available to the decoder (due to, for example, loss or layer switch - down), the decoder will not know whether the pic_output_flag of that picture is equal to 1 or 0. If it is equal to 1, it makes sense to output the lower - layer picture, but if it is equal to 0, from the perspective of user experience, outputting the lower - layer picture may be worse because the encoder (content provider) has set this value to 0 for some reason. For example, for this particular OLS, this AU should not have a picture output.

[0298] 5. List of embodiments and techniques

[0299] To solve the above - mentioned problems and other problems, the methods outlined below are disclosed. The listed items should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. In addition, these items can be applied individually or combined in any way.

[0300] Solutions to Problems 1 to 5

[0301] 1) To solve Problem 1, one or both of the maximum values of chroma_format_idc and bit_depth_minus8 of all pictures of all layers can be signaled in the VPS.

[0302] 2) To solve Problem 2, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be at least based on one or both of the maximum picture width and height of all pictures of all layers that can be signaled in the VPS.

[0303] 3) To solve Problem 3, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be at least based on one or both of the maximum values of chroma_format_idc and bit_depth_minus8 of all pictures of all layers that can be signaled in the VPS.

[0304] 4) To solve Problem 4, the setting of the value of the variable NoOutputOfPriorPicsFlag can be specified to be independent of the value of separate_colour_plane_flag.

[0305] 5) To solve Problem 5, both the semantics of no_output_of_prior_pics_flag and its use in the setting of NoOutputOfPriorPicsFlag can be specified in an AU - specific manner.

[0306] a. In one example, when present, for all pictures in an AU, the value of no_output_of_prior_pics_flag should be the same, and the value of no_output_of_prior_pics_flag for the AU is considered to be the value of no_output_of_prior_pics_flag for the pictures of the AU.

[0307] b. Alternatively, in one example, when irap_or_gdr_au_flag is equal to 1, no_output_of_prior_pics_flag can be removed from the PH syntax and can be signaled in the AUD syntax.

[0308] i. For a single-layer bitstream, since AUD is optional, when AUD is not present for an IRAP or GDR AU, the value of no_output_of_prior_pics_flag can be inferred to be equal to 1 (which means that if the encoder wants to signal a value of 0 for no_output_of_prior_pics_flag for an IRAP or GDR AU in a single-layer bitstream, it must signal AUD for that AU in the bitstream).

[0309] c. Alternatively, in one example, the value of no_output_of_prior_pics_flag for an AU can be considered equal to 0 if and only if the value of no_output_of_prior_pics_flag for each picture of the AU is equal to 0, otherwise the value of no_output_of_prior_pics_flag for the AU can be considered equal to 1.

[0310] i. The disadvantage of this method is that the setting of NoOutputOfPriorPicsFlag and the output of pictures of the CVSS AU need to wait for all pictures in the AU to arrive.

[0311] Solutions to problems 6 to 10

[0312] 6) To solve problem 6, the setting of PictureOutputFlag for the current picture can be specified to be at least based on the pic_output_flag (instead of PictureOutputFlag) of pictures in the same AU as the current picture and in a higher layer than the current picture.

[0313] 7) To solve problems 7 to 9, whenever the current picture does not belong to the output layer, the value of PictureOutputFlag for the current picture is set to be equal to 0.

[0314] a. Alternatively, to solve problems 7 and 8, when there is only one output layer and there is no output layer for the AU (which must be the top layer when there is only one output layer), for the picture among all pictures of the AU available to the decoder that has the highest value of nuh_layer_id and pic_output_flag is equal to 1, PictureOutputFlag is set to be equal to 1, and for all other pictures of the AU available to the decoder, PictureOutputFlag is set to be equal to 0.

[0315] 8) To solve problem 10, the value of pic_output_flag of the output layer picture of the AU can be signaled in the AUD or SEI message in the AU, or in the PH of one or more other pictures in the AU.

[0316] 6. Embodiment

[0317] The following are some example embodiments of some aspects of the present invention summarized in section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-Q2001-vE / v15. Most of the relevant parts that have been added or modified are highlighted in bold italics, and some deleted parts are marked with double brackets (e.g., [[a]] means deleting the character "a"). There are also some other changes that are essentially editorial and thus not highlighted.

[0318] 6.1. First Embodiment

[0319] This embodiment is for items 1, 2, 3, 4, 5, and 5a.

[0320] 7.3.2.2 Video Parameter Set Syntax

[0321] ...

[0323] 7.4.3.2 Video Parameter Set RBSP Semantics ...

[0325] ols_dpb_pic_width[i] specifies the width in luma samples of each picture storage buffer of the i-th OLS.

[0326] ols_dpb_pic_height[i] specifies the height in luma samples of each picture storage buffer of the i-th OLS.

[0327]

[0328] ols_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure that applies to the i-th OLS when NumLayersInOls[i] is greater than 1 (index into the list of dpb_parameters() syntax structures in the VPS). When present, the value of ols_dpb_params_idx[i] shall be in the range of 0 to vps_num_dpb_params-1, inclusive. When ols_dpb_params_idx[i] is not present, the value of ols_dpb_params_idx[i] is inferred to be equal to 0.

[0329] When NumLayersInOls[i] is equal to 1, the dpb_parameters() syntax structure that applies to the i-th OLS is present in the SPS referenced by the layers in the i-th OLS. ...

[0331] 7.4.3.3 Sequence parameter set RBSP semantics ...

[0333] gdr_enabled_flag equal to 1 specifies that GDR pictures may be present in the CLVS referencing the SPS. gdr_enabled_flag equal to 0 specifies that GDR pictures are not present in the CLVS referencing the SPS.

[0334] chroma_format_idc specifies the chroma samples relative to the luma samples as specified in clause 6.2.

[0335] ...

[0337] bit_depth_minus8 specifies the bit depth BitDepth of the samples of the luma and chroma arrays and the value of the luma and chroma quantization parameter range offset QpBdOffset as follows:

[0338] BitDepth=8+bit_depth_minus8 (45)

[0339] QpBdOffset=6*bit_depth_minus8 (46)

[0340] bit_depth_minus8 should be in the range of 0 to 8 (inclusive).

[0341] ...

[0343] 7.4.3.7 Semantics of Picture Header Structure ...

[0345] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding the bitstream specified in Appendix C in the bitstream specified in Appendix C after that.

[0346]

[0347] ...

[0349] C.1 General ...

[0351] For each bitstream conformance test, the CPB size (in bits) is CpbSize[Htid][ScIdx] as specified in Clause 7.4.6.3, where ScIdx and the HRD parameters are specified in that clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found or derived from the dpb_parameters() syntax structure applied to the target OLS as follows:

[0352] -- If the dpb_parameters() syntax structure is found in the SPS, which is referred to as the layer in the target OLS

[0353] -- Otherwise (the target OLS contains more than one layer), dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsIdx] found in the VPS ...

[0355] C.3.2 Removing Pictures from the DPB before Decoding the Current Picture

[0356] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (which contains the current picture) and is as follows:

[0357] --Call the decoding process constructed from the reference picture list specified in Clause 8.3.2 and call the decoding process of the reference picture marking specified in Clause 8.3.3.

[0358] --When the current AU is a CVSS AU other than AU 0, apply the following ordered steps:

[0359] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:

[0360] --If derived or the value of max_dec_pic_buffering_minus1[Htid] is different from the previous derived or the value of max_dec_pic_buffering_minus1[Htid], then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag.

[0361] Note--Although under these conditions, it is best to set NoOutputOfPriorPicsFlag to be equal to no_output_of_prior_pics_flag, in this case, the decoder under test is allowed to set NoOutputOfPriorPicsFlag to 1.

[0362] --Otherwise, NoOutputOfPriorPicsFlag is set to be equal to no_output_of_prior_pics_flag.

[0363] 2. The value of NoOutputOfPriorPicsFlag derived for the decoder under test is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are emptied without outputting the pictures they contain, and the DPB fullness is set to be equal to 0.

[0364] --When for any picture k in the DPB, the following two conditions are true, all such pictures k in the DPB will be removed from the DPB:

[0365] -- Picture k is marked as "not used for reference".

[0366] -- The PictureOutputFlag of Picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current Picture n; that is, DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].

[0367] -- For each picture removed from the DPB, the DPB fullness is decremented by 1.

[0368] C.5.2.2 Output and Removal of Pictures from the DPB

[0369] Output and removal of pictures from the DPB occurs instantaneously when removing the first DU of the AU containing the current picture from the CPB (but after parsing the slice header of the first slice of the current picture) and proceeds as follows:

[0370] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.

[0371] -- If the current is not CLVSS as 0 then apply the following ordered steps:

[0372] 1. The variable NoOutputOfPriorPicsFlag is derived for the DUT as follows:

[0373] -- If the derived or the value of max_dec_pic_buffering_minus1[Htid] is different from the value for the previous or the value of max_dec_pic_buffering_minus1[Htid], then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the DUT, regardless of the value of no_output_of_prior_pics_ .

[0374] Note--Although under these conditions, it is preferable that NoOutputOfPriorPicsFlag be set to be equal to no_output_of_prior_pics_flag, but in this case, the DUT decoder is allowed to set NoOutputOfPriorPicsFlag to 1.

[0375] -- Otherwise, NoOutputOfPriorPicsFlag is set to be equal to no_output_of_prior_pics_flag.

[0376] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT decoder is applied to the HRD as follows:

[0377] -- If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set to be equal to 0.

[0378] -- Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output) by repeatedly calling the "collision" procedure specified in Clause C.5.2.4, and all non-empty picture storage buffers in the DPB are cleared, and the DPB fullness is set to be equal to 0.

[0379] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are cleared (without output). For each picture storage buffer that is cleared, the DPB fullness is decremented by 1. When one or more of the following conditions are true, the "collision" procedure specified in Clause C.5.2.4 is repeatedly called, and for each additional picture storage buffer that is cleared, the DPB fullness is further decremented by 1 until none of the following conditions are true:

[0380] -- The number of pictures marked as "required for output" in the DPB is greater than max_num_reorder_pics[Htid].

[0381] -- max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "required for output" and its associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].

[0382] -- The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid] + 1.

[0383] 6.2. Second Embodiment

[0384] This embodiment is for Items 1, 2, 3, 4, 5, and 5c, where the text varies with respect to the text of the first embodiment.

[0385] 7.4.3.7 Picture Header Structure Semantics ...

[0387] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the DPB after decoding a picture in a CVSS AU that is not the first AU in the bitstream specified in Appendix C.

[0388] [[The requirement for bitstream consistency is that when present, the value of the no_output_of_prior_pics_flag should be the same for all pictures of an AU.

[0389] When the no_output_of_prior_pics_flag is present in the PH of a picture of an AU, the value of the no_output_of_prior_pics_flag of the AU is the value of the no_output_of_prior_pics_flag of the picture of the AU.]] ...

[0391] C.3.2 Removing Pictures from the DPB before Decoding the Current Picture

[0392] Removing pictures from the DPB before decoding the current picture (but after parsing the slice header of the first slice of the current picture) occurs at the CPB removal time instant of the first DU of AU n (including the current picture) and is as follows:

[0393] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.

[0394] -- When the current AU is a CVSS AU other than AU 0, apply the following ordered steps:

[0395] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:

[0396] -- If the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous AU in decoding order, then NoOutputOfPriorPicsFlag may (but shall not) be set to 1 by the DUT, regardless of [[the value of]] no_output_of_prior_pics_flag [[for the current AU]]

[0397] Note -- [[Although]] under these conditions, it is preferable that NoOutputOfPriorPicsFlag be set to equal 0 [[equal to the value of no_output_of_prior_pics_flag for the current AU]], but in such a case, the DUT is allowed to set NoOutputOfPriorPicsFlag to 1.

[0398] -- Otherwise, [[NoOutputOfPriorPicsFlag is set to equal the value of no_output_of_prior_pics_flag for the current AU]]

[0399] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT is applied to the HRD such that when the value of NoOutputOfPriorPicsFlag equals 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set to equal 0.

[0400] -- All such pictures k in the DPB will be removed from the DPB when the following two conditions are true for any picture k in the DPB:

[0401] -- Picture k is marked as "not used for reference".

[0402] -- The PictureOutputFlag of picture k is equal to 0, or its DPB output time is less than or equal to the CPB removal time of the first DU of the current picture n (denoted as DU m); that is, DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m].

[0403] -- For each picture removed from the DPB, the DPB fullness is decremented by 1.

[0404] C.5.2.2 Output and removal of pictures from the DPB

[0405] Output and removal of pictures from the DPB occur instantaneously when the first DU of the AU containing the current picture is removed from the CPB, before the decoding of the current picture (but after parsing the slice header of the first slice of the current picture) and are as follows:

[0406] -- Invoke the decoding process for reference picture list construction specified in Clause 8.3.2 and the decoding process for reference picture marking specified in Clause 8.3.3.

[0407] -- If the current AU is a CLVSS AU other than AU 0, apply the following ordered steps:

[0408] 1. The variable NoOutputOfPriorPicsFlag is derived for the decoder under test as follows:

[0409] -- If the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the current AU are different from the values of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8, or max_dec_pic_buffering_minus1[Htid] derived for the previous AU in decoding order, then NoOutputOfPriorPicsFlag may (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag of [[the current AU]]

[0410] Note-- [[Although]] under these conditions, it is preferable that NoOutputOfPriorPicsFlag be set equal to 0 [[equal to the no_output_of_prior_pics_flag of the current AU]], in this case, the DUT decoder is allowed to set NoOutputOfPriorPicsFlag to 1.

[0411] -- Otherwise, then NoOutputOfPriorPicsFlag will be [[equal to the no_output_of_prior_pics_flag of the current AU]].

[0412]

[0413] 2. The value of NoOutputOfPriorPicsFlag derived for the DUT decoder is applied to the HRD as follows:

[0414] -- If NoOutputOfPriorPicsFlag is equal to 1, all picture storage buffers in the DPB are cleared without outputting the pictures they contain, and the DPB fullness is set equal to 0.

[0415] -- Otherwise (NoOutputOfPriorPicsFlag is equal to 0), all picture storage buffers containing pictures marked as "not to be output" and "not used for reference" are cleared (without output) by repeatedly calling the "collision" process specified in Clause C.5.2.4, and all non-empty picture storage buffers in the DPB are cleared, and the DPB fullness is set equal to 0.

[0416] -- Otherwise (the current picture is not a CLVSS picture or the CLVSS picture is picture 0), all picture storage buffers containing pictures marked as "not to be output" and "not used for reference" are cleared (without output). For each picture storage buffer cleared, the DPB fullness is decremented by 1. When one or more of the following conditions are true, the "collision" process specified in Clause C.5.2.4 is repeatedly called, and for each additional picture storage buffer cleared, the DPB fullness is further decremented by 1 until none of the following conditions are true:

[0417] -- The number of pictures marked as "to be output" in the DPB is greater than max_num_reorder_pics[Htid].

[0418] --max_latency_increase_plus 1[Htid] is not equal to 0, and at least one picture in the DPB is marked as "to be output", and its related variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].

[0419] --The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid] + 1.

[0420] 6.3. Third Embodiment

[0421] This embodiment is for Item 6, Item 7 (the changed text does not include the notes added in Clause 8.1.2), and Item 7a (the notes added in Clause 8.1.2).

[0422] 7.4.3.7 Picture Header Structure Semantics ...

[0424] recovery_poc_cnt specifies the recovery point of the decoded pictures in the output order.

[0425]

[0426]

[0427] If the current picture is a [[PH - associated]] GDR picture, and there is a picture picA in the CLVS that follows the current GDR picture in the decoding order, and its PicOrderCntVal is equal to [[the value of PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt]], then the picture picA is called a recovery - point picture. Otherwise, the first picture in the output order whose PicOrderCntVal is greater than [[the value of PicOrderCntVal of the current picture plus the value of recovery_poc_cnt]] is called a recovery - point picture. The recovery - point picture should not be before the current GDR picture in the decoding order. The value of recovery_poc_cnt should be in the range (including the end values) from 0 to MaxPicOrderCntLsb - 1.

[0428] [[When the current picture is a GDR picture, the variable RpPicOrderCntVal is derived as follows:

[0429] RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt(81)]

[0430] Note 2--When gdr_enabled_flag is equal to 1 and the PicOrderCntVal of the current picture is greater than or equal to [[RpPicOrderCntVal]] of the relevant GDR picture, the current and subsequent decoded pictures in the output order exactly match the corresponding pictures generated by decoding the process starting from the previous IRAP picture, and when present, are before the relevant GDR picture in the decoding order. ...

[0432] 8.1.2 Decoding Process of Encoded and Decoded Pictures ...

[0434]

[0435] -- [[PictureOutputFlag is set as follows:

[0436] -- If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:

[0437] -- The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0438] -- gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0439] -- gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0440] -- sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0441] -- The PictureOutputFlag of PicA is equal to 1.

[0442] -- The nuh_layer_id of PicA, nuhLid, is greater than the nuh_layer_id of the current picture.

[0443] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).

[0444] -- sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0445] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag. ...

[0447] 6.4. Fourth Embodiment

[0448] This embodiment is used for Project 6 and Project 7a.

[0449] 8.1.2 Decoding Process of Encoded / Decoded Pictures ...

[0451] -- PictureOutputFlag is set as follows:

[0452] -- If one of the following conditions is true, then PictureOutputFlag is set to be equal to 0:

[0453] -- The current picture is a RASL picture, and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0454] -- gdr_enabled_flag is equal to 1, and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0455] -- gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0456]

[0457] --The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0458] --The PictureOutputFlag of PicA is equal to 1.

[0459] --The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.

[0460] --PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).]]

[0461] --The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0462] --Otherwise, the PictureOutputFlag is set to be equal to pic_output_flag. ...

[0464] 6.5. Fifth Embodiment

[0465] This embodiment is only used for Project 6.

[0466] 8.1.2 Decoding Process of Encoded / Decoded Pictures ...

[0468] --The PictureOutputFlag is set as follows:

[0469] --If one of the following conditions is true, the PictureOutputFlag is set to be equal to 0:

[0470] --The current picture is a RASL picture and the NoOutputBeforeRecoveryFlag of the associated IRAP picture is equal to 1.

[0471] --The gdr_enabled_flag is equal to 1 and the current picture is a GDR picture with NoOutputBeforeRecoveryFlag equal to 1.

[0472] -- The gdr_enabled_flag is equal to 1, the current picture is associated with a GDR picture with NoOutputBeforeRecoveryFlag equal to 1, and the PicOrderCntVal of the current picture is less than the RpPicOrderCntVal of the associated GDR picture.

[0473] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 0, and the current AU contains a picture picA that satisfies all of the following conditions:

[0474] -- For PicA [[PictureOutputFlag]] is equal to 1.

[0475] -- The nuh_layer_id nuhLid of PicA is greater than the nuh_layer_id of the current picture.

[0476] -- PicA belongs to the output layer of OLS (i.e., OutputLayerIdInOls[TargetOlsIdx][0] is equal to nuhLid).

[0477] -- The sps_video_parameter_set_id is greater than 0, the ols_mode_idc is equal to 2, and ols_output_layer_flag[TargetOlsIdx][GeneralLayerIdx[nuh_layer_id]] is equal to 0.

[0478] -- Otherwise, PictureOutputFlag is set to be equal to pic_output_flag.

[0479] Figure 1 FIG. 24 is a block diagram showing an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0480] System 1900 may include a codec component 1904 that may implement various codec or encoding methods described in this document. The codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Codec techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 may be stored or transmitted via a communication connection as represented by component 1906. The stored or communicatively transmitted bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or a displayable video to be sent to the display interface 1910. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Further, although certain video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0481] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0482] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 may be configured to implement one or more methods described in this document. The memory (multiple memories) 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry.

[0483] Figure 4 is a block diagram showing an example video codec system 100 that may utilize the techniques of the present disclosure.

[0484] As Figure 4As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, where the source device 110 may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, where the destination device 120 may be referred to as a video decoding device.

[0485] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0486] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0487] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0488] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.

[0489] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or additional standards.

[0490] Figure 5 is a block diagram showing an example of a video encoder 200, which may be Figure 4 the video encoder 114 in the system 100 shown.

[0491] Video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 the example, video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0492] The functional components of video encoder 200 may include a splitting unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206), a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0493] In other examples, video encoder 200 may include more, fewer, or different functional components. In an example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0494] In addition, some components such as motion estimation unit 204 and motion compensation unit 205 may be highly integrated, but are shown separately in the Figure 5 example for explanatory purposes.

[0495] Splitting unit 201 may split a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support various video block sizes.

[0496] Mode selection unit 203 may select one of the coding / decoding modes (e.g., intra or inter) based on error results, and provide the resulting intra-coded / decoded block or inter-coded / decoded block to residual generation unit 207 to generate residual block data, and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 may select a combination of intra and inter prediction modes (CIIP), where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 may also select the resolution of the motion vector of the block (e.g., sub-pixel or integer-pixel precision).

[0497] To perform inter prediction on the current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0498] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0499] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search for a reference picture in list 0 or list 1 of reference video blocks for the current video block. The motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 1, which includes the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0500] In other examples, the motion estimation unit 204 can perform bi-directional prediction on the current video block. The motion estimation unit 204 can search for a reference video block of the current video block in the reference pictures in list 0 and can also search for another reference video block of the current video block in list 1. The motion estimation unit 204 can then generate a reference index indicating the reference pictures in list 0 and list 1 that include the reference video blocks and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0501] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoding process of the decoder.

[0502] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.

[0503] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and the value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0504] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0505] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0506] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0507] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0508] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0509] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0510] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0511] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0512] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0513] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0514] Figure 6 is a block diagram showing an example of a video decoder 300, and the video decoder 300 can be Figure 4 the video decoder 114 in the system 100 shown.

[0515] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 6 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0516] In Figure 6 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described for the video encoder 200 ( Figure 5 ).

[0517] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes.

[0518] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0519] The motion compensation unit 302 may use an interpolation filter such as that used by the video encoder 200 during encoding of a video block to compute an interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine, based on received syntax information, the interpolation filter used by the video encoder 200 and may use that interpolation filter to generate a prediction block.

[0520] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode (a) frame(s) and / or (a) slice(s) of an encoded video sequence, partitioning information describing how each macroblock of a picture describing the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0521] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in a bitstream. The inverse quantization unit 303 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301, i.e., dequantizes. The inverse transform unit 303 applies an inverse transform.

[0522] The reconstruction unit 306 may add a residual block to a corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307 to provide a reference block for subsequent motion compensation / intra prediction and also to produce decoded video for presentation on a display device.

[0523] Next, a list of preferred examples of some embodiments is provided.

[0524] The first set of clauses illustrates example embodiments of the techniques discussed in the previous chapter. The following clauses illustrate example embodiments of the techniques discussed in the previous chapter (e.g., item 1).

[0525] 1. A video processing method (e.g., Figure 3 method 3000 as illustrated), comprising: performing a conversion (3002) between a video and a codec representation of the video having one or more video layers including one or more video pictures; wherein the codec representation includes a video parameter set that indicates a maximum value of a chroma format indicator and / or a maximum value of a bit depth for pixels representing the video.

[0526] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0527] 2. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies a maximum picture width and / or a maximum picture height of video pictures of all video layers to control a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer.

[0528] 3. The method according to clause 2, wherein the variable is signaled in a video parameter set.

[0529] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).

[0530] 4. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies a maximum value of a chroma format indicator and / or a maximum bit depth of pixels representing the video to control a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer.

[0531] 5. The method according to clause 4, wherein the variable is signaled in a video parameter set.

[0532] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., item 4).

[0533] 6. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies that a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is independent of whether separate color planes are used to encode the video.

[0534] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., item 5).

[0535] 7. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies that a value of a variable indicating whether a picture in a decoder buffer is output before being removed from the decoder buffer is included in the encoded / decoded representation at an access unit (AU) level.

[0536] 8. The method according to clause 7, wherein the format rule specifies that the value is the same for all AUs in the encoded / decoded representation.

[0537] 9. The method according to any one of clauses 7 - 8, wherein the variable is indicated in the picture header.

[0538] 10. The method according to any one of clauses 7 - 8, wherein the variable is indicated in the access unit delimiter.

[0539] The following clauses show example embodiments of the technology discussed in the previous section (e.g., item 6).

[0540] 11. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies that a picture output flag of a video picture in an access unit is determined based on a pic_output_flag variable of another video picture in the access unit.

[0541] The following clauses show example embodiments of the technology discussed in the previous section (e.g., item 7).

[0542] 12. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to a format rule that specifies a value of a picture output flag for a video picture that does not belong to an output layer.

[0543] 13. The method according to clause 12, wherein the format rule specifies that the value of the picture output flag of the video picture is set to zero.

[0544] 14. The method according to clause 12, wherein the video includes only one output layer, and wherein an access unit that does not include an output layer is encoded / decoded by setting the picture output flag value to logical 1 for a picture with the highest layer id value and setting the picture output flag value to logical 0 for all other pictures.

[0545] The following clauses show example embodiments of the technology discussed in the previous section (e.g., item 8).

[0546] 15. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the access unit delimiter.

[0547] 16. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the supplementary enhancement information field.

[0548] 17. The method according to any one of clauses 1 - 14, wherein the picture output flag is included in the picture header of one or more pictures.

[0549] 18. The method according to any one of clauses 1 to 17, wherein the conversion includes encoding the video into a codec representation.

[0550] 19. The method according to any one of clauses 1 to 17, wherein the conversion includes decoding the codec representation to generate pixel values of the video.

[0551] 20. A video decoding apparatus, comprising a processor configured to implement the method according to one or more of clauses 1 to 19.

[0552] 21. A video encoding apparatus, comprising a processor configured to implement the method according to one or more of clauses 1 to 19.

[0553] 22. A computer program product storing computer code, which when executed by a processor causes the processor to implement the method according to any one of clauses 1 to 19.

[0554] 23. A method, apparatus or system described in this document.

[0555] The second set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., items 1 - 4).

[0556] 1. A method for video processing (e.g., the method 710 as Figure 7A shown), comprising: performing a conversion 712 between a video and a bitstream of the video according to format rules, wherein the bitstream includes one or more output layer sets (OLSs), each OLS includes one or more codec layer video sequences, and wherein the format rules specify that a video parameter set indicates, for each of the one or more OLSs, a maximum allowable value of a chrominance format indicator for representing pixels of the video and / or a maximum allowable value of a bit depth.

[0557] 2. The method according to clause 1, wherein the maximum allowable value of the chrominance format indicator of the OLS applies to all sequence parameter sets referred to by one or more codec layer video sequences in the OLS.

[0558] 3. The method according to clause 1 or 2, wherein the maximum allowable value of the bit depth of the OLS applies to all sequence parameter sets referred to by one or more codec layer video sequences in the OLS.

[0559] 4. The method according to any one of clauses 1 to 3, wherein, in order to perform conversion on an OLS containing a video sequence with more than one coding / decoding layer and having an OLS index i, the rule specifies allocating memory for the decoded picture buffer according to the value of at least one of the syntax elements including ols_dpb_pic_width[i] indicating the width of each picture storage buffer of the i-th OLS, ols_dpb_pic_height[i] indicating the height of each picture storage buffer of the i-th OLS, a syntax element indicating the maximum allowable value of the chroma format indicator of the i-th OLS, and a syntax element indicating the maximum allowable value of the bit depth of the i-th OLS.

[0560] 5. The method according to any one of clauses 1 to 4, wherein the video parameter set is included in the bitstream.

[0561] 6. The method according to any one of clauses 1 to 4, wherein the video parameter set is indicated separately from the bitstream.

[0562] 7. A method for video processing (e.g., method 720 as shown in Figure 7B ), including: performing conversion 722 between a video with one or more video layers and the bitstream of the video according to format rules, and wherein the format rules specify a variable value that controls whether pictures in the decoded picture buffer before the current picture in decoding order in the bitstream are output before the pictures are removed from the decoded picture buffer for the maximum picture width and / or maximum picture height of the video pictures of all video layers.

[0563] 8. The method according to clause 7, wherein the variable is derived based on at least one or more syntax elements included in the video parameter set.

[0564] 9. The method according to clause 7, wherein, in the case where the value of the maximum width per picture, the maximum height per picture, the maximum allowable value of the chroma format indicator, or the maximum allowable value of the bit depth derived for each picture of the current access unit is different from the value of the maximum width per picture, the maximum height per picture, the maximum allowable value of the chroma format indicator, or the maximum allowable value of the bit depth derived for the previous access unit in decoding order, the value of the variable is set to 1.

[0565] 10. The method according to clause 9, wherein a value of the variable equal to 1 indicates that pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures are removed from the decoded picture buffer.

[0566] 11. The method according to any one of clauses 7 to 10, wherein the value of the variable is further based on the maximum allowable value of the chrominance format indicator and / or the maximum allowable value of the bit depth for the pixels representing the video.

[0567] 12. The method according to any one of clauses 7 to 11, wherein the video parameter set is included in the bitstream.

[0568] 13. The method according to any one of clauses 7 to 11, wherein the video parameter set is indicated separately from the bitstream.

[0569] 14. A method for video processing (e.g., as shown in Figure 7C 730), comprising: performing a conversion between a video having one or more video layers and a bitstream of the video according to format rules 732, and wherein the format rules specify the maximum allowable value of the chrominance format indicator and / or the maximum allowable value of the bit depth for the pixels representing the video to control the value of a variable indicating whether a picture in the decoded picture buffer before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer.

[0570] 15. The method according to clause 14, wherein the variable is derived based at least on one or more syntax elements signaled in the video parameter set.

[0571] 16. The method according to clause 14, wherein, in a case where the value of the maximum width of each picture, the maximum height of each picture, the maximum allowable value of the chrominance format indicator, or the maximum allowable value of the bit depth for each picture derived for the current access unit is different from the value of the maximum width of each picture, the maximum height of each picture, the maximum allowable value of the chrominance format indicator, or the maximum allowable value of the bit depth for each picture derived for the previous access unit in decoding order, the value of the variable is set to 1.

[0572] 17. The method according to clause 16, wherein the value of the variable being equal to 1 indicates that a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is removed from the decoded picture buffer.

[0573] 18. The method according to any one of clauses 14 to 17, wherein the value of the variable is further based on the maximum picture width and / or the maximum picture height of the video pictures of all video layers.

[0574] 19. The method according to any one of clauses 14 to 18, wherein the video parameter set is included in the bitstream.

[0575] 20. The method according to any one of clauses 14 to 18, wherein the video parameter set is indicated separately from the bitstream.

[0576] 21. A method for video processing (e.g., the method 740 as Figure 7D shown), comprising: performing a conversion 742 between a video having one or more video layers and a bitstream of the video according to a rule, and wherein the rule specifies that the value of a variable indicating whether a picture in the decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer is independent of whether separate color planes are used for encoding the video.

[0577] 22. The method according to clause 21, wherein, in the case where separate color planes are not used for encoding the video, the rule specifies that the video picture is decoded only once, or in the case where separate color planes are used for encoding the video, the rule specifies that the picture decoding is called three times.

[0578] 23. The method according to any one of clauses 1 to 22, wherein the conversion includes encoding the video into a bitstream.

[0579] 24. The method according to any one of clauses 1 to 22, wherein the conversion includes decoding the video from the bitstream.

[0580] 25. The method according to clauses 1 to 22, wherein the conversion includes generating a bitstream from the video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.

[0581] 26. A video processing apparatus, comprising a processor configured to implement the method according to any one or more of clauses 1 to 25.

[0582] 27. A method for storing a bitstream of a video, comprising the method according to any one of clauses 1 to 25, and further including storing the bitstream in a non-transitory computer-readable recording medium.

[0583] 28. A computer-readable medium storing program code, which when executed causes a processor to implement the method according to any one or more of clauses 1 to 25.

[0584] 29. A computer-readable medium storing a bitstream generated according to any one of the above methods.

[0585] 30. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 25.

[0586] The third set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., item 5).

[0587] 1. A method for video processing (e.g., the method asFigure 8A The method (810) shown includes: performing a conversion (812) between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a flag indicating whether a picture that has been previously decoded and stored in a decoded picture buffer is removed from the decoded picture buffer when decoding an access unit of a certain type is included in the bitstream.

[0588] 2. The method according to clause 1, wherein the format rules specify that the value is the same for all pictures in the access unit.

[0589] 3. The method according to clause 1 or 2, wherein the format rules specify that a value of a variable indicating whether a picture in the decoded picture buffer that is in decoding order before the current picture in the bitstream is output before the picture is removed from the decoded picture buffer is based on the value of the flag.

[0590] 4. The method according to any one of clauses 1 to 3, wherein the flag is indicated in a picture header.

[0591] 5. The method according to any one of clauses 1 to 3, wherein the flag is indicated in a slice header.

[0592] 6. A method for video processing (e.g., the method 820 as Figure 8B shown), includes: performing a conversion (822) between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that a value of a first flag indicating whether a picture that has been previously decoded and stored in a decoded picture buffer is removed from the decoded picture buffer when decoding a particular type of access unit is not indicated in a picture header.

[0593] 7. The method according to clause 6, wherein the first flag is indicated in an access unit delimiter.

[0594] 8. The method according to clause 6, wherein a second flag indicating an IRAP (Intra Random Access Point picture) or GDR (Gradual Decoding Refresh) access unit has a certain value.

[0595] 9. The method according to clause 6, wherein in a case where there is no access unit delimiter for an IRAP (Intra Random Access Point picture) or GDR (Gradual Decoding Refresh) access unit, the value of the first flag is inferred to be equal to 1.

[0596] 10. A method for video processing (e.g., as Figure 8CThe method shown (830) includes: performing a conversion (832) between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that the value of a flag associated with an access unit indicating whether to remove a previously decoded and stored picture from a decoded picture buffer depends on the value of a flag of each picture of the access unit.

[0597] 11. The method according to clause 10, wherein the formatting rules specify that, in the case where the flag of each picture of the access unit is equal to 0, the value of the flag of the access unit is considered to be equal to 0, otherwise, the value of the flag of the access unit is considered to be equal to 1.

[0598] 12. The method according to any one of clauses 1 to 11, wherein the conversion includes encoding the video into a bitstream.

[0599] 13. The method according to any one of clauses 1 to 11, wherein the conversion includes decoding the video from the bitstream.

[0600] 14. The method according to any one of clauses 1 to 11, wherein the conversion includes generating a bitstream from the video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.

[0601] 15. A video processing apparatus, including a processor configured to implement the method according to any one or more of clauses 1 to 14.

[0602] 16. A method of storing a bitstream of a video, including the method according to any one of clauses 1 to 14, and further including storing the bitstream into a non-transitory computer-readable recording medium.

[0603] 17. A computer-readable medium storing program code that, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 14.

[0604] 18. A computer-readable medium storing a bitstream generated according to any one of the above methods.

[0605] 19. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 14.

[0606] The fourth set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., items 6-8).

[0607] 1. A method of video processing (e.g., as Figure 9AThe method 910) shown, includes: performing a conversion 912 between a video having one or more video layers and a bitstream of the video according to formatting rules, and wherein the formatting rules specify that a value of a variable indicating whether to output a picture in an access unit is determined based on a flag indicating whether to output another picture in the access unit.

[0608] 2. The method according to clause 1, wherein the other picture is in a higher layer than the picture.

[0609] 3. The method according to clause 1 or 2, wherein the flag controls a decoded picture output and removal process.

[0610] 4. The method according to clause 1 or 2, wherein the flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a slice group header.

[0611] 5. The method according to any one of clauses 1 to 4, wherein the value of the variable is further based on at least one of the following: i) a flag specifying a value of an identifier of a video parameter set (VPS), ii) whether the current video layer is an output layer, ii) whether the current picture is a random access skipped previous picture, a progressive decoded refresh picture, a recovery picture of a progressive decoded refresh picture, or iii) whether a picture in a decoded picture buffer before the current picture in decoding order is output before the picture is recovered.

[0612] 6. The method according to clause 5, wherein the value of the variable is set to be equal to 0 when i) the flag specifying the value of the identifier of the VPS is greater than 0 and the current layer is not an output layer, or ii) one of the following conditions is true:

[0613] The current picture is a random access skipped previous picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the intra random access point picture is recovered; or

[0614] The current picture is a progressive decoded refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is recovered, or the current picture is a recovery picture of a progressive decoded refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture is recovered.

[0615] 7. The method according to clause 6, wherein the value of the variable is set to be equal to the value of the flag when neither i) nor ii) is satisfied.

[0616] 8. The method according to clause 6, wherein a value of the variable being equal to 0 indicates not to output a picture in the access unit.

[0617] 9. The method according to clause 1, wherein the variable is PictureOutputFlag, and the flag is pic_output_flag.

[0618] 10. The method according to any one of clauses 1 to 9, wherein the flag is included in the access unit delimiter.

[0619] 11. The method according to any one of clauses 1 to 9, wherein the flag is included in the supplementary enhancement information field.

[0620] 12. The method according to any one of clauses 1 to 9, wherein the flag is included in the picture header of one or more pictures.

[0621] 13. A method for video processing (e.g., the method 920 as Figure 9B shown), comprising: performing a conversion 922 between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when a picture in the access unit does not belong to the output layer, a value of a variable indicating whether to output the picture is set to be equal to a certain value.

[0622] 14. The method according to clause 13, wherein the certain value is 0.

[0623] 15. A method for video processing (e.g., the method 930 as Figure 9C shown), comprising: performing a conversion 932 between a video having one or more video layers and a bitstream of the video according to format rules, and wherein the format rules specify that when the video includes only one output layer, an access unit that does not include the output layer is encoded and decoded by setting a variable indicating whether to output a picture in the access unit to a first value for a picture having the highest layer ID (identification) value and a second value for all other pictures.

[0624] 16. The method according to clause 15, wherein the first value is 1, and the second value is 0.

[0625] 17. The method according to any one of clauses 1 to 16, wherein the conversion includes encoding the video into a bitstream.

[0626] 18. The method according to any one of clauses 1 to 16, wherein the conversion includes decoding the video from the bitstream.

[0627] 19. The method according to clauses 1 to 16, wherein the conversion includes generating a bitstream from a video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.

[0628] 20. A video processing apparatus, including a processor configured to implement the method according to any one or more of clauses 1 to 19.

[0629] 21. A method for storing a bitstream of a video, including the method according to any one of clauses 1 to 19, and further including storing the bitstream in a non-transitory computer-readable recording medium.

[0630] 22. A computer-readable medium storing program code that, when executed, causes a processor to implement the method according to any one or more of clauses 1 to 19.

[0631] 23. A computer-readable medium storing a bitstream generated according to any one of the above methods.

[0632] 24. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 19.

[0633] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from the pixel representation of a video to the corresponding bitstream representation, a video compression algorithm may be applied, and vice versa. As defined by the syntax, the bitstream representation of the current video block may, for example, correspond to bits co-located or scattered in different places within the bitstream. For example, a macroblock may be encoded according to the transform and coding / decoding error residual values and also using bits in the headers and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or exclude a particular syntax field and generate the coded / decoded representation accordingly by including the syntax field or excluding the syntax field from the coded / decoded representation.

[0634] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal that is generated to encode information for transmission to a suitable receiver apparatus, e.g., a machine-generated electrical, optical, or electromagnetic signal.

[0635] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers distributed across one site or multiple sites and interconnected by a communication network.

[0636] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0637] Processors suitable for executing computer programs include, for example, any one or more processors of general and special purpose microprocessors, as well as any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to the one or more mass storage devices, or both receive data from and transfer data to them. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0638] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features that are described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0639] Similarly, although operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as required in all embodiments.

[0640] Only some embodiments and examples have been described, and other embodiments, enhancements, and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A method for video processing, comprising: Performing a conversion between a video having one or more video layers and a bitstream of the video according to a first format rule, wherein the first format rule specifies that when the video includes only one output layer, a first access unit that does not include the output layer is encoded and decoded by setting a first variable indicating whether to output a picture in the first access unit to a first value for a picture having the highest layer identifier (ID) value and a second value for all other pictures, the first value being 1 and the second value being 0, and wherein the conversion is further performed according to a third format rule, wherein the third format rule specifies that the value of a third variable indicating whether to output a picture is determined based on at least one of the following: a flag of a value of an identifier specifying a video parameter set (VPS), whether the current video layer is an output layer, whether the current picture is a random access skipped previous picture, a progressive decoding refresh picture, or a recovery picture of a progressive decoding refresh picture, or whether a picture in a decoded picture buffer before the current picture in decoding order is output before the picture in the decoded picture buffer is recovered, wherein the value of the third variable is set to be equal to 0 when i) the flag of the value of the identifier specifying the VPS is greater than 0 and the current video layer is not an output layer, or ii) one of the following conditions is true: The current picture is a random access skipped previous picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the associated intra random access point picture is recovered; or The current picture is a progressive decoding refresh picture, wherein pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures in the decoded picture buffer are recovered, or the current picture is a recovery picture of a progressive decoding refresh picture, wherein pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures in the decoded picture buffer are recovered.

2. The method according to claim 1, wherein, The conversion is performed according to a second format rule, wherein the second format rule specifies that the value of a second variable indicating whether to output a picture in a second access unit is determined based on a flag indicating whether to output another picture in the second access unit.

3. The method according to claim 2, wherein The other picture is in a layer higher than the picture in the second access unit.

4. The method according to claim 2, wherein The flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a slice group header.

5. The method according to claim 1, wherein, When neither i) nor ii) is satisfied, the value of the third variable is set to be equal to the value of the flag.

6. The method according to claim 1, wherein, The value of the third variable being equal to 0 indicates not to output the picture in the access unit.

7. The method according to claim 5, wherein The third variable is PictureOutputFlag, and the flag is pic_output_flag.

8. The method according to claim 5, wherein, The flag controls the decoded picture output and removal process.

9. The method according to claim 5, wherein The flag is included in an access unit delimiter, a supplementary enhancement information field, or a picture header of one or more pictures.

10. The method according to claim 1, wherein, The conversion is performed according to a fourth formatting rule, wherein the fourth formatting rule specifies that when a picture in a fourth access unit is not included in an output layer, a value of a fourth variable indicating whether to output the picture is set to be equal to a certain value.

11. The method according to claim 10, wherein, The certain value is 0.

12. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.

13. The method according to claim 1, wherein The conversion includes decoding the video from the bitstream.

14. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: perform a conversion between a video having one or more video layers and a bitstream of the video according to a first formatting rule, Among them, wherein the first formatting rule specifies that when the video includes only one output layer, a first access unit not including the output layer is encoded and decoded by setting a first variable indicating whether to output a picture in the first access unit to a first value for a picture having a highest layer identifier (ID) value and a second value for all other pictures, the first value being 1, and the second value being 0, and wherein the conversion is further performed according to a third formatting rule, wherein the third formatting rule specifies that a value of a third variable indicating whether to output a picture is determined based on at least one of: a flag of a value of an identifier specifying a video parameter set (VPS), whether the current video layer is an output layer, whether the current picture is a random access skipped preceding picture, a progressive decoding refresh picture, or a recovery picture of a progressive decoding refresh picture, or whether a picture in a decoded picture buffer before the current picture in decoding order is output before the picture in the decoded picture buffer is recovered, wherein the value of the third variable is set to be equal to 0 when i) the flag of the value of the identifier specifying the VPS is greater than 0 and the current video layer is not an output layer, or ii) one of the following conditions is true: the current picture is a random access skipped preceding picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the associated intra random access point picture is recovered; or the current picture is a progressive decoding refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture in the decoded picture buffer is recovered, or the current picture is a recovery picture of a progressive decoding refresh picture, wherein a picture in the decoded picture buffer before the current picture in decoding order is not output before the picture in the decoded picture buffer is recovered.

15. The apparatus according to claim 14, wherein, The conversion is performed according to a second formatting rule, wherein the second formatting rule specifies that a value of a second variable indicating whether to output a picture in a second access unit is determined based on a flag indicating whether to output another picture in the second access unit.

16. The apparatus according to claim 15, wherein, The another picture is in a layer higher than the picture in the second access unit.

17. The device according to claim 15, wherein, The flag is a syntax element included in a Sequence Parameter Set (SPS), a Video Parameter Set (VPS), a Picture Parameter Set (PPS), a picture header, a slice header, or a slice group header.

18. The device according to claim 14, wherein When both i) and ii) are not satisfied, the value of the third variable is set to be equal to the value of the flag.

19. The apparatus according to claim 14, wherein, The value of the third variable being equal to 0 indicates that the picture in the access unit is not output.

20. The apparatus according to claim 18, wherein, The third variable is PictureOutputFlag, and the flag is pic_output_flag.

21. The device according to claim 18, wherein, The flag controls the decoding picture output and removal process.

22. The device according to claim 18, wherein, The flag is included in an access unit delimiter, an auxiliary enhancement information field, or a picture header of one or more pictures.

23. The apparatus according to claim 14, wherein, The conversion is performed according to a fourth formatting rule, wherein the fourth formatting rule specifies that when the picture in the fourth access unit is not included in the output layer, the value of a fourth variable indicating whether to output the picture is set to be equal to a certain value.

24. The apparatus according to claim 23, wherein The certain value is 0.

25. The apparatus according to claim 14, wherein, The conversion includes encoding the video into the bitstream.

26. The apparatus according to claim 14, wherein, The conversion includes decoding the video from the bitstream.

27. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video having one or more video layers and a bitstream of the video according to a first formatting rule, Among them, wherein the first formatting rule specifies that when the video includes only one output layer, a first access unit not including the output layer is encoded and decoded by setting a first variable indicating whether to output the picture in the first access unit to a first value for a picture having a highest layer identifier (ID) value and a second value for all other pictures, the first value is 1, and the second value is 0, and wherein the conversion is further performed according to a third formatting rule, wherein the third formatting rule specifies that the value of a third variable indicating whether to output a picture is determined based on at least one of: a flag with a value of an identifier specifying a Video Parameter Set (VPS), whether the current video layer is an output layer, whether the current picture is a random access skipped previous picture, a progressive decoding refresh picture, or a recovery picture of a progressive decoding refresh picture, or whether a picture in a decoded picture buffer before the current picture in decoding order is output before the picture in the decoded picture buffer is recovered, wherein when i) the flag with a value of the identifier specifying the VPS is greater than 0 and the current video layer is not an output layer, or ii) one of the following conditions is true, the value of the third variable is set to be equal to 0: the current picture is a random access skipped previous picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the associated intra random access point picture is recovered; or The current picture is a progressively decoded and refreshed picture, where pictures in the decoded picture buffer that are before the current picture in decoding order are not output until the pictures in the decoded picture buffer are restored, or the current picture is a restored picture of a progressively decoded and refreshed picture, where pictures in the decoded picture buffer that are before the current picture in decoding order are not output until the pictures in the decoded picture buffer are restored.

28. The non-transitory computer-readable storage medium according to claim 27, wherein, The transformation is performed according to a second format rule. Wherein, the second format rule specifies that the value of a second variable indicating whether to output the picture in the second access unit is determined based on a flag indicating whether to output another picture in the second access unit.

29. The non-transitory computer-readable storage medium according to claim 28, wherein, The other picture is in a layer higher than the picture in the second access unit.

30. The non-transitory computer-readable storage medium according to claim 28, wherein, The flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a slice group header.

31. The non-transitory computer-readable storage medium according to claim 27, wherein, When neither i) nor ii) is satisfied, the value of the third variable is set to be equal to the value of the flag.

32. The non-transitory computer-readable storage medium according to claim 27, wherein, The value of the third variable being equal to 0 indicates not to output the picture in the access unit.

33. The non-transitory computer-readable storage medium according to claim 31, wherein, The third variable is PictureOutputFlag, and the flag is pic_output_flag.

34. The non-transitory computer-readable storage medium according to claim 31, wherein, The flag controls the decoded picture output and removal process.

35. The non-transitory computer-readable storage medium according to claim 31, wherein, The flag is included in an access unit delimiter, an auxiliary enhancement information field, or a picture header of one or more pictures.

36. The non-transitory computer-readable storage medium according to claim 27, wherein, The transformation is performed according to a fourth format rule. Wherein, the fourth format rule specifies that when the picture in the fourth access unit is not included in the output layer, the value of a fourth variable indicating whether to output the picture is set to be equal to a certain value.

37. The non-transitory computer-readable storage medium according to claim 36, wherein, The certain value is 0.

38. The non-transitory computer-readable storage medium according to claim 27, wherein, The transformation includes encoding the video into the bitstream.

39. The non-transitory computer-readable storage medium according to claim 27, wherein, The transformation includes decoding the video from the bitstream.

40. A method for storing a bitstream of a video, comprising: generating the bitstream of the video having one or more video layers according to a first format rule; and storing the bitstream in a non-transitory computer-readable recording medium, wherein, the first format rule specifies that when the video includes only one output layer, the first access unit that does not include the output layer is encoded and decoded by setting the value of a first variable indicating whether to output the picture in the first access unit to a first value for the picture having the highest layer identifier (ID) value and a second value for all other pictures, the first value is 1, and the second value is 0, and wherein, the generating is also performed according to a third format rule. Among them, the value of a third variable specified by the third formatting rule for indicating whether to output a picture is determined based on at least one of the following: a flag specifying the value of an identifier of a video parameter set (VPS), whether the current video layer is an output layer, whether the current picture is a random access skipped previous picture, a progressive decoding refresh picture, or a recovery picture of a progressive decoding refresh picture, or whether a picture in the decoded picture buffer before the current picture in decoding order is output before the picture in the decoded picture buffer is recovered. Among them, the value of the third variable is set to be equal to 0 when i) the flag specifying the value of the identifier of the VPS is greater than 0 and the current video layer is not an output layer, or ii) one of the following conditions is true: The current picture is a random access skipped previous picture, and an associated intra random access point picture in the decoded picture buffer before the current picture in decoding order is not output before the associated intra random access point picture is recovered; or The current picture is a progressive decoding refresh picture, where pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures in the decoded picture buffer are recovered, or the current picture is a recovery picture of a progressive decoding refresh picture, where pictures in the decoded picture buffer before the current picture in decoding order are not output before the pictures in the decoded picture buffer are recovered.

41. The method according to claim 40, wherein, The generation is performed according to the second formatting rule. Among them, the value of a second variable specified by the second formatting rule for indicating whether to output a picture in a second access unit is determined based on a flag indicating whether to output another picture in the second access unit.

42. The method according to claim 41, wherein, The another picture is in a layer higher than the picture in the second access unit.

43. The method according to claim 41, wherein, The flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a slice header, or a slice group header.

44. The method according to claim 40, wherein When neither i) nor ii) is satisfied, the value of the third variable is set to be equal to the value of the flag.

45. The method according to claim 40, wherein The value of the third variable being equal to 0 indicates not to output the picture in the access unit.

46. The method according to claim 44, wherein, The third variable is PictureOutputFlag, and the flag is pic_output_flag.

47. The method according to claim 44, wherein, The flag controls the decoded picture output and removal process.

48. The method according to claim 44, wherein, The flag is included in an access unit delimiter, an auxiliary enhancement information field, or a picture header of one or more pictures.

49. The method according to claim 40, wherein, The generation is performed according to the fourth formatting rule. Among them, the fourth formatting rule specifies that when a picture in a fourth access unit is not included in the output layer, the value of a fourth variable for indicating whether to output the picture is set to be equal to a certain value.

50. The method according to claim 49, wherein, The certain value is 0.

Citation Information

Patent Citations

  • Improved inference of no output of prior pictures flag in video coding

    CN105900426A

  • Signaling change in output layer sets

    US20200077107A1