USE OF VIDEO PARAMETER SET IN VIDEO CODING

MX430945BActive Publication Date: 2026-02-25BYTEDANCE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
MX2022011208
Authority / Receiving Office
MX · MX
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-17
Filing Date
2022-09-08
Publication Date
2026-02-25
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

The existing scalability design in VVC text has issues with specifying maximum image width, height, chroma format, and bit depth for all layers, affecting DPB memory allocation and image output flags, leading to inefficiencies and incorrect image handling.

Method used

Specify maximum chroma format and bit depth values for all layers in the VPS, and adjust DPB parameters based on these values, ensuring consistent and correct handling of image output flags across all layers.

Benefits of technology

Ensures efficient and accurate DPB memory allocation and image output handling, resolving issues with layer-specific and AU-specific flag settings, improving decoder performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure MX430945B0
    Figure MX430945B0
Patent Text Reader

Abstract

Methods and devices for video processing, including encoding and decoding, are described. An example video processing method includes performing a conversion between a video and a bitstream of the video according to a format rule, wherein the bitstream includes one or more output layer sets (OLS), each OLS comprising one or more encoded layer video sequences, and wherein the format rule specifies that a set of video parameters indicates, for each of the one or more OLS, a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video.
Need to check novelty before this filing date? Find Prior Art

Description

USE OF VIDEO PARAMETER SET IN VIDEO CODING Cross-reference to Related Applications Under patent law and / or applicable rules pursuant to the Paris Convention, this application is made to timely claim priority and benefits of U.S. Provisional Patent Application No. 62 / 990,749, filed on March 17, 2020. For all purposes under the law, all disclosures of the foregoing application are incorporated by reference as part of the disclosure of this application. Field of invention This patent document relates to the encoding and decoding of images and video. Background of the invention Digital video accounts for the largest use of bandwidth on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth for digital video use is expected to continue growing. Brief description of the invention This document discloses techniques that can be used by video encoders and decoders to process coded video representations using control information useful for decoding the coded representation. In one example, a video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers comprising one or more video frames and an encoded representation of the video; wherein the encoded representation includes a set of video parameters indicating a maximum value for a chroma format indicator and / or a maximum bit depth used to represent pixels of the video. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and an encoded representation of the video, where the encoded representation conforms to a format rule that specifies a maximum image width and / or a maximum image height for video images from all video layers. This rule controls the value of a variable that indicates whether images in a decoder buffer are produced before being removed from the decoder buffer. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to an encoded representation of the video. The encoded representation conforms to a format rule that specifies the maximum value of a chroma format indicator and / or the maximum bit depth used to represent video control pixels. This value indicates whether images in a decoder buffer are produced before being removed from the decoder buffer. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to an encoded representation of the video. The encoded representation conforms to a format rule that specifies that the value of a variable indicating whether images in a decoder buffer are produced before being removed from the decoder buffer is independent of whether separate color planes are used to encode the video. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to an encoded representation of the video. The encoded representation conforms to a format rule that specifies that a variable indicating whether images in a decoder buffer are produced before being removed from the decoder buffer is included in the encoded representation at the access unit (AU) level. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to an encoded representation of the video. The encoded representation conforms to a format rule that specifies that an image output flag for a video image on an access unit is determined based on a `pic_output_flag` variable from another video image on the access unit. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to an encoded representation of the video. The encoded representation conforms to a format rule that specifies that, for a video image not belonging to an output layer, an output image flag value is used. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video and a video bitstream according to a format rule, wherein the bitstream includes one or more output layer sets (OLS), each OLS comprising one or more encoded layer video sequences, and wherein the format rule specifies that a set of video parameters indicates, for each of the one or more OLS, a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent video pixels. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies a maximum image width and / or a maximum image height for video images from all video layers. This rule controls the value of a variable that indicates whether images in a decoded image buffer prior to a current image in the decoding order in the bitstream are processed before images are removed from the decoded image buffer. In another example, another video processing method is disclosed. The method includes performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and where the format rule specifies a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video control a value of a variable that indicates whether the images in a decoded image buffer prior to a current image in decoding order in the bit sequence are produced before the images are removed from the decoded image buffer. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a rule, and wherein the rule specifies that a value of a variable indicating whether the images in a decoded image buffer prior to a current image in the decoding order in the bitstream occur before the images are removed from the decoded image buffer is independent of whether separate color planes are used to encode the video. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a flag value indicating whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer when a certain type of access unit is decoded is included in the bitstream. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a first flag value indicating whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer when decoding an access unit of a particular type is not indicated in an image header. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that the value of a flag associated with an access unit, which indicates whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer, depends on a flag value for each image in the access unit. In another example, a different video processing method is disclosed. The method involves performing a conversion between a video that has one or more video layers and a video bitstream according to a format rule, and where the format rule specifies that the value of a variable indicating whether an image is to be produced in an access unit is determined based on a flag indicating whether another image is to be produced in the access unit. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers into a video bitstream according to a format rule. The format rule specifies that a variable indicating whether an image should be produced in an access unit is set to a certain value if the image does not belong to an output layer. In another example, a different video processing method is disclosed. This method involves converting a video with one or more video layers to a video bitstream according to a format rule. The format rule specifies that if the video comprises only one output layer, an access unit that does not include an output layer is encoded by setting a variable that indicates whether an image should be produced in the access unit. This variable assigns a first value to the image with the highest layer ID and a second value to all other images. In yet another example, a video encoder is disclosed. The video encoder comprises a processor configured to implement the methods described above. In yet another example, a video decoder device is disclosed. The video decoder comprises a processor configured to implement the methods described above. In yet another example, a computer-readable medium containing stored code is disclosed. The code incorporates one of the methods described herein in the form of processor-executable code. These and other features are described throughout this document. Brief description of the drawings Figure 1 is a block diagram of an example video processing system. Figure 2 is a block diagram of a video processing device. Figure 3 is a flowchart for an example video processing method. Figure 4 is a block diagram illustrating a video coding system according to some modalities of the present disclosure. Figure 5 is a block diagram illustrating an encoder according to some modalities of the present disclosure. Figure 6 is a block diagram illustrating a decoder according to some modalities of the present disclosure. Figures 7A to 7D show flowcharts for example video processing methods based on some implementations of the disclosed technology. Figures 8A to 8C show flowcharts for example video processing methods based on some implementations of the disclosed technology. Figures 9A to 9C show flowcharts for example video processing methods based on some implementations of the disclosed technology. Detailed description of the invention The section headings used herein are for ease of understanding and do not limit the applicability of the techniques and methods disclosed in each section to that section only. Furthermore, H.266 terminology is used in some descriptions only for ease of understanding and not to limit the scope of the techniques disclosed. As such, the techniques described herein are also applicable to other video codec protocols and designs. 1. Initial Analysis This patent document relates to video coding technologies. Specifically, it concerns the signaling of decoded image buffer (DPB) parameters for DPB memory allocation, as well as specifying the output of decoded images in scalable video coding, where a video bitstream can contain more than one layer. The ideas can be applied individually or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding, for example, the Versatile Video Coding (VVC) currently under development. 2. Abbreviations APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding CLVS Encoded Layer Video Sequence CPB Encoded image buffer CRA Clean Random Access CTU Coding Tree Unit CVS Encoded video sequence DCI Decoding Capability Information DPB Decoded image buffer EOB End of bit stream EOS End of sequence GDR Gradual Decoding Update HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instant Decoding Update JEM Joint Exploration Model MCTS Set of tiles with movement restriction NAL Network Abstraction Layer OLS Output Layer Set PH image header PPS Image Parameter Set PTL Profile, Grade and Level PU Imaging Unit RAP Random Access Point RBSP Raw Byte Sequence Payload SEI Supplementary Improvement Information SPS Sequence parameter set ινΐΛ / a / zuzz / uii ¿uo iviA / a / zuzz / uii zuo SVC Scalable Video Coding VCL Video Encoding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding 3. Introduction to video coding Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes both time prediction and transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly founded the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into the reference software called the Joint Exploration Model (JEM).The JVET meeting is held quarterly, and the new encoding standard aims for a 50% reduction in bitrate compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the April 2018 JVET meeting, and the first version of the VVC Test Model (VTM) was released at that time. As there is ongoing effort to standardize VVC, new encoding techniques are being adopted according to the VVC standard at each JVET meeting. The VVC Working Draft and the VTM Test Model are updated after each meeting. The VVC project is now aiming for Technical Finalization (FDIS) at the July 2020 meeting. 3.1. Scalable Video Coding (SVC) in general and in VVC Scalable video coding (SVC, sometimes also called scalability in video coding) refers to video coding that uses a base layer (BL), sometimes called a reference layer (RL), and one or more scalable enhancement layers (ELs). In SVC, the base layer can carry video data with a basic level of quality. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to a previously coded layer. For example, a lower layer can serve as a BL, while a higher layer can serve as an EL. Middle layers can serve as either an EL or an RL, or both.For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL (Effective Layer) for the layers below it, such as the base layer or any intermediate enhancement layers, and simultaneously serve as an RL (Relative Layer) for one or more enhancement layers above it. Similarly, in the Multiview or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to encode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancies). In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the encoding level (e.g., video level, sequence level, image level, sector level, etc.) at which they can be used. For example, parameters that can be used by one or more encoded video sequences at different layers in the bitstream can be included in a video parameter set (VPS), and parameters used by one or more images in an encoded video sequence can be included in a sequence parameter set (SPS). Similarly, parameters used by one or more sectors in an image can be included in an image parameter set (PPS), and other parameters specific to a single sector can be included in a sector header.Similarly, the indication of which set(s) of parameters a particular layer is using at any given time can be provided at various encoding levels. Thanks to Reference Picture Resampling (RPR) support in VVC, supporting a bitstream containing multiple layers—for example, two layers with SD and HD resolutions—in VVC can be done without any additional signal processing-level coding tools, since the extra sampling required for spatial scalability support can simply be done using the RPR upsampling filter. However, high-level syntax changes (compared to not supporting scalability) are required for scalability support. Scalability support is specified in VVC version 1. Unlike scalability support in any previous video coding standard, including AVC and HEVC extensions, VVC scalability design has been made as easy as possible for single-layer decoder designs.The decoding capability for multilayer bitstreams is specified as if there were only one layer in the bitstream. For example, the decoding capability, such as the DPB size, is specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, a decoder designed for single-layer bitstreams doesn't require much modification to decode multilayer bitstreams. Compared to the multilayer extension designs of AVC and HEVC, HLS has been significantly simplified at the expense of some flexibility. For example, an IRAP ALI is required to hold an image for each of the layers present in the CVS. 3.2. Random access and its supports in HEVC and VVC Random access refers to the initial access and decoding of a bitstream starting from an image that is not the first image in the bitstream in the decoding order. To support tuning and channel switching in broadcast / multicast and multiparty videoconferencing, search in playback and live streaming, as well as live-stream adaptation within live streaming, the bitstream must include frequent random access points, which are typically intracoded images but can also be intercoded images (for example, in the case of gradual decoding updates). HEVC includes intra-random access point (IRAP) image signaling at the NAL unit header, across NAL unit types. Three types of IRAP images are supported: instant decoder update (IDR) images, clean random access (CRA) images, and broken link access (BLA) images. IDR images restrict the inter-image prediction structure to not reference any images prior to the current image group (GOP), conventionally known as closed GOP random access points. CRA images are less restrictive, allowing certain images to reference images prior to the current GOP, all of which are discarded in the event of random access. CRA images are conventionally known as open GOP random access points.BLA images typically originate from splicing two bitstreams or parts thereof into a CRA image, for example, during flow switching. To enable better use of IRAP imaging systems, six different NAL units are defined to signal IRAP image properties. These can be used to better match the flow access point types as defined in the ISO Base Media File Format (ISOBMFF), which is used for Dynamic Adaptive Live Streaming (DASH) support. VVC supports three IRAP image types, two IDR image types (one with and one without associated RADL images), and one CRA image type. These are essentially the same as in HEVC. BLA image types, as used in HEVC, are not included in VVC, primarily for two reasons: i) The basic functionality of BLA images can be achieved using CRA images plus the NAL sequence unit terminator, the presence of which indicates that the subsequent image initiates a new CVS in a single-layer bitstream; ii) During VVC's development, there was a desire to specify fewer NAL unit types than HEVC, as evidenced by the use of five bits instead of six for the NAL unit type field in the NAL unit header. Another key difference in random access support between VVC and HEVC is VVC's more normative support for GDR. In GDR, decoding a bitstream can begin with an intercoded image, and although the entire image region cannot be correctly decoded initially, after a certain number of images, the complete image region will be correctly decoded. AVC and HEVC also support GDR, using the SEI recovery point message to signal GDR random access points and recovery points. In VVC, a new type of NAL unit is specified for indicating GDR images, and the recovery point is signaled in the image header syntax structure. A CVS and a bitstream are both allowed to begin with a GDR image. This means that a complete bitstream can contain only intercoded images without a single intracoded image.The main benefit of specifying GDR support in this way is to provide conformance behavior for GDR. GDR allows encoders to smooth the bitrate of a bitstream by distributing intracoded sectors or blocks across multiple frames instead of intracoding entire frames. This enables a significant reduction in end-to-end latency, which is considered more important today than ever before, as ultra-low latency applications such as wireless display, online gaming, and drain-based applications become more popular. Another feature related to GDR in VVC is virtual boundary signaling. The boundary between the updated region (i.e., the correctly decoded region) and the unupdated region in an image between a GDR image and its retrieval point can be signaled as a virtual boundary. When signaled, no loop filtering is applied across the boundary, thus preventing decoding mismatch for some samples at or near the boundary. This can be useful when the application determines to display correctly decoded regions during the GDR process. IRAP images and GDR images can be collectively referred to as random access point (RAP) images. 3.3. Parameter Sets AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all versions of AVC, HEVC, and VVC. VPS was introduced in HEVC and is included in both HEVC and VVC. APS was not included in AVC or HEVC but is included in the latest draft version of VVC. SPS was designed to carry sequence-level header information, and PPS was designed to carry infrequently variable image-level header information. With SPS and PPS, it is not necessary to repeat the infrequently variable information for each sequence or image, thus avoiding redundant signaling of this information. Furthermore, the use of SPS and PPS allows for out-of-band transmission of important header information, thereby not only eliminating the need for redundant transmissions but also improving fault resilience. VPS was introduced to carry sequence-level header information that is common to all layers in multilayer bitstreams. APS was introduced to carry this information at the image or sector level, which needs a lot of bits to encode, can be shared through multiple images, and in a sequence there can be quite a few different variations. 3.4. Related Definitions in VVC The related definitions in the latest VVC text (in JVET-Q2001-vE / v15) are as follows. associated IRAP image (of a particular image): The previous IRAP image in decoding order (when present) that has the same nuhjayerjd value as the particular image. Coded video sequence (CVS): An AU sequence consisting, in decoding order, of a CVSS AU, followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AUs that are CVSS AUs. Coded Video Sequence Start (CVSS) AU: An AU in which there is a PU for each layer in the CVS and the encoded image in each PU is a CLVSS image. Gradual Decoding Update (GDR) AU: An AU in which the encoded image in each PU present is a GDR image. Gradual Decoding Update PU (GDR): A PU in which the encoded image is a GDR image. ινΐΛ / a / zuzz / uii ¿uo Gradual Decoding Update (GDR) image: An image for which every VCL NAL unit has nal_unit_type equal to GDRNUT. Intra-random access point (IRAP) AU: An AU in which there is a PU for each layer in the CVS and the encoded image in each PU is an IRAP image. Intra-random access point image (IRAP): An encoded image for which all VCL NAL units have the same nal_unit_type value in the range from IDR_W_RADL to CRA_NUT, inclusive. front image: An image that is in the same layer as the associated IRAP image and precedes the associated IRAP image in output order. rear image: A non-IRAP image that follows the associated IRAP image in output order and is not an STSA image. NOTE - Subsequent images associated with an IRAP image also follow the IRAP image in decoding order. Images that follow the associated IRAP image in output order and precede the associated IRAP image in decoding order are not allowed. 3.5. Syntax and semantics of VPS in VVC VVC supports scalability, also known as scalable video coding, where multiple layers can be encoded into one encoded video bitstream. In the latest VVC text (in JVET-Q2001-vE / v15), scalability information is signaled in the VPS, for which the syntax and semantics are as follows. 7.3.2.2 Video Parameter Set Syntax video parameter set rbsp() { Descriptor vps video parameter set id u(4) vps max layers minusl u(6) vps max sublayers minusl u(3) if(vps max layers minusl > 0 && vps max sublayers minusl > 0) vps all layers same num sublayers flag u(1) if(vps max layers minusl > 0) vps all independen! layers flag u(1) for(¡ = 0; i <= vps max layers minusl; i++) { vps layer id[¡] u(6) if(i > 0 && Ivps all independent layers flag) { vps independent layer flag[¡] u(1) if(!vps independent layer flagfi]) f for(¡ = 0; ¡ < i; j++) vps direct ref layer flag[i][j] u(1) max tid ref present flag[i] u(1) if(max tid ref present flagfi]) max tid il ref pies pluslfi] u(3)}} if(vps max layers minusl > 0) { if(vps all independent layers flag) each layer is an oís flag u(1) if(!each layer is an oís flag) { if(!vps all independent layers flag) oís mode ide u(2) if(ols mode ¡de = = 2) { num output layer sets minusl u(8) for(¡ = 1; i <= num output layer sets minusl; i ++) for(¡ = 0; ¡ <= vps max layers minusl;¡++) oís output layer flag[i][j] u(1)}} vps num ptls minusl u(8) for(¡ = 0; i <= vps num ptls minusl; i++) { if(¡ > 0) pt present flag [i] u(1) if(vps_max_sublayers_minus1 >0 && Ivps all layers same num sublayers flag) ptl max temporal id[i] u(3) j while(!byte alignedO) vps ptl alignment zero bit / * igual a 0 7 f(D for(¡ = 0; i <= vps num ptls minusl; i++) profile tier level(pt present flag[i], ptl max temporal id[i]) for(¡ = 0; i < TotaINumOIss; i++) if(vps num ptls minusl > 0) oís ptl idxfi] u(8) if(!vps all independent layers flag) vps num dpb params ue(v) if(vps. num. dpb. .params > 0 && vps..max. sublayers..minusl >0); vps sublayer dpb params present flag u(1) for(¡ = 0; i < vps num dpb params; i++) { Íf(vps_max_sublayers_minus1 >0 && !vps all layers same num sublayers flag) dpb max temporal idfi] u(3) dpb_parameters(dpb_max_temporal_id[i], vps sublayer dpb params present flag)} for(¡ = 0; i < TotaINumOIss; I++) { if(NumLayerslnOls[i] > 1) { oís dpb pie widthfi] ue(v) oís dpb pie heightfi] ue(v) if(vps num dpb params >1) oís dpb params idxfi] ue(v)}} if(!each layer is an oís flag) vps general hrd params present flag u(1) if(vps general hrd params present flag) { general hrd parameters() if(vps max sublayers minusl > 0) vps sublayer cpb params present flag U(1) num oís hrd params minusl ue(v) for(i = 0; i <= num oís hrd params minusl;i++) { if(vps_max_sublayers_minus1 >0 && Ivps all layers same num sublayers flag) hrd max tid[i] u(3) firstSubLayer = vps_sublayer_cpb_params_present_flag ? 0 hrd max tidfi] oís hrd parametersffirstSubLayer, hrd max tidfi]) if(num_ols_hrd_params_minus1 + 1 != TotaINumOIss && num oís hrd params minusl > 0) for(¡ = 1; i < TotaINumOIss; I++) if(NumLayerslnOls[i] > 1) ois hrd idxfi] ue(v)} vps extension flag u(1) if(vps extension flag) whilefmore rbsp data()) vps extension data flag u(1) rbsp trailing bits()}; 7.4.3.2 RBSP Semantics of Video Parameter Set An RBSP VPS will be available for the decoding process before it is referenced, included in at least one AU with Temporalld equal to 0 or provided through external means. All NAL VPS units with a particular vps_video_parameter_set_id value in a CVS must have the same content. `vps_video_parameter_set_id` provides an identifier for the VPS for reference by other syntax elements. The value of `vps_video_parameter_set_id` must be greater than 0. vps_max_layers_minus1 plus 1 specifies the maximum number of layers allowed on each CVS with reference to the VPS. vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporary sublayers that can be present in a layer in each CVS with reference to the VPS. The value of vps_max_sublayers_minus1 must be in the range of 0 and 6, inclusive. `vps_all_layers_same_num_sublayers_flag` equal to 1 specifies that the number of temporary sublayers is the same for all layers in each CVS that refer to the VPS. `vps_all_layers_same_num_sublayers_flag` equal to 0 specifies that the layers in each CVS that refer to the VPS may or may not have the same number of temporary sublayers. When not present, the value of `vps_all_layers_same_num_sublayers_flag` is assumed to be 1. `vps_all_independent_layers_flag` equal to 1 specifies that all layers in the CVS are independently encoded without using interlayer prediction. `vps_all_independent_layers_flag` equal to 0 specifies that one or more of the layers in the CVS may use interlayer prediction. When not present, the value of `vps_all_independent_layers_flag` is inferred to be 1. vps_layer_id[i] specifies the value of nuh layerjd of the i-th layer. For any two non-negative integer values ​​of m and n, when m is less than n, the value of vps_layer_id[m] must be less than vps_layer_id[n]. `vps_independent_layer_flag[i]` equal to 1 specifies that the layer with index i does not use interlayer prediction. `vps_independent_layer_flag[i]` equal to 0 specifies that the layer with index i can use interlayer prediction and the syntax elements `vps_direct_ref_layer_flag[i][j]` for j in the range 0 to i - 1 inclusive are present in VPS. When not present, it is inferred that the value of `vps_independent_layer_flag[i]` is equal to 1. `vps_direct_ref_layer_flag[i][j]` equal to 0 specifies that the layer with index j is not a direct reference layer for the layer with index i. `vps_direct_ref_layer_flag[i][j]` equal to 1 specifies that the layer with index j is a direct reference layer for the layer with index i. When `vps_direct_ref_layer_flag[i][j]` is not present for i and j in the range from 0 to `vps_max_layers_minus1`, inclusive, it is inferred to be equal to 0. When `vps_independent_layer_flag[i]` is equal to 0, there will be at least one value of j in the range from 0 to `i - 1`, inclusive, such that the value of `vps_direct_ref_layer_flag[i][j]` is equal to 1. The variables NumDirectRefLayers[¡], DirectRefLayerldx[i][d], NumRefLayers[i],RefLayerldx [i][r] and LayerUsedAsRefLayerFlag[j] are derived as follows: for(¡ = 0; i <= vps_max_layers_minus1; i++) { for(j = 0; j <= vps_max_layers_minus1; j++) { dependentFlag[i][j] = vps_direct_ref_layer_flag[i][j] for(k = 0; k < i; k++) if(vps_direct_ref_layer_flag[i][k] && dependencyFlag[k][j]) dependency Flag[i][j] = 1} LayerUsedAsRefLayerFlag[¡] = 0 for(¡ = 0; i <= vps_max_layers_minus1; i++) { for(j = 0, d = 0, r = 0; j <= vps_max_layers_minus1; j++) { (37) if(vps_direct_ref_layer_flag[¡][j]) { DirectRef Layerldx[i][d++] = j LayerllsedAsRefLayerFlag[j] = 1} if (dependency Flag[i][j]) RefLayerldx[i][r++] = j} NumDirectRefLayers[¡] = d NumRefLayers[¡] = r} La variable GeneralLayerldx[i], que especifica el índice de capa de la capa con nuhjayerjd igual a vps_layer_¡d[¡], se deriva como sigue: for(¡ = 0; i <= vps_max_layers_minus1; i++) (38) GeneralLayerldx[vps_layer_id[¡]] = i For any two different values ​​of i and j, both in the range from 0 to vps_max_layers_minus1 inclusive, when dependencyFlag[i][j] equals 1, it is a bitstream conformance requirement that the values ​​of chroma_format_idc and bit_depth_minus8 that apply to the i-th layer be equal to the values ​​of chroma_format_idc and bit_depth_minus8, respectively, that apply to the j-th layer. max_tid_ref_present_flag[i] equal to 1 specifies that the syntax element max_tid_il_ref_pics_plus1[i] is present. max_tid_ref_present_flag[i] equal to 0 specifies that the syntax element max_tid_il_ref_pics_plus1[i] is not present. max_tid_il_ref_pics_plus1[i] equal to 0 specifies that cross-layer prediction is not used by non-IRAP images in the i-th layer. max_tid_il_ref_pics_plus1[i] greater than 0 specifies that, to decode images in the i-th layer, no image with Temporalld greater than max_tid_il_ref_pics_plus1[i] - 1 is used as an ILRP. When not present, it is inferred that max_tid_il_ref_pics_plus1[i] is equal to 7. `each_layer_is_an_ols_flag` equal to 1 specifies that each OLS contains only one layer, and each layer itself in a CVS that refers to the VPS is an OLS with its individual included layer, which is the only output layer. `each_layer_is_an_ols_flag` equal to 0 indicates that an OLS can contain more than one layer. If `vps_max_layers_minus1` equals 0, it is inferred that the value of `each_layer_is_an_ols_flag` is equal to 1. Otherwise, when `vps_all_independent_layers_flag` equals 0, it is inferred that the value of `each_layer_is_an_ols_flag` is equal to 0. ols_mode_idc equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_maxjayers_minus1 + 1, the i-th OLS includes layers with layer indices from 0 to i inclusive, and for each OLS only the highest layer in the OLS is generated. ols_mode_idc equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes layers with layer indices from 0 to i inclusive, and for each OLS all layers in the OLS are generated. ols_mode_idc equal to 2 specifies that the total number of OLSs specified by the VPS is explicitly indicated and for each OLS the output layers are explicitly signaled and other layers are the layers that are direct or indirect reference layers of the OLS output layers. The value of ols_mode_idc must be in the range of 0 to 2, inclusive. The value 3 of ols_mode_idc is reserved for future use by ITU-T | ISO / IEC. When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, it is inferred that the value of ols_mode_idc is equal to 2. num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode idc equals 2. The variable TotaINumOIss, specifying the total number of OLSs specified by the VPS, is derived as follows: if(vps_max_layers_minus1 == 0) TotaINumOIss = 1 else if(each_layer_is_an_ols_flag | | ols_mode_idc = = 0 | | ols_mode_idc = = 1) TotaINumOIss = vps_maxlayers_minus1 + 1 (39) else if(ols_modejdc = = 2) TotaINumOIss = num_outputjayer_sets_minus1 + 1 ols_outputjayerjlag[i][j] equal to 1 specifies that the layer with nuhjayerjd equal to vps_layer_id[j] is an output layer of the i-th OLS when ols_mode_idc equals 2. ols_output_layer_flag[i][j] equal to 0 specifies that the layer with nuhjayerjd equal to vpsjayer id[j] is not an output layer of the i-th OLS when ols_modejdc equals 2. The variable NumOutputLayerslnOls[i], which specifies the number of output layers in the i-th OLS, the variable NumSubLayersInLayerlnOLS[i][j], which specifies the number of sublayers in the j-th layer in the i-th OLS, the variable OutputLayerldlnOls[i]|JL, which specifies the value of the j-th output layer in the i-th OLS, and the variable LayerUsedAsOutputLayerFlag[k], which specifies whether the k-th layer is used as an output layer in at least one OLS, are derived as follows: NumOutputLayerslnOls[0] = 1 OutputLayerldlnOls[0][0] = vps_layerJd[0] NumSubLayerslnLayerlnOLS[0][0] = vps_rnax_sub_layers_minus1 + 1 LayerUsedAsOutputLayerFlag[0] = 1 for(¡ = 1, i <= vps_max_layers_minus1; i++) { if(each_layer_is_an_ols_flag | | ols_mode_idc < 2) LayerUsedAsOutputLayerFlag[¡] = 1 else / *(!each_layer_is_an_ols_flag && oís modejdc == 2)* / LayerUsedAsOutputLayerFlag[¡] = 0 for(i = 1; i < TotaINumOIss; i++) if(each_layer_is_an_ols_flag | | ols_mode_idc = = 0) { NumOutputLayerslnOls[¡] = 1 OutputLayerldlnOls[¡][0] = vps_layer_id[¡] for(j = 0; j < i && (ols_mode_idc = = 0); j++) NumSubLayerslnLayerlnOLS[i][j] = max_tid_il_ref_pics_plus1[i] NumSubLayerslnLayerlnOLS[¡][¡] = vps_max_sub_layers_minus1 + 1} else if(ols_mode_idc = = 1) { NumOutputLayerslnOls[¡] = i + 1 for(j = 0; j < NumOutputLayerslnOls[¡]; j++) { OutputLayerldl nOls[i][j] = vps_layer_id[j] NumSubLayerslnLayerlnOLS[i][j] = vps_max_sub_layers_minus1 + 1} } else if(ols_mode_idc = = 2) { for(j = 0; j <= vps_max_layers_minus1; j++) { layerlncludedlnOlsFlag[i][j] = 0 NumSubLayerslnLayerlnOLS[i][j] = 0} for(k = 0, j = 0; k <= vps_max_layers_minus1; k++) (40) if(ols_output_layer_flag[i][k]) { layerlncludedlnOlsFlag[i][k] = 1 LayerUsedAsOutputLayerFlag[k] = 1 OutputLayerldx[i][j] = k OutputLayerldl nOls[i][j++] = vps_layer_id[k] NumSubLayerslnLayerlnOLS[i][j] = vps_max_sub_layers_minus1 + 1 NumOutputLayerslnOls[i] = j for(j = 0; j < NumOutputLayerslnOls[¡]; j++) { idx = OutputLayer ldx[i][j] for(k = 0; k < NumRefLayers[idx]; k++) { layerlncludedlnOlsFlag[i][Ref Layerldx[idx][k]j = 1 ¡f(NumSubLayerslnLayerlnOLS[i][RefLayerldx[idx][k]] < max_tid_il_ref_pics_plus1 [OutputLayerldlnOls[i][j]]) NumSubLayerslnLayerlnOLS[i][RefLayerldx[idx][k]] = max_tid_il_ref_pics_plus1 [OutputLayerIdlnOls[i][j]]}}} For each value of i in the range from 0 to vps_maxjayers_minus1 inclusive, the values ​​of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] must not be equal to 0. In other words, there will not be any layer that is an output layer of at least one OLS or a direct reference layer of any other layer. For each OLS, there will be at least one layer that is an output layer. In other words, for any value of i in the range from 0 to TotaINumOIss - 1, inclusive, the value of NumOutputLayersInOls[i] will be greater than or equal to 1. The variable NumLayerslnOls[i], which specifies the number of layers in the i-th OLS, and the variable LayerldlnOls[i][j], which specifies the nuhjayerjd value of the j-th layer in the i-th OLS, are derived as follows: NumLayerslnOls[0] = 1 LayerldlnOls[0][0] = vps_layer_id[O] for(¡ = 1; i < TotaINumOIss; i++) { if(each_layer_is_an_ols_flag) { NumLayerslnOls[¡] = 1 LayerldlnOls[¡][0] = vpsjayerjd[¡] (41)} else if(ols_mode_idc = = 0 | | ols_mode_idc = = 1) { NumLayerslnOls[¡] = i + 1 for(j = 0; j < NumLayerslnOls[¡]; j++) LayerldlnOls[i][j] = vps_layer_id[j]} else if(ols_mode_idc = = 2) { for(k = 0, j = 0; k <= vps_max_layers_minus1; k++) if(layerlncludedlnOlsFlag[i][k]) Layerld lnOls[i][j++] = vpsjayerid[k] NumLayerslnOls[¡] = j} } iviA / a / zuzz / u i i zuo NOTA 1 - El OLS 0 contiene solo la capa más baja (es decir, la capa con nuhjayerjd igual a vpsJayerJd[0]) y para el OLS 0 se produce la única capa incluida. La variable OlsLayerldx[i][j], especificando el índice de capa OLS de la capa con nuhjayerjd igual a LayerldlnOls [i][j], se deriva como sigue: for(¡ = 0; i < TotaINumOIss; i++) for j = 0; j < NumLayerslnOls[i]; j++) (42) OlsLayerldx[i][Layerldl nOls[i][j]] = j The lowest layer in each OLS will be an independent layer. In other words, for each i in the range from 0 to TotaINumOIss - 1, inclusive, the value of vpsJndependentJayerJlag[GeneralLayerldx [LayerldlnOls [i][0]]] will be equal to 1. Each layer will be included in at least one OLS specified by the VPS. In other words, for each layer with a particular value of nuhjayerjd nuhLayerld, equal to one of vpsjayerjd[k] for k in the range of 0 to vps_maxjayers_minus1, inclusive, there will be at least one pair of iyj values, where i is in the range of 0 to TotaINumOIss - 1, inclusive, and j is in the range of NumLayersInOls [i] - 1, inclusive, such that the value of LayerldlnOls [i][j] is equal to nuhLayerld. vps_num_ptls_minus1 plus 1 specifies the number of profilejierjevel() syntax structures in the VPS. The value of vps_num_ptls_minus1 must be less than TotalInumOIss. `pt_presentjlag[i]` equal to 1 specifies that profile, level, and general constraint information are present in the i-th `profilejierjevel()` syntax structure on the VPS. `pt_presentjlag[i]` equal to 0 specifies that profile, level, and general constraint information are not present in the i-th `profileJierJevel()` syntax structure on the VPS. It is inferred that the value of `pt_presentJlag[0]` is equal to 1. When `pt_presentJlag[i]` is equal to 0, it is inferred that the profile, level, and general constraint information for the i-th `profilejierlevel()` syntax structure on the VPS is the same as for the (i-1)-th `profilejierjevel()` syntax structure on the VPS. ptl_maxjemporaljd[i] specifies the Temporald of the highest sublayer representation for which level information is present in the i-th profilejierjevel() syntax structure in the VPS. The value of ptl_maxjemporaljd[i] must be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, it is inferred that the value of ptl_maxjemporaljd[i] is equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_alljayers_same_num_sublayersjlag is equal to 1, it is inferred that the value of ptl_maxjemporaljd[i] is equal to vps_max_sublayers_minus1. vps_ptl_alignment_zero_bit must be equal to 0. ols_ptljdx[i] specifies the index, in the list of profilejierjevel() syntax structures in the VPS, of the profilejierjevel() syntax structure that applies to the i-th OLS. When present, the value of olsptljdx[i] will be in the range of 0 to vps_num_ptls_minus1, inclusive. When vps_num_ptls_minus1 is equal to 0, it is inferred that the value of ols_ptljdx[i] is equal to 0. When NumLayersInOls[i] equals 1, the profiletierlevel() syntax structure that applies to the i-th OLS is also present in the SPS referenced by the layer in the i-th OLS. It is a bitstream conformance requirement that, when NumLayersInOls[i] equals 1, the profiletierlevel() syntax structures signaled in the VPS and in the SPS for the i-th OLS will be identical. `vps_num_dpb_params` specifies the number of `dpb_parameters()` syntax structures in the VPS. The value of `vps_num_dpb_params` must be between 0 and 16, inclusive. When not present, the value of `vps_num_dpb_params` is assumed to be 0. The `vps_sublayer_dpb_params_present_flag` flag is used to check for the presence of the syntax elements `max_dec_pic_buffering_minus1[], max_num_reorder_pics[],` and `max_latency_increase_plus1[]` in the `dpb_parameters()` syntax structures in the VPS. When it is not present, it is inferred that `vps_sub_dpb_params_info_present_flag` is equal to 0. `dpb_max_temporal_id[!]` specifies the Temporald of the highest sublayer representation for which DPB parameters can be present in the i-th `dpb_parameters()` syntax structure in the VPS. The value of `dpb_max_temporal_id[!]` must be in the range of 0 to `vps_max_sublayers_minus1`, inclusive. When `vps_max_sublayers_minus1` is equal to 0, it is inferred that the value of `dpb_max_temporal_id[!]` is equal to 0. When `vps_max_sublayers_minus1` is greater than 0 and `vps_all_layers_same_num_sublayers_flag` is equal to 1, it is inferred that the value of `dpb_max_temporal_id[!]` is equal to `vps_max_sublayers_minus1`. ols_dpb_pic_width[¡] specifies the width, in luma sample units, of each image storage buffer for the i-th OLS. ols_dpb_pic_height[¡] specifies the height, in luma sample units, of each image storage buffer for the i-th OLS. ols_dpb_params_idx[i] specifies the index, to the list of dpb_parameters() syntax structures in the VPS, of the dpb_parameters() syntax structure that is applied to the i-th OLS when NumLayerslnOls[i] is greater than 1. When present, the value of ols_dpb_params_idx[i] must be in the range of 0 to vps_num_dpb_params - 1, inclusive. When ols_dpb_params_idx[i] is not present, it is inferred that the value of ols_dpb_params_idx[i] is equal to 0. When NumLayerslnOls[¡] is equal to 1, the dpb_parameters() syntax structure that applies to the i-th OLS is present in the SPS referenced by the layer in the i-th OLS. The `vps general hrd params present flag` flag set to 1 specifies that the VPS contains a `general_hrd_parameters()` syntax structure and other HRD parameters. A `vps_general_hrd_params_present_flag` flag set to 0 specifies that the VPS does not contain a `general_hrd_parameters()` syntax structure or other HRD parameters. When not present, the value of `vps_general_hrd_params_present_flag` is assumed to be 0. When NumLayerslnOls[¡] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure that apply to the i-th OLS are present in the SPS referenced by the layer in the i-th OLS. `vps_sublayer_cpb_params_present_flag` equal to 1 specifies that the i-th `ols_hrd_parameters()` syntax structure in the VPS contains HRD parameters for sublayer representations with Temporalld in the range of 0 to hrd_max_tid[i], inclusive. `vps_sublayer_cpb_params_present_flag` equal to 0 specifies that the i-th `ols_hrd_parameters()` syntax structure in the VPS contains HRD parameters for the sublayer representation with Temporalld equal to hrd_max_tid[i] only. When `vps_max_sublayers_minus1` is equal to 0, it is inferred that the value of `vps_sublayer_cpb_params_present_flag` is equal to 0. When vps_sublayer_cpb_params_present_flag is equal to 0, it is inferred that the HRD parameters for sublayer representations with Temporalld in the range of 0 to hrd_max_tid[i] - 1, inclusive, are the same as for the sublayer representation with Temporalld equal to hrd_max_tid[i]. These include the HRD parameters starting from the syntax element fixed_pic_rate_general_flag[i] up to the syntax structure sublayer hrd parameters(i) immediately under the condition if(general_vcl_hrd_params_present_flag) in the syntax structure ols_hrd_parameters. num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present on the VPS when vps_general_hrd_params_present_flag equals 1. The value of num_ols_hrd_params_minus1 must be in the range of 0 to TotaINumOIss - 1, inclusive. hrd_max_tid[i] specifies the Temporalld of the upper sublayer representation for which the HRD parameters are contained in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] must be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, it is inferred that the value of hrd_max_tid[i] is equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, it is inferred that the value of hrd_max_tid[i] is equal to vps_max_sublayers_minus1. ols_hrd_idx[i] specifies the index, to the list of ols_hrd_parameters() syntax structures in the VPS, of the ols_hrd_parameters() syntax structure that is applied to the i-th OLS when NumLayerslnOls[i] is greater than 1. The value of ols_hrd_idx[[i] must be in the range of 0 to num_ols_hrd_params_minus1, inclusive. When NumLayerslnOls[¡] is equal to 1, the syntax structure ols_hrd_parameters() that applies to the i-th OLS is present in the SPS referenced by the layer in the i-th OLS. If the value of num_ols_hrd_param_minus1 + 1 is equal to TotaINumOIss, it is inferred that the value of ols_hrd_idx[i] is equal to i. Otherwise, when NumLayersInOls [i] is greater than 1 and num_ols_hrd_params_minus1 is equal to 0, it is inferred that the value of ols_hrd_idx[[i]] is equal to 0. vps_extension_flag equal to 0 specifies that there are no vps_extension_data_flag syntax elements present in the VPS RBSP syntax structure. vps_extension_flag equal to 1 specifies that there are vps_extension_data_flag syntax elements present in the VPS RBSP syntax structure. `vps_extension_data_flag` can have any value. Its presence and value do not affect the decoder's conformance to the profiles specified in this version of this specification. Decoders that conform to this version of this specification will ignore all elements of `vps_extension_data_flag` syntax. 3.6. Syntax and semantics of SPS in VVC In the latest VVC text (in JVET-Q2001-vE / v15), the syntax and semantics of SPS that are most relevant to the inventions herein are as follows. 7.3.2.3 Sequence Parameter Set RBSP Syntax seq parameter set rbsp() { Descriptor ue(v) gdr enabled flag u(1) chroma format idc u(2) ue(v) a) bit depth minus8 ue(v) ue(v)} 7.4.3.3 RBSP Semantics of Sequence Parameter Set gdr_enabled_flag equal to 1 specifies that GDR images may be present in CLVS that refer to the SPS. gdr_enabled_flag equal to 0 specifies that GDR images are not present in CLVS that refer to the SPS. chroma_format_idc specifies chroma sampling relative to luma sampling as specified in clause 6.2. bit_depth_minus8 specifies the bit depth of the luma and chroma array samples, BitDepth, and the value of the luma and chroma quantization parameter interval offset, QpBdOffset, as follows: BitDepth = 8 + bit_depth_minus8 (45) QpBdOffset = 6 * bit_depth_minus8 (46) bit_depth_minus8 must be between 0 and 8, inclusive. 3.7. Syntax and semantics of image header structure in VVC In the latest VVC text (in JVET-Q2001-vE / v15), the image header structure syntax and semantics that are most relevant to the inventions herein are as follows. 7.3.2.7 Image Header Structure Syntax picture header structure() { Descriptor gdr or irap pie flag u(1) if(gdr or irap pie flag) gdr pie flag u(1) ph pie order cnt Isb u(v) if(gdr or irap pie flag) no output of prior pies flag u(1) if(gdr pie flag) recovery poc cnt ue(v) ue(v)} 7.4.3.7 Image Header Structure Semantics The PH syntax structure contains information that is common to all sectors of the encoded image associated with the PH syntax structure. gdr_or_irap_pic_flag equal to 1 specifies that the current image is a GDR or IRAP image. gdr_or_irap_pic_flag equal to 0 specifies that the current image may or may not be a GDR or IRAP image. A `gdr_pic_flag` value of 1 specifies that the image associated with the PH is a GDR image. A `gdr_pic_flag` value of 0 specifies that the image associated with the PH is not a GDR image. When it is not present, the value of `gdr_pic_flag` is assumed to be 0. When `gdr_enabled_flag` is 0, the value of `gdr_pic_flag` must also be 0. NOTE 1: When gdr or irap pic flag is equal to 1 and gdr pic flag is equal to 0, the image associated with the PH is an IRAP image. `ph_pic_order_cnt_lsb` specifies the image order count modulo `MaxPicOrderCntLsb` for the current image. The length of the syntax element `ph_pic_order_cnt_lsb` is `Iog2_max_pic_order_cnt_lsb_minus4` + 4 bits. The value of `ph_pic_order_cnt_lsb` must be in the range of 0 to `MaxPicOrderCntLsb - 1`, inclusive. no_output_of_prior_pics_flag affects the output of previously decoded images in the DPB after decoding a CLVSS image that is not the first image in the bitstream as specified in Annex C. `recovery_poc_cnt` specifies the recovery point of decoded images in output order. If the current image is a GDR image associated with the PH, and there is a `picA` image following the current GDR image in decoding order in the CLVS with a `PicOrderCntVal` equal to the `PicOrderCntVal` of the current GDR image plus the value of `recovery_poc_cnt`, then the `picA` image is known as the recovery point image. Otherwise, the first image in output order with a `PicOrderCntVal` greater than the `PicOrderCntVal` of the current image plus the value of `recovery_poc_cnt` is known as the recovery point image. The recovery point image must not precede the current GDR image in decoding order. The value of `recovery_poc_cnt` must be in the range of 0 to `MaxPicOrderCntLsb` - 1, inclusive. When the current image is a GDR image, the RpPicOrderCntVal variable is derived as follows: RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81) NOTE 2: When gdr_enabled_flag is equal to 1 and PicOrderCntVal of the current image is greater than or equal to RpPicOrderCntVal of the GDR's associated image, the current and subsequent decoded images in output order exactly match the corresponding images produced when starting the decoding process of the previous IRAP image, when present, before the GDR's associated image in decoding order. 3.8. PictureOutputFIag Setting In the latest VVC text (in JVET-Q2001-vE / v15), the specification for adjusting the value of the PictureOutputFIag variable is as follows (as part of the process of decoding clause 8.1.2 for an encoded picture). 8.1.2 Decoding process for an encoded image The decoding processes specified in this clause are applied to each encoded image, called the current image and denoted by the variable CurrPic, in BitstreamToDecode. Depending on the value of chroma_format_idc, the number of sample arrays of the current image is as follows: - If chroma_format_idc is equal to 0, the current image consists of 1 array of Sl samples. - Otherwise (chroma_format_idc is not equal to 0), the current image consists of 3 sample arrays Sl, Scb, Ser. The decoding process for the current image takes as inputs the syntax elements and uppercase variables in clause 7. When interpreting the semantics of each syntax element in each NAL unit, and in the remaining parts of clause 8, the term the bitstream (or part of it, e.g., a CVS of the bitstream) refers to BitstreamToDecode (or part of it). Depending on the value of separate_colour_plane_flag, the decoding process is structured as follows: - If separate_colour_plane_flag is equal to 0, the decoding process is invoked only once with the current image being the output. Otherwise (separate_colour_plane_flag equals 1), the decoding process is invoked three times. The inputs for the decoding process are all NAL units from the encoded image with the same color_plane_id value. The process of decoding NAL units with a particular color_plane_id value is specified as if only one monochrome color-formatted CVS with that particular color_plane_id value were present in the bitstream. The output of each of the three decoding processes is mapped to one of the three sample arrays of the current image, with the NAL units having color_plane_ids of 0, 1, and 2 being mapped to Sl, Scb, and Ser, respectively. NOTE: The ChromaArrayType variable is derived as equal to 0 when separate_colour_plane_flag is equal to 1 and chroma_format_idc is equal to 3. In the decoding process, the value of this variable is evaluated, resulting in operations identical to those of monochrome images (when chroma_format_idc is equal to 0). The decoding process operates as follows for the current CurrPic image: 1. The decoding of NAL units is specified in clause 8.2. 2. The processes in clause 8.3 specify the following decoding processes using syntax elements in and above the sector header layer: - The variables and functions related to the image order count are derived as specified in clause 8.3.1. This should only be invoked for the first sector of an image. - At the beginning of the decoding process for each sector of a non-IDR image, the decoding process for building reference image lists specified in clause 8.3.2 is invoked for derivation of reference image list 0 (RefPicList [0]) and reference image list 1 (RefPicList [1]). The decoding process for marking reference images is invoked in clause 8.3.3, where reference images can be marked as either “not used for reference” or “used for long-term reference.” This should only be invoked for the first sector of an image. - When the current image is a CRA image with NoOutputBeforeRecoveryFlag equal to 1 or a GDR image with NoOutputBeforeRecoveryFlag equal to 1, the decoding process is invoked to generate unavailable reference images specified in subclause 8.3.4, which should be invoked only for the first sector of an image. - PictureOutputFIag is set as follows: - If one of the following conditions is met, PictureOutputFIag is set to 0: - the current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - If gdr_enabled_flag is equal to 1, the current image is associated with a GDR image with NoOutputBeforeRecoveryFlag equal to 1, and PicOrderCntVal of the current image is less than RpPicOrderCntVal of the associated GDR image. - spsvideoparametersetid is greater than 0, ols_mode_idc is equal to 0 and the current AU contains a picA image that meets all of the following conditions: - PicA has PictureOutputFIag equal to 1. - PicA has a nuh_layer_id nuhLid greater than that of the current image. - PicA belongs to the output layer of the OLS (i.e., OutputLayerldlnOls[TargetOlsldx][Q] is equal to nuhLid). - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2 and ols_output_layer_flag[TargetOlsldx][GeneralLayerldx[nuh_layer_id]] is equal to 0. - Otherwise, PictureOutputFIag is set equal to pic_output_flag. 3. The processes in clauses 8.4, 8.5, 8.6, 8.7, and 8.8 specify the decoding processes using syntax elements at all layers of the syntax structure. It is a bitstream conformance requirement that the encoded sections of the image contain section data for iviA / a / zuzz / uii ¿uo each CTU of the image, so that the division of the image into sectors and the division of the sectors into CTUs form a partition of the image. 4. After all sectors of the current image have been decoded, the current decoded image is marked as used for short-term reference, and each ILRP entry in RefPicList[0] or RefPicList[1] is marked as used for short-term reference. 3.9. Adjusting DPB parameters for HRD operations In the latest VVC text (in JVET-Q2001-vE / v15), the specification for adjusting DPB parameters for HRD operations is as follows (as part of clause C.1). C.1 General For each bitstream conformance test, the CPB size (number of bits) is CpbSize[Htid][Scldx] as specified in clause 7.4.6.3, where Scldx and the HRD parameters are specified earlier in this clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found in or derived from the dpb_parameters() syntax structure that applies to the target OLS as follows: - If the target OLS contains only one layer, the dpb_parameters() syntax structure is found in the SPS that is referenced as the layer in the target OLS. - Otherwise (the target OLS contains more than one layer), dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsldx] located on the VPS. 3.10. Setting NoOutputOfPriorPicsFlag In the latest VVC text (in JVET-Q2001-vE / v15), the specifications for adjusting the value of the NoOutputOfPriorPicsFlag variable are as follows (as part of the specifications for removing images from the DPB). C.3.2 Removal of DPB images before decoding the current image The removal of images from the DPB before decoding the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously at the time of CPB removal of the first DU of AU n (which contains the current image) and proceeds as follows: - The decoding process is invoked for building the reference image list as specified in clause 8.3.2 and the decoding process is invoked for marking reference images as specified in clause 8.3.3. - When the current AU is a CVSS AU that is not AU 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: If the value of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1 [Htid] derived for any image in the current AU is different from the value of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1 [Htid], respectively, derived for the previous image in the same CLVS, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of the value of iviA / a / zuzz / uii ¿uo no_output_of_prior_pics_flag. NOTE: Although it is preferred to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test is applied to the HRD, so that when the NoOutputOfPriorPicsFlag value is equal to 1, all image storage buffers in the DPB are emptied without output of the images they contain, and the fullness of the DPB is set to 0. - When both conditions are met for any k image in the DPB, all k images in the DPB are eliminated: - Image k is marked as not used for reference. - image k has PictureOutputFIag equal to 0 or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current image n; i.e., DpbOutputTime[k] is less than or equal to DuCpbRemovalTimefm]. - For each image that is removed from the DPB, the DPB fullness is decreased by one. C.5.2.2 DPB Image Output and Deletion The removal and deletion of images from the DPB before the decoding of the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously when the first DU of the AU containing the current image is removed from the CPB and proceeds as follows: - The decoding process for building a list of reference images is invoked as specified in clause 8.3.2 and the decoding process for marking reference images is invoked as specified in clause 8.3.3. - If the current image is a CLVSS image other than image 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: - If the value of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid] derived for any image in the current AU is different from the value of pic_width_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc, separate_colour_plane_flag, bit_depth_minus8, or max_dec_pic_buffering_minus1[Htid], respectively, for the previous image in the same CLVS, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag. NOTE: Although it is preferred to set NoOutputOfPriorPicsFlag equal to no_output_of_prior_pics_flag under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test applies to iviA / a / zuzz / uii zuo the HRD as follows: - If NoOutputOfPriorPicsFlag equals 1, all image storage buffers in the DPB are emptied without output of the images they contain and the DPB fullness is set to 0. - Otherwise (NoOutputOfPriorPicsFlag equals 0), all image storage buffers containing an image that is marked as not needed for output and not used for reference are emptied (no output) and all non-empty image storage buffers in the DPB are emptied by repeatedly invoking the offset process specified in clause C.5.2.4 and the fullness of DPB is set to 0. Otherwise (the current image is not a CLVSS image or the CLVSS image is image 0), all image buffers containing an image marked as not needed for output and not used for reference are flushed (no output). For each image buffer that is flushed, the DPB fullness is decreased by one. When one or more of the following conditions are met, the offset process specified in clause C.5.2.4 is repeatedly invoked while further decreasing the DPB fullness by one for each additional image buffer that is flushed, until none of the following conditions are met: - The number of images in the DPB that are marked as required for output is greater than max_num_reorder_pics[Htid]. - max_latencyjncrease_plus1 [Htid] is not equal to 0 and there is at least one image in the DPB that is marked as required for output for which the associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid]. - The number of images in the DPB is greater than or equal to max_dec_pic_buffering_minus1 [Htid] + 1. 4. Technical problems solved by the disclosed technical solutions The existing scalability design in the latest VVC text (in JVET-Q2001-vE / v15) has the following problems: 1) Currently, the maximum image width and height values ​​for all images across all layers are signaled in the VPS to allow the decoder to correctly allocate memory for the DPB. Like image width and height, the chroma format and bit depth, currently specified by the SPS syntax elements chroma_format_idc and bit_depth_minus8, respectively, also affect the size of an image buffer in the DPB. However, the maximum values ​​for chroma_format_idc and bit_depth_minus8 for all images across all layers are not signaled. 2) Currently, adjusting the value of the NoOutputOfPriorPicsFlag variable involves changing the values ​​of p¡c_width_max_jn_luma_samples and p¡c_height_maxjn_luma_samples. However, the maximum image width and height values ​​should be used for all images in all layers. 3) Currently, adjusting NoOutputOfPriorPicsFlag involves changing the value of chroma_format_idc or bit_depth_minus8. However, the maximum chroma format and bit depth values ​​should be used for all images in all layers. 4) Currently, the NoOutputOfPriorPicsFlag setting involves changing the separate_colour_plane_flag value. However, the separate_colour_plane_flag is only present and used when chroma_format_idc is equal to 3, which specifies the 4:4:4 chroma format. For the 4:4:4 chroma format, a separate_colour_plane_flag value of 0 or 1 does not affect the buffer size required to store a decoded image. Therefore, the NoOutputOfPriorPicsFlag setting should not involve changing the separate_colour_plane_flag value. 5) Currently, the no_output_of_prior_pics_flag is signaled in the PH for IRAP and GDR images, and both the semantics of this flag and the process for setting NoOutputOfPriorPicsFlag are specified in a layer-specific or PU-specific manner. However, since the DPB operation is OLS-specific or AU-specific, both the semantics of no_output_of_prior_pics_flag and the use of this flag in setting NoOutputOfPriorPicsFlag must be specified in an AU-specific manner. 6) The current text for adjusting the value of the PictureOutputFIag variable for a current image implies using PictureOutputFIag from an image in the same AU as the current image and on a higher layer than the current image. However, for a picA image that has a higher nuh_layer_id than that of the current image, when PictureOutputFIag of the current image is derived, PictureOutputFIag of picA has not yet been derived. 7) The current text for adjusting the PictureOutputFIag variable has a problem, as described below. There are two layers in the OLS bitstream, and only the top layer is an output layer. In a particular AU, auA, the top-layer image has a pic_output_flag of 0. On the decoder side, the top-layer image of auA is not present (due, for example, to loss or a down-layer change), while the bottom-layer image of auA is present and has a pic_output_flag of 1. Therefore, the PictureOutputFIag value of the bottom-layer image of auA would be set to 1. However, when an OLS has only one output layer and an output-layer image has a pic_output_flag of 0, it should be interpreted that the encoder (or content provider) did not intend for an image to be produced for the AU containing that image. 8) The current text for adjusting the PictureOutputFlag variable has a problem, as described below. There are three or more layers in the OLS bitstream, and only the top layer is an output layer. On the decoder side, the top-layer image for the current AU is not present (due, for example, to loss or a down-layer change), while two or more images from the lower layers for the current AU are present, and these images have pic_output_flag equal to 1. Therefore, more than one image would be produced for this AU. However, this is problematic because there is only one output layer for the OLS, so the encoder or content provider expected the output of a single image. 9) The current text for adjusting the value of the PictureOutputFIag variable has a problem as described below. OLS mode 2 (when ols_mode_idc is equal to 2) can also specify only an output layer like mode 0, but the behavior of outputting a lower layer image for an AU when the output layer image (which is also the highest layer image) is not present is only specified for mode 0. 10) For an OLS containing only one output layer, when the output layer image (which is also the top layer image) is unavailable (due, for example, to loss or a down-layer change) to the decoder, the decoder will not be able to tell if the `pic_output_flag` of that image was equal to 1 or 0. If it was equal to 1, then it makes sense to produce a lower-layer image, but if it was equal to 0, it could be worse from a user experience perspective to produce a lower-layer image since the encoder (content provider) set the value to 0 for a reason—for example, there shouldn't be an image output for this AU for this particular OLS. 5. A list of modalities and techniques To solve the problems mentioned above, and others, methods are presented and summarized below. The listed elements should be considered as examples to illustrate the general concepts and should not be interpreted restrictively. Furthermore, these elements can be applied individually or combined in any way. Solutions to solving problems 1 to 5 1) To solve problem 1, one or both of the maximum values ​​of chroma_format_idc and bit_depth_minus8 for all images of all layers can be signaled in the VPS. 2) To solve problem 2, you can specify that the setting of the NoOutputOfPriorPicsFlag variable value is based on at least one or both of the maximum image width and height for all images from all layers that can be signaled on the VPS. 3) To solve problem 3, you can specify that the setting of the NoOutputOfPriorPicsFlag variable value is based on at least one or both of the maximum values ​​of chroma_format_idc and blt_depth_minus8 for all images of all layers that can be signaled on the VPS. 4) To solve problem 4, you can specify that the adjustment of the NoOutputOfPriorPicsFlag variable value is independent of the separate_colour_plane_flag value. 5) To solve problem 5, both the semantics of no_output_of_prior_pics_flag and the use of this flag in the setting of NoOutputOfPriorPicsFlag can be specified in an AU-specific way. a. In one example, it may be required that, when present, the value of no_output_of_prior_pics_flag be the same for all images in an AU, and the value of no_output_of_prior_pics_flag of the AU is considered to be the value of no_output_of_prior_pics_flag of the images in the AU. b. Alternatively, in one example, the no_output_of_prior_pics_flag can be removed from the PH syntax and the AUD syntax can be signaled when irap_or_gdr_au_flag equals 1. i. For single-layer bitstreams, since AUD is optional, it can be inferred that the value of no_output_of_prior_pics_flag is equal to 1 when AUD is not present for an IRAP or GDR AU (that is, if the encoder wants to signal a value of 0 for no_output_of_prior_pics_flag for an IRAP or GDR AU in a single-layer bitstream, it has to signal AUD for that AU in the bitstream). c. Alternatively, in one example, the no_output_of_prior_pics_flag value of an AU can be considered equal to 0 if and only if no_output_of_prior_pics_flag for each image in the AU is equal to 0, and otherwise the no_output_of_prior_pics_flag value of the AU can be considered equal to 1. i. The drawback of this approach is that the NoOutputOfPriorPicsFlag setting and the image output from a CVSS AU must wait for all images to arrive at the AU. Solutions for solving problems 6 to 10 6) To solve problem 6, you can specify that the PictureOutputFIag setting for a current image is based at least on pic_output_flag (instead of PictureOutputFIag) of an image in the same AU as the current image and on a higher layer than the current image. 7) To solve problems 7 through 9, the PictureOutputFIag value for a current picture is set to 0 whenever the current picture does not belong to an output layer. a. Alternatively, to solve problems 7 and 8, when there is only one output layer, and the output layer (which must be the top layer when there is only one output layer) is not present for an AU, then PictureOutputFIag is set equal to 1 for the picture that has the highest nuhjayerjd value among all the pictures in the AU available to the decoder and that have pic_outputjlag equal to 1, and is set equal to 0 for all other pictures in the AU available to the decoder. 8) To solve problem 10, the pic_outputjlag value of an output layer image in an AU can be signaled in the AUD or an SEI message in the AU, or signaled in the PH of one or more additional images in the AU. 6. Modalities The following are some example embodiments for some of the aspects of the invention summarized above in Section 5, which can be applied to the VVC specification. The amended texts are based on the latest VVC text in JVET-Q2001-vE / v15. Most of the relevant parts that have been added or amended are highlighted in italics and bold, and some of the deleted parts are set off by double square brackets (e.g., [[a]] denotes the deletion of the character 'a'). There are some other changes that are editorial in nature and are therefore not highlighted. 6.1. First modality This modality is for elements 1, 2, 3, 4, 5, and 5a. i ¿uo 7.3.2.2 Video Parameter Set Syntax video parameter set rbsp() { Descriptor for(i = 0; i < TotaINumOIss; I++) { if(NumLayerslnOls[i] > 1) { oís dpb pie widthfi] ue(v) oís dpb pie height[¡] ue(v) oís dpb chroma format[¡] u(2) oís dpb bitdepth mlnus8[i] ue(v) if(vps num dpb params>1) oís dpb params idx[¡] ue(v)}} uo 7.4.3.2 RBSP Semantics of Video Parameter Set ols_dpb_pic_width[¡] specifies the width, in luma sample units, of each picture storage buffer for the i-th OLS. ols_dpb_pic_height[i] specifies the height, in luma sample units, of each image storage buffer for the i-th OLS. ols_dpb_chroma_format[i] specifies the maximum allowed value of chroma_format_idc for all SPS referenced by CL VS in the CVS for the i-th OLS. ols_dpb_bitdepth_minus8[i] specifies the maximum allowed value of bit_depth_minus8 for all SPS referenced by the CL VS in the CVS for the i-th OLS. NOTE 2 - To decode an OLS containing more than one layer and having an OLS index i, the decoder can safely allocate memory for the DPB according to the values ​​of the syntax elements ols_dpb_pic_width[i], ols_dpb_pic_height[i], ols_dpb_chroma_format[i] and ols_dpb_bitdepth_minus8[i]. ols_dpb_params_idx[i] specifies the index, to the list of dpb_parameters() syntax structures in the VPS, of the dpb_parameters() syntax structure that is applied to the i-th OLS when NumLayerslnOls[i] is greater than 1. When present, the value of ols_dpb_params_idx[i] must be in the range of 0 to vps_num_dpb_params - 1, inclusive. When ols_dpb_params_idx[i] is not present, it is inferred that the value of ols_dpb_params_idx[i] is equal to 0. When NumLayerslnOls[i] is equal to 1, the dpb_parameters() syntax structure that applies to the i-th OLS is present in the SPS referenced by the layer in the i-th OLS. 7.4.3.3 RBSP Semantics of Sequence Parameter Set gdr_enabled_flag equal to 1 specifies that GDR images may be present in CLVS that refer to the SPS. gdr_enabled_flag equal to 0 specifies that GDR images are not present in CLVS that refer to the SPS. chroma_format_idc specifies chroma sampling relative to luma sampling as specified in clause 6.2. When sps_video_parameter_set_id is greater than 0, it is a bitstream conformance requirement that, for any OLS with OLS index i containing one or more layers that refer to the SPS, the value of chroma_format_idc be less than or equal to the value of ols_dpb_chroma_format[i]. bit_depth_minus8 specifies the bit depth of the luma and chroma array samples, BitDepth, and the value of the luma and chroma quantization parameter interval offset, QpBdOffset, as follows: BitDepth = 8 + bit_depth_minus8 (45) QpBdOffset - 6* bit_depth_minus8 (46) bit_depth_minus8 must be between 0 and 8, inclusive. When sps_video_parameter_set_id is greater than 0, it is a bitstream conformance requirement that, for any OLS with OLS index i containing one or more layers that refer to the SPS, the value of bit_depth_minus8 be less than or equal to the value of ols_dpb_bitdepth_minus8[!]. 7.4.3.7 Image header structure semantics no_output_of_prior_pics_flag affects the output of previously decoded images in the DPB after decoding an image in a CVSS AU that is not the first AU in the bitstream as specified in Annex C. It is a bitstream conformance requirement that, when present, the value of no_output_of_prior_pics_flag be the same for all images in an AU. When no_output_of_prior_pics_flag is present in the PH of the images of an AU, the no_output_of_prior_pics_flag value of the AU is the no_output_of_prior_pics_flag value of the images of the AU. C.1 General For each bitstream conformance test, the CPB size (number of bits) is CpbSize[Htid][Scldx] as specified in clause 7.4.6.3, where Scldx and the HRD parameters are specified earlier in this clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found in or derived from the dpb_parameters() syntax structure that is applied to the target OLS as follows: - If NumLayersInOlsjTargetOlsIdx] is equal to 1, the dpb_parameters() syntax structure is located in the SPS referenced as the layer in the target OLS, and the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat and MaxBitDepthMinusS are set to pic_w¡dth_max_in_luma_samples, pic_height_max_in_luma_samples, chroma_format_idc and bit_depth_minus8 respectively, which are located in the SPS referenced as the layer in the target OLS. - Otherwise (the target OLS contains more than one layer), dpb_parameters() is identified by ols_dpb_params_idx[TargetOlsldx] located on the VPS, and the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat and MaxBitDepthMinusS are set equal to ols_dpb_pic_width[TargetOlsldx], ols_dpb_pic_height[TargetOlsldx], ols_dpb_chroma_format[TargetOlsldx] and ols_dpb_bitdepth_minus8, respectively, located on the VPS. C.3.2 Removal of DPB images before decoding the current image The removal of images from the DPB before decoding the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously at the time of CPB removal of the first DU of AU n (which contains the current image) and proceeds as follows: - The decoding process is invoked for building the reference image list as specified in clause 8.3.2 and the decoding process is invoked for marking reference images as specified in clause 8.3.3. - When the current AU is a CVSS AU that is not AU 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: - If the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid], respectively, derived for the previous AU in decoding order, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag of the current AU. NOTE: Although it is preferred to set NoOutputOfPriorPicsFlag equal to the current AU's no_output_of_prior_pics_flag under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag of the current AU. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test is applied to the HRD, so that when the NoOutputOfPriorPicsFlag value is equal to 1, all image storage buffers in the DPB are emptied without output of the images they contain, and the fullness of the DPB is set to 0. - When both conditions are met for any k image in the DPB, all k images in the DPB are eliminated: - Image k is marked as not used for reference. - image k has PictureOutputFIag equal to 0 or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current image n; i.e., DpbOutputTime[k] is less than or equal to DuCpbRemovalTime[m]. - For each image that is removed from the DPB, the DPB fullness is decreased by one. C.5.2.2 DPB Image Output and Deletion The removal and deletion of images from the DPB before the decoding of the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously when the first DU of the AU containing the current image is removed from the CPB and proceeds as follows: - The decoding process for building a list of reference images is invoked as specified in clause 8.3.2 and the decoding process for marking reference images is invoked as specified in clause 8.3.3. - If the current AU is a CVSS AU that is not AU 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: - If the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid], respectively, derived for the previous AU in decoding order, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of the value of no_output_of_prior_pics_flag of the current AU. NOTE: Although it is preferred to set NoOutputOfPriorPicsFlag equal to the current AU's no_output_of_prior_pics_flag under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag of the current AU. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test is applied to the HRD as follows: - If NoOutputOfPriorPicsFlag equals 1, all image storage buffers in the DPB are emptied without output of the images they contain and the DPB fullness is set to 0. - Otherwise (NoOutputOfPriorPicsFlag equals 0), all image storage buffers containing an image that is marked as not needed for output and not used for reference are emptied (no output) and all non-empty image storage buffers in the DPB are emptied by repeatedly invoking the offset process specified in clause C.5.2.4 and the fullness of DPB is set to 0. Otherwise (the current image is not a CLVSS image or the CLVSS image is image 0), all image buffers containing an image marked as not needed for output and not used for reference are flushed (no output). For each image buffer that is flushed, the DPB fullness is decreased by one. When one or more of the following conditions are met, the offset process specified in clause C.5.2.4 is repeatedly invoked while further decreasing the DPB fullness by one for each additional image buffer that is flushed, until none of the following conditions are met: - The number of images in the DPB that are marked as required for output is greater than max_num_reorder_pics[Htid]. - max_latency_increase_plus1 [Htid] is not equal to 0 and there is at least one image in the DPB that is marked as required for output for which the associated variable PicLatencyCount is greater than or equal to iviA / a / zuzz / uii zuo MaxLatencyPictures[Htid]. - The number of images in the DPB is greater than or equal to max_dec_pic_buffering_minus1 [Htid] + 1. 6.2. Second modality This modality is for elements 1, 2, 3, 4, 5, and 5c, with text changes compared to the text of the first modality. 7.4.3.7 Image header structure semantics no_output_of_prior_pics_flag affects the output of previously decoded images in the DPB after decoding an image in a CVSS AU that is not the first AU in the bitstream as specified in Annex C. [[It is a bitstream conformance requirement that, when present, the value of no_output_of_prior_pics_flag be the same for all images in an AU. When no_output_of_prior_pics_flag is present in the PH of the images of an AU, the value of no_output_of_prior_pics_flag of the AU is the value of no_output_of_prior_pics_flag of the images of the AU. C.3.2 Removal of DPB images before decoding the current image The removal of images from the DPB before decoding the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously at the time of CPB removal of the first DU of AU n (which contains the current image) and proceeds as follows: - The decoding process is invoked for building the reference image list as specified in clause 8.3.2 and the decoding process is invoked for marking reference images as specified in clause 8.3.3. - When the current AU is a CVSS AU that is not AU 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: - If the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid], respectively, derived for the previous AU in decoding order, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of whether the value of no_output_of_prior_pics_flag [[of the current AU]] is equal to 0 for each image in the current AU. NOTE: When no_output_of_prior_pics_flag is equal to 0 for each image in the current AU, [[Although]] although it is preferred to set NoOutputOfPriorPicsFlag equal to 0 [[equal to no_output_of_prior_pics_flag of the current AU]] under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, [[NoOutputOfPriorPicsFlag is set equal to no_output_of_prior_pics_flag of the current AU]] if no_output_of_prior_pics_flag is equal to 0 for each image in the current AU, NoOutputOfPriorPicsFlag is set equal to 0. iviA / a / zuzz / u ι ί ¿uo - Otherwise, NoOutputOfPriorPicsFlag is set to 1. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test is applied to the HRD, so that when the NoOutputOfPriorPicsFlag value is equal to 1, all image storage buffers in the DPB are emptied without output of the images they contain, and the fullness of the DPB is set to 0. - When both conditions are met for any k image in the DPB, all k images in the DPB are eliminated: - Image k is marked as not used for reference. - image k has PictureOutputFIag equal to 0 or its DPB output time is less than or equal to the CPB removal time of the first DU (denoted as DU m) of the current image n; i.e., DpbOutputTime[k] is less than or equal to DuCpbRemovalTimefm]. - For each image that is removed from the DPB, the DPB fullness is decreased by one. C.5.2.2 DPB Image Output and Deletion The removal and deletion of images from the DPB before the decoding of the current image (but after analyzing the sector header of the first sector of the current image) occurs instantaneously when the first DU of the AU containing the current image is removed from the CPB and proceeds as follows: - The decoding process for building a list of reference images is invoked as specified in clause 8.3.2 and the decoding process for marking reference images is invoked as specified in clause 8.3.3. - If the current AU is a CVSS AU that is not AU 0, the following steps apply in order: 1. The NoOutputOfPriorPicsFlag variable is derived for the decoder under test as follows: - If the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid] derived for the current AU is different from the value of PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, MaxBitDepthMinus8 or max_dec_pic_buffering_minus1 [Htid], respectively, derived for the previous AU in decoding order, NoOutputOfPriorPicsFlag can (but should not) be set to 1 by the decoder under test, regardless of whether the value of no_output_of_prior_pics_flag [[of the current AU]] is equal to 0 for each image in the current AU. NOTE: When no_output_of_prior_pics_flag is equal to 0 for each image in the current AU, although it is preferred to set NoOutputOfPriorPicsFlag equal to 0 [equal to no_output_of_prior_pics_flag of the current AU] under these conditions, the decoder under test may set NoOutputOfPriorPicsFlag to 1 in this case. - Otherwise, if no_output_of_prior_pics_flag is equal to 0 for each image in the current AU, NoOutputOfPriorPicsFlag is set to 0 [[equal to no output of prior pics flag of the current AU]]. - Otherwise, NoOutputOfPriorPicsFlag is set to 1. 2. The NoOutputOfPriorPicsFlag value derived for the decoder under test is applied to the HRD as follows: i ¿uo - If NoOutputOfPriorPicsFlag equals 1, all image storage buffers in the DPB are emptied without output of the images they contain and the DPB fullness is set to 0. - Otherwise (NoOutputOfPriorPicsFlag equals 0), all image storage buffers containing an image that is marked as not needed for output and not used for reference are emptied (no output) and all non-empty image storage buffers in the DPB are emptied by repeatedly invoking the offset process specified in clause C.5.2.4 and the fullness of DPB is set to 0. Otherwise (the current image is not a CLVSS image or the CLVSS image is image 0), all image buffers containing an image marked as not needed for output and not used for reference are flushed (no output). For each image buffer that is flushed, the DPB fullness is decreased by one. When one or more of the following conditions are met, the offset process specified in clause C.5.2.4 is repeatedly invoked while further decreasing the DPB fullness by one for each additional image buffer that is flushed, until none of the following conditions are met: - The number of images in the DPB that are marked as needed for output is greater than max_num_reorder_pics[Ht¡d]. - max_latency_increase_plus1 [Htid] is not equal to 0 and there is at least one image in the DPB that is marked as required for output for which the associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid]. - The number of images in the DPB is greater than or equal to max_dec_pic_buffering_minus1 [Htid] + 1. 6.3. Third modality This modality is for element 6, element 7 (the modified texts except for the NOTE added in clause 8.1.2) and element 7a (the NOTE added in clause 8.1.2). 7.4.3 .7 Image header structure semantics recovery_poc_cnt specifies the recovery point of decoded images in output order. When the current image is a GDR image, the RecoveryPointPocVal variable is derived as follows: recoveryPointPocVal = PicOrderCntVal + recovery_poc_cnt(81) If the current image is a GDR image associated with the PH, and there is a picA image following the current GDR image in the CLVS decoding order that has PicOrderCntVal equal to recoveryPointPocVal [[the PicOrderCntVal of the current GDR image plus the recovery_poc_cnt value]], the picA image is known as the recovery point image. Otherwise, the first image in output order that has PicOrderCntVal greater than recoveryPointPocVal [[the PicOrderCntVal of the current image plus the recovery_poc_cnt value]] in the CLVS is known as the recovery point image. The recovery point image must not precede the current GDR image in decoding order. Images associated with the current GDR image that have PicOrderCntVal less than recoveryPointPocVal are known as the recovery images of the GDR image. The value of recovery_poc_cnt must be in the range of 0 to MaxPicOrderCntLsb - 1, both inclusive. [[When the current image is a GDR image, the RpPicOrderCntVal variable is derived as follows:] RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt (81)]] NOTE 2: When gdr_enabled_flag is equal to 1 and PicOrderCntVal of the current image is greater than or equal to recoveryPointPocVal [[RpPicOrderCntVal]] of the associated GDR image, the current and subsequent decoded images in output order exactly match the corresponding images produced when starting the decoding process of the previous IRAP image, when present, before the associated GDR image in decoding order. 8.1.2 Decoding process for an encoded image - The PictureOutputFIag variable of the current image is set as follows: - If sps_video_parameter_set_id is greater than 0 and the current layer is not an output layer (i.e., nuh_layer_id is not equal to OutputLayerldlnOls[TargetOlsldx][i] for any value of i in the range from Oa NumOutputLayersinOls[TargetOlsIdx] - 1, inclusive), or one of the following conditions is true, PictureOutputFIag is set equal to 0: - The current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - The current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1 or a recovery image of a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - Otherwise, PictureOutputFIag is set equal to pic_output_flag. NOTE: In one implementation, the decoder may output an image that does not belong to an output layer. For example, when there is only one output layer and the output layer image is unavailable in a AU due to a loss or downgrade, the decoder may set PictureOutputFIag equal to 1 for the image with the highest nuh_layer_id value among all images in the AU available to the decoder that have pic_output_flag equal to 1, and set PictureOutputFIag equal to 0 for all other images in the AU available to the decoder. - [[PictureOutputFIag is set as follows: - If one of the following conditions is met, PictureOutputFIag is set to 0: - the current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - gdr_enabled_flag is equal to 1, the current image is associated with a GDR image with NoOutputBeforeRecoveryFlag equal to 1, and PicOrderCntVal of the current image is less than uo RpPicOrderCntVal of the associated GDR image. - sps_video_parameter_setjd is greater than 0, ols_modejdc is equal to 0 and the current AU contains a picaA image that meets all of the following conditions: - PicA has PictureOutputFIag equal to 1. - PicA has a nuhjayerjd nuhLid greater than that of the current image. - PicA belongs to the output layer of the OLS (i.e., OutputLayerldlnOls[TargetOlsldx][0] is equal to nuhLid). - sps_video_parameter_set_id is greater than 0, olsmodejdc is equal to 2 and ols_outputJayerJlag[TargetOlsldx][GeneralLayerldx[nuhJayerid]] is equal to 0. - Otherwise, PictureOutputFIag is set equal to pic_output_flag.]] 6.4. Fourth Modality This modality is for element 6 and 7a. 8.1.2 Decoding process for an encoded image - PictureOutputFIag is set as follows: - If one of the following conditions is met, PictureOutputFIag is set to 0: - the current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - If gdr_enabled_flag is equal to 1, the current image is associated with a GDR image with NoOutputBeforeRecoveryFlag equal to 1, and PicOrderCntVal of the current image is less than RpPicOrderCntVal of the associated GDR image. - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0, the current layer is not the target output layer (i.e., nuh_layer_id is less than OutputLayerldlnOls[TargetOlsldx][0]), and an image with pic_output_flag equal to 1 and nuh_layer_id greater than that of the current image is present in the current AU. - [[sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0 and the current AU contains a picaA image that meets all of the following conditions: - PicA has PictureOutputFIag equal to 1. - PicA has a nuhjayerjd nuhLid greater than that of the current image. - PicA belongs to the output layer of the OLS (i.e., OutputLayerldlnOls[TargetOlsldx][0] is equal to nuhLid).]] - sps_video_parameter_setjd is greater than 0, ols_modejdc is equal to 2 and ols_output layerJlag[TargetOlsldx][GeneralLayerldx[nuhJayer id]] is equal to 0. - Otherwise, PictureOutputFIag is set equal to pic_outputjlag. 6.5. Fifth Modality This option is only for item 6. 8.1.2 Decoding process for an encoded image i ¿uo - PictureOutputFIag is set as follows: - If one of the following conditions is met, PictureOutputFIag is set to 0: - the current image is a RASL image and the NoOutputBeforeRecoveryFlag of the associated IRAP image is equal to 1. - gdr_enabled_flag is equal to 1 and the current image is a GDR image with NoOutputBeforeRecoveryFlag equal to 1. - If gdr_enabled_flag is equal to 1, the current image is associated with a GDR image with NoOutputBeforeRecoveryFlag equal to 1, and PicOrderCntVal of the current image is less than RpPicOrderCntVal of the associated GDR image. - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 0 and the current AU contains a picaA image that meets all of the following conditions: - PicA has pic_output_flag [[PictureOutputFIag]] equal to 1. - PicA has a nuhjayerjd nuhLid greater than that of the current image. - PicA belongs to the output layer of the OLS (i.e., OutputLayerldlnOls[TargetOlsldx][Q] is equal to nuhLid). - sps_video_parameter_set_id is greater than 0, ols_mode_idc is equal to 2 and ols_output_layer_flag[TargetOlsldx][GeneralLayerldx[nuh_layer_id]] is equal to 0. - Otherwise, PictureOutputFIag is set equal to pic_output_flag. Figure 1 is a block diagram showing a sample 1900 video processing system in which different techniques disclosed herein can be implemented. Various implementations may include some or all of the 1900 system components. The 1900 system may include input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, for example, 8- or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces. The 1900 system may include an encoding component 1904 that can implement the various encoding methods described herein. The encoding component 1904 can reduce the average bitrate of the video from input 1902 to the output of the encoding component 1904 to produce an encoded representation of the video. Therefore, encoding techniques are sometimes called video compression techniques or video transcoding techniques. The output of the encoding component 1904 can be stored or transmitted over a connected communication, as represented by component 1906. The stored or transmitted (or encoded) bitstream representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or viewable video that is sent to a display interface 1910.The process of generating user-viewable video from a bitstream representation is sometimes called video decompression. Furthermore, while certain video processing operations are referred to as encoding operations or tools, it's worth noting that encoding tools or operations are used in an encoder, and the corresponding decoding tools or operations that reverse the encoding results are performed by a decoder. Examples of peripheral bus interfaces or display interfaces include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), DisplayPort, and others. Examples of storage interfaces include SATA (Serial Advanced Technology Junction), PCI, IDE, and similar interfaces. The techniques described herein can be incorporated into various electronic devices such as mobile phones, laptops, smartphones, and other devices capable of digital data processing and / or video display. Figure 2 is a block diagram of a 3600 video processing device. The 3600 device can be used to implement one or more of the methods described herein. The 3600 device can be incorporated into a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The 3600 device can include one or more 3602 processors, one or more 3604 memories, and 3606 video processing hardware. The 3602 processor(s) can be configured to implement one or more of the methods described herein. The 3604 memory(s) can be used to store data and the code used to implement the methods and techniques described herein. The 3606 video processing hardware can be used to implement, in hardware circuitry, some of the techniques described herein. Figure 4 is a block diagram illustrating an example video coding system that can utilize the techniques in this disclosure. As shown in Figure 4, the video encoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and can be referred to as a video decoding device. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116. Video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may comprise one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded images and associated data. The encoded image is an encoded representation of an image. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter.Encoded video data can be transmitted directly to the destination device 120 via I / O interface 116a over network 130a. Alternatively, encoded video data can be stored on a storage medium / server 130b for access by the destination device 120. The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can acquire encoded video data from the source device 110 or the storage / server medium 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the target device 120, or it can be external to the target device 120 and configured to interface with an external display device. The 114 video encoder and 124 video decoder can operate in accordance with a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or additional standards. Figure 5 is a block diagram illustrating an example of video encoder 200, which can be video encoder 114 in system 100 illustrated in Figure 4. The Video Encoder 200 can be configured to perform any or all of the techniques in this disclosure. In the example in Figure 5, the Video Encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the Video Encoder 200. In some examples, a single processor can be configured to perform any or all of the techniques described in this disclosure. The functional components of the video encoder 200 may include a partitioning unit 201, a predication unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intraprediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213 and an entropic encoding unit 214. In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the preaching unit 202 may include an intrablock copy (IBC) unit. The IBC unit can perform preaching in an IBC mode in which at least one reference image is an image where the current video block is located. Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are represented separately in the example in Figure 5 for explanatory purposes. The 201 partition unit can partition an image into one or more video blocks. The 200 video encoder and 300 video decoder can support various video block sizes. ivix / a / zuzz / uii ¿uo The mode selection unit 203 can select one of the encoding modes, intra or inter, for example, based on the error results, and provide the resulting intracoded or intercoded block to a residual generation unit 207 to generate residual block data and to a reconstruction unit 212 to reconstruct the encoded block for use as a reference image. In some examples, the mode selection unit 203 can select a combination of intrapredication and interpredication (CIIP) modes in which the predication is based on an interpredication signal and an intrapredication signal. The mode selection unit 203 can also select a resolution for a motion vector (for example, subpixel or whole-pixel precision) for the block in the case of interpredication. To perform interprediction on a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 can then determine a predicted video block for the current video block based on the motion information and decoded image samples from buffer 213 other than the image associated with the current video block. The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, for example, depending on whether the current video block is in an I sector, a P sector, or a B sector. In some examples, the Motion Estimation Unit 204 can perform unidirectional prediction for the current video block. The Motion Estimation Unit 204 can search for reference images in List 0 or List 1 for a reference video block. The Motion Estimation Unit 204 can then generate a reference index indicating the reference image in List 0 or List 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The Motion Estimation Unit 204 can generate the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block.The 205 motion compensation unit can generate the predicted video block of the current block based on the reference video block indicated by the motion information of the current video block. In other examples, the motion estimation unit 204 can perform bidirectional prediction for the current video block. It can search for reference images in list 0 for a reference video block for the current video block and also search for reference images in list 1 for another reference video block for the current video block. The motion estimation unit 204 can then generate reference indices indicating the reference images in lists 0 and 1 containing the reference video blocks, and motion vectors indicating spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 can generate the reference indices and motion vectors for the current video block as motion information for that block.The 205 motion compensation unit can generate the predicted video block from the current video block based on the reference video blocks indicated by the motion information of the current video block. In some examples, the motion estimation unit 204 can generate a complete set of motion information for a decoder's decoding processing. In some examples, motion estimation unit 204 may not generate a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block. In one example, the motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that tells the video decoder 300 that the current video block has the same motion information as the other video block. In another example, motion estimation unit 204 can identify, within a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the specified video block. Video decoder 300 can use the motion vector of the specified video block and the motion vector difference to determine the motion vector of the current video block. As discussed earlier, the Video Encoder 200 can predictively signal the motion vector. Two examples of predictive signaling techniques that can be implemented using the Video Encoder 200 include advanced motion vector predication (AMVP) and fusion mode signaling. The intraprediction unit 206 can perform intraprediction on the current video block. When intraprediction unit 206 performs intraprediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same image. The prediction data for the current video block can include a predicted video block and various syntax elements. The residual generation unit 207 can generate residual data for the current video block by subtracting (for example, indicated by the minus sign) the predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks that correspond to different sample components of the samples in the current video block. In other examples, there may be no residual data for the current video block, for example, in a skip mode, and the residual generation unit 207 may not perform the subtraction operation. The 208 transformation processing unit can generate one or more video blocks of transformation coefficient for the current video block by applying one or more transformations to a residual video block associated with the current video block. After the transformation processing unit 208 generates a transformation coefficient video block associated with the current video block, the quantization unit 209 can quantize the transformation coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block. The inverse quantization unit 210 and the inverse transformation unit 211 can apply inverse quantization and inverse transformations to the transformation coefficient video block, respectively, to reconstruct a residual video block from the transformation coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current block for storage in the buffer 213. After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video blocking artifacts in the video block. The entropic encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropic encoding unit 214 receives the data, the entropic encoding unit 214 can perform one or more entropic encoding operations to generate entropy-encoded data and generate a bitstream that includes the entropy-encoded data. Figure 6 is a block diagram illustrating an example of a video decoder 300, which can be a video decoder 114 in the system 100 illustrated in Figure 4. The Video Decoder 300 can be configured to perform any or all of the techniques described in this disclosure. In the example in Figure 6, the Video Decoder 300 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the Video Decoder 300. In some examples, a single processor can be configured to perform any or all of the techniques described in this disclosure. In the example in Figure 6, the video decoder 300 includes an entropic decoding unit 301, a motion compensation unit 302, an intraprediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. The video decoder 300 can, in some examples, perform a decoding pass that is generally the reciprocal of the encoding pass described with respect to the video encoder 200 (Figure 5). The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector accuracy, reference image list indices, and other motion information. The motion compensation unit 302 can, for example, determine this information when performing merging mode and AMVP. The 302 motion compensation unit can produce motion-compensated blocks, possibly by performing interpolation based on interpolation filters. Identifiers for interpolation filters to be used with subpixel precision can be included in the syntax elements. The Motion Compensation Unit 302 can use interpolation filters, as used by the Video Encoder 200 during video block encoding, to calculate interpolated values ​​for sub-integer pixels in a reference block. The Motion Compensation Unit 302 can determine the interpolation filters used by the Video Encoder 200 based on the received syntax information and use those interpolation filters to produce predictive blocks. The 302 motion compensation unit can use some of the syntax information to determine the block sizes used to encode frames and / or sectors of the encoded video sequence, partitioning information that describes how each macroblock of an image in the encoded video sequence is divided, modes that indicate how each partition is encoded, one or more reference frames (and reference frame lists) for each intercoded block, and other information to decode the encoded video sequence. The intraprediction unit 303 can use intraprediction modes, for example, received in the bitstream, to form a prediction block from spatially adjacent blocks. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropic decoding unit 301. The inverse transformation unit 303 applies an inverse transformation. The reconstruction unit 306 can sum the residual blocks with the corresponding prediction blocks generated by the motion compensation unit 202 or the intraprediction unit 303 to form decoded blocks. If desired, an unblocking filter can also be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are stored in buffer 307, which provides reference blocks for post-motion compensation / intrapredication and also produces decoded video for display on a screen. Below is a list of examples preferred by some modalities. The first set of clauses shows examples of the types of techniques analyzed in the previous section. The following clauses show examples of the types of techniques analyzed in the previous section (e.g., element 1). 1. A video processing method (for example, method 3000 shown in Figure 3), comprising: performing (3002) a conversion between a video having one or more video layers comprising one or more video images and an encoded representation of the video; wherein the encoded representation includes a set of video parameters indicating a maximum value of a chroma format indicator and / or a maximum bit depth value used to represent pixels iviA / a / zuzz / uii zuo of the video. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 2). 2. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule that specifies a maximum image width and / or a maximum image height for video images from all video layers, controlling a value of a variable that indicates whether images in a decoder buffer are produced before being removed from the decoder buffer. 3. The method of clause 2, where the variable is signaled in a set of video parameters. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 3). 4. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that the maximum value of a chroma format indicator and / or a maximum bit depth value used to represent video control pixels is a value of a variable indicating whether images in a decoder buffer are produced before being cleared from the decoder buffer. 5. The method of clause 4, where the variable is signaled in a set of video parameters. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 4). 6. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that a value of a variable indicating whether images in a decoder buffer are produced before being removed from the decoder buffer is independent of whether separate color planes are used to encode the video. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 5). 7. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that a value of a variable indicating whether the images in a decoder buffer are produced before being removed from the decoder buffer is included in the encoded representation at an access unit (AU) level. iviA / a / zuzz / uii ¿uo 8. The method of clause 7, where the format rule specifies that the value is the same for all AUs in the encoded representation. 9. The method of any of clauses 7-8, where the variable is indicated in an image header. 10. The method of any of clauses 7-8, wherein the variable is indicated in an access unit delimiter. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 6). 11. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that an image output flag for a video image in an access unit is determined based on a pic_output_flag variable of another video image in the access unit. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 7). 12. A video processing method, comprising: performing a conversion between a video having one or more video layers and an encoded representation of the video, wherein the encoded representation conforms to a format rule specifying that, for a video image not belonging to an output layer, a value of an output image flag. 13. The method of clause 12, where the format rule specifies that the value of the picture output flag for a video picture is set to zero. 14. The method of clause 12, wherein the video comprises only one output layer, and wherein an access unit not including the output layer is encoded by setting the image output flag value to logic 1 for an image having a higher layer identification value and logic 0 for all other images. The following clauses show examples of types of techniques analyzed in the previous section (e.g., element 8). 15. The method of any of clauses 1-14, wherein the image exit flag is included in a delimited access unit. 16. The method of any of clauses 1-14, wherein the image exit flag is included in a supplemental enhancement information field. 17. The method of any of clauses 1-14, wherein the image exit flag is included in image headers of one or more images. 18. The method of any of clauses 1 to 17, wherein the conversion comprises encoding the video into the encoded representation. 19. The method of any of clauses 1 to 17, wherein the conversion comprises decoding the encoded representation to generate pixel values ​​from the video. 20. A video decoding apparatus comprising a processor configured to implement a method in accordance with one or more of clauses 1 to 19. 21. A video encoding apparatus comprising a processor configured to implement a method in accordance with one or more of clauses 1 to 19. 22. A computer program product that has computer code stored therein, the code, when executed by a processor, causes the processor to implement a method mentioned in any of clauses 1 to 19. 23. A method, apparatus or system described in this document. The second set of clauses shows examples of modalities of techniques analyzed in the previous section (e.g., items 1-4). 1. A video processing method (for example, method 710 as shown in Figure 7A) comprising: performing 712 a conversion between a video and a video bitstream according to a format rule, wherein the bitstream includes one or more output layer sets (OLS), each OLS comprising one or more encoded layer video sequences, and wherein the format rule specifies that a set of video parameters indicates, for each of the one or more OLS, a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video. 2. The method of clause 1, wherein the maximum allowable value of the chroma format indicator for an OLS is applicable to all sets of sequence parameters to which the one or more layer-encoded video sequences in the OLS refer. 3. The method of clause 1 or 2, wherein the maximum allowed value of the bit depth for an OLS is applicable to all sets of sequence parameters referenced by the one or more layer video sequences encoded in the OLS. 4. The method of any of clauses 1 to 3, wherein, to perform the conversion for an OLS containing more than one encoded layer video sequence and having an OLS index, i, the rule specifies allocating memory for a decoded picture buffer according to values ​​of at least one of the syntax elements including ols_dpb_pic_width[i] indicating a width for each picture buffer for an i-th OLS, ols_dpb_pic_height[i] indicating a height for each picture buffer for the i-th OLS, a syntax element indicating the highest allowed value of the chroma format flag for the i-th OLS, and a syntax element indicating the highest allowed value of a bit depth for the i-th OLS. 5. The method of any of clauses 1 to 4, wherein the set of video parameters is included in the bitstream. 6. The method of any of clauses 1 to 4, wherein the set of video parameters is specified separately from the bitstream. 7. A video processing method (for example, method 720 as shown in Figure 7B) comprising: performing 722 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a maximum image width and / or a maximum image height for video images from all video layers controls a value of a variable indicating whether images in a decoded image buffer prior to a current image in decoding order in the bitstream are produced before images are removed from the decoded image buffer. 8. The method of clause 7, where the variable is derived based on at least one or more syntax elements included in a set of video parameters. 9. The method of clause 7, wherein the value of the variable is set to 1 in case a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a current access unit is different from a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a previous access unit in decoding order. 10. The method in clause 9, wherein the value of the variable equal to 1 indicates images in a decoded image buffer before a current image in decoding order are not produced before the images are removed from the decoded image buffer. 11. The method of any of clauses 7 to 10, wherein the value of the variable is further based on a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video. 12. The method of any of clauses 7 to 11, wherein the set of video parameters is included in the bitstream. 13. The method of any of clauses 7 to 11, wherein the set of video parameters is indicated separately from the bitstream. 14. A video processing method (for example, 730 as shown in Figure 7C), comprising: performing 732 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video control a value of a variable indicating whether images in a decoded image buffer prior to a current image in decoding order in the bitstream are produced before images are removed from the decoded image buffer. 15. The method of clause 14, wherein the variable is derived based on at least one or more syntax elements signaled in a set of video parameters. 16. The method of clause 14, wherein the value of the variable is set to 1 in case a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a current access unit is different from a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a previous access unit in decoding order. 17. The method of clause 16, wherein the value of the variable that is equal to 1 indicates images in a decoded image buffer before a current image in decoding order are not produced before the images are removed from the decoded image buffer. 18. The method of any of clauses 14 to 17, wherein the value of the variable is further based on a maximum image width and / or a maximum image height for video images of all video layers. 19. The method of any of clauses 14 to 18, wherein the set of video parameters is included in the bitstream. 20. The method of any of clauses 14 to 18, wherein the set of video parameters is indicated separately from the bitstream. 21. A video processing method (for example, method 740 as shown in Figure 7D) comprising: performing 742 a conversion between a video having one or more video layers and a video bitstream according to a rule, and wherein the rule specifies that a value of a variable indicating whether images in a decoded image buffer prior to a current image in decoding order in the bitstream occur before images are removed from the decoded image buffer is independent of whether separate color planes are used to encode the video. 22. The method of clause 21, wherein, in the event that separate color planes are not used to encode the video, the rule specifies that decoding is performed only once for a video image, or, in the event that separate color planes are used to encode the video, the rule specifies that image decoding is invoked three times. 23. The method of any of clauses 1 to 22, wherein the conversion includes encoding the video in the bitstream. 24. The method of any of clauses 1 to 22, wherein the conversion includes decoding the video from the bitstream. 25. The method of clauses 1 to 22, wherein the conversion includes generating the bitstream from the video, and the method further comprises: storing the bitstream on a non-transient, computer-readable recording medium. 26. A video processing apparatus comprising a processor configured to implement a method mentioned in any or more of clauses 1 to 25. 27. A method for storing a video bitstream, comprising a method mentioned in any of clauses 1 to 25, and further including storing the bitstream on a non-transient, computer-readable recording medium. 28. A computer-readable medium that stores program code that, when executed, causes a processor to implement a method in accordance with any one or more of the clauses in 25. 29. A computer-readable medium that stores a bit stream generated according to any of the methods described above. 30. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method in accordance with any one or more of clauses 1 to 25. The third set of clauses shows examples of modalities of techniques analyzed in the previous section (e.g., element 5). 1. A video processing method (for example, method 810 as shown in Figure 8A), comprising: performing 812 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a flag value indicating whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer when a certain type of access unit is decoded is included in the bitstream. 2. The method of clause 1, where the format rule specifies that the value is the same for all images on an access unit. 3. The method of clause 1 or 2, wherein the format rule specifies that a value of a variable indicating whether images in the decoded image buffer prior to a current image in decoding order in the bitstream occur before images are removed from the decoded image buffer is based on the value of the flag. 4. The method of any of clauses 1 to 3, where the flag is indicated in an image header. 5. The method of any of clauses 1 to 3, wherein the flag is indicated in a sector header. 6. A video processing method (for example, method 820 as shown in Figure 8B) comprising: performing 822 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a value of a first flag indicating whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer when decoding an access unit of a particular type is not indicated in an image header. 7. The method of clause 6, wherein the first flag is indicated in an access unit delimiter. 8. The method of clause 6, wherein a second flag indicating an IRAP (intra-random access point image) or GDR (gradual decoding update) access unit has a certain value. 9. The method of clause 6, where it is inferred that the value of the first flag is equal to 1 iviA / a / zuzz / uii zuo in case there is no access unit delimiter present for an IRAP (intra-random access point image) or GDR (gradual decoding update) access unit. 10. A video processing method (for example, method 830 as shown in Figure 8C), comprising: performing 832 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a value of a flag associated with an access unit indicating whether previously decoded images stored in a decoded image buffer should be removed from the decoded image buffer depends on a flag value for each image in the access unit. 11. The method of clause 10, wherein the format rule specifies that the value of the access unit flag is considered to be 0 if the flag for each image of an access unit is equal to 0, and otherwise the value of the access unit flag is considered to be 1. 12. The method of any of clauses 1 to 11, wherein the conversion includes encoding the video in the bitstream. 13. The method of any of clauses 1 to 11, wherein the conversion includes decoding the video from the bitstream. 14. The method of any of clauses 1 to 11, wherein the conversion includes generating the bitstream from the video, and the method further comprises: storing the bitstream on a non-transient, computer-readable recording medium. 15. A video processing apparatus comprising a processor configured to implement a method mentioned in any or more of clauses 1 to 14. 16. A method for storing a video bitstream, comprising a method mentioned in any of clauses 1 to 14, and further including storing the bitstream on a non-transient, computer-readable recording medium. 17. A computer-readable medium that stores program code that, when executed, causes a processor to implement a method in accordance with any one or more of clauses 1 to 14. 18. A computer-readable medium that stores a bit stream generated according to any of the methods described above. 19. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method in accordance with any one or more of clauses 1 to 14. The fourth set of clauses shows examples of modalities of techniques analyzed in the previous section (e.g., items 6-8). 1. A video processing method (for example, method 910 as shown in Figure 9A), comprising: 912 performing a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that a value of a variable indicating whether an image is to be produced in an access unit is determined based on a flag indicating whether another image is to be produced in the access unit. 2. The method of clause 1, where the other image is on a higher layer than the image. 3. The method of clause 1 or 2, where the flag controls the processes of deletion and output of decoded images. 4. The method of clause 1 or 2, wherein the flag is a syntax element included in a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a sector header, or a tile group header. 5. The method of any of clauses 1 to 4, wherein the value of the variable is further based on at least one of i) a flag specifying a value of an identifier for the video parameter set (VPS), ii) whether a current video layer is an output layer, iii) whether a current picture is a skipped random-access front picture, a stepped-decode update picture, a retrieved picture from a stepped-decode update picture, or iii) whether pictures in the decoded picture buffer prior to the current picture in decoding order are produced before pictures are retrieved. 6. The method in clause 5, wherein the value of the variable is set to 0 if i) the flag specifying the identifier value for the VPS is greater than 0 and the current layer is not an output layer, or i) one of the following conditions is met: The current image is a skipped random access front image and the associated intra-random access point image in the decoded image buffer before the current image in decoding order does not occur before the intra-random access point image is retrieved; or the current image is a gradual decoding update image with images in the decoded image buffer before the current image in decoding order does not occur before the images are retrieved or a gradual decoding update image retrieval image with images in the decoded image buffer before the current image in decoding order does not occur before the images are retrieved. 7. The method of clause 6, where the value of the variable is set equal to a flag value in case both i) and i¡) are not satisfied. 8. The method of clause 6, where the value of the variable being equal to 0 indicates not to produce an image on an access unit. 9. The method in clause 1, where the variable is PictureOutputFIag and the flag is pic_output_flag. 10. The method of any of clauses 1 to 9, wherein the flag is included in an access unit delimiter. ivix / a / zuzz / uii ¿uo 11. The method of any of clauses 1 to 9, where the flag is included in a supplementary improvement information field. 12. The method of any of clauses 1 to 9, where the flag is included in image headers of one or more images. 13. A video processing method (for example, method 920 as shown in Figure 9B), comprising: performing 922 a conversion between a video having one or more video layers and a bitstream of the video according to a format rule, and wherein the format rule specifies that a value of a variable indicating whether an image is to be produced in an access unit is set equal to a certain value in case the image does not belong to an output layer. 14. The method of clause 13, where the determined value is zero. 15. A video processing method (for example, method 930 as shown in Figure 9C), comprising: performing 932 a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies that in the event the video comprises only one output layer, an access unit that does not include an output layer is encoded by setting a variable indicating whether an image is to be produced in the access unit to a first value for the image having the highest layer identification (ID) value and a second value for all other images. 16. The method of clause 15, where the first value is 1 and the second value is 0. 17. The method of any of clauses 1 to 16, wherein the conversion includes encoding the video in the bitstream. 18. The method of any of clauses 1 to 16, wherein the conversion includes decoding the video from the bitstream. 19. The method of clauses 1 to 16, wherein the conversion includes generating the bitstream from the video, and the method further comprises: storing the bitstream on a non-transient, computer-readable recording medium. 20. A video processing apparatus comprising a processor configured to implement a method mentioned in any or more of clauses 1 to 19. 21. A method for storing a video bitstream, comprising a method mentioned in any of clauses 1 to 19, and further including storing the bitstream on a non-transient, computer-readable recording medium. 22. A computer-readable medium that stores program code that, when executed, causes a processor to implement a method in accordance with any one or more of clauses 1 to 19. 23. A computer-readable medium that stores a bit stream generated according to any of the methods described above. 24. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method in accordance with any one or more of clauses 1 to 19. iviA / a / zuzz / uii ¿uo In this document, the term “video processing” may refer to video encoding, video decoding, video compression, or video decompression. For example, video compression algorithms can be applied during the conversion of a video's pixel representation to a corresponding bitstream representation, or vice versa. The bitstream representation of a video block may, for example, correspond to bits that are contiguous or spread across different locations within the bitstream, as defined by the syntax. For instance, a macroblock may be encoded in terms of transformed and encoded error residual values, and also using bits in headers and other fields within the bitstream.Furthermore, during conversion, a decoder can analyze a bitstream knowing that certain fields may be present or absent, based on the determination described in the previous solutions. Similarly, an encoder can determine whether certain syntax fields will be included or excluded and generate the encoded representation accordingly by including or excluding those syntax fields. The disclosed and other solutions, examples, modalities, modules, and functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other modalities may be implemented as one or more computer program products, that is, one or more computer program instruction modules encoded in a computer-readable medium for execution by means of, or to control the operation of, a data processing device.A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that produces a machine-readable propagated signal, or a combination of one or more of these. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may include, in addition to hardware, code for creating an execution environment for the computer program in question, for example, the code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.A propagated signal is a signal that is artificially generated, for example, a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiving device. A computer program (also known as a program, software, sequence, software application, or code) can be written in any programming language, including compiled or interpreted languages, and can be developed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file containing other programs or data (for example, one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (for example, files that store one or more modules, subprograms, or code snippets).A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network. The processes and logic flows described in this document can be implemented using one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. These processes and logic flows can also be implemented using, and the apparatus can also be implemented as, special-purpose logic circuitry, such as an FPGA (Field-Programmable Gate Array) or an ASIO (Application-Specific Integrated Circuit). The processors suitable for executing a computer program include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from read-only memory or random-access memory (RAM). The essential elements of a computer are a processor to carry out instructions and one or more memory devices to store instructions and data. Generally, a computer will include, or be operatively coupled to receive or transfer data, one or more mass storage devices to store data, such as magnetic, magneto-optical, or optical disks. However, a computer does not necessarily need to have these devices.Computer-legal media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry. While this patent document contains many details, these should not be interpreted as limitations on the scope of any subject matter described or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular techniques. Certain features described in this patent document may also be implemented in the context of separate embodiments in combination with a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.Furthermore, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features of a claimed combination may in some cases be separated from the combination, and the claimed combination may be directed towards a subcombination or variation of a subcombination. Similarly, while the operations are illustrated in the drawings in a particular order, this should not be construed as requiring that such operations be carried out in the specific order shown or in a sequential order, or that all the illustrated operations be performed in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments. Only a few implementations and examples are described, and other implementations, improvements, and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method comprising: performing a conversion between a video and a video bitstream according to a format rule, wherein the bitstream includes one or more output layer sets (OLS), each OLS comprising one or more encoded layer video sequences, and wherein the format rule specifies that a set of video parameters indicates, for each of the one or more OLS, a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent video pixels.

2. The method according to claim 1, wherein the maximum allowable value of the chroma format indicator for an OLS is applicable to all sets of sequence parameters to which the one or more layer video sequences encoded in the OLS refer.

3. The method according to claim 1 or 2, wherein the maximum allowable value of the bit depth for an OLS is applicable to all sets of sequence parameters referenced by the one or more layer video sequences encoded in the OLS.

4. The method according to any one of claims 1 to 3, wherein, to perform the conversion for an OLS containing more than one encoded layer video sequence and having an OLS index, i, the rule specifies allocating memory for a decoded picture buffer according to values ​​of at least one of the syntax elements including ols_dpb_pic_width[i] indicating a width of each picture buffer for an i-th OLS, ols_dpb_pic_height[i] indicating a height of each picture buffer for the i-th OLS, a syntax element indicating the highest allowed value of the chroma format flag for the i-th OLS, and a syntax element indicating the highest allowed value of a bit depth for the i-th OLS.

5. The method according to any of claims 1 to 4, wherein the set of video parameters is included in the bitstream.

6. The method according to any of claims 1 to 4, wherein the set of video parameters is specified separately from the bitstream.

7. A video processing method comprising: performing a conversion between a video having one or more video layers and a video bitstream according to a format rule, wherein the format rule specifies a maximum image width and / or a maximum image height for video images from all video layers, controlling a value of a variable that indicates whether images in a decoded image buffer prior to a current image in decoding order in the bitstream are produced before images are removed from the decoded image buffer.

8. The method according to claim 7, wherein the variable is derived based on at least one or more syntax elements included in a set of video parameters.

9. The method according to claim 7, wherein the value of the variable is set to 1 iviA / a / zuzz / uii ¿uo in case a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a current access unit is different from a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a previous access unit in decoding order.

10. The method according to claim 9, wherein the value of the variable equal to 1 indicates images in a decoded image buffer prior to a current image in decoding order are not produced before the images are removed from the decoded image buffer.

11. The method according to any of claims 7 to 10, wherein the value of the variable is further based on a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent video pixels.

12. The method according to any of claims 7 to 11, wherein the set of video parameters is included in the bitstream.

13. The method according to any of claims 7 to 11, wherein the set of video parameters is specified separately from the bitstream.

14. A video processing method comprising: performing a conversion between a video having one or more video layers and a video bitstream according to a format rule, and wherein the format rule specifies a maximum allowable value of a chroma format indicator and / or a maximum allowable value of a bit depth used to represent pixels of the video control a value of a variable indicating whether images in a decoded image buffer prior to a current image in decoding order in the bitstream are produced before images are removed from the decoded image buffer.

15. The method according to claim 14, wherein the variable is derived based on at least one or more syntax elements signaled in a set of video parameters.

16. The method according to claim 14, wherein the value of the variable is set to 1 in the event that a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a current access unit is different from a value of a maximum width of each image, a maximum height of each image, a maximum allowable value of a chroma format indicator, or a maximum allowable value of a bit depth derived for a previous access unit in decoding order.

17. The method according to claim 16, wherein the value of the variable equal to 1 indicates images in a decoded image buffer prior to a current image in decoding order are not produced before the images are removed from the decoded image buffer.

18. The method according to any of claims 14 to 17, wherein the value of the variable is further based on a maximum image width and / or a maximum image height for video images of all video layers.

19. The method according to any of claims 14 to 18, wherein the set of video parameters is included in the bitstream.

20. The method according to any of claims 14 to 18, wherein the set of video parameters is specified separately from the bitstream.

21. A video processing method comprising: performing a conversion between a video having one or more video layers and a video bitstream according to a rule, wherein the rule specifies that a value of a variable indicating whether images in a decoded image buffer prior to a current image in decoding order in the bitstream occur before images are removed from the decoded image buffer is independent of whether separate color planes are used to encode the video.

22. The method according to claim 21, wherein, if separate color planes are not used to encode the video, the rule specifies that decoding is performed only once for a video image, or, if separate color planes are used to encode the video, the rule specifies that image decoding is invoked three times.

23. The method according to any of claims 1 to 22, wherein the conversion includes encoding the video in the bitstream.

24. The method according to any of claims 1 to 22, wherein the conversion includes decoding the video from the bitstream.

25. The method according to claims 1 to 22, wherein the conversion includes generating the bitstream from the video, and the method further comprises: storing the bitstream on a non-transient, computer-readable recording medium.

26. A video processing apparatus comprising a processor configured to implement a method mentioned in any or more of claims 1 to 25.

27. A method for storing a video bitstream, comprising a method mentioned in any one of claims 1 to 25, and further including storing the bitstream on a non-transient, computer-readable recording medium.

28. A computer-readable medium that stores program code that, when executed, causes a processor to implement a method in accordance with any one or more of claims 1 to 25.

29. A computer-readable medium that stores a bit stream generated according to any of the methods described above.

30. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method according to any one or more of claims 1 to 25.