Signaling of Decoded Picture Buffer Parameters in Hierarchical Video

By modifying the syntax elements and signaling rules in the VVC draft text, the problems of inflexible PTL information signaling notification and HRD parameter duplication signaling in VVC were solved, improving encoding and decoding efficiency and flexibility.

CN114868158BActive Publication Date: 2025-08-01DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080090437.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-24
Publication Date
2025-08-01
Estimated Expiration
2040-12-24

AI Technical Summary

Technical Problem

The existing VVC draft text has insufficient flexibility in the constraints on slice_type, which makes it impossible to effectively signal PTL information in multi-layer bitstreams, and the HRD parameter repeats unnecessary signaling notification, affecting encoding and decoding efficiency.

Method used

By modifying the syntax elements and signaling rules in the VVC draft text, more flexible signaling notifications are allowed. For example, the value of ols_ptl_idx[i] is inferred to be 0 when it does not exist, unnecessary layer _output_dpb_params_idx[i] signaling is avoided, the representation of vps_num_dpb_params is adjusted, PTL information is allowed to be repeated in VPS, and HRD parameters are signaled only when needed.

Benefits of technology

It improves encoding and decoding efficiency, reduces unnecessary signaling overhead, ensures complete signaling notification of PTL information in multi-layer bitstreams, and enhances the flexibility and efficiency of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114868158B_ABST
    Figure CN114868158B_ABST
Patent Text Reader

Abstract

A video processing method includes performing a conversion between a video and a bitstream of the video. The bitstream includes one or more output layer sets, each including one or more video layers. The bitstream conforms to format rules, wherein the format rules specify that when each output layer set includes a single video layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to zero; or when it is not true that each output layer set includes a single layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to one plus the value of a syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is the U.S. national phase entry of International Patent Application No. PCT / US2020 / 067019, filed on December 24, 2020, which claims the priority of U.S. Provisional Application No. 62 / 953,854, filed on December 26, 2019, and U.S. Provisional Application No. 62 / 955,185, filed on December 30, 2019. The entire disclosure of the above applications is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image encoding and decoding and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by a video encoder and a video decoder to perform video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more scalable video layers and a bitstream of the video. The video includes one or more video pictures, and the video pictures include one or more slices. The bitstream conforms to format rules. The format rules specify that, when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the slice does not use inter - layer prediction, the value of a field indicating the slice type of the slice is set to indicate the type of an intra - slice.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including multiple video layers and a bitstream of the video, where the bitstream includes multiple output layer sets (OLSs), each output layer set includes one or more of the multiple scalable video layers, and the bitstream conforms to format rules, where the format rules specify that, for an OLS having a single layer, the profile - tier - level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0008] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein the format rules specify a relationship between a plurality of profile-tier-level (PTL) syntax structures that appear in a video parameter set of the bitstream and a byte alignment syntax field in the video parameter set; wherein each PTL syntax structure indicates the profile, tier, and level of one or more of the plurality of OLSs.

[0009] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of scalable video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of scalable video layers, and the bitstream conforms to format rules, wherein the format rules specify that during encoding, when the value of an index of a syntax structure that describes the profile, tier, and level of one or more of the plurality of OLSs is zero, a syntax element that indicates the index is excluded from a video parameter set of the bitstream, or during decoding, when there is no syntax element in the bitstream, the value is inferred to be zero.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein the format rules specify that for layer i, where i is an integer, the bitstream includes a first set of syntax elements that indicates a first variable that indicates whether layer i is included in at least one of the plurality of OLSs.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets, each output layer set includes one or more video layers; wherein the bitstream conforms to format rules, wherein the format rules specify that when each output layer set includes a single video layer, the number of decoded picture buffer parameter syntax structures included in a video parameter set of the bitstream is equal to zero; or when it is not true that each output layer set includes a single layer, the number is equal to one plus the value of a syntax element.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, where the bitstream includes a coded video sequence (CVS), and the coded video sequence includes one or more coded video pictures of one or more video layers; and where the bitstream conforms to a format rule that specifies that one or more sequence parameter sets (SPSs) indicating conversion parameters referenced by one or more of the coded pictures of the CVS have the same reference video parameter set (VPS) identifier that indicates a reference VPS.

[0013] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, where the bitstream includes one or more output layer sets (OLSs), each output layer set including one or more video layers, where the bitstream conforms to a format rule; where the format rule specifies whether or how a first syntax element is included in a video parameter set (VPS) of the bitstream, and the first syntax element indicates whether a first syntax structure describing parameters of a hypothetical reference decoder (HRD) is used for the conversion.

[0014] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, where the bitstream includes one or more output layer sets (OLSs), each output layer set including one or more video layers, where the bitstream conforms to a format rule; where the format rule specifies whether or how a first syntax structure describing general HRD parameters and a plurality of second syntax structures describing OLS-specific HRD parameters are included in a video parameter set (VPS) of the bitstream.

[0015] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.

[0016] In another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.

[0017] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0018] In yet another example aspect, a method of writing a bitstream generated according to one of the above methods to a computer-readable medium is disclosed.

[0019] In another exemplary aspect, a computer-readable medium storing a bitstream of a video generated according to the above method is disclosed.

[0020] These features and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a block diagram showing a video codec system according to some embodiments of the present disclosure.

[0022] Figure 2 is a block diagram of an exemplary hardware platform for video processing.

[0023] Figure 3 is a flowchart of an exemplary method for video processing.

[0024] Figure 4 is a block diagram showing an exemplary video codec system.

[0025] Figure 5 is a block diagram showing an encoder according to some embodiments of the present disclosure.

[0026] Figure 6 is a block diagram showing a decoder according to some embodiments of the present disclosure.

[0027] Figures 7A - 7I is a flowchart of examples of various video processing methods. DETAILED DESCRIPTION

[0028] The use of section headings in this document is for ease of understanding and does not limit the techniques and embodiments disclosed in each section to only that section. Additionally, the use of H.266 terminology in some of the descriptions is for ease of understanding and not to limit the scope of the disclosed techniques. Thus, the techniques described herein also apply to other video codec protocols and designs.

[0029] 1. Summary

[0030] This document is related to video codec technology. Specifically, it is about various improvements to scalable video coding, where a video bitstream can contain more than one layer. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Coding (VVC) being developed.

[0031] 2. Abbreviations

[0032] APS Adaptive Parameter Set

[0033] AU Access Unit

[0034] AUD Access Unit Delimiter

[0035] AVC Advanced Video Coding

[0036] CLVS Coding Layer Video Sequence

[0037] CPB Coding Picture Buffer

[0038] CRA Clean Random Access

[0039] CTU Coding Tree Unit

[0040] CVS Coding Video Sequence

[0041] DPB Decoding Picture Buffer

[0042] DPS Decoding Parameter Set

[0043] EOB End of Bitstream

[0044] EOS End of Sequence

[0045] GDR Gradual Decoding Refresh

[0046] HEVC High Efficiency Video Coding

[0047] HRD Hypothetical Reference Decoder

[0048] IDR Instantaneous Decoding Refresh

[0049] JEM Joint Exploration Model

[0050] MCTS Motion Constrained Tile Set

[0051] NAL Network Abstraction Layer

[0052] OLS Output Layer Set

[0053] PH Picture Header

[0054] PPS Picture Parameter Set

[0055] PTL Profile, Tier and Level

[0056] PU Picture Unit

[0057] RBSP Raw Byte Sequence Payload

[0058] SEI Supplemental Enhancement Information

[0059] SPS Sequence Parameter Set

[0060] SVC Scalable Video Coding

[0061] VCL Video Coding Layer

[0062] VPS Video Parameter Set

[0063] VTM VVC Test Model

[0064] VUI Video Usability Information

[0065] VVC Versatile Video Coding

[0066] 3. Preliminary Discussion

[0067] Video coding standards have mainly evolved through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard as well as the H.265 / HEVC [1] standard. Since H.262, video coding standards have been based on the hybrid video coding structure, in which temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and applied to the reference software called the Joint Exploration Model (JEM) [2]. JVET meetings are held quarterly simultaneously, and the goal of the new coding standard is to reduce the bit rate by 50% compared to HEVC. At the JVET meeting in April 2018, the new video coding standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at that time. With continuous efforts dedicated to VVC standardization, new coding technologies are adopted for the VCC standard at each JVET meeting. Then the VVC working draft and the test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Feature Complete (FDIS) at the meeting in July 2020.

[0068] 3.1. Scalable Video Coding (SVC)

[0069] Scalable Video Coding (SVC) refers to video coding in which a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (ELs) are used. In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial levels, temporal levels, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can act as the BL, and the top layer can act as the EL. Intermediate layers can act as ELs or RLs, or both simultaneously. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below it (e.g., the base layer or any intermediate enhancement layer) and simultaneously act as an RL for one or more enhancement layers above it. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and the information of one view can be used to code (e.g., encode or decode) the information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).

[0070] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding levels at which they can be used (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be used by one or more coded video sequences of different layers in a bitstream can be included in the Video Parameter Set (VPS), and the parameters used by one or more pictures in a coded video sequence can be included in the Sequence Parameter Set (SPS). Similarly, the parameters used by one or more slices in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single slice can be included in the slice header. Similarly, an indication of which parameter set a given layer uses at a given time can be provided at various coding levels.

[0071] 3.2. Parameter Sets

[0072] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, VPS, and DPS. All AVC, HEVC, and VVC support SPS and PPS. VPS was introduced since HEVC and is included in HEVC and VVC. APS and DPS are not included in AVC or HEVC but are included in the latest VVC draft text.

[0073] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry picture-level header information that does not change frequently. Using the SPS and PPS, it is not necessary to repeat the information that does not change frequently for each sequence or picture, so the redundant signaling of this information can be avoided. In addition, the use of the SPS and PPS enables the out-of-band transmission of important header information, so not only is the need for redundant transmission avoided, but also the error recovery ability is improved.

[0074] The VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0075] The APS is introduced to carry picture-level information or slice-level information that requires a considerable number of bits to encode and decode, can be shared by multiple pictures, and can have a considerable number of different variants in a sequence.

[0076] The DPS is introduced to carry bitstream-level information that indicates the highest capabilities required to decode the entire bitstream.

[0077] 3.3. VPS Syntax and Semantics in VVC

[0078] VVC supports scalability, also known as scalable video coding, where multiple layers can be encoded in a single coded video bitstream.

[0079] In the latest VVC text, scalability information is signaled in the VPS, and the syntax and semantics are as follows.

[0080] 7.3.2.2 Video Parameter Set Syntax

[0081]

[0082]

[0083]

[0084] 7.4.3.2 Video Parameter Set RBSP Semantics

[0085] The VPS RBSP should be available for the decoding process before being referenced, including in at least one AU where the TemporalId is equal to 0 or provided externally.

[0086] All VPS NAL units with a specific value of vps_video_parameter_set_id in the CVS should have the same content.

[0087] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.

[0088] vps_max_layers_minus1 plus 1 specifies the maximum allowed number of layers in each CVS that refers to the VPS.

[0089] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in each CVS that refers to the VPS. The value of vps_max_sublayers_minus1 shall be in the range from 0 to 6 (inclusive of 0 and 6).

[0090] vps_all_layers_same_num_sublayers_flag being equal to 1 specifies that for all layers in each CVS that refers to the VPS, the number of temporal sublayers is the same. vps_all_layers_same_num_sublayers_flag being equal to 0 specifies that the layers in each CVS that refers to the VPS may or may not have the same number of temporal sublayers. When absent, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.

[0091] vps_all_independent_layers_flag being equal to 1 specifies that all layers in the CVS are independently encoded and decoded without using inter-layer prediction. vps_all_independent_layers_flag being equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. When absent, the value of vps_all_independent_layers_flag is inferred to be equal to 1. When vps_all_independent_layers_flag is equal to 1, the value of vps_independent_layer_flag[i] is inferred to be equal to 1. When vps_all_independent_layers_flag is equal to 0, the value of vps_independent_layer_flag[0] is inferred to be equal to 1.

[0092] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] shall be less than the value of vps_layer_id[n].

[0093] When vps_independent_layer_flag[i] equals 1, it is specified that the layer with index i does not use inter-layer prediction. When vps_independent_layer_flag[i] equals 0, it is specified that the layer with index i can use inter-layer prediction, and there exists a syntax element vps_direct_ref_layer_flag[i][j] (for j in the range from 0 to i - 1, inclusive of 0 and i - 1) in the VPS. When it does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1.

[0094] When vps_direct_ref_layer_flag[i][j] equals 0, it is specified that the layer with index j is not a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[i][j] equals 1, it is specified that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[i][j] does not exist (for i and j in the range from 0 to vps_max_layers_minus1, inclusive of 0 and vps_max_layers_minus1), it is inferred to be equal to 0. When vps_independent_layer_flag[i] equals 0, there should exist at least one value of j in the range from 0 to i - 1, inclusive of 0 and i - 1, such that the value of vps_direct_ref_layer_flag[i][j] equals 1.

[0095] Derive the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] as follows:

[0096]

[0097]

[0098] Derive the variable GeneralLayerIdx[i] as follows, which specifies the layer index of the layer where nuh_layer_id equals vps_layer_id[i]:

[0099]

[0100] When each_layer_is_an_ols_flag equals 1, it is specified that each output layer set contains only one layer, and each layer in the bitstream is itself an output layer set, where the single included layer is the only output layer. When each_layer_is_an_ols_flag equals 0, it is specified that the output layer set may contain multiple layers. If vps_max_layers_minus1 equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag equals 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.

[0101] When ols_mode_idc equals 0, it is specified that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i (including 0 and i), and for each OLS, only the highest layer in the OLS is output.

[0102] When ols_mode_idc equals 1, it is specified that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i (including 0 and i), and for each OLS, all layers in the OLS are output.

[0103] When ols_mode_idc equals 2, it is specified that the total number of OLSs specified by the VPS is signaled explicitly, and for each OLS, the output layers are signaled explicitly, while the other layers are direct or indirect reference layers of the output layers of the OLS.

[0104] The value of ols_mode_idc shall be in the range of 0 to 2 (including 0 and 2). The value 3 of ols_mode_idc is reserved for future use by ITU-T|ISO / IEC.

[0105] When vps_all_independent_layers_flag equals 1 and each_layer_is_an_ols_flag equals 0, the value of ols_mode_idc is inferred to be equal to 2.

[0106] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc equals 2.

[0107] The variable TotalNumOlss is derived as follows, which specifies the total number of OLSs specified by the VPS:

[0108]

[0109] ols_output_layer_flag[i][j] is equal to 1, which specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. ols_output_layer_flag[i][j] is equal to 0, which specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.

[0110] Derive the variable NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS, and the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer in the i-th OLS in the following way:

[0111]

[0112]

[0113] For each OLS, there should be at least one layer as the output layer. In other words, for any i value in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.

[0114] Derive the variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS in the following way:

[0115]

[0116]

[0117] Note 1 - The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, only the contained layer is output.

[0118] Derive the variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j] in the following way:

[0119]

[0120] The lowest layer in each OLS shall be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]] shall be equal to 1.

[0121] Each layer shall be included in at least one OLS specified by the VPS. In other words, for each layer with a specific value nuhLayerId of nuh_layer_id (equal to one of vps_layer_id[k] for k in the range from 0 to vps_max_layers_minus1 (including 0 and vps_max_layers_minus1)), there shall exist at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), and j is in the range up to and including NumLayersInOls[i] - 1, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.

[0122] vps_num_ptls specifies the number of profile_tier_level() syntax structures in the VPS.

[0123] That pt_present_flag[i] is equal to 1 specifies the existence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. That pt_present_flag[i] is equal to 0 specifies the non-existence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 0. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (i - 1)-th profile_tier_level() syntax structure in the VPS.

[0124] ptl_max_temporal_id[i] specifies the TemporalId represented by the highest sublayer with level information in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0125] vps_ptl_byte_alignment_zero_bit shall be equal to 0.

[0126] ols_ptl_idx[i] specifies the index of the profile_tier_level() syntax structure applied to the i-th OLS in the list of profile_tier_level() syntax structures in the VPS. When present, the value of ols_ptl_idx[i] shall be in the range of 0 to vps_num_ptls - 1, inclusive.

[0127] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS exists in the SPS referred to by the layer in the i-th OLS.

[0128] vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params shall be in the range of 0 to 16, inclusive. When not present, the value of vps_num_dpb_params is inferred to be equal to 0.

[0129] When same_dpb_size_output_or_nonoutput_flag equals 1, it is specified that the syntax element layer_nonoutput_dpb_params_idx[i] does not exist in the VPS. When same_dpb_size_output_or_nonoutput_flag equals 0, it is specified that the syntax element layer_nonoutput_dpb_params_idx[i] may or may not exist in the VPS.

[0130] The vps_sublayer_dpb_params_present_flag is used to control the existence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[] and max_latency_increase_plus1[] in the dpb_parameters() syntax structure in the VPS. When they do not exist, the vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.

[0131] When dpb_size_only_flag[i] equals 1, it is specified that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] do not exist in the i-th dpb_parameters() syntax structure in the VPS. When dpb_size_only_flag[i] equals 1, it is specified that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] may exist in the i-th dpb_parameters() syntax structure in the VPS.

[0132] dpb_max_temporal_id[i] specifies the TemporalId represented by the highest sublayer in the i-th dpb_parameters() syntax structure where DPB parameters may be present in the VPS. The value of dpb_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0133] layer_output_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure list in the VPS that is applied to the i-th layer when it is an output layer in the OLS. When present, the value of layer_output_dpb_params_idx[i] shall be in the range of 0 to vps_num_dpb_params - 1, inclusive.

[0134] If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is an output layer, the dpb_parameters() syntax structure applied to the i-th layer is the dpb_parameter() syntax structure present in the SPS that the layer references.

[0135] Otherwise (vps_independent_layer_flag[i] is equal to 0), the following applies:

[0136] - When vps_num_dpb_params is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0137] - A requirement for bitstream consistency is that the value of layer_output_dpb_params_idx[i] shall be such that dpb_size_only_flag[layer_output_dpb_params_idx[i]] is equal to 0.

[0138] layer_nonoutput_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure list in the VPS that is applied to the i-th layer as a non-output layer in the OLS. When present, the value of layer_nonoutput_dpb_params_idx[i] shall be in the range from 0 to vps_num_dpb_params-1, inclusive (including 0 and vps_num_dpb_params-1).

[0139] If same_dpb_size_output_or_nonoutput_flag is equal to 1, the following applies:

[0140] - If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is a non-output layer, the dpb_parameters() syntax structure applied to that layer is the dpb_parameters() syntax structure that exists in the SPS referred to by that layer.

[0141] - Otherwise (vps_independent_layer_flag[i] is equal to 0), the value of layer_nonoutput_dpb_params_idx[i] is inferred to be equal to layer_output_dpb_params_idx[i].

[0142] Otherwise (same_dpb_size_output_or_nonoutput_flag is equal to 0), when vps_num_dpb_params is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0143] vps_general_hrd_params_present_flag being equal to 1 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are present in the VPS RBSP syntax structure. vps_general_hrd_params_present_flag being equal to 0 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are not present in the VPS RBSP syntax structure.

[0144] The vps_sublayer_cpb_params_present_flag being equal to 1 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId ranges from 0 to hrd_max_tid[i] (inclusive of 0 and hrd_max_tid[i]). The vps_sublayer_cpb_params_present_flag being equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId is equal to hrd_max_tid[i] only. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.

[0145] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented at the sublayer (where the TemporalId ranges from 0 to hrd_max_tid[i] - 1 (inclusive of 0 and hrd_max_tid[i] - 1)) are inferred to be the same as the HRD parameters represented at the sublayer (where the TemporalId is equal to hrd_max_tid[i]). These include HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element up to the sublayer_hrd_parameters(i) syntax structure under the "if(general_vcl_hrd_params_present_flag)" condition in the ols_hrd_parameters syntax structure.

[0146] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure. The value of num_ols_hrd_params_minus1 shall range from 0 to 63 (inclusive of 0 and 63). When TotalNumOlss is greater than 1, the value of num_ols_hrd_params_minus1 is inferred to be equal to 0.

[0147] hrd_max_tid[i] specifies the TemporalId of the highest sublayer representation for which the HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive (including 0 and vps_max_sublayers_minus1). When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0.

[0148] ols_hrd_idx[i] specifies the index of the ols_hrd_parameters() syntax structure applied to the i-th OLS. The value of ols_hrd_idx[[i] shall be in the range of 0 to num_ols_hrd_params_minus1, inclusive (including 0 and num_ols_hrd_params_minus1). When it does not exist, the value of ols_hrd_idx[[i] is inferred to be equal to 0.

[0149] vps_extension_flag being equal to 0 specifies that there is no vps_extension_data_flag syntax element in the VPS RBSP syntax structure. vps_extension_flag being equal to 1 specifies that there is a vps_extension_data_flag syntax element in the VPS RBSP syntax structure.

[0150] vps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of this specification. Decoders compliant with this version of this specification shall ignore all vps_extension_data_flag syntax elements.

[0151] 3.4. SPS Syntax and Semantics in VVC

[0152] In the latest VVC draft text in JVET-P2001-v14, the SPS syntax and semantics most relevant to the invention of this article are as follows.

[0153] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0154]

[0155] 7.4.3.3 Sequence Parameter Set RBSP Semantics

[0156] The SPS RBSP shall be available for the decoding process before it is referenced, included in at least one AU with TemporalId equal to 0, or provided by an external means.

[0157] All SPS NAL units with a specific value of sps_seq_parameter_set_id in the CVS shall have the same content.

[0158] When sps_decoding_parameter_set_id is greater than 0, it specifies the value of dps_decoding_parameter_set_id of the DPS referenced by the SPS. When sps_decoding_parameter_set_id is equal to 0, the SPS does not reference the DPS, and the DPS is not referenced when decoding each CLVS that references the SPS. The value of sps_decoding_parameter_set_id shall be the same in all SPSs referenced by coded pictures in the bitstream.

[0159] When sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS.

[0160] When sps_video_parameter_set_id is equal to 0, the following applies:

[0161] - The SPS does not reference the VPS.

[0162] - When decoding each CLVS that references the SPS, the VPS is not referenced.

[0163] - The value of vps_max_layers_minus1 is inferred to be equal to 0.

[0164] - The CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same nuh_layer_id value).

[0165] - The value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.

[0166] - The value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.

[0167] When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referred to by the CLVS with a specific nuh_layer_id value nuhLayerId shall have a nuh_layer_id equal to nuhLayerId.

[0168] sps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in the CLVS of each reference SPS. The value of sps_max_sublayers_minus1 shall be in the range from 0 to vps_max_sublayers_minus1 (including 0 and vps_max_sublayers_minus1).

[0169] sps_reserved_zero_4bits shall be equal to 0 in the bitstream of this version that complies with this specification. Other values of sps_reserved_zero_4bits are reserved for future use by ITU-T|ISO / IEC.

[0170] sps_ptl_dpb_hrd_params_present_flag being equal to 1 specifies that there are the profile_tier_level() syntax structure and the dpb_parameters() syntax structure in the SPS, and there may also be the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure in the SPS. sps_ptl_dpb_hrd_params_present_flag being equal to 0 specifies that these syntax structures do not exist in the SPS. The value of sps_ptl_dpb_hrd_params_present_flag shall be equal to vps_independent_layer_flag[nuh_layer_id].

[0171] If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the variable MaxDecPicBuffMinus1 shall be set to max_dec_pic_buffering_minus1[sps_max_sublayers_minus1] in the dpb_parameters() syntax structure in the SPS. Otherwise, MaxDecPicBuffMinus1 is set to be equal to max_dec_pic_buffering_minus1[sps_max_sublayers_minus1] in the dpb_parameters() syntax structure at the layer_nonoutput_dpb_params_idx[GeneralLayerIdx[nuh_layer_id]]-th in the VPS.

[0172] 3.5. Syntax and Semantics of Slice Headers in VVC

[0173] In the latest VVC draft text in JVET-P2001-v14, the slice header syntax and semantics most relevant to the present invention are as follows.

[0174] 7.3.7.1 General Slice Header Syntax

[0175]

[0176] 7.4.8.1 General Slice Header Semantics

[0177] The slice_type specifies the coding type of the slice according to Table 9.

[0178] Table 9 - Names Associated with slice_type

[0179] slice_type Name of slice_type 0 B (B-band) 1 P (P-band) 2 I (I-band)

[0180] When the nal_unit_type is a value of nal_unit_type within the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and the current picture is the first picture in the access unit, the slice_type shall be equal to 2.

[0181] 4. Technical Problems Solved by the Described Technical Solutions

[0182] The existing scalable designs in VVC have the following problems:

[0183] 1) The latest VVC draft text includes the following constraints on slice_type:

[0184] When the nal_unit_type is a nal_unit_type value within the range of IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and the current picture is the first picture in an AU, slice_type shall be equal to 2.

[0185] For a slice, a slice_type value equal to 2 means that the slice is intra-coded without using inter-prediction from reference pictures.

[0186] However, in an AU, not only the first picture that is located in the AU and is an IRAP picture needs to contain only intra-coded slices, but also all IRAP pictures in all independent layers need to contain only intra-coded slices. Therefore, the above constraints do need to be updated.

[0187] 2) When the syntax element ols_ptl_idx[i] does not exist, this value still needs to be used. However, when ols_ptl_idx[i] does not exist, there is a lack of inference about its value.

[0188] 3) When the i-th layer is not used as an output layer in any OLS, signaling of the syntax element layer_output_dpb_params_idx[i] is unnecessary.

[0189] 4) The value of the syntax element vps_num_dpb_params can be equal to 0. However, when vps_all_independent_layers_flag is equal to 0, there needs to be at least one dpb_parameters() syntax structure in the VPS.

[0190] 5) In the latest VVC draft text, the PTL information of the OLS only contains one layer, which is an independent coding layer that does not refer to any other layers and is only signaled in the SPS. However, for the purpose of session negotiation, it is desired to signal the PTL information for all OLSs in the bitstream in the VPS.

[0191] 6) When the number of PTL syntax structures signaled in the VPS is zero, signaling of the vps_ptl_byte_alignment_zero_bit syntax element is unnecessary.

[0192] 7) In the semantics of sps_video_parameter_set_id, there are the following constraints:

[0193] When sps_video_parameter_set_id is equal to 0, the CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same nuh_layer_id value).

[0194] However, this constraint does not allow including an independent layer that does not refer to the VPS in a multi-layer bitstream. Since sps_video_parameter_set_id being equal to 0 means that the SPS (and the layer) does not refer to the VPS.

[0195] 8) When each_layer_is_an_ols_flag is equal to 1, the value of vps_general_hrd_params_present_flag may be equal to 1. However, when each_layer_is_an_ols_flag is equal to 1, the HRD parameters are signaled only in the SPS, so the value of vps_general_hrd_params_present_flag should not be equal to 1. In some cases, when zero old_hrd_parameters() syntax structures are signaled in the VPS, signaling the general_hrd_parameter() syntax structure may not make sense.

[0196] 9) Signal the HRD parameters of an OLS that contains only one layer in both the VPS and the SPS. However, for an OLS that contains only one layer, repeating the HRD parameters in the VPS is useless.

[0197] 5. Example Embodiments and Techniques

[0198] To solve the above problems and other problems, the following summarized methods are disclosed. These inventions should be regarded as examples for explaining general concepts and should not be narrowly construed. In addition, these inventions can be applied alone or in any combination.

[0199] 1) To solve the first problem, the following constraint is specified:

[0200] When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type shall be equal to 2.

[0201] 2) To solve the second problem, for each possible value of i, when the syntax element does not exist, the value of ols_ptl_idx[i] is inferred to be equal to 0.

[0202] 3) To solve the third problem, a variable LayerUsedAsOutputLayerFlag[i] is defined to indicate whether the i-th layer is used as an output layer in any OLS, and when this variable equals 0, signaling of the variable layer_output_dpb_params_idx[i] is avoided.

[0203] a. Additionally, the following constraint can be further defined: for each value of i in the range from 0 to vps_max_layers_minus1 (including 0 and vps_max_layers_minus1), the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] should not both be equal to 0. In other words, there should not be a layer that is neither a direct reference layer of any other layer nor an output layer of at least one OLS.

[0204] 4) To solve the fourth problem, vps_num_dpb_params is changed to vps_num_dpb_params_minus1, and VpsNumDpbParams is defined as follows:

[0205] If (!vps_all_independent_layers_flag)

[0206] VpsNumDpbParams = vps_num_dpb_params_minus1 + 1

[0207] Otherwise

[0208] VpsNumDpbParams = 0

[0209] And vps_num_dpb_params used in the syntax condition and semantics is replaced with VpsNumDpbParams.

[0210] a. Additionally, the following constraint is further defined: when it does not exist, the value of same_dpb_size_output_or_nonoutput_flag is inferred to be equal to 1.

[0211] 5) To solve the fifth problem, for the purpose of session negotiation, it is allowed to repeat the PTL information of an OLS that contains only one layer in the VPS. This can be achieved by changing vps_num_ptls to vps_num_ptls_minus1.

[0212] a. Alternatively, keep vps_num_ptls (without making it vps_num_ptls_minus1), but remove "NumLayersInOls[i]>1&&" from the syntax condition of ols_ptl_idx[i].

[0213] a. Alternatively, keep vps_num_ptls (without making it vps_num_ptls_minus1), but set the vps_num_ptls condition to "if (!each_layer_is_an_ols_flag)" or "if (vps_max_layers_minus1>0 &&!vps_all_independent_layers_flag)".

[0214] b. Alternatively, also allow the repetition of DPB parameter information (only DPB size or all DPB parameters) for an OLS that contains only one layer in the VPS.

[0215] 6) To solve the sixth problem, as long as the number of PTL syntax structures signaled in the VPS is zero, the vps_ptl_byte_alignment_zero_bit syntax element is not signaled. This can be achieved by setting the vps_ptl_byte_alignment_zero_bit condition to "if (vps_num_ptls>0)", or by changing vps_num_ptls to vps_num_ptls_minus1, which effectively does not allow the number of PTL syntax structures signaled in the VPS to be equal to zero.

[0216] 7) To solve the seventh problem, remove the following constraint:

[0217] When sps_video_parameter_set_id is equal to 0, the CVS should contain only one layer (i.e., all VCL NAL units in the CVS should have the same nuh_layer_id value).

[0218] And add the following constraint:

[0219] The value of sps_video_parameter_set_id should be the same in all SPSs referenced by coded pictures in the CVS, and sps_video_parameter_set_id is greater than 0.

[0220] a. Alternatively, keep the following constraint:

[0221] When sps_video_parameter_set_id is equal to 0, the CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same nuh_layer_id value).

[0222] And the following constraints are specified:

[0223] The value of sps_video_parameter_set_id shall be the same in all SPSs referred to by the coded pictures in the CVS.

[0224] 8) To solve the eighth problem, when vps_general_hrd_params_present_flag is equal to 1, the syntax element vps_general_hrd_params_present_flag is not signaled, and when it is absent, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.

[0225] a. Alternatively, when each_layer_is_an_ols_flag is equal to 1, the value of vps_general_hrd_params_present_flag is restricted to be equal to 0.

[0226] b. In addition, since the syntax condition of the syntax element num_ols_hrd_params_minus1, i.e., "if (TotalNumOlss > 1)", is not required, it will be removed. This is because when TotalNumOlss is equal to 1, the value of each_layer_is_an_ols_flag will be equal to 1, and then the value of vps_general_hrd_params_present_flag will be equal to 0, and then the HRD parameters will not be signaled in the VPS.

[0227] In some cases, when the ols_hrd_parameters() syntax structure does not exist in the VPS, the general_hrd_parameters() syntax structure is not signaled in the VPS.

[0228] 9) To solve the ninth problem, the HRD parameters of the OLS with only one layer are signaled only in the SPS and not in the VPS.

[0229] 6. Embodiment

[0230] The following are some example embodiments of the aspects summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-P2001-v14. Most of the relevant parts that have been added or modified are highlighted as Underline Bold and some of the deleted parts are highlighted in italic bold. There are also some other changes that are editorial in nature and thus not highlighted.

[0231] 6.1. First Embodiment

[0232] 6.1.1. VPS Syntax and Semantics

[0233] 7.3.2.2 Video Parameter Set Syntax

[0234]

[0235]

[0236]

[0237] 7.4.3.2 Video Parameter Set RBSP Semantics

[0238] ols_output_layer_flag[i][j] being equal to 1 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. ols_output_layer_flag[i][j] being equal to 0 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.

[0239] The variables NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS and the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer in the i-th OLS are derived as follows:

[0240]

[0241]

[0242]

[0243] For each OLS, there should be at least one layer that serves as the output layer. In other words, for any value of i in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.

[0244] Derive the variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS in the following way:

[0245]

[0246]

[0247] Note - The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and only the contained layer is output for the 0-th OLS.

[0248] Derive the variable OlsLayeIdx[i][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j] in the following way:

[0249]

[0250] The lowest layer in each OLS should be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]] should be equal to 1.

[0251] Each layer shall be included in at least one OLS specified by the VPS. In other words, for each layer for which the value of nuh_layer_id, nuhLayerId, is equal to one of vps_layer_id[k] (for k in the range from 0 to vps_max_layers_minus1, inclusive of 0 and vps_max_layers_minus1), there shall exist at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1, inclusive of 0 and TotalNumOlss - 1, and j is in the range from 0 to NumLayersInOls[i] - 1, inclusive of NumLayersInOls[i] - 1, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.

[0252] Specify the number of profile_tier_level() syntax structures in the VPS.

[0253] pt_present_flag[i] being equal to 1 specifies the presence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] being equal to 0 specifies the absence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as the tier, layer, and general constraint information in the (i - 1)-th profile_tier_level() syntax structure in the VPS.

[0254] ptl_max_temporal_id[i] specifies the TemporalId for the highest sublayer representation for which level information exists in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[i] shall be in the range from 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0255] vps_ptl_byte_alignment_zero_bit shall be equal to 0.

[0256] ols_ptl_idx[i] specifies the index of the profile_tier_level() syntax structure in the list of profile_tier_level() syntax structures in the VPS that applies to the i-th OLS. When present, the value of ols_ptl_idx[i] shall be in the range from 0 to inclusive.

[0257] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure that applies to the i-th OLS exists in the SPS referenced by the layer in the i-th OLS.

[0258] plus 1 specifies the number of dpb_parameters() syntax structures in the VPS. When present, the value shall be in the range from 0 to 15, inclusive.

[0259]

[0260] When same_dpb_size_output_or_nonoutput_flag equals 1, it specifies that there is no layer_nonoutput_dpb_params_idx[i] syntax element in the VPS. When same_dpb_size_output_or_nonoutput_flag equals 0, it specifies that there may or may not be a layer_nonoutput_dpb_params_idx[i] syntax element in the VPS.

[0261] The vps_sublayer_dpb_params_present_flag is used to control the presence of the max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] syntax elements in the dpb_parameters() syntax structure in the VPS. When they are absent, the vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.

[0262] When dpb_size_only_flag[i] equals 1, it specifies that there are no max_num_reorder_pics[] and max_latency_increase_plus1[] syntax elements in the i-th dpb_parameters() syntax structure in the VPS. When dpb_size_only_flag[i] equals 0, it specifies that there may be max_num_reorder_pics[] and max_latency_increase_plus1[] syntax elements in the i-th dpb_parameters() syntax structure in the VPS.

[0263] dpb_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation for which DPB parameters may be present in the i-th dpb_parameters() syntax structure in the VPS. The value of dpb_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1 (including 0 and vps_max_sublayers_minus1). When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0264] layer_output_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure in the list of dpb_parameters() syntax structures in the VPS that is applied to the i-th layer as the output layer in the OLS. When present, the value of layer_output_dpb_params_idx[i] shall be in the range of 0 to (including 0 and VpsNumDpbParams-1).

[0265] If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is the output layer, the dpb_parameters() syntax structure applied to the i-th layer is the dpb_parameters() syntax structure present in the SPS that the layer references.

[0266] Otherwise (vps_independent_layer_flag[i] is equal to 0), the following applies:

[0267] - When is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0268] - A requirement for bitstream compliance is that the value of layer_output_dpb_params_idx[i] shall be such that dpb_size_only_flag[layer_output_dpb_params_idx[i]] is equal to 0.

[0269] layer_nonoutput_dpb_params_idx[i] specifies the index in the list of dpb_parameters() syntax structures in the VPS of the dpb_parameters() syntax structure applied to the i-th layer as a non-output layer in the OLS. When present, the value of layer_nonoutput_dpb_params_idx[i] shall be in the range from 0 to VpsNumDpbParams - 1, inclusive (including 0 and VpsNumDpbParams - 1).

[0270] If same_dpb_size_output_or_nonoutput_flag is equal to 1, the following applies:

[0271] - If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is a non-output layer, the dpb_parameters() syntax structure applied to that layer is the dpb_parameters() syntax structure present in the SPS that the layer refers to.

[0272] - Otherwise (vps_independent_layer_flag[i] is equal to 0), the value of layer_nonoutput_dpb_params_idx[i] is inferred to be equal to layer_output_dpb_params_idx[i].

[0273] Otherwise (same_dpb_size_output_or_nonoutput_flag is equal to 0), when VpsNumDpbParams is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0274] vps_general_hrd_params_present_flag being equal to 1 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are present in the VPS RBSP syntax structure. vps_general_hrd_params_present_flag being equal to 0 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are not present in the VPS RBSP syntax structure.

[0275] The vps_sublayer_cpb_params_present_flag being equal to 1 specifies that the ith ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId ranges from 0 to hrd_max_tid[i] (inclusive of 0 and hrd_max_tid[i]). The vps_sublayer_cpb_params_present_flag being equal to 0 specifies that the ith ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId is equal to only hrd_max_tid[i]. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.

[0276] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented at the sublayer with TemporalId in the range from 0 to hrd_max_tid[i] - 1 (inclusive of 0 and hrd_max_tid[i] - 1) are inferred to be the same as the HRD parameters represented at the sublayer with TemporalId equal to hrd_max_tid[i]. These include such HRD parameters: starting from the fixed_pic_rate_general_flag[i] syntax element up to the sublayer_hrd_parameters(i) syntax structure immediately under the "if(general_vcl_hrd_params_present_flag)" condition in the ols_hrd_parameters syntax structure.

[0277] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure. The value of num_ols_hrd_params_minus1 shall be in the range from 0 to 63 (inclusive of 0 and 63).

[0278] hrd_max_tid[i] specifies the TemporalId of the highest sublayer representation for which HRD parameters are contained in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[i] is inferred to be equal to vps_max_sublayers_minus1.

[0279] ols_hrd_idx[i] specifies The ols_hrd_parameters() syntax structure applied to the i-th OLS When present, the value of ols_hrd_idx[i] shall be in the range of 0 to num_ols_hrd_params_minus1, inclusive. When , the value of ols_hrd_idx[[i]] is inferred to be equal to 0.

[0280] vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element is not present in the VPS RBSP syntax structure. vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element is present in the VPS RBSP syntax structure.

[0281] The vps_extension_data_flag can have any value. Its presence and value do not affect decoder conformance to the profiles specified in this version of this specification. Decoders conforming to this version of this specification should ignore all vps_extension_data_flag syntax elements.

[0282] 6.1.2.SPS Semantics

[0283] 7.4.3.3 Sequence Parameter Set RBSP Semantics

[0284] The SPS RBSP shall be available for the decoding process before being referenced, included in at least one AU (where TemporalId equals 0), or provided by an external means.

[0285] All SPS NAL units with a specific value of sps_seq_parameter_set_id in the CVS shall have the same content.

[0286] When sps_decoding_parameter_set_id is greater than 0, it specifies the value of dps_decoding_parameter_set_id of the DPS referenced by the SPS. When sps_decoding_parameter_set_id equals 0, the SPS does not reference the DPS, and no DPS is referenced when decoding each CLVS that references the SPS. The value of sps_decoding_parameter_set_id shall be the same for all SPSs decoded pictures referenced in the bitstream.

[0287] When sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS.

[0288] When sps_video_parameter_set_id equals 0, the following applies:

[0289] - The SPS does not reference the VPS.

[0290] - When decoding each CLVS that references the SPS, the VPS is not referenced.

[0291] - The value of vps_max_layers_minus1 is inferred to be equal to 0.

[0292]

[0293] - The value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.

[0294] - The value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.

[0295] When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referred to by the CLVS with a specific nuh_layer_id value nuhLayerId shall have a nuh_layer_id equal to nuhLayerId.

[0296] 6.1.3. Syntax of slice header

[0297] 7.4.8.1 Syntax of general slice header ...

[0299] The slice_type specifies the coding / decoding type of the slice according to Table 9.

[0300] Table 9 - Names associated with slice_type

[0301] slice_type Name of slice_type 0 B (B-band) 1 P (P-band) 2 I (I-band)

[0302] When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), the slice_type shall be equal to 2.

[0303] Figure 1 FIG. shows a block diagram of an exemplary video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or it can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, passive optical network (PON), etc.), and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0304] System 1900 may include an encoding / decoding component 1904, which may implement various encoding or coding methods described in this document. The encoding / decoding component 1904 may reduce the average bit rate of a video from the input 1902 to the output of the encoding / decoding component 1904 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 1904 may be stored or transmitted via communication over a connection represented by component 1906. A stored or transmitted (or encoded / decoded) representation of the video received at the input 1902 may be used by component 1908 to generate pixel values or a displayable video to be sent to a display interface 1910. The process of generating a user-viewable video from the bit stream is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding / decoding" operations or tools, it will be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations to reverse the encoding result will be performed by the decoder.

[0305] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0306] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 may be configured to implement one or more methods described in this document. One or more memories 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 3606 may be partially or entirely located within the (multiple) processors 3602 (e.g., a graphics processor).

[0307] Figure 4 is a block diagram showing an example video encoding / decoding system 100 that may utilize the techniques of the present disclosure.

[0308] As Figure 4As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data that may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the source device 110 may be referred to as a video decoding device.

[0309] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0310] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0311] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0312] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be located external to the destination device 120, and the destination device 120 is configured to interface with an external display device.

[0313] The video encoder 114 and the video decoder 124 may operate according to video compression standards (e.g., the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards).

[0314] Figure 5 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may be Figure 4 the video encoder 114 in the system 100 shown.

[0315] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0316] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0317] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0318] In addition, some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be highly integrated, but are shown separately in the Figure 5 example for purposes of explanation.

[0319] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0320] The mode selection unit 203 may select one of the coding / decoding modes (intra or inter) based on, for example, error results, and provide the resulting intra or inter coded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra prediction and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel accuracy or integer pixel accuracy).

[0321] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures other than the picture associated with the current video block from buffer 213.

[0322] For example, the motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0323] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0324] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0, and may also search for another reference video block of the current video block in the reference pictures of list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector that indicates the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0325] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.

[0326] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to the motion information of another video block to

[0327] Notify the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0328] In one example, the motion estimation unit 204 may indicate, in the syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0329] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0330] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0331] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0332] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0333] In other examples, for the current video block, there may be no residual data for the current video block. For example, in the skip mode, the residual generation unit 207 may not perform the subtraction operation.

[0334] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0335] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0336] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct the residual video block based on the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0337] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0338] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0339] Figure 6 is a block diagram showing an example of a video decoder 300, and the video decoder 300 can be Figure 4 the video decoder 114 in the system 100 shown.

[0340] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 5 an example, the video decoder 300 includes multiple functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0341] In Figure 6 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding pass that is generally opposite to the encoding pass ( Figure 5 ) described with respect to the video encoder 200.

[0342] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data, and the motion compensation unit 302 can determine motion information based on the entropy decoded video data, and the motion information includes motion vectors, motion vector precision, reference picture list indices, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge mode.

[0343] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. The syntax elements may include an identifier for the interpolation filter to be used with sub-pixel precision.

[0344] The motion compensation unit 302 may use the interpolation filter used by the video encoder 20 during video block coding to calculate the interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information, and use the interpolation filter to generate a prediction block.

[0345] The motion compensation unit 302 may use some syntax information to determine the size of the blocks for encoding the frames and / or slices of the coded video sequence, the partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.

[0346] The intra prediction unit 303 may form a prediction block according to spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0347] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, and the buffer 307 provides a reference block for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.

[0348] A list of preferred solutions for some embodiments is provided below.

[0349] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0350] 1. A video processing method (e.g., Figure 3 method 600 in), comprising: performing a conversion between a video including one or more scalable video layers and a coded representation of the video, wherein the coded representation conforms to format rules, and wherein the format rules specify that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the segment does not use inter-layer prediction, the value of the field indicating the segment type of the segment is set to indicate the type of intra segment.

[0351] 2. The method of Solution 1, wherein the stripe type of the stripe is set to 2.

[0352] 3. The method of Solution 1, wherein the predetermined range is from IDR_W_RADL to CRA_NUT, both inclusive.

[0353] 4. A method for video processing, comprising: performing a conversion between a video including one or more scalable video layers and a coded representation of the video, wherein the coded representation conforms to format rules, and wherein the format rules specify that in the case where an index of a syntax structure describing the profile, layer, and level of one or more scalable video layers is excluded from the coded representation, the indication of the index is inferred to be equal to 0.

[0354] 5. The method of Solution 4, wherein the indication includes the ols_ptl_idx[i] syntax element, where i is an integer.

[0355] 6. A method for video processing, comprising: performing a conversion between a video including one or more scalable video layers and a coded representation of the video, wherein the coded representation includes a plurality of output layer sets (OLSs), wherein each OLS is a set of layers in the coded representation, and wherein the set of layers is specified to be output, wherein the coded representation conforms to format rules, and wherein the format rules specify that in the case where an OLS includes a single layer, the set of video parameters of the coded representation is allowed to have duplicate information regarding the profile, layer, and level of the OLS.

[0356] 7. The method of Solution 6, wherein the format rules specify a signaling digital field that indicates the total number of sets of information regarding the profile, layer, and level of a plurality of OLSs, where the total number is at least one.

[0357] 8. The method of any one of Solutions 1 - 7, wherein performing the conversion includes encoding the video to generate the coded representation.

[0358] 9. The method of any one of Solutions 1 - 7, wherein performing the conversion includes parsing and decoding the coded representation to generate the video.

[0359] 10. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 9.

[0360] 11. A video encoding apparatus, comprising a processor configured to implement the method described in one or more of Solutions 1 to 9.

[0361] 12. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method recited in any one of Solutions 1 to 9.

[0362] 13. The method, apparatus or system described in this document.

[0363] Regarding Figures 7A through 7I , the following listed solutions can be preferably implemented in some embodiments.

[0364] For example, the following solution can be implemented according to Item 1 in the previous section.

[0365] 1. A method for video processing (e.g., Figure 7A the method 710 described in

[0366] ), including performing (712) a conversion between a video including one or more scalable video layers and a bitstream of the video, where the video includes one or more video pictures including one or more strips; where the bitstream conforms to format rules, and where the format rules stipulate that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the strip does not use inter-layer prediction, the value of the field indicating the strip type of the strip is set to the type indicating an intra strip.

[0367] 2. The method of Solution 1, where the strip type of the strip is set to 2.

[0368] 3. The method of Solution 1, where the predetermined range is from IDR_W_RADL to CRA_NUT, both inclusive, where IDR_W_RADL indicates a strip having an instantaneous decoding refresh type, and where CRA_NUT indicates a strip having a clean random access type.

[0369] 4. The method of any one of Solutions 1 - 3, where the corresponding video layer flag corresponds to vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]], where nuh_layer_id corresponds to the identifier of the video layer, GeneralLayerIdx[] is the index of the video layer, and vps_independent_layer_flag corresponds to the video layer flag.

[0370] 5. The method of any one of Solutions 1 - 4, where performing the conversion includes encoding the video into the bitstream.

[0371] 7. A method for storing a video bitstream, comprising: generating a bitstream according to a video including one or more scalable video layers; and storing the bitstream in a non-transitory computer-readable storage medium; wherein the video includes one or more video pictures, and the video pictures include one or more slices; wherein the bitstream conforms to format rules, and the format rules stipulate that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the slice does not use inter-layer prediction, the value of the field indicating the slice type of the slice is set to indicate the type of I slice.

[0372] For example, the following solution can be implemented according to item 5 in the previous section.

[0373] 1. A video processing method (e.g., Figure 7B the method 720 shown), comprising: performing (722) a conversion between a video including multiple video layers and a bitstream of the video, wherein the bitstream includes multiple output layer sets (OLSs), each output layer set (OLS) includes one or more of the multiple scalable video layers, and the bitstream conforms to format rules, and the format rules stipulate that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, layer, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream.

[0374] 2. The method of solution 1, wherein the format rules stipulate that the syntax element is included in the video parameter set indicating the multiple PTL syntax structures decoded in the video parameter set.

[0375] 3. The method of solution 2, wherein the number of PTL syntax structures decoded in the video parameter set is equal to the value of the syntax element plus one.

[0376] 4. The method of solution 2, wherein the number of PTL syntax structures decoded in the video parameter set is equal to the value of the syntax element.

[0377] 5. The method of solution 4, wherein the format rules stipulate that when it is determined that one or more of the multiple OLSs contain more than one video layer, the syntax element is decoded in the video parameter set during encoding or parsed from the video parameter set during decoding.

[0378] 6. The method of solution 4, wherein the format rules stipulate that when it is determined that the number of video layers is greater than 0 and one or more video layers use inter-layer prediction, the syntax element is decoded in the video parameter set during encoding or parsed from the video parameter set during decoding.

[0379] 7. The method of any one of Solutions 1 to 6, wherein the formatting rules further stipulate that for an OLS with a single layer, the decoded picture buffer size parameter of the OLS is included in the video parameter set of the bitstream, and the decoded picture buffer size parameter of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0380] 8. The method of any one of Solutions 1 to 6, wherein the formatting rules further stipulate that for an OLS with a single layer, the parameter information of the decoded picture buffer of the OLS is included in the video parameter set of the bitstream, and the parameter information of the decoded picture buffer of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0381] For example, the following solutions can be implemented according to Item 6 in the previous section.

[0382] 9. A video processing method (e.g., Figure 7C the method 730 shown), comprising: performing (732) a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of video layers, and the bitstream conforms to formatting rules, wherein the formatting rules stipulate the relationship between a plurality of profile-tier-level (PTL) syntax structures appearing in the video parameter set of the bitstream and the byte alignment syntax fields in the video parameter set; wherein each PTL syntax structure indicates the profile, tier, and level of one or more of the plurality of OLSs.

[0383] 10. The method of Solution 9, wherein the formatting rules stipulate that the video parameter set includes at least one PTL syntax structure, and due to including at least one PTL syntax structure, one or more instances of the byte alignment syntax fields are included in the video parameter set.

[0384] 11. The method of Solution 10, wherein the byte alignment syntax field is one bit.

[0385] 12. The method of Solution 11, wherein each value of one or more instances of the byte alignment syntax field has a value of 0.

[0386] For example, the following solutions can be implemented according to Item 2 in the previous section.

[0387] 13. A video processing method (e.g., Figure 7DThe method shown (740) includes: performing (742) a conversion between a video including a plurality of scalable video layers and a bitstream of the video, where the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of scalable video layers, and the bitstream conforms to format rules, where the format rules specify that during encoding, in the case where the value of an index is zero, a syntax element indicating an index of a syntax structure that describes the profile, layer, and level of one or more of the plurality of OLSs is excluded from a set of video parameters of the bitstream, or during decoding, in the case where the syntax element does not exist in the bitstream, the value is inferred to be zero.

[0388] 14. The method of solution 13, where the index points to a syntax structure that indicates the profile, layer, and level of at least one OLS.

[0389] 15. A method for storing a bitstream of a video, including: generating a bitstream from a video including a plurality of video layers and storing the bitstream in a non-transitory computer-readable storage medium; where the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, where the format rules specify that for an OLS with a single layer, a profile-tier-level (PTL) syntax structure indicating the profile, layer, and level of the OLS is included in a set of video parameters of the bitstream, and the PTL syntax structure of the OLS is also included in a set of sequence parameters encoded and decoded in the bitstream.

[0390] 16. The method of any one of solutions 1-14, where performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0391] For example, the following solutions can be implemented according to item 3 in the previous section.

[0392] 1. A video processing method (e.g., Figure 7E the method shown 750) includes: performing (752) a conversion between a video including a plurality of video layers and a bitstream of the video, where the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of video layers, and the bitstream conforms to format rules, where the format rules specify that for layer i, where i is an integer, the bitstream includes a set of first syntax elements indicating a first variable that indicates whether layer i is included in at least one of the plurality of OLSs.

[0393] 2. The method of Solution 1, wherein the formatting rule stipulates that in the case where the first variable of layer i is equal to zero, it means that layer i is not included in any of the multiple OLSs, and the bitstream excludes the second set of syntax elements indicating the decoded picture buffer parameters of layer i.

[0394] 3. The method of any one of Solutions 1 - 2, wherein the formatting rule further stipulates that the bitstream includes a third set of syntax elements indicating a second variable, the second variable indicating whether layer i is used as a reference layer for at least one of the multiple video layers, and wherein the formatting rule does not allow the first variable and the second variable to have zero values.

[0395] 4. The method of Solution 3, wherein the formatting rule does not allow the values of the first variable and the second variable to both be equal to 0, which indicates that no layer is neither a direct reference layer of any other layer nor an output layer of at least one OLS.

[0396] 5. The method of any one of Solutions 1 - 4, wherein the first variable is a one - bit flag, denoted as LayerUsedAsOutputLayerFlag.

[0397] 6. The method of Solution 5, wherein the first variable is determined based on repeatedly checking the value of a third variable for each layer in the multiple video layers, the value indicating the relationship between the multiple layers included in the multiple OLSs.

[0398] 7. The method of Solution 6, wherein the third variable indicating the relationship between the multiple layers included in the multiple OLSs is allowed to have values 0, 1, or 2.

[0399] 8. The method of any one of Solutions 1 - 7, wherein performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non - transitory computer - readable storage medium.

[0400] 9. A method for storing a video bitstream, including: generating a bitstream according to a video including multiple video layers, and storing the bitstream in a non - transitory computer - readable storage medium; wherein the bitstream includes multiple output layer sets (OLSs), each output layer set including one or more of the multiple video layers, and the bitstream conforms to formatting rules, wherein the formatting rule stipulates that for layer i (where i is an integer), the bitstream includes a set of first syntax elements indicating a first variable, the first variable indicating whether layer i is included in at least one of the multiple OLSs.

[0401] For example, the following solutions can be implemented according to Item 4 in the previous section.

[0402] 1. A video processing method (for example, Figure 7FThe method shown (760) includes: performing (762) a conversion between a video and a bitstream of the video, where the bitstream includes one or more sets of output layers, and each set of output layers includes one or more video layers; where the bitstream conforms to format rules, and the format rules stipulate that in the case where each set of output layers includes a single video layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: zero; or in the case where it is not true that each set of output layers includes a single layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: one plus the value of a syntax element.

[0403] The method of Solution 1, where the syntax element corresponds to the vps_num_dpb_params_minus1 syntax element.

[0404] The method of any one of Solutions 1-2, where the format rules stipulate that in the case where there is no other syntax element in the video parameter set (the syntax element indicates whether the same dimensions are used to indicate the decoded picture buffer syntax structures of video layers that are included and not included in one or more sets of output layers), the value of the other syntax element is inferred to be equal to 1.

[0405] The method of any one of Solutions 1-3, where performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0406] A method for storing a bitstream of a video, including: generating a bitstream according to the video; and storing the bitstream in a non-transitory computer-readable storage medium; where the bitstream includes one or more sets of output layers, and each set of output layers includes one or more video layers; where the bitstream conforms to format rules, and the format rules stipulate that in the case where each set of output layers includes a single video layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: zero; or in the case where it is not true that each set of output layers includes a single layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: one plus the value of a syntax element.

[0407] For example, the following solutions can be implemented according to Item 7 in the previous section.

[0408] 1. A video processing method (for example, Figure 7GThe method shown (770) includes: performing (772) a conversion between a video and a bitstream of the video, where the bitstream includes a coded video sequence (CVS), the coded video sequence including one or more coded video pictures of one or more video layers; and where the bitstream conforms to format rules that specify that one or more sequence parameter sets (SPSs) that indicate conversion parameters referenced by one or more coded pictures of the CVS have the same reference video parameter set (VPS) identifier that indicates a reference VPS.

[0409] The method of Solution 1, where the format rules further specify that the same reference VPS identifier has a value greater than 0.

[0410] The method of any one of Solutions 1-2, where the format rules further specify that, in response to and only when the CVS includes a single video layer, the value zero of the SPS identifier is used.

[0411] The method of any one of Solutions 1-3, where performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0412] A method for storing a bitstream of a video, including: generating a bitstream according to the video and storing the bitstream in a non-transitory computer-readable storage medium; where the video includes one or more video pictures, the video pictures including one or more slices; where the bitstream includes a coded video sequence (CVS), the coded video sequence including one or more coded video pictures of one or more video layers; and where the bitstream conforms to format rules that specify that one or more sequence parameter sets (SPSs) that indicate conversion parameters referenced by one or more coded pictures of the CVS have the same reference video parameter set (VPS) identifier that indicates a reference VPS.

[0413] For example, the following solutions can be implemented according to item 8 in the previous section.

[0414] 1. A video processing method (e.g., Figure 7H the method shown 780) includes: performing (782) a conversion between a video and a bitstream of the video, where the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) including one or more video layers, where the bitstream conforms to format rules; where the format rules specify whether or how a first syntax element is included in a video parameter set (VPS) of the bitstream, the first syntax element indicating whether a first syntax structure that describes parameters of a hypothetical reference decoder (HRD) is used for the conversion.

[0415] 2. The method of Solution 1, wherein the first syntax structure includes a set of common HRD parameters.

[0416] 3. The method of Solutions 1-2, wherein the formatting rules specify that since each of one or more OLSs includes one or more video layers, when the first syntax element is not present in the VPS, the first syntax element is ignored from the VPS and is inferred to have a zero value, and wherein each of the one or more OLSs includes a single video layer.

[0417] 4. The method of Solutions 1-2, wherein the formatting rules specify that since each of one or more OLSs includes one or more video layers, when the first syntax element appears in the VPS, the first syntax element has a zero value in the VPS, and wherein each of the one or more OLSs includes a single video layer.

[0418] 5. The method of any one of Solutions 1-4, wherein the formatting rules further specify whether or how the VPS includes a second syntax element that indicates a plurality of syntax structures that describe OLS-specific HRD parameters.

[0419] 6. The method of Solution 5, wherein the formatting rules further specify that when the first syntax element has a value of 1, the second syntax element is included in the VPS regardless of whether the total number of OLSs in the one or more OLSs is greater than 1.

[0420] For example, the following solutions can be implemented according to Items 9 and 10 in the previous section.

[0421] 7. A video processing method (e.g., Figure 7I the method 790 shown), comprising: performing (792) a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) includes one or more video layers, and wherein the bitstream conforms to formatting rules; wherein the formatting rules specify whether or how the video parameter set (VPS) of the bitstream includes a first syntax structure that describes common hypothetical reference decoder (HRD) parameters and a plurality of second syntax structures that describe OLS-specific HRD parameters.

[0422] 8. The method of Solution 7, wherein the formatting rules specify that the first syntax structure is omitted from the VPS in the case where no second syntax structure is included in the VPS.

[0423] 9. The method of Solution 7, wherein, for an OLS that includes only one video layer, the formatting rules exclude the first syntax structure and the second syntax structure from being included in the VPS, and wherein the formatting rules allow the first syntax structure and the second syntax structure to be included in an OLS that includes only one video layer.

[0424] 10. A method for storing a bitstream of a video, comprising: generating a bitstream according to the video; storing the bitstream in a non-transitory computer-readable storage medium; wherein the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) includes one or more video layers, wherein the bitstream conforms to formatting rules; wherein the formatting rules specify whether or how a first syntax element is included in a video parameter set (VPS) of the bitstream, and the first syntax element indicates whether a first syntax structure describing parameters of a hypothetical reference decoder (HRD) is used for transformation.

[0425] The solutions listed above may further include:

[0426] In some embodiments, in the solutions listed above, performing the transformation includes encoding the video into a bitstream.

[0427] In some embodiments, in the solutions listed above, performing the transformation includes parsing and decoding the video according to the bitstream.

[0428] In some embodiments, in the solutions listed above, performing the transformation includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0429] In some embodiments, a video decoding device includes a processor configured to implement the method described in one or more of the solutions listed above.

[0430] In some embodiments, a video encoding device includes a processor configured to implement the method described in one or more of the solutions listed above.

[0431] In some embodiments, a non-transitory computer-readable storage medium may store instructions that cause a processor to implement the method described in one or more of the solutions listed above.

[0432] In some embodiments, a non-transitory computer-readable storage medium stores a bitstream of a video, and the bitstream is generated by the method described in one or more of the solutions listed above.

[0433] In some embodiments, the encoding method described above may be implemented by a device, and the device may further write the bitstream generated by implementing the method to a computer-readable medium.

[0434] The disclosed solutions and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in: digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of the foregoing. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, such as including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver apparatus.

[0435] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can also be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0436] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0437] For example, processors suitable for executing a computer program include general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to receive data from and transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks). However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including for example semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0438] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject or of the claimed subject matter, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features that are described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in a particular combination and even initially claimed as such, in some cases, one or more features from the claimed combination can be deleted, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.

[0439] Similarly, although these operations are described in a particular order in the figures, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood to require such separation in all embodiments.

[0440] Only some implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets, and each output layer set includes one or more video layers; wherein the bitstream conforms to format rules, wherein the format rules specify that when the value of the third syntax element vps_max_layers_minus1 is equal to 0, the value of the first syntax element vps_all_independent_layers_flag, which indicates whether all layers specified by a set of video parameters for the bitstream are independently encoded and decoded without using inter-layer prediction, is equal to 1, the value of the second syntax element each_layer_is_an_ols_flag, which indicates whether each output layer set includes a single video layer, in the set of video parameters is equal to 1, and the value of the variable VpsNumDpbParams, which indicates the number of decoded picture buffer parameter syntax structures included in the set of video parameters, is equal to 0, wherein adding 1 to the third syntax element specifies the number of layers, and the number of layers is the maximum allowed number of layers in each coded video sequence with reference to the set of video parameters, wherein the format rules specify that a fifth syntax element is conditionally included in the set of video parameters, and the fifth syntax element indicates an index of a decoded picture buffer parameter syntax structure applied to an output layer set relative to a list of decoded picture buffer parameter syntax structures included in the set of video parameters, wherein when the fifth syntax element does not exist and the number of decoded picture buffer parameter syntax structures is equal to 1, the value of the fifth syntax element is inferred to be equal to 0.

2. The method according to claim 1, wherein The format rules specify that when it is not true that each output layer set includes a single video layer, the variable VpsNumDpbParams indicates that the number of decoded picture buffer parameter syntax structures is equal to one plus the value of a fourth syntax element, and the fourth syntax element corresponds to the vps_num_dpb_params_minus1 syntax element.

3. The method according to claim 1, wherein, When the fifth syntax element exists, the value of the fifth syntax element is in the range from 0 to the number of decoded picture buffer parameter syntax structures minus 1.

4. The method according to claim 1, wherein, The format rules specify that when a sixth syntax element does not exist in the set of video parameters, the value of the sixth syntax element is inferred to be equal to 1, and the sixth syntax element indicates whether the same size is used to indicate a decoded picture buffer syntax structure for a video layer allowed to be included in the one or more output layer sets.

5. The method according to claim 1, wherein Performing the conversion includes encoding the video into the bitstream.

6. The method according to claim 1, wherein, Performing the conversion includes decoding the video from the bitstream.

7. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to: Perform a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets, and each output layer set includes one or more video layers; wherein the bitstream conforms to format rules, Wherein, the format rule stipulates that when the value of the third syntax element vps_max_layers_minus1 is equal to 0, the value of the first syntax element vps_all_independent_layers_flag, which indicates whether all layers specified by the video parameter set for the bitstream are independently coded and decoded without using inter-layer prediction, is equal to 1, the value of the second syntax element each_layer_is_an_ols_flag, which indicates whether each output layer set in the video parameter set includes a single video layer, is equal to 1, and the value of the variable VpsNumDpbParams, which indicates the number of decoded picture buffer parameter syntax structures included in the video parameter set, is equal to 0. When the value of the third syntax element vps_max_layers_minus1 is equal to 0, where adding 1 to the third syntax element specifies the number of layers, and the number of layers is the maximum allowable number of layers in each coded and decoded video sequence with reference to the video parameter set. Wherein, the format rule stipulates that the fifth syntax element is conditionally included in the video parameter set, and the fifth syntax element indicates an index of the decoded picture buffer parameter syntax structure applied to the output layer set relative to the list of decoded picture buffer parameter syntax structures included in the video parameter set. Wherein, when the fifth syntax element does not exist and the number of decoded picture buffer parameter syntax structures is equal to 1, the value of the fifth syntax element is inferred to be equal to 0.

8. The apparatus according to claim 7, wherein, The format rule stipulates that when it is not true that each output layer set includes a single video layer, the variable VpsNumDpbParams indicates that the number of decoded picture buffer parameter syntax structures is equal to one plus the value of the fourth syntax element, and wherein, the fourth syntax element corresponds to the vps_num_dpb_params_minus1 syntax element.

9. The apparatus according to claim 7, wherein, When the fifth syntax element exists, the value of the fifth syntax element ranges from 0 to the number of decoded picture buffer parameter syntax structures minus 1.

10. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video and a bitstream of the video, Among them, the bitstream including one or more output layer sets, each output layer set including one or more video layers; wherein, the bitstream conforms to the format rule. Among them, the format rule stipulates that when the value of the third syntax element vps_max_layers_minus1 is equal to 0, the value of the first syntax element vps_all_independent_layers_flag, which indicates whether all the layers specified by the video parameter set for the bitstream are independently encoded and decoded without using inter-layer prediction, is equal to 1, the value of the second syntax element each_layer_is_an_ols_flag in the video parameter set, which indicates whether each output layer set includes a single video layer, is equal to 1, and the value of the variable VpsNumDpbParams, which indicates the number of decoded picture buffer parameter syntax structures included in the video parameter set, is equal to 0, where adding 1 to the third syntax element specifies the number of layers, and the number of layers is the maximum allowable number of layers in each encoded and decoded video sequence with reference to the video parameter set. Among them, the format rule stipulates that the fifth syntax element is conditionally included in the video parameter set, and the fifth syntax element indicates the index of the decoded picture buffer parameter syntax structure applied to the output layer set relative to the list of decoded picture buffer parameter syntax structures included in the video parameter set. Among them, when the fifth syntax element does not exist and the number of decoded picture buffer parameter syntax structures is equal to 1, the value of the fifth syntax element is inferred to be equal to 0.

11. The non-transitory computer-readable storage medium according to claim 10, Among them, When the fifth syntax element exists, the value of the fifth syntax element is in the range from 0 to the number of decoded picture buffer parameter syntax structures minus 1.

12. A non-transitory computer-readable storage medium storing a bitstream of a video, the bitstream being generated by a method executed by a video processing device, wherein, The method includes: Generating a bitstream of the video, Among them, the bitstream includes one or more output layer sets, and each output layer set includes one or more video layers; Among them, the bitstream conforms to the format rule, [[ID=⑨]]Among them, the format rule stipulates that when the value of the third syntax element vps_max_layers_minus1 is equal to 0, the value of the first syntax element vps_all_independent_layers_flag, which indicates whether all the layers specified by the video parameter set for the bitstream are independently encoded and decoded without using inter-layer prediction, is equal to 1, the value of the second syntax element each_layer_is_an_ols_flag in the video parameter set, which indicates whether each output layer set includes a single video layer, is equal to 1, and the value of the variable VpsNumDpbParams, which indicates the number of decoded picture buffer parameter syntax structures included in the video parameter set, is equal to 0, where adding 1 to the third syntax element specifies the number of layers, and the number of layers is the maximum allowable number of layers in each encoded and decoded video sequence with reference to the video parameter set. Wherein, the format rule specifies that a fifth syntax element is conditionally included in the video parameter set, and the fifth syntax element indicates an index of a decoded picture buffer parameter syntax structure applied to an output layer set relative to a list of decoded picture buffer parameter syntax structures included in the video parameter set. Wherein, when the fifth syntax element does not exist and the number of decoded picture buffer parameter syntax structures is equal to 1, the value of the fifth syntax element is inferred to be equal to 0.

13. The non-transitory computer-readable storage medium according to claim 12, Among them, When the fifth syntax element exists, the value of the fifth syntax element is in the range from 0 to the number of decoded picture buffer parameter syntax structures minus 1.

14. A method for storing a bitstream of a video, comprising: Generating the bitstream of the video, Storing the bitstream in a non-transitory computer-readable storage medium, Wherein, the bitstream includes one or more output layer sets, and each output layer set includes one or more video layers; Wherein, the bitstream conforms to a format rule, Wherein, the format rule specifies that in the case where the value of a third syntax element vps_max_layers_minus1 is equal to 0, the value of a first syntax element vps_all_independent_layers_flag indicating whether all layers specified by a video parameter set for the bitstream are independently coded and decoded without using inter-layer prediction is equal to 1, the value of a second syntax element each_layer_is_an_ols_flag indicating whether each output layer set includes a single video layer in the video parameter set is equal to 1, and the value of a variable VpsNumDpbParams indicating the number of decoded picture buffer parameter syntax structures included in the video parameter set is equal to 0, wherein the third syntax element plus 1 specifies the number of layers, and the number of layers is the maximum allowed number of layers in each coded and decoded video sequence with reference to the video parameter set. Wherein, the format rule specifies that a fifth syntax element is conditionally included in the video parameter set, and the fifth syntax element indicates an index of a decoded picture buffer parameter syntax structure applied to an output layer set relative to a list of decoded picture buffer parameter syntax structures included in the video parameter set. Wherein, when the fifth syntax element does not exist and the number of decoded picture buffer parameter syntax structures is equal to 1, the value of the fifth syntax element is inferred to be equal to 0.