Profile, Tier, and Level Indication in Video Coding and Decoding

The refinement of VVC standard signaling rules for slice types, layer dependencies, and HRD parameters addresses inefficiencies in multi-layer video coding, enhancing decoding consistency and reducing unnecessary signaling.

CN114902674BActive Publication Date: 2025-07-15DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080090659.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-30
Filing Date
2020-12-24
Publication Date
2025-07-15
Estimated Expiration
2040-12-24

AI Technical Summary

Technical Problem

There is insufficient flexibility in the existing VVC draft text on slice_type, resulting in unnecessary signaling notifications and information duplication in some cases and failure to effectively support independent layer and inter-layer prediction in multi-layer bitstreams.

Method used

By adjusting the syntax elements and signaling rules in the VVC draft text, more flexible slice_type definitions are allowed, missing syntax element values are inferred, signaling of DPB parameters and HRD parameters is optimized, and signaling of DPB parameters and HRD parameters is ensured that PTL information is reasonably signaled in VPS and SPS, and independent layer and inter-layer predictions are supported in multi-layer bitstreams.

Benefits of technology

It improves the efficiency and flexibility of the VVC encoding and decoding process, reduces unnecessary signaling notifications and information duplication, supports independent layer and inter-layer prediction in multi-layer bitstreams, and improves the negotiation ability of the encoding and decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902674B_ABST
    Figure CN114902674B_ABST
Patent Text Reader

Abstract

A video processing method includes performing a conversion between a video including multiple video layers and a bitstream of the video, where the bitstream includes multiple output layer sets (OLSs), each including one or more of the multiple video layers, and the bitstream conforms to format rules, where the format rules stipulate that for an OLS having a single layer, a profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application is the U.S. national phase entry of International Patent Application No. PCT / US2020 / 067015, filed on December 24, 2020, which claims the priority of U.S. Provisional Application No. 62 / 953,854, filed on December 26, 2019, and U.S. Provisional Application No. 62 / 954,907, filed on December 30, 2019. The entire disclosure of the above applications is incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image encoding and decoding as well as video encoding and decoding. Background Art

[0004] Digital video accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is expected that the bandwidth demand for digital video will continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by a video encoder and a video decoder to perform video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more scalable video layers and a bitstream of the video. The video includes one or more video pictures, and the video pictures include one or more slices. The bitstream conforms to format rules. The format rules specify that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the slice does not use inter - layer prediction, the value of the field indicating the slice type of the slice is set to indicate the type of an intra - slice.

[0007] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including multiple video layers and a bitstream of the video, where the bitstream includes multiple output layer sets (OLSs), each output layer set includes one or more of the multiple scalable video layers, and the bitstream conforms to format rules, where the format rules specify that for an OLS having a single layer, the profile - tier - level (PTL) syntax structure indicating the profile, layer, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0008] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein the format rules specify a relationship between the occurrence of a plurality of profile-tier-level (PTL) syntax structures in a video parameter set of the bitstream and a byte alignment syntax field in the video parameter set; wherein each PTL syntax structure indicates the profile, tier, and level of one or more of the plurality of OLSs.

[0009] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of scalable video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of scalable video layers, and the bitstream conforms to format rules, wherein the format rules specify that during encoding, in the case where the value of an index of a syntax structure that describes the profile, tier, and level of one or more of the plurality of OLSs is zero, a syntax element indicating the index is excluded from a video parameter set of the bitstream, or during decoding, in the case where the syntax element does not exist in the bitstream, the value is inferred to be zero.

[0010] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein the format rules specify that for layer i, where i is an integer, the bitstream includes a first set of syntax elements indicating a first variable that indicates whether layer i is included in at least one of the plurality of OLSs.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets, each output layer set includes one or more video layers; wherein the bitstream conforms to format rules, wherein the format rules specify that in the case where each output layer set includes a single video layer, the number of decoded picture buffer parameter syntax structures included in a video parameter set of the bitstream is equal to zero; or in the case where it is not true that each output layer set includes a single layer, the number is equal to one plus the value of a syntax element.

[0012] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes a coded video sequence (CVS), the coded video sequence including one or more coded video pictures of one or more video layers; and wherein the bitstream conforms to a format rule that specifies that one or more sequence parameter sets (SPSs) indicating conversion parameters referred to by one or more of the coded pictures of the CVS have the same reference video parameter set (VPS) identifier that indicates a reference VPS.

[0013] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets (OLSs), each output layer set including one or more video layers, wherein the bitstream conforms to a format rule; wherein the format rule specifies whether or how a first syntax element is included in a video parameter set (VPS) of the bitstream, the first syntax element indicating whether a first syntax structure describing parameters of a hypothetical reference decoder (HRD) is used for the conversion.

[0014] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets (OLSs), each output layer set including one or more video layers, wherein the bitstream conforms to a format rule; wherein the format rule specifies whether or how a first syntax structure describing parameters of a general hypothetical reference decoder (HRD) and a plurality of second syntax structures describing OLS-specific HRD parameters are included in a video parameter set (VPS) of the bitstream.

[0015] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the method described above.

[0016] In another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the method described above.

[0017] In yet another example aspect, a computer-readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in a form of processor-executable code.

[0018] In yet another example aspect, a method of writing a bitstream generated according to one of the methods described above to a computer-readable medium is disclosed.

[0019] In another exemplary aspect, a computer-readable medium storing a bitstream of a video generated according to the above method is disclosed.

[0020] These features and other features are described in this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a block diagram showing a video codec system according to some embodiments of the present disclosure.

[0022] Figure 2 is a block diagram of an exemplary hardware platform for video processing.

[0023] Figure 3 is a flowchart of an exemplary method for video processing.

[0024] Figure 4 is a block diagram showing an exemplary video codec system.

[0025] Figure 5 is a block diagram showing an encoder according to some embodiments of the present disclosure.

[0026] Figure 6 is a block diagram showing a decoder according to some embodiments of the present disclosure.

[0027] Figures 7A - 7I is a flowchart of examples of various video processing methods. DETAILED DESCRIPTION

[0028] The use of section headings in this document is for ease of understanding and does not limit the techniques and embodiments disclosed in each section to only that section. Additionally, the use of H.266 terminology in some of the descriptions is for ease of understanding and not to limit the scope of the disclosed techniques. Thus, the techniques described herein also apply to other video codec protocols and designs.

[0029] 1. Summary

[0030] This document is related to video codec technology. Specifically, it is about various improvements to scalable video coding, where a video bitstream can contain more than one layer. These ideas can be applied alone or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video coding, such as the Versatile Video Coding (VVC) being developed.

[0031] 2. Abbreviations

[0032] APS Adaptive Parameter Set

[0033] AU Access Unit

[0034] AUD Access Unit Delimiter

[0035] AVC Advanced Video Coding

[0036] CLVS Coding Layer Video Sequence

[0037] CPB Coding Picture Buffer

[0038] CRA Clean Random Access

[0039] CTU Coding Tree Unit

[0040] CVS Coding Video Sequence

[0041] DPB Decoding Picture Buffer

[0042] DPS Decoding Parameter Set

[0043] EOB End of Bitstream

[0044] EOS End of Sequence

[0045] GDR Gradual Decoding Refresh

[0046] HEVC High Efficiency Video Coding

[0047] HRD Hypothetical Reference Decoder

[0048] IDR Instantaneous Decoding Refresh

[0049] JEM Joint Exploration Model

[0050] MCTS Motion Constrained Tile Set

[0051] NAL Network Abstraction Layer

[0052] OLS Output Layer Set

[0053] PH Picture Header

[0054] PPS Picture Parameter Set

[0055] PTL Profile, Tier and Level

[0056] PU Picture Unit

[0057] RBSP Raw Byte Sequence Payload

[0058] SEI Supplemental Enhancement Information

[0059] SPS Sequence Parameter Set

[0060] SVC Scalable Video Coding

[0061] VCL Video Coding Layer

[0062] VPS Video Parameter Set

[0063] VTM VVC Test Model

[0064] VUI Video Usability Information

[0065] VVC Versatile Video Coding

[0066] 3. Preliminary Discussion

[0067] Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard as well as the H.265 / HEVC [1] standard. Since H.262, video coding standards have been based on a hybrid video coding structure, in which temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and applied to a reference software called the Joint Exploration Model (JEM) [2]. JVET meetings are held quarterly simultaneously, and the goal of the new coding standard is to reduce the bit rate by 50% compared to HEVC. At the JVET meeting in April 2018, the new video coding standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at that time. With continuous efforts dedicated to VVC standardization, each JVET meeting adopts new coding technologies for the VCC standard. Then the VVC working draft and the test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Feature Complete (FDIS) at the meeting in July 2020.

[0068] 3.1. Scalable Video Coding (SVC)

[0069] Scalable Video Coding (SVC) refers to such video coding and decoding: wherein, a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (EL) are used. In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial levels, temporal levels, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can act as the BL, while the top layer can act as the EL. Intermediate layers can act as ELs or RLs, or both simultaneously. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below it (e.g., the base layer or any intermediate enhancement layer) and simultaneously act as an RL for one or more enhancement layers above it. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and the information of one view can be used to code (e.g., encode or decode) the information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).

[0070] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the coding and decoding levels at which they can be used (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be used by one or more coded video sequences of different layers in the bitstream can be included in the Video Parameter Set (VPS), and the parameters used by one or more pictures in the coded video sequence can be included in the Sequence Parameter Set (SPS). Similarly, the parameters used by one or more slices in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single slice can be included in the slice header. Similarly, an indication of which parameter set a given layer uses at a given time can be provided at various coding and decoding levels.

[0071] 3.2. Parameter Sets

[0072] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, VPS, and DPS. All AVC, HEVC, and VVC support SPS and PPS. VPS was introduced since HEVC and is included in HEVC and VVC. APS and DPS are not included in AVC or HEVC but are included in the latest VVC draft text.

[0073] The SPS is designed to carry sequence-level header information, and the PPS is designed to carry picture-level header information that does not change frequently. Using the SPS and PPS, it is not necessary to repeat the information that does not change frequently for each sequence or picture, so the redundant signaling of this information can be avoided. In addition, the use of the SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving the error recovery ability.

[0074] The VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0075] The APS is introduced to carry picture-level information or slice-level information that requires a considerable number of bits to encode and decode, can be shared by multiple pictures, and can have a considerable number of different variants in a sequence.

[0076] The DPS is introduced to carry bitstream-level information that indicates the highest capabilities required to decode the entire bitstream.

[0077] 3.3. VPS Syntax and Semantics in VVC

[0078] VVC supports scalability, also known as scalable video coding, in which multiple layers can be encoded in a single coded video bitstream.

[0079] In the latest VVC text, scalability information is signaled in the VPS, and the syntax and semantics are as follows.

[0080] 7.3.2.2 Video Parameter Set Syntax

[0081]

[0082]

[0083]

[0084] 7.4.3.2 Video Parameter Set RBSP Semantics

[0085] The VPS RBSP shall be available for the decoding process before being referenced, including in at least one AU where the TemporalId is equal to 0 or provided externally.

[0086] All VPS NAL units with a specific value of vps_video_parameter_set_id in the CVS shall have the same content.

[0087] The vps_video_parameter_set_id provides an identifier for the VPS for reference by other syntax elements. The value of vps_video_parameter_set_id shall be greater than 0.

[0088] vps_max_layers_minus1 plus 1 specifies the maximum number of allowed layers in each CVS that refers to the VPS.

[0089] vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in each CVS that refers to the VPS. The value of vps_max_sublayers_minus1 shall be in the range from 0 to 6 (inclusive of 0 and 6).

[0090] vps_all_layers_same_num_sublayers_flag being equal to 1 specifies that for all layers in each CVS that refers to the VPS, the number of temporal sublayers is the same. vps_all_layers_same_num_sublayers_flag being equal to 0 specifies that the layers in each CVS that refers to the VPS may or may not have the same number of temporal sublayers. When not present, the value of vps_all_layers_same_num_sublayers_flag is inferred to be equal to 1.

[0091] vps_all_independent_layers_flag being equal to 1 specifies that all layers in the CVS are independently encoded and decoded without using inter-layer prediction. vps_all_independent_layers_flag being equal to 0 specifies that one or more layers in the CVS may use inter-layer prediction. When not present, the value of vps_all_independent_layers_flag is inferred to be equal to 1. When vps_all_independent_layers_flag is equal to 1, the value of vps_independent_layer_flag[i] is inferred to be equal to 1. When vps_all_independent_layers_flag is equal to 0, the value of vps_independent_layer_flag[0] is inferred to be equal to 1.

[0092] vps_layer_id[i] specifies the nuh_layer_id value of the i-th layer. For any two non-negative integer values of m and n, when m is less than n, the value of vps_layer_id[m] shall be less than the value of vps_layer_id[n].

[0093] When vps_independent_layer_flag[i] is equal to 1, it is specified that the layer with index i does not use inter-layer prediction. When vps_independent_layer_flag[i] is equal to 0, it is specified that the layer with index i can use inter-layer prediction and there exists a syntax element vps_direct_ref_layer_flag[i][j] (for j in the range from 0 to i - 1, inclusive of 0 and i - 1) in the VPS. When it does not exist, the value of vps_independent_layer_flag[i] is inferred to be equal to 1.

[0094] When vps_direct_ref_layer_flag[i][j] is equal to 0, it is specified that the layer with index j is not a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[i][j] is equal to 1, it is specified that the layer with index j is a direct reference layer of the layer with index i. When vps_direct_ref_layer_flag[i][j] does not exist (for i and j in the range from 0 to vps_max_layers_minus1, inclusive of 0 and vps_max_layers_minus1), it is inferred to be equal to 0. When vps_independent_layer_flag[i] is equal to 0, there shall exist at least one value of j in the range from 0 to i - 1, inclusive of 0 and i - 1, such that the value of vps_direct_ref_layer_flag[i][j] is equal to 1.

[0095] Derive the variables NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j] as follows:

[0096]

[0097] Derive the variable GeneralLayerIdx[i] as follows, which specifies the layer index of the layer where nuh_layer_id is equal to vps_layer_id[i]:

[0098] for (i = 0; i <= vps_max_layers_minus1;

[0099] i++)(38)

[0100] GeneralLayerIdx[vps_layer_id[i]] = i

[0101] each_layer_is_an_ols_flag being equal to 1 specifies that each output layer set contains only one layer, and each layer in the bitstream itself is an output layer set, where the single included layer is the only output layer. each_layer_is_an_ols_flag being equal to 0 specifies that the output layer set may contain multiple layers. If vps_max_layers_minus1 is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 1. Otherwise, when vps_all_independent_layers_flag is equal to 0, the value of each_layer_is_an_ols_flag is inferred to be equal to 0.

[0102] ols_mode_idc being equal to 0 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i (including 0 and i), and for each OLS, only the highest layer in the OLS is output.

[0103] ols_mode_idc being equal to 1 specifies that the total number of OLSs specified by the VPS is equal to vps_max_layers_minus1 + 1, the i-th OLS includes the layers with layer indices from 0 to i (including 0 and i), and for each OLS, all layers in the OLS are output.

[0104] ols_mode_idc being equal to 2 specifies that the total number of OLSs specified by the VPS is signaled explicitly, and for each OLS, the output layers are signaled explicitly, while the other layers are direct or indirect reference layers of the output layers of the OLS.

[0105] The value of ols_mode_idc shall be in the range of 0 to 2 (including 0 and 2). The value 3 of ols_mode_idc is reserved for future use by ITU-T|ISO / IEC.

[0106] When vps_all_independent_layers_flag is equal to 1 and each_layer_is_an_ols_flag is equal to 0, the value of ols_mode_idc is inferred to be equal to 2.

[0107] num_output_layer_sets_minus1 plus 1 specifies the total number of OLSs specified by the VPS when ols_mode_idc is equal to 2.

[0108] Derive the variable TotalNumOlss in the following way, which specifies the total number of OLSs specified by VPS:

[0109]

[0110] ols_output_layer_flag[i][j] being equal to 1 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. ols_output_layer_flag[i][j] being equal to 0 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.

[0111] Derive the variable NumOutputLayersInOls[i] that specifies the number of output layers in the i-th OLS, and the variable OutputLayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th output layer in the i-th OLS in the following way:

[0112]

[0113] For each OLS, there should be at least one layer as the output layer. In other words, for any i value in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.

[0114] Derive the variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS in the following way:

[0115]

[0116] Note 1 - The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and for the 0-th OLS, only the contained layer is output.

[0117] Derive the variable OlsLayerIdx[i][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j] in the following way:

[0118]

[0119] The lowest layer in each OLS shall be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss-1 (including 0 and TotalNumOlss-1), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]] shall be equal to 1.

[0120] Each layer shall be included in at least one OLS specified by the VPS. In other words, for each layer with a specific value nuhLayerId of nuh_layer_id (equal to one of vps_layer_id[k] for k in the range from 0 to vps_max_layers_minus1 (including 0 and vps_max_layers_minus1)), there shall exist at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss-1 (including 0 and TotalNumOlss-1), and j is in the range up to and including NumLayersInOls[i]-1, such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.

[0121] vps_num_ptls specifies the number of profile_tier_level() syntax structures in the VPS.

[0122] pt_present_flag[i] being equal to 1 specifies the presence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] being equal to 0 specifies the absence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 0. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (i-1)-th profile_tier_level() syntax structure in the VPS.

[0123] ptl_max_temporal_id[i] specifies the TemporalId of the highest sublayer with level information in the profile_tier_level() syntax structure of the i-th profile_tier_level() in the VPS. The value of ptl_max_temporal_id[i] shall be in the range from 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0124] vps_ptl_byte_alignment_zero_bit shall be equal to 0.

[0125] ols_ptl_idx[i] specifies the index of the profile_tier_level() syntax structure applied to the i-th OLS in the list of profile_tier_level() syntax structures in the VPS. When present, the value of ols_ptl_idx[i] shall be in the range from 0 to vps_num_ptls - 1, inclusive.

[0126] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS is present in the SPS referred to by the layer in the i-th OLS.

[0127] vps_num_dpb_params specifies the number of dpb_parameters() syntax structures in the VPS. The value of vps_num_dpb_params shall be in the range from 0 to 16, inclusive. When not present, the value of vps_num_dpb_params is inferred to be equal to 0.

[0128] When same_dpb_size_output_or_nonoutput_flag equals 1, it specifies that the syntax element layer_nonoutput_dpb_params_idx[i] does not exist in the VPS. When same_dpb_size_output_or_nonoutput_flag equals 0, it specifies that the syntax element layer_nonoutput_dpb_params_idx[i] may or may not exist in the VPS.

[0129] The vps_sublayer_dpb_params_present_flag is used to control the presence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure in the VPS. When they are absent, the vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.

[0130] When dpb_size_only_flag[i] equals 1, it specifies that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] do not exist in the i-th dpb_parameters() syntax structure in the VPS. When dpb_size_only_flag[i] equals 1, it specifies that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] may exist in the i-th dpb_parameters() syntax structure in the VPS.

[0131] dpb_max_temporal_id[i] specifies the TemporalId represented by the highest sublayer in the i-th dpb_parameters() syntax structure where DPB parameters may be present in the VPS. The value of dpb_max_temporal_id[i] shall be in the range from 0 to vps_max_sublayers_minus1, inclusive (including 0 and vps_max_sublayers_minus1). When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0132] layer_output_dpb_params_idx[i] specifies the index of the list of dpb_parameters() syntax structures in the VPS that applies to the i-th layer when it is the output layer in the OLS. When present, the value of layer_output_dpb_params_idx[i] shall be in the range from 0 to vps_num_dpb_params - 1, inclusive (including 0 and vps_num_dpb_params - 1).

[0133] If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is the output layer, the dpb_parameters() syntax structure applied to the i-th layer is the dpb_parameter() syntax structure present in the SPS that the layer references.

[0134] Otherwise (vps_independent_layer_flag[i] is equal to 0), the following applies:

[0135] - When vps_num_dpb_params is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0136] - A requirement for bitstream consistency is that the value of layer_output_dpb_params_idx[i] shall be such that dpb_size_only_flag[layer_output_dpb_params_idx[i]] is equal to 0.

[0137] layer_nonoutput_dpb_params_idx[i] specifies the index of the dpb_parameters() syntax structure list in the VPS that is applied to the i-th layer as a non-output layer in the OLS. When present, the value of layer_nonoutput_dpb_params_idx[i] shall be in the range from 0 to vps_num_dpb_params - 1, inclusive of 0 and vps_num_dpb_params - 1.

[0138] If same_dpb_size_output_or_nonoutput_flag is equal to 1, the following applies:

[0139] - If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is a non-output layer, the dpb_parameters() syntax structure applied to that layer is the dpb_parameters() syntax structure that exists in the SPS referred to by that layer.

[0140] - Otherwise (vps_independent_layer_flag[i] is equal to 0), the value of layer_nonoutput_dpb_params_idx[i] is inferred to be equal to layer_output_dpb_params_idx[i].

[0141] Otherwise (same_dpb_size_output_or_nonoutput_flag is equal to 0), when vps_num_dpb_params is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0142] vps_general_hrd_params_present_flag being equal to 1 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are present in the VPS RBSP syntax structure. vps_general_hrd_params_present_flag being equal to 0 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are not present in the VPS RBSP syntax structure.

[0143] The vps_sublayer_cpb_params_present_flag being equal to 1 specifies that the ith ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId is in the range from 0 to hrd_max_tid[i] (inclusive of 0 and hrd_max_tid[i]). The vps_sublayer_cpb_params_present_flag being equal to 0 specifies that the ith ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented at the sublayer, where the TemporalId is equal to hrd_max_tid[i] only. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.

[0144] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented at the sublayer (where the TemporalId is in the range from 0 to hrd_max_tid[i] - 1 (inclusive of 0 and hrd_max_tid[i] - 1)) are inferred to be the same as the HRD parameters represented at the sublayer (where the TemporalId is equal to hrd_max_tid[i]). These include HRD parameters starting from the fixed_pic_rate_general_flag[i] syntax element up to the sublayer_hrd_parameters(i) syntax structure under the "if(general_vcl_hrd_params_present_flag)" condition in the ols_hrd_parameters syntax structure.

[0145] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure. The value of num_ols_hrd_params_minus1 shall be in the range from 0 to 63 (inclusive of 0 and 63). When TotalNumOlss is greater than 1, the value of num_ols_hrd_params_minus1 is inferred to be equal to 0.

[0146] hrd_max_tid[i] specifies the TemporalId of the highest sublayer representation for which the HRD parameters are included in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0.

[0147] ols_hrd_idx[i] specifies the index of the ols_hrd_parameters() syntax structure applied to the i-th OLS. The value of ols_hrd_idx[[i] shall be in the range of 0 to num_ols_hrd_params_minus1, inclusive. When not present, the value of ols_hrd_idx[[i] is inferred to be equal to 0.

[0148] vps_extension_flag being equal to 0 specifies that there is no vps_extension_data_flag syntax element in the VPS RBSP syntax structure. vps_extension_flag being equal to 1 specifies that there is a vps_extension_data_flag syntax element in the VPS RBSP syntax structure.

[0149] vps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of this specification. Decoders compliant with this version of this specification shall ignore all vps_extension_data_flag syntax elements.

[0150] 3.4. SPS Syntax and Semantics in VVC

[0151] In the latest VVC draft text in JVET-P2001-v14, the SPS syntax and semantics most relevant to the present invention are as follows.

[0152] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0153]

[0154]

[0155] 7.4.3.3 Sequence Parameter Set RBSP Semantics

[0156] The SPS RBSP shall be available for the decoding process before being referenced, included in at least one AU where TemporalId equals 0, or provided by an external means.

[0157] All SPS NAL units with a specific value of sps_seq_parameter_set_id in the CVS shall have the same content.

[0158] When sps_decoding_parameter_set_id is greater than 0, it specifies the value of dps_decoding_parameter_set_id of the DPS referenced by the SPS. When sps_decoding_parameter_set_id equals 0, the SPS does not reference the DPS, and the DPS is not referenced when decoding each CLVS that references the SPS. The value of sps_decoding_parameter_set_id shall be the same in all SPSs referenced by coded pictures in the bitstream.

[0159] When sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS.

[0160] When sps_video_parameter_set_id equals 0, the following applies:

[0161] - The SPS does not reference the VPS.

[0162] - When decoding each CLVS that references the SPS, the VPS is not referenced.

[0163] - The value of vps_max_layers_minus1 is inferred to be equal to 0.

[0164] - The CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same value of nuh_layer_id).

[0165] - The value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.

[0166] - The value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.

[0167] When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referred to by the CLVS with a specific nuh_layer_id value nuhLayerId shall have a nuh_layer_id equal to nuhLayerId.

[0168] sps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may be present in the CLVS of each reference SPS. The value of sps_max_sublayers_minus1 shall be in the range from 0 to vps_max_sublayers_minus1, inclusive (including 0 and vps_max_sublayers_minus1).

[0169] sps_reserved_zero_4bits shall be equal to 0 in the bitstream of this version that conforms to this specification. Other values of sps_reserved_zero_4bits are reserved for future use by ITU-T|ISO / IEC.

[0170] sps_ptl_dpb_hrd_params_present_flag being equal to 1 specifies that the profile_tier_level() syntax structure and the dpb_parameters() syntax structure are present in the SPS, and the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure may also be present in the SPS. sps_ptl_dpb_hrd_params_present_flag being equal to 0 specifies that these syntax structures are not present in the SPS. The value of sps_ptl_dpb_hrd_params_present_flag shall be equal to vps_independent_layer_flag[nuh_layer_id].

[0171] If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the variable MaxDecPicBuffMinus1 is set to max_dec_pic_buffering_minus1[sps_max_sublayers_minus1] in the dpb_parameters() syntax structure in the SPS. Otherwise, MaxDecPicBuffMinus1 is set to be equal to max_dec_pic_buffering_minus1[sps_max_sublayers_minus1] in the dpb_parameters() syntax structure at the layer_nonoutput_dpb_params_idx[GeneralLayerIdx[nuh_layer_id]]-th in the VPS.

[0172] 3.5. Syntax and Semantics of Slice Headers in VVC

[0173] In the latest VVC draft text in JVET-P2001-v14, the slice header syntax and semantics most relevant to the present invention are as follows.

[0174] 7.3.7.1 General Slice Header Syntax

[0175]

[0176] 7.4.8.1 General Slice Header Semantics

[0177] The slice_type specifies the coding type of the slice according to Table 9.

[0178] Table 9 - Names Associated with slice_type

[0179] slice_type Name of slice_type 0 B (B band) 1 P (P band) 2 I (I band)

[0180] When the nal_unit_type is a value of nal_unit_type within the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and the current picture is the first picture in the access unit, the slice_type shall be equal to 2.

[0181] 4. Technical Problems Solved by the Described Technical Solutions

[0182] The existing scalable designs in VVC have the following problems:

[0183] 1) The latest VVC draft text includes the following constraints on slice_type:

[0184] When the nal_unit_type is a nal_unit_type value within the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and the current picture is the first picture in an AU, the slice_type shall be equal to 2.

[0185] For a slice, a slice_type value equal to 2 means that the slice is intra-coded without using inter prediction from reference pictures.

[0186] However, in an AU, not only the first picture in the AU that is an IRAP picture needs to contain only intra-coded slices, but also all IRAP pictures in all independent layers need to contain only intra-coded slices. Therefore, the above constraints do need to be updated.

[0187] 2) When the syntax element ols_ptl_idx[i] does not exist, the value still needs to be used. However, when ols_ptl_idx[i] does not exist, there is a lack of inference of its value.

[0188] 3) When the i-th layer is not used as an output layer in any OLS, the signaling of the syntax element layer_output_dpb_params_idx[i] is unnecessary.

[0189] 4) The value of the syntax element vps_num_dpb_params can be equal to 0. However, when vps_all_independent_layers_flag is equal to 0, there needs to be at least one dpb_parameters() syntax structure in the VPS.

[0190] 5) In the latest VVC draft text, the PTL information of OLS only contains one layer, which is an independent coding layer that does not refer to any other layers and is only signaled in the SPS. However, for the purpose of session negotiation, it is desirable to signal the PTL information for all OLSs in the bitstream in the VPS.

[0191] 6) When the number of PTL syntax structures signaled in the VPS is zero, the signaling of the vps_ptl_byte_alignment_zero_bit syntax element is unnecessary.

[0192] 7) In the semantics of sps_video_parameter_set_id, there are the following constraints:

[0193] When sps_video_parameter_set_id is equal to 0, the CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same value of nuh_layer_id).

[0194] However, this constraint does not allow including a stand-alone layer that does not refer to the VPS in a multi-layer bitstream. Since sps_video_parameter_set_id being equal to 0 means that the SPS (and layer) does not refer to the VPS.

[0195] 8) When each_layer_is_an_ols_flag is equal to 1, the value of vps_general_hrd_params_present_flag may be equal to 1. However, when each_layer_is_an_ols_flag is equal to 1, the HRD parameters are signaled only in the SPS, so the value of vps_general_hrd_params_present_flag shall not be equal to 1. In some cases, signaling the general_hrd_parameter() syntax structure may not make sense when signaling the zero old_hrd_parameters() syntax structure in the VPS.

[0196] 9) Signal the HRD parameters of an OLS that contains only one layer in both the VPS and the SPS. However, for an OLS that contains only one layer, repeating the HRD parameters in the VPS is useless.

[0197] 5. Example Embodiments and Techniques

[0198] To solve the above problems and other problems, the following summarized methods are disclosed. These inventions should be regarded as examples for explaining general concepts and should not be narrowly construed. In addition, these inventions can be applied alone or in any combination.

[0199] 1) To solve the first problem, the following constraint is specified:

[0200] When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type shall be equal to 2.

[0201] 2) To solve the second problem, for each possible value of i, when the syntax element does not exist, the value of ols_ptl_idx[i] is inferred to be equal to 0.

[0202] 3) To solve the third problem, it is stipulated that the variable LayerUsedAsOutputLayerFlag[i] is used to indicate whether the i-th layer is used as an output layer in any OLS, and when this variable is equal to 0, signaling of the variable layer_output_dpb_params_idx[i] is avoided.

[0203] a. In addition, the following constraint can be further stipulated: for each value of i in the range from 0 to vps_max_layers_minus1 (including 0 and vps_max_layers_minus1), the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] should not both be equal to 0. In other words, there should not be a layer that is neither a direct reference layer of any other layer nor an output layer of at least one OLS.

[0204] 4) To solve the fourth problem, vps_num_dpb_params is changed to vps_num_dpb_params_minus1, and VpsNumDpbParams is stipulated as follows:

[0205] If (!vps_all_independent_layers_flag)

[0206] VpsNumDpbParams = vps_num_dpb_params_minus1 + 1

[0207] Otherwise

[0208] VpsNumDpbParams = 0

[0209] And vps_num_dpb_params used in the syntax condition and semantics is replaced with VpsNumDpbParams.

[0210] a. In addition, the following constraint is further stipulated: when it does not exist, the value of same_dpb_size_output_or_nonoutput_flag is inferred to be equal to 1.

[0211] 5) To solve the fifth problem, for the purpose of session negotiation, it is allowed to repeat the PTL information of an OLS that contains only one layer in the VPS. This can be achieved by changing vps_num_ptls to vps_num_ptls_minus1.

[0212] a. Alternatively, keep vps_num_ptls (without making it vps_num_ptls_minus1), but remove “NumLayersInOls[i]>1&&” from the syntax condition of ols_ptl_idx[i].

[0213] a. Alternatively, keep vps_num_ptls (without making it vps_num_ptls_minus1), but set the vps_num_ptls condition to “if (!each_layer_is_an_ols_flag)” or “if (vps_max_layers_minus1>0 &&!vps_all_independent_layers_flag)”.

[0214] b. Alternatively, also allow the repetition of DPB parameter information (only DPB size or all DPB parameters) for an OLS that contains only one layer in the VPS.

[0215] 6) To solve the sixth problem, the vps_ptl_byte_alignment_zero_bit syntax element shall not be signaled as long as the number of PTL syntax structures signaled in the VPS is zero. This can be achieved by setting the vps_ptl_byte_alignment_zero_bit condition to “if (vps_num_ptls>0)”, or by changing vps_num_ptls to vps_num_ptls_minus1, which effectively does not allow the number of PTL syntax structures signaled in the VPS to be equal to zero.

[0216] 7) To solve the seventh problem, remove the following constraint:

[0217] When sps_video_parameter_set_id is equal to 0, the CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same nuh_layer_id value).

[0218] And add the following constraint:

[0219] The value of sps_video_parameter_set_id shall be the same in all SPSs referenced by coded pictures in the CVS, and sps_video_parameter_set_id shall be greater than 0.

[0220] a. Alternatively, keep the following constraint:

[0221] When sps_video_parameter_set_id is equal to 0, the CVS shall contain only one layer (i.e., all VCL NAL units in the CVS shall have the same nuh_layer_id value).

[0222] And the following constraints are specified:

[0223] The value of sps_video_parameter_set_id shall be the same in all SPSs referenced by coded pictures in the CVS.

[0224] 8) To solve the eighth problem, when vps_general_hrd_params_present_flag is equal to 1, the syntax element vps_general_hrd_params_present_flag is not signaled, and when it is absent, the value of vps_general_hrd_params_present_flag is inferred to be equal to 0.

[0225] a. Alternatively, when each_layer_is_an_ols_flag is equal to 1, the value of vps_general_hrd_params_present_flag is restricted to be equal to 0.

[0226] b. In addition, since the syntax condition for the syntax element num_ols_hrd_params_minus1, i.e., "if (TotalNumOlss > 1)", is not required, it will be removed. This is because when TotalNumOlss is equal to 1, the value of each_layer_is_an_ols_flag will be equal to 1, and then the value of vps_general_hrd_params_present_flag will be equal to 0, and then the HRD parameters will not be signaled in the VPS.

[0227] In some cases, when the ols_hrd_parameters() syntax structure is not present in the VPS, the general_hrd_parameters() syntax structure is not signaled in the VPS.

[0228] 9) To solve the ninth problem, the HRD parameters for an OLS with only one layer are signaled only in the SPS and not in the VPS.

[0229] 6. Embodiment

[0230] The following are some example embodiments of the aspects summarized in Section 5 above, which can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-P2001-v14. Most of the relevant parts added or modified are highlighted with Underline Bold highlighted, and some of the deleted parts are highlighted in italic bold. There are also some other changes that are editorial in nature and thus not highlighted.

[0231] 6.1. First Embodiment

[0232] 6.1.1. VPS Syntax and Semantics

[0233] 7.3.2.2 Video Parameter Set Syntax

[0234]

[0235]

[0236]

[0237] 7.4.3.2 Video Parameter Set RBSP Semantics ...

[0239] ols_output_layer_flag[i][j] being equal to 1 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. ols_output_layer_flag[i][j] being equal to 0 specifies that when ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.

[0240] The variables NumOutputLayersInOls[i] that specify the number of output layers in the i-th OLS and the variables OutputLayerIdInOls[i][j] that specify the nuh_layer_id value of the j-th output layer in the i-th OLS are derived as follows:

[0241]

[0242]

[0243]

[0244] For each OLS, there should be at least one layer that serves as the output layer. In other words, for any value of i in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of NumOutputLayersInOls[i] should be greater than or equal to 1.

[0245] Derive the variable NumLayersInOls[i] that specifies the number of layers in the i-th OLS, and the variable LayerIdInOls[i][j] that specifies the nuh_layer_id value of the j-th layer in the i-th OLS in the following manner:

[0246]

[0247]

[0248] Note - The 0-th OLS only contains the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and only the included layer is output for the 0-th OLS.

[0249] Derive the variable OlsLayeIdx[i][j] that specifies the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j] in the following manner:

[0250]

[0251] The lowest layer in each OLS should be an independent layer. In other words, for each i in the range from 0 to TotalNumOlss - 1 (including 0 and TotalNumOlss - 1), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]] should be equal to 1.

[0252] Each layer shall be included in at least one OLS specified by the VPS. In other words, for each layer where the value of nuh_layer_id, nuhLayerId, is equal to one of vps_layer_id[k] (for k in the range from 0 to vps_max_layers_minus1, inclusive of 0 and vps_max_layers_minus1), there shall exist at least one pair of values of i and j, where i is in the range from 0 to TotalNumOlss - 1 (inclusive of 0 and TotalNumOlss - 1), and j is in the range from 0 to NumLayersInOls[i] - 1 (inclusive of NumLayersInOls[i] - 1), such that the value of LayerIdInOls[i][j] is equal to nuhLayerId.

[0253] Specify the number of profile_tier_level() syntax structures in the VPS.

[0254] pt_present_flag[i] being equal to 1 specifies the existence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. pt_present_flag[i] being equal to 0 specifies the non-existence of tier, layer, and general constraint information in the i-th profile_tier_level() syntax structure in the VPS. The value of pt_present_flag[0] is inferred to be equal to 1. When pt_present_flag[i] is equal to 0, the tier, layer, and general constraint information of the i-th profile_tier_level() syntax structure in the VPS is inferred to be the same as that of the (i - 1)-th profile_tier_level() syntax structure in the VPS.

[0255] ptl_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation for which there is level information in the i-th profile_tier_level() syntax structure in the VPS. The value of ptl_max_temporal_id[i] shall be in the range from 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of ptl_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of ptl_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0256] vps_ptl_byte_alignment_zero_bit shall be equal to 0.

[0257] ols_ptl_idx[i] specifies the index of the profile_tier_level() syntax structure of the i-th OLS in the list of profile_tier_level() syntax structures in the VPS. When present, the value of ols_ptl_idx[i] shall be in the range from 0 to inclusive.

[0258] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS is present in the SPS referenced by the layer in the i-th OLS.

[0259] plus 1 specifies the number of dpb_parameters() syntax structures in the VPS. When present, the value shall be in the range from 0 to 15, inclusive.

[0260]

[0261] When same_dpb_size_output_or_nonoutput_flag equals 1, it specifies that the syntax element layer_nonoutput_dpb_params_idx[i] does not exist in the VPS. When same_dpb_size_output_or_nonoutput_flag equals 0, it specifies that the syntax element layer_nonoutput_dpb_params_idx[i] may or may not exist in the VPS.

[0262] The vps_sublayer_dpb_params_present_flag is used to control the presence of the syntax elements max_dec_pic_buffering_minus1[], max_num_reorder_pics[], and max_latency_increase_plus1[] in the dpb_parameters() syntax structure in the VPS. When they are absent, the vps_sub_dpb_params_info_present_flag is inferred to be equal to 0.

[0263] When dpb_size_only_flag[i] equals 1, it specifies that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] do not exist in the i-th dpb_parameters() syntax structure in the VPS. When dpb_size_only_flag[i] equals 0, it specifies that the syntax elements max_num_reorder_pics[] and max_latency_increase_plus1[] may exist in the i-th dpb_parameters() syntax structure in the VPS.

[0264] dpb_max_temporal_id[i] specifies the TemporalId of the highest sublayer representation for which DPB parameters may be present in the i-th dpb_parameters() syntax structure in the VPS. The value of dpb_max_temporal_id[i] shall be in the range of 0 to vps_max_sublayers_minus1 (including 0 and vps_max_sublayers_minus1). When vps_max_sublayers_minus1 is equal to 0, the value of dpb_max_temporal_id[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of dpb_max_temporal_id[i] is inferred to be equal to vps_max_sublayers_minus1.

[0265] layer_output_dpb_params_idx[i] specifies the index in the list of dpb_parameters() syntax structures in the VPS of the dpb_parameters() syntax structure applied to the i-th layer as the output layer in the OLS. When present, the value of layer_output_dpb_params_idx[i] shall be in the range of 0 to (including 0 and VpsNumDpbParams - 1).

[0266] If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is the output layer, the dpb_parameters() syntax structure applied to the i-th layer is the dpb_parameters() syntax structure present in the SPS referred to by that layer.

[0267] Otherwise (vps_independent_layer_flag[i] is equal to 0), the following applies:

[0268] - When is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0269] - A requirement for bitstream consistency is that the value of layer_output_dpb_params_idx[i] shall be such that dpb_size_only_flag[layer_output_dpb_params_idx[i]] is equal to 0.

[0270] layer_nonoutput_dpb_params_idx[i] specifies the index in the list of dpb_parameters() syntax structures in the VPS of the dpb_parameters() syntax structure applied to the i-th layer as a non-output layer in the OLS. When present, the value of layer_nonoutput_dpb_params_idx[i] shall be in the range from 0 to VpsNumDpbParams - 1, inclusive of 0 and VpsNumDpbParams - 1.

[0271] If same_dpb_size_output_or_nonoutput_flag is equal to 1, the following applies:

[0272] - If vps_independent_layer_flag[i] is equal to 1, then when the i-th layer is a non-output layer, the dpb_parameters() syntax structure applied to that layer is the dpb_parameters() syntax structure present in the SPS that the layer refers to.

[0273] - Otherwise (vps_independent_layer_flag[i] is equal to 0), the value of layer_nonoutput_dpb_params_idx[i] is inferred to be equal to layer_output_dpb_params_idx[i].

[0274] Otherwise (same_dpb_size_output_or_nonoutput_flag is equal to 0), when VpsNumDpbParams is equal to 1, the value of layer_output_dpb_params_idx[i] is inferred to be equal to 0.

[0275] vps_general_hrd_params_present_flag being equal to 1 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are present in the VPS RBSP syntax structure. vps_general_hrd_params_present_flag being equal to 0 specifies that the syntax structure general_hrd_parameters() and other HRD parameters are not present in the VPS RBSP syntax structure.

[0276] The vps_sublayer_cpb_params_present_flag being equal to 1 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented by sublayers, where the TemporalId is in the range from 0 to hrd_max_tid[i] (including 0 and hrd_max_tid[i]). The vps_sublayer_cpb_params_present_flag being equal to 0 specifies that the i-th ols_hrd_parameters() syntax structure in the VPS contains HRD parameters represented by sublayers, where the TemporalId is equal to hrd_max_tid[i] only. When vps_max_sublayers_minus1 is equal to 0, the value of vps_sublayer_cpb_params_present_flag is inferred to be equal to 0.

[0277] When vps_sublayer_cpb_params_present_flag is equal to 0, the HRD parameters represented by sublayers in the range of TemporalId from 0 to hrd_max_tid[i] - 1 (including 0 and hrd_max_tid[i] - 1) are inferred to be the same as the HRD parameters represented by the sublayer with TemporalId equal to hrd_max_tid[i]. These include such HRD parameters: starting from the fixed_pic_rate_general_flag[i] syntax element until the sublayer_hrd_parameters(i) syntax structure immediately under the "if(general_vcl_hrd_params_present_flag)" condition in the ols_hrd_parameters syntax structure.

[0278] num_ols_hrd_params_minus1 plus 1 specifies the number of ols_hrd_parameters() syntax structures present in the general_hrd_parameters() syntax structure. The value of num_ols_hrd_params_minus1 shall be in the range from 0 to 63 (including 0 and 63).

[0279] hrd_max_tid[i] specifies the TemporalId of the highest sublayer representation for which the HRD parameters are contained in the i-th ols_hrd_parameters() syntax structure. The value of hrd_max_tid[i] shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. When vps_max_sublayers_minus1 is equal to 0, the value of hrd_max_tid[i] is inferred to be equal to 0. When vps_max_sublayers_minus1 is greater than 0 and vps_all_layers_same_num_sublayers_flag is equal to 1, the value of hrd_max_tid[i] is inferred to be equal to vps_max_sublayers_minus1.

[0280] ols_hrd_idx[i] specifies The syntax of ols_hrd_parameters() for the ith OLS When present, the value of ols_hrd_idx[[i]] shall be in the range of 0 to num_ols_hrd_params_minus1, inclusive. , the value of ols_hrd_idx[[i]] is inferred to be equal to 0.

[0281] vps_extension_flag equal to 0 specifies that the vps_extension_data_flag syntax element is not present in the VPS RBSP syntax structure. vps_extension_flag equal to 1 specifies that the vps_extension_data_flag syntax element is present in the VPS RBSP syntax structure.

[0282] The vps_extension_data_flag may have any value. Its presence and value do not affect the conformance of a decoder to the profiles specified in this version of this specification. Decoders conforming to this version of this specification shall ignore all vps_extension_data_flag syntax elements.

[0283] 6.1.2.SPS Semantics

[0284] 7.4.3.3 Sequence parameter set RBSP semantics

[0285] The SPS RBSP shall be available for the decoding process before it is referenced, included in at least one AU (where TemporalId equals 0), or provided by an external means.

[0286] All SPS NAL units with a specific value of sps_seq_parameter_set_id in the CVS shall have the same content.

[0287] When sps_decoding_parameter_set_id is greater than 0, it specifies the value of dps_decoding_parameter_set_id of the DPS referenced by the SPS. When sps_decoding_parameter_set_id equals 0, the SPS does not reference the DPS, and no DPS is referenced when decoding each CLVS that references the SPS. The value of sps_decoding_parameter_set_id shall be the same for all SPSs decoded pictures referenced in the bitstream.

[0288] When sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS.

[0289] When sps_video_parameter_set_id equals 0, the following applies:

[0290] - The SPS does not reference the VPS.

[0291] - When decoding each CLVS that references the SPS, the VPS is not referenced.

[0292] - The value of vps_max_layers_minus1 is inferred to be equal to 0.

[0293] -

[0294] - The value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.

[0295] - The value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.

[0296] When vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referred to by the CLVS with a specific nuh_layer_id value nuhLayerId shall have a nuh_layer_id equal to nuhLayerId.

[0297] 6.1.3. Syntax of slice header

[0298] 7.4.8.1 Syntax of general slice header ...

[0300] The slice_type specifies the coding type of the slice according to Table 9.

[0301] Table 9 - Names associated with slice_type

[0302] slice_type Name of slice_type 0 B (B band) 1 P (P band) 2 I (I band)

[0303] When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), the slice_type shall be equal to 2.

[0304] Figure 1 FIG. shows a block diagram of an exemplary video processing system 1900, in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, passive optical network (PON), etc.), and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0305] System 1900 may include an encoding / decoding component 1904, which may implement various encoding or coding methods described in this document. The encoding / decoding component 1904 may reduce the average bit rate of a video from the input 1902 to the output of the encoding / decoding component 1904 to produce an encoded / decoded representation of the video. Thus, encoding / decoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding / decoding component 1904 may be stored or transmitted via communication over a connection represented by component 1906. The stored or transmitted bitstream (or encoded / decoded) representation of the video received at the input 1902 may be used by component 1908 to generate pixel values or a displayable video that is sent to the display interface 1910. The process of generating a user-viewable video from the bitstream is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "encoding / decoding" operations or tools, it will be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding result will be performed by the decoder.

[0306] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include Serial Advanced Technology Attachment (SATA), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0307] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The (multiple) processors 3602 may be configured to implement one or more methods described in this document. One or more memories 3604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the hardware 3606 may be partially or wholly located within the (multiple) processors 3602 (e.g., a graphics processor).

[0308] Figure 4 is a block diagram showing an example video encoding / decoding system 100 that may utilize the techniques of the present disclosure.

[0309] As Figure 4As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data that may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, and the source device 110 may be referred to as a video decoding device.

[0310] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0311] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to the destination device 120 via the I / O interface 116 over a network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0312] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0313] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be located external to the destination device 120, and the destination device 120 is configured to interface with an external display device.

[0314] The video encoder 114 and the video decoder 124 may operate according to video compression standards (such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards).

[0315] Figure 5 is a block diagram showing an example of a video encoder 200, and the video encoder 200 may be Figure 4 the video encoder 114 in the system 100 shown.

[0316] The video encoder 200 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0317] The functional components of the video encoder 200 may include a splitting unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0318] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0319] In addition, some components (e.g., the motion estimation unit 204 and the motion compensation unit 205) may be highly integrated, but are shown separately in Figure 5 the example for purposes of explanation.

[0320] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0321] The mode selection unit 203 may select one of the coding / decoding modes (intra or inter) based on, for example, error results, and provide the resulting intra or inter coded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block to be used as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra prediction and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel accuracy or integer-pixel accuracy).

[0322] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures other than the picture associated with the current video block from buffer 213.

[0323] For example, the motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0324] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0325] In other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search for a reference video block of the current video block in the reference pictures of list 0, and may also search for another reference video block of the current video block in the reference pictures of list 1. Then, the motion estimation unit 204 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector that indicates the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0326] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.

[0327] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to the motion information of another video block to

[0328] Notify the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0329] In one example, the motion estimation unit 204 may indicate in a syntax structure associated with the current video block a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0330] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0331] As discussed above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0332] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0333] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., denoted by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0334] In other examples, for the current video block, there may be no residual data for the current video block. For example, in the skip mode, the residual generation unit 207 may not perform a subtraction operation.

[0335] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0336] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0337] The inverse quantization unit 210 and the inverse transform unit 211 may respectively apply inverse quantization and inverse transform to the transformed coefficient video block to reconstruct the residual video block based on the transformed coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0338] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0339] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0340] Figure 6 is a block diagram illustrating an example of a video decoder 300, and the video decoder 300 may be Figure 4 the video decoder 114 in the system 100 shown.

[0341] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 5 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0342] In Figure 6 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform a decoding pass that is generally opposite to the encoding pass ( Figure 5 ) described with respect to the video encoder 200.

[0343] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy encoded video data, and the motion compensation unit 302 may determine motion information based on the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. For example, the motion compensation unit 302 may determine such information by performing AMVP and merge modes.

[0344] The motion compensation unit 302 may generate motion-compensated blocks and may perform interpolation based on an interpolation filter. The syntax elements may include an identifier for the interpolation filter to be used with sub-pixel precision.

[0345] The motion compensation unit 302 may use the interpolation filter used by the video encoder 20 during video block encoding to calculate the interpolation of sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information and use the interpolation filter to generate a prediction block.

[0346] The motion compensation unit 302 may use some syntax information to determine the size of the blocks for encoding frames and / or slices of an encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0347] The intra prediction unit 303 may form a prediction block based on spatially neighboring blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0348] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, and the buffer 307 provides a reference block for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.

[0349] A list of preferred solutions for some embodiments is provided below.

[0350] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0351] 1. A video processing method (e.g., Figure 3 method 600 therein), comprising: performing a conversion between a video including one or more scalable video layers and an encoded / decoded representation of the video, wherein the encoded / decoded representation conforms to format rules, and wherein the format rules specify that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to a segment does not use inter-layer prediction, the value of the field indicating the segment type of the segment is set to indicate the type of an intra segment.

[0352] 2. The method of Solution 1, wherein the stripe type of the stripe is set to 2.

[0353] 3. The method of Solution 1, wherein the predetermined range is from IDR_W_RADL to CRA_NUT, both inclusive.

[0354] 4. A method for video processing, comprising: performing a conversion between a video including one or more scalable video layers and an encoded and decoded representation of the video, wherein the encoded and decoded representation conforms to format rules, and wherein the format rules stipulate that in the case where an index of a syntax structure describing the profile, layer, and level of one or more scalable video layers is excluded from the encoded and decoded representation, the indication of the index is inferred to be equal to 0.

[0355] 5. The method of Solution 4, wherein the indication includes the ols_ptl_idx[i] syntax element, where i is an integer.

[0356] 6. A method for video processing, comprising: performing a conversion between a video including one or more scalable video layers and an encoded and decoded representation of the video, wherein the encoded and decoded representation includes a plurality of output layer sets (OLSs), wherein each OLS is a set of layers in the encoded and decoded representation, and wherein the set of layers is specified to be output, wherein the encoded and decoded representation conforms to format rules, and wherein the format rules stipulate that in the case where an OLS includes a single layer, the set of video parameters of the encoded and decoded representation is allowed to have duplicate information regarding the profile, layer, and level of the OLS.

[0357] 7. The method of Solution 6, wherein the format rules stipulate a signaling digital field that indicates the total number of sets of information regarding the profile, layer, and level of the plurality of OLSs, where the total number is at least one.

[0358] 8. The method of any one of Solutions 1-7, wherein performing the conversion includes encoding the video to generate the encoded and decoded representation.

[0359] 9. The method of any one of Solutions 1-7, wherein performing the conversion includes parsing and decoding the encoded and decoded representation to generate the video.

[0360] 10. A video decoding device, comprising a processor configured to implement the method described in one or more of Solutions 1 to 9.

[0361] 11. A video encoding device, comprising a processor configured to implement the method described in one or more of Solutions 1 to 9.

[0362] 12. A computer program product having computer code stored thereon which, when executed by a processor, causes the processor to implement the method recited in any one of Solutions 1 to 9.

[0363] 13. The method, apparatus or system described in this document.

[0364] Regarding Figures 7A to 7I , the following listed solutions can be preferably implemented in some embodiments.

[0365] For example, the following solution can be implemented according to Item 1 in the previous section.

[0366] 1. A method for video processing (e.g., Figure 7A the method 710 described in

[0367] ), including performing (712) a conversion between a video including one or more scalable video layers and a video bitstream, where the video includes one or more video pictures including one or more strips; where the bitstream conforms to format rules, where the format rules stipulate that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the strip does not use inter-layer prediction, the value of the field indicating the strip type of the strip is set to the type indicating an intra strip.

[0368] 2. The method of Solution 1, where the strip type of the strip is set to 2.

[0369] 3. The method of Solution 1, where the predetermined range is from IDR_W_RADL to CRA_NUT, both inclusive, where IDR_W_RADL indicates a strip with an instantaneous decoding refresh type, and where CRA_NUT indicates a strip with a clean random access type.

[0370] 4. The method of any one of Solutions 1 - 3, where the corresponding video layer flag corresponds to vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]], where nuh_layer_id corresponds to the identifier of the video layer, GeneralLayerIdx[] is the index of the video layer, and vps_independent_layer_flag corresponds to the video layer flag.

[0371] 5. The method of any one of Solutions 1 - 4, where performing the conversion includes encoding the video into the bitstream.

[0372] 7. A method for storing a video bitstream, comprising: generating a bitstream according to a video including one or more scalable video layers; and storing the bitstream in a non-transitory computer-readable storage medium; wherein the video includes one or more video pictures, and the video pictures include one or more slices; wherein the bitstream conforms to format rules, and the format rules specify that when the corresponding network abstraction layer unit type is within a predetermined range and the corresponding video layer flag indicates that the video layer corresponding to the slice does not use inter-layer prediction, the value of the field indicating the slice type of the slice is set to indicate the type of I slice.

[0373] For example, the following solutions can be implemented according to item 5 in the previous section.

[0374] 1. A video processing method (e.g., Figure 7B the method 720 shown), comprising: performing (722) a conversion between a video including multiple video layers and a bitstream of the video, wherein the bitstream includes multiple output layer sets (OLSs), each output layer set (OLS) includes one or more of the multiple scalable video layers, and the bitstream conforms to format rules, and the format rules specify that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, layer, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0375] 2. The method of solution 1, wherein the format rules specify that a syntax element is included in the video parameter set indicating multiple PTL syntax structures encoded and decoded in the video parameter set.

[0376] 3. The method of solution 2, wherein the number of PTL syntax structures encoded and decoded in the video parameter set is equal to the value of the syntax element plus one.

[0377] 4. The method of solution 2, wherein the number of PTL syntax structures encoded and decoded in the video parameter set is equal to the value of the syntax element.

[0378] 5. The method of solution 4, wherein the format rules specify that when it is determined that one or more of the multiple OLSs contain more than one video layer, the syntax element is encoded and decoded in the video parameter set during encoding, or the syntax element is parsed from the video parameter set during decoding.

[0379] 6. The method of solution 4, wherein the format rules specify that when it is determined that the number of video layers is greater than 0 and one or more video layers use inter-layer prediction, the syntax element is encoded and decoded in the video parameter set during encoding, or the syntax element is parsed from the video parameter set during decoding.

[0380] 7. A method according to any one of Solutions 1 to 6, wherein the formatting rule further provides that for an OLS with a single layer, the decoded picture buffer size parameter of the OLS is included in the video parameter set of the bitstream, and the decoded picture buffer size parameter of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0381] 8. A method according to any one of Solutions 1 to 6, wherein the formatting rule further provides that for an OLS with a single layer, the parameter information of the decoded picture buffer of the OLS is included in the video parameter set of the bitstream, and the parameter information of the decoded picture buffer of the OLS is also included in the sequence parameter set encoded and decoded in the bitstream.

[0382] For example, the following solutions can be implemented according to Item 6 in the previous section.

[0383] 9. A video processing method (e.g., Figure 7C the method 730 shown), comprising: performing (732) a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of video layers, and the bitstream conforms to a formatting rule, wherein the formatting rule specifies a relationship between the occurrence of a plurality of profile-tier-level (PTL) syntax structures in the video parameter set of the bitstream and a byte alignment syntax field in the video parameter set; wherein each PTL syntax structure indicates the profile, tier, and level of one or more of the plurality of OLSs.

[0384] 10. The method of Solution 9, wherein the formatting rule provides that the video parameter set includes at least one PTL syntax structure, and due to including at least one PTL syntax structure, one or more instances of the byte alignment syntax field are included in the video parameter set.

[0385] 11. The method of Solution 10, wherein the byte alignment syntax field is one bit.

[0386] 12. The method of Solution 11, wherein each value of one or more instances of the byte alignment syntax field has a value of 0.

[0387] For example, the following solutions can be implemented according to Item 2 in the previous section.

[0388] 13. A video processing method (e.g., Figure 7DThe method shown (740) includes: performing (742) a conversion between a video including a plurality of scalable video layers and a bitstream of the video, where the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of scalable video layers, and the bitstream complies with format rules, where the format rules specify that during encoding, in the case where the value of an index is zero, a syntax element indicating an index of a syntax structure that describes the profile, layer, and level of one or more of the plurality of OLSs is excluded from a set of video parameters of the bitstream, or during decoding, in the case where the syntax element does not exist in the bitstream, the value is inferred to be zero.

[0389] 14. The method of solution 13, where the index points to a syntax structure that indicates the profile, layer, and level of at least one OLS.

[0390] 15. A method for storing a bitstream of a video, including: generating a bitstream from a video including a plurality of video layers and storing the bitstream in a non-transitory computer-readable storage medium; where the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream complies with format rules, where the format rules specify that for an OLS having a single layer, a profile-tier-level (PTL) syntax structure indicating the profile, layer, and level of the OLS is included in a set of video parameters of the bitstream, and the PTL syntax structure of the OLS is also included in a set of sequence parameters encoded and decoded in the bitstream.

[0391] 16. The method of any one of solutions 1-14, where performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0392] For example, the following solutions can be implemented according to item 3 in the previous section.

[0393] 1. A video processing method (e.g., Figure 7E the method shown 750) includes: performing (752) a conversion between a video including a plurality of video layers and a bitstream of the video, where the bitstream includes a plurality of output layer sets (OLSs), each output layer set (OLS) includes one or more of the plurality of video layers, and the bitstream complies with format rules, where the format rules specify that for layer i, where i is an integer, the bitstream includes a set of first syntax elements indicating a first variable that indicates whether layer i is included in at least one of the plurality of OLSs.

[0394] 2. The method of Solution 1, wherein the formatting rule stipulates that in the case where the first variable of layer i is equal to zero, it means that layer i is not included in any of the multiple OLSs, and the bitstream excludes the second set of syntax elements indicating the decoded picture buffer parameters of layer i.

[0395] 3. The method of any one of Solutions 1-2, wherein the formatting rule further stipulates that the bitstream includes a third set of syntax elements indicating a second variable, which indicates whether layer i is used as a reference layer for at least one of the multiple video layers, and wherein the formatting rule does not allow the first variable and the second variable to have zero values.

[0396] 4. The method of Solution 3, wherein the formatting rule does not allow the values of the first variable and the second variable to both be equal to 0, which indicates that no layer is neither a direct reference layer of any other layer nor an output layer of at least one OLS.

[0397] 5. The method of any one of Solutions 1-4, wherein the first variable is a one-bit flag, denoted as LayerUsedAsOutputLayerFlag.

[0398] 6. The method of Solution 5, wherein the first variable is determined based on repeatedly checking the value of a third variable for each layer in the multiple video layers, and this value indicates the relationship between the multiple layers included in the multiple OLSs.

[0399] 7. The method of Solution 6, wherein the third variable indicating the relationship between the multiple layers included in the multiple OLSs is allowed to have values of 0, 1, or 2.

[0400] 8. The method of any one of Solutions 1-7, wherein performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0401] 9. A method for storing a video bitstream, including: generating a bitstream according to a video including multiple video layers, and storing the bitstream in a non-transitory computer-readable storage medium; wherein the bitstream includes multiple output layer sets (OLSs), each output layer set includes one or more of the multiple video layers, and the bitstream conforms to formatting rules, wherein the formatting rule stipulates that for layer i (where i is an integer), the bitstream includes a set of first syntax elements indicating a first variable, and this first variable indicates whether layer i is included in at least one of the multiple OLSs.

[0402] For example, the following solutions can be implemented according to Item 4 in the previous section.

[0403] 1. A video processing method (for example, Figure 7FThe method shown (760) includes: performing (762) a conversion between a video and a bitstream of the video, where the bitstream includes one or more sets of output layers, and each set of output layers includes one or more video layers; where the bitstream conforms to format rules, and where the format rules specify that in the case where each set of output layers includes a single video layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: zero; or in the case where it is not true that each set of output layers includes a single layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: one plus the value of a syntax element.

[0404] The method of Solution 1, wherein the syntax element corresponds to the vps_num_dpb_params_minus1 syntax element.

[0405] The method of any one of Solutions 1 - 2, wherein the format rules specify that in the case where there is no other syntax element in the video parameter set (the syntax element indicating whether the same dimensions are used to indicate the decoded picture buffer syntax structures of video layers that are included and not included in one or more sets of output layers), the value of the other syntax element is inferred to be equal to 1.

[0406] The method of any one of Solutions 1 - 3, wherein performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non - transitory computer - readable storage medium.

[0407] A method for storing a bitstream of a video, including: generating a bitstream according to the video; and storing the bitstream in a non - transitory computer - readable storage medium; where the bitstream includes one or more sets of output layers, and each set of output layers includes one or more video layers; where the bitstream conforms to format rules, and where the format rules specify that in the case where each set of output layers includes a single video layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: zero; or in the case where it is not true that each set of output layers includes a single layer, the number of decoded picture buffer parameter syntax structures included in the video parameter set of the bitstream is equal to: one plus the value of a syntax element.

[0408] For example, the following solutions can be implemented according to item 7 in the previous section.

[0409] 1. A video processing method (for example, Figure 7GThe method (e.g., method 770) shown includes: performing (772) a conversion between a video and a bitstream of the video, where the bitstream includes a coded video sequence (CVS), and the coded video sequence includes one or more coded video pictures of one or more video layers; and where the bitstream conforms to a format rule that specifies that one or more sequence parameter sets (SPSs) indicating conversion parameters referred to by one or more of the coded pictures of the CVS have the same reference video parameter set (VPS) identifier indicating a reference VPS.

[0410] The method of Solution 1, where the format rule further specifies that the same reference VPS identifier has a value greater than 0.

[0411] The method of any one of Solutions 1 - 2, where the format rule further specifies that, in response to and only when the CVS includes a single video layer, the value zero of the SPS identifier is used.

[0412] The method of any one of Solutions 1 - 3, where performing the conversion includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non - transitory computer - readable storage medium.

[0413] A method for storing a bitstream of a video, including: generating a bitstream according to the video and storing the bitstream in a non - transitory computer - readable storage medium; where the video includes one or more video pictures, and the video pictures include one or more slices; where the bitstream includes a coded video sequence (CVS), and the coded video sequence includes one or more coded video pictures of one or more video layers; and where the bitstream conforms to a format rule that specifies that one or more sequence parameter sets (SPSs) indicating conversion parameters referred to by one or more of the coded pictures of the CVS have the same reference video parameter set (VPS) identifier indicating a reference VPS.

[0414] For example, the following solutions can be implemented according to item 8 in the previous section.

[0415] 1. A video processing method (e.g., Figure 7H the method 780) shown includes: performing (782) a conversion between a video and a bitstream of the video, where the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) includes one or more video layers, where the bitstream conforms to a format rule; where the format rule specifies whether or how a first syntax element is included in a video parameter set (VPS) of the bitstream, and the first syntax element indicates whether a first syntax structure describing parameters of a hypothetical reference decoder (HRD) is used for the conversion.

[0416] 2. The method of Solution 1, wherein the first syntax structure includes a set of common HRD parameters.

[0417] 3. The method of Solutions 1-2, wherein the formatting rules specify that, since each of one or more OLSs includes one or more video layers, when the first syntax element is not present in the VPS, the first syntax element is ignored from the VPS and is inferred to have a zero value, and wherein each of the one or more OLSs includes a single video layer.

[0418] 4. The method of Solutions 1-2, wherein the formatting rules specify that, since each of one or more OLSs includes one or more video layers, when the first syntax element appears in the VPS, the first syntax element has a zero value in the VPS, and wherein each of the one or more OLSs includes a single video layer.

[0419] 5. The method of any one of Solutions 1-4, wherein the formatting rules further specify whether or how the VPS includes a second syntax element that indicates a plurality of syntax structures describing OLS-specific HRD parameters.

[0420] 6. The method of Solution 5, wherein the formatting rules further specify that when the first syntax element has a value of 1, the second syntax element is included in the VPS regardless of whether the total number of OLSs in the one or more OLSs is greater than 1.

[0421] For example, the following solutions can be implemented according to Items 9 and 10 in the previous section.

[0422] 7. A video processing method (e.g., Figure 7I the method 790 shown), comprising: performing (792) a conversion between a video and a bitstream of the video, wherein the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) includes one or more video layers, and wherein the bitstream conforms to formatting rules; wherein the formatting rules specify whether or how the video parameter set (VPS) of the bitstream includes a first syntax structure describing common hypothetical reference decoder (HRD) parameters and a plurality of second syntax structures describing OLS-specific HRD parameters.

[0423] 8. The method of Solution 7, wherein the formatting rules specify that the first syntax structure is omitted from the VPS in the case where no second syntax structure is included in the VPS.

[0424] 9. The method of Solution 7, wherein, for an OLS including only one video layer, the formatting rules exclude the first and second syntax structures from being included in the VPS, and wherein the formatting rules allow the first and second syntax structures to be included in an OLS including only one video layer.

[0425] 10. A method for storing a bitstream of a video, comprising: generating a bitstream according to the video; storing the bitstream in a non-transitory computer-readable storage medium; wherein the bitstream includes one or more output layer sets (OLSs), each output layer set (OLS) includes one or more video layers, wherein the bitstream conforms to formatting rules; wherein the formatting rules specify whether or how a first syntax element is included in the video parameter set (VPS) of the bitstream, and the first syntax element indicates whether a first syntax structure describing parameters of a hypothetical reference decoder (HRD) is used for transformation.

[0426] The solutions listed above may further include:

[0427] In some embodiments, in the solutions listed above, performing the transformation includes encoding the video into a bitstream.

[0428] In some embodiments, in the solutions listed above, performing the transformation includes parsing and decoding the video according to the bitstream.

[0429] In some embodiments, in the solutions listed above, performing the transformation includes encoding the video into a bitstream; and the method further includes storing the bitstream in a non-transitory computer-readable storage medium.

[0430] In some embodiments, a video decoding device includes a processor configured to implement the method described in one or more of the solutions listed above.

[0431] In some embodiments, a video encoding device includes a processor configured to implement the method described in one or more of the solutions listed above.

[0432] In some embodiments, a non-transitory computer-readable storage medium may store instructions that cause a processor to implement the method described in one or more of the solutions listed above.

[0433] In some embodiments, a non-transitory computer-readable storage medium stores a bitstream of a video, which is generated by the method described in one or more of the solutions listed above.

[0434] In some embodiments, the encoding method described above may be implemented by a device, and the device may further write the bitstream generated by implementing the method to a computer-readable medium.

[0435] The disclosed solutions and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in: digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of the foregoing. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution or control of the operation by a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, such as including programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code for creating an execution environment for the computer programs under discussion, e.g., code constituting processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.

[0436] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can also be deployed in any form, including as a stand-alone program or module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0437] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0438] For example, processors suitable for executing a computer program include general and special purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for executing the instructions and one or more memory devices for storing the instructions and data. Generally, a computer will also include or be operatively coupled to receive data from and transfer data to one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks). However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including for example semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0439] Although this patent document contains many details, these details should not be construed as limiting the scope of any subject matter or of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features that are described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in a particular combination and even initially claimed as such, in some cases, one or more features may be deleted from the claimed combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.

[0440] Similarly, although these operations are depicted in the drawings in a particular order, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood to require such separation in all embodiments.

[0441] Only some implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein the format rules stipulate that for an OLS with a single layer, a profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream; wherein the format rules stipulate that a number syntax element is included in the video parameter set, and the number syntax element indicates the number of PTL syntax structures decoded in the video parameter set; wherein the number of the PTL syntax structures is equal to the value of the number syntax element plus one.

2. The method according to claim 1, wherein, The format rules further stipulate the relationship between the occurrences of a plurality of PTL syntax structures in the video parameter set of the bitstream and a byte alignment syntax field in the video parameter set.

3. The method according to claim 2, wherein The format rules stipulate that at least one PTL syntax structure is included in the video parameter set, and due to the inclusion of the at least one PTL syntax structure, one or more instances of the byte alignment syntax field are included in the video parameter set.

4. The method according to claim 3, wherein The byte alignment syntax field is one bit.

5. The method according to claim 4, wherein, Each value of one or more instances of the byte alignment syntax field is 0.

6. The method according to claim 1, wherein, The format rules further stipulate that: During encoding, in the case where the value of the number syntax element is equal to zero, an index syntax element indicating the index of the PTL syntax structure is excluded from the video parameter set of the bitstream, and During decoding, in the case where the index syntax element does not exist in the bitstream, the value of the index syntax element is inferred to be zero.

7. The method according to claim 1, wherein Performing the conversion includes encoding the video into the bitstream.

8. The method according to claim 1, wherein, Performing the conversion includes decoding the video from the bitstream.

9. The method according to claim 1, wherein The format rules further stipulate that for an OLS with a single layer, the decoded picture buffer size parameter of the OLS is included in the video parameter set of the bitstream, and the decoded picture buffer size parameter of the OLS is also included in the sequence parameter set decoded in the bitstream.

10. The method according to claim 1, wherein, The format rules further stipulate that for an OLS with a single layer, the parameter information of the decoded picture buffer of the OLS is included in the video parameter set of the bitstream, and the parameter information of the decoded picture buffer of the OLS is also included in the sequence parameter set decoded in the bitstream.

11. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to: Perform a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, Among them, the format rule stipulates that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream; Among them, the format rule stipulates that a number of syntax elements are included in the video parameter set, and the number of syntax elements indicates the number of PTL syntax structures decoded in the video parameter set; Among them, the number of the PTL syntax structures is equal to the value of the number syntax element plus one.

12. The apparatus according to claim 11, wherein The format rule also stipulates the relationship between the appearance of multiple PTL syntax structures in the video parameter set of the bitstream and the byte alignment syntax field in the video parameter set, Among them, the format rule stipulates that at least one PTL syntax structure is included in the video parameter set, and due to the inclusion of the at least one PTL syntax structure, one or more instances of the byte alignment syntax field are included in the video parameter set, Among them, the byte alignment syntax field is one bit, and Among them, the value of each of the one or more instances of the byte alignment syntax field is 0.

13. The apparatus according to claim 11, wherein, The format rule also stipulates that: During encoding, when the value of the number syntax element is equal to zero, the index syntax element indicating the index of the PTL syntax structure is excluded from the video parameter set of the bitstream, and During decoding, when the index syntax element does not exist in the bitstream, the value of the index syntax element is inferred to be zero.

14. A non-transitory computer-readable storage medium storing instructions that cause a processor to: Perform the conversion between a video including a plurality of video layers and the bitstream of the video, Among them, The bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to the format rule, Among them, the format rule stipulates that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream; Among them, the format rule stipulates that a number of syntax elements are included in the video parameter set, and the number of syntax elements indicates the number of PTL syntax structures decoded in the video parameter set; Among them, the number of the PTL syntax structures is equal to the value of the number syntax element plus one.

15. The non-transitory computer-readable storage medium according to claim 14, wherein, The format rule also stipulates the relationship between the appearance of multiple PTL syntax structures in the video parameter set of the bitstream and the byte alignment syntax field in the video parameter set, Among them, the format rule stipulates that at least one PTL syntax structure is included in the video parameter set, and due to the inclusion of the at least one PTL syntax structure, one or more instances of the byte alignment syntax field are included in the video parameter set, wherein, the byte alignment syntax field is one bit, and wherein, the value of each of one or more instances of the byte alignment syntax field is 0.

16. A non-transitory computer-readable storage medium storing a bitstream of video, the bitstream being generated by a method executed by a video processing device, wherein, The method includes: generating a bitstream of a video including a plurality of video layers, wherein, the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein, the format rules specify that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream; wherein, the format rules specify that a number syntax element is included in the video parameter set, and the number syntax element indicates the number of PTL syntax structures decoded in the video parameter set; wherein, the number of the PTL syntax structures is equal to the value of the number syntax element plus one.

17. The non-transitory computer-readable storage medium according to claim 16, wherein, The format rules further specify the relationship between the occurrences of a plurality of PTL syntax structures in the video parameter set of the bitstream and the byte alignment syntax field in the video parameter set, wherein, the format rules specify that at least one PTL syntax structure is included in the video parameter set, and due to including the at least one PTL syntax structure, one or more instances of the byte alignment syntax field are included in the video parameter set, wherein, the byte alignment syntax field is one bit, and wherein, the value of each of one or more instances of the byte alignment syntax field is 0.

18. A method for storing a bitstream of a video, including: generating a bitstream of a video including a plurality of video layers, storing the bitstream into a non-transitory computer-readable storage medium, wherein, the bitstream includes a plurality of output layer sets (OLSs), each output layer set includes one or more of the plurality of video layers, and the bitstream conforms to format rules, wherein, the format rules specify that for an OLS with a single layer, the profile-tier-level (PTL) syntax structure indicating the profile, tier, and level of the OLS is included in the video parameter set of the bitstream, and the PTL syntax structure of the OLS is also included in the sequence parameter set decoded in the bitstream; wherein, the format rules specify that a number syntax element is included in the video parameter set, and the number syntax element indicates the number of PTL syntax structures decoded in the video parameter set; wherein, the number of the PTL syntax structures is equal to the value of the number syntax element plus one.

19. A video decoding device, including a processor, the processor being configured to implement the method according to any one of claims 1-10.

20. A video encoding device, including a processor, the processor being configured to implement the method according to any one of claims 1-10.

21. A video processing device for storing a bitstream, wherein, The video processing device is configured to implement the method according to any one of claims 1-10 by a processor.

Citation Information

Patent Citations

  • Profile, tier, level for the 0-th output layer set in video coding

    CN106464919A

  • Method and an apparatus and a computer program for encoding media content

    CN109155861A