Video Coding and Decoding Using Parameter Sets
By adding syntax elements to the video encoding and decoding standard to control the generation of strip types and reference picture lists in the video area, the problems of insufficient control of B striping and redundant parameter transmission in the prior art are solved, and more efficient video encoding and decoding and lower bandwidth requirements are achieved.
Patent Information
- Application Number
- CN202180026286.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-06
- Filing Date
- 2021-04-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-04-01
AI Technical Summary
The existing video codec standards have shortcomings in controlling the types of stripes allowed in videos and the use of bidirectional predicted B-bands, resulting in the decoder being unable to efficiently process videos containing only P and I bands, and there is redundant parameter transmission.
By adding syntax elements in the video parameter set and general constraint information syntax, indicating whether only specific types of stripes are allowed in the video area, and controlling the generation of reference picture lists, ensuring that only allowed strip types are signaled and used.
It realizes more efficient video encoding and decoding, reduces the transmission of unnecessary parameters, improves the control ability of video content types, and is suitable for video transmission in low-bandwidth environments.
Smart Images

Figure CN115428454B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit and priority of U.S. Patent Application No. 63 / 006,054, filed on April 6, 2020, and International Patent Application No. PCT / US2021 / 025351, filed on April 1, 2021. All of the above - mentioned patent applications are hereby incorporated by reference in their entirety. Technical Field
[0003] This patent document relates to image and video encoding, decoding, and transcoding. Background Art
[0004] In the Internet and other digital communication networks, digital video occupies the largest bandwidth. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to grow. Summary of the Invention
[0005] Techniques are disclosed herein that can be used by video encoders and decoders to process coded representations of video using control information useful for decoding the coded representations.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video including one or more layers each including one or more video regions and a coded representation of the video according to formatting rules, where the formatting rules specify that one or more syntax elements are included in the coded representation at one or more video region levels corresponding to allowed slice types of the corresponding video regions.
[0007] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more layers each including one or more video slices and a coded representation of the video according to formatting rules, where the one or more layers include one or more video pictures each including one or more video slices, and the formatting rules specify that syntax elements related to enabling or using a coded mode at the slice level are included at most once between a picture header and a slice header according to a second rule.
[0008] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video pictures each including one or more video slices and a coded representation of the video according to formatting rules, where the formatting rules specify whether a reference picture list for allowed slice types in a video picture is signaled in the coded representation or generated from the coded representation.
[0009] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between videos including one or more video pictures each including one or more sub-pictures, where the codec representation conforms to format rules, and where the format rules specify the processing of non-coded sub-pictures of the video pictures.
[0010] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify that, in response to one or more conditions being met, a syntax element indicating whether a first syntax structure providing profile, layer, and level information and a second syntax structure providing decoded picture buffer information exist in a sequence parameter set is set to be equal to 1 to indicate that the first syntax structure and the second syntax structure exist in the sequence parameter set.
[0011] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more coded layers and a bitstream of the video according to format rules, and where the format rules specify that one or more parameter sets and / or general constraint information syntax structures include one or more syntax elements indicating the allowed slice types in pictures of the coded layer video sequence.
[0012] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more layers and a bitstream of the video according to format rules, where the one or more layers include one or more pictures each including one or more slices, and the format rules specify that syntax elements are included in a picture header or a slice header to indicate whether a bi-predictive B slice is allowed for the corresponding picture or slice of the video or whether the bi-predictive B slice is used for the corresponding picture or slice of the video.
[0013] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more layers and a bitstream of the video according to format rules, where the one or more layers include one or more pictures each including one or more slices, and the format rules specify that one or more syntax elements related to the enabling or use of a coded mode at the slice level are included at most once between picture headers or slice headers according to a second rule.
[0014] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify setting the value of a variable indicating whether a picture in a decoded picture buffer that is before the current picture in decoding order in the bitstream is output before the picture is removed from the decoded picture buffer based on the picture order count value of the current picture.
[0015] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify enabling control of picture type and layer independence i) whether a syntax element indicating allowance of an inter-frame strip or a B strip or a P strip is included in the picture and / or prediction information, and / or ii) an indication of the presence of prediction information.
[0016] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify that the use of a reference picture list during the conversion of a coded layer video sequence depends on the strip type allowed in the pictures of the video corresponding to the coded layer video sequence.
[0017] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more video sequences and a bitstream of the video according to format rules, where the format rules specify whether or under which conditions two sets of adaptive parameters in the video sequence or the bitstream are allowed to have the same adaptive parameter set identifier.
[0018] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video according to format rules, where the format rules specify that a first parameter set and a second parameter set are dependent on each other such that whether or how a syntax element is included in the second parameter set is based on the first parameter set.
[0019] In another example aspect, another video processing method is disclosed. The method includes performing a conversion between a video including one or more pictures and a bitstream of the video, each picture including one or more sub-pictures, where the format rules specify the processing of non-coded sub-pictures of the picture.
[0020] In yet another example aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the above method.
[0021] In yet another example aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the above method.
[0022] In yet another example aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0023] These and other features will be described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a block diagram of an example video processing system;
[0025] Figure 2 is a block diagram of a video processing apparatus;
[0026] Figure 3 is a flowchart of an example method of video processing;
[0027] Figure 4 is a block diagram showing a video codec system according to some embodiments of the present disclosure;
[0028] Figure 5 is a block diagram showing an encoder according to some embodiments of the present disclosure;
[0029] Figure 6 is a block diagram showing a decoder according to some embodiments of the present disclosure; and
[0030] Figures 7A to 7J is a flowchart of an example method of video processing based on some implementations of the disclosed technology. Detailed Description
[0031] The section headings are used herein for ease of understanding and are not intended to limit the applicability of the technologies and embodiments disclosed in each section to only that section. Additionally, the use of H.266 terminology in some descriptions is for ease of understanding only and is not intended to limit the scope of the disclosed technologies. Accordingly, the technologies described herein are also applicable to other video codec protocols and designs. In this document, certain embodiments are shown as changes to the current VVC specification, where new text is added, shown in bold italic, and deleted text is marked with double brackets (e.g., [[a]] indicates deletion of the character 'a').
[0032] 1. Introduction
[0033] This document relates to video codec technology. Specifically, it pertains to improvements in signaling of allowed slice types and associated codec tools applicable only to bi - directionally predicted slices, as well as support for non - coded sub - pictures. These ideas can be applied alone or in various combinations to any video codec standard or non - standard video codec that supports multi - layer video coding, such as the Versatile Video Coding (VVC) currently under development.
[0034] 2. Abbreviations
[0035] ALF Adaptive Loop Filter
[0036] APS Adaptive Parameter Set
[0037] AU Access Unit
[0038] AUD Access Unit Delimiter
[0039] AVC Advanced Video Coding
[0040] CLVS Coding Layer Video Sequence
[0041] CPB Coding Picture Buffer
[0042] CRA Complete Random Access
[0043] CTU Coding Tree Unit
[0044] CVS Coding Video Sequence
[0045] DCI Decoding Capability Information
[0046] DPB Decoding Picture Buffer
[0047] DU Decoding Unit
[0048] EOB End of Bitstream
[0049] EOS End of Sequence
[0050] GDR Gradual Decoding Refresh
[0051] HEVC High Efficiency Video Coding
[0052] HRD Hypothetical Reference Decoder
[0053] IDR Instantaneous Decoding Refresh
[0054] JEM Joint Exploration Model
[0055] LMCS Luminance Mapping and Chrominance Scaling
[0056] MCTS Motion Constrained Tile Set
[0057] NAL Network Abstraction Layer
[0058] OLS Output Layer Set
[0059] PH Picture Header
[0060] PPS Picture Parameter Set
[0061] PTL Profile, Tier and Level
[0062] PU Picture Unit
[0063] RADL Random Access Decodable Leading (picture)
[0064] RAP Random Access Point
[0065] RASL Random Access Skip Leading (Picture)
[0066] RBSP Raw Byte Sequence Payload
[0067] RPL Reference Picture List
[0068] SAO Sample Adaptive Offset
[0069] SEI Supplemental Enhancement Information
[0070] SPS Sequence Parameter Set
[0071] STSA Stepwise Temporal Sub-layer Access
[0072] SVC Scalable Video Coding
[0073] VCL Video Coding Layer
[0074] VPS Video Parameter Set
[0075] VTM VVC Test Model
[0076] VUI Video Usability Information
[0077] VVC Versatile Video Coding
[0078] 3. Preliminary Discussion
[0079] Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 video, and the two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction plus transform coding is used. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to the continuous efforts in VVC standardization, new coding technologies have been incorporated into the VVC standard at each JVET meeting. The working draft and test model VTM of VVC are updated after each meeting. The latest VVC working draft JVET-Q2001_vE can be downloaded from the following address:
[0080] http: / / phenix.it-
[0081] sudparis.eu / jvet / doc_end_user / documents / 17_Brussels / wg11 / JVET-Q2001-
[0082] v15.zip
[0083] The VVC project now aims to be technically completed (FDIS) at the meeting in July 2020.
[0084] 3.1. Parameter Sets
[0085] AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced starting from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.
[0086] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, information that does not change frequently does not need to be repeated for each sequence or picture, thus avoiding redundant signaling of this information. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.
[0087] VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0088] APS is introduced to carry such picture-level or slice-level information that requires a significant number of bits to encode and decode, can be shared by multiple pictures, and can have a significant number of different variations in the sequence.
[0089] 3.1.1. Video Parameter Set (VPS)
[0090] The syntax tables and semantics of multiple syntax elements in the latest VVC draft text (JVET-Q2001-vE / v15) are defined as follows:
[0091] 7.3.2.2 Video Parameter Set RBSP Syntax
[0092]
[0093] 3.1.2. Sequence Parameter Set (SPS)
[0094] The syntax tables and semantics of multiple syntax elements in the latest VVC draft text (JVET-Q2001-vE / v15) are defined as follows:
[0095] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0096]
[0097]
[0098] 3.1.3. General Constraint Flag
[0099] 7.3.3.2 General Constraint Information Syntax
[0100]
[0101]
[0102] When no_bdof_constraint_flag is equal to 1, it specifies that sps_bdof_enabled_flag shall be equal to 0. When no_bdof_constraint_flag is equal to 0, such a constraint is not imposed.
[0103] When no_dmvr_constraint_flag is equal to 1, it specifies that sps_dmvr_enabled_flag shall be equal to 0. When no_dmvr_constraint_flag is equal to 0, such a constraint is not imposed.
[0104] When no_bcw_constraint_flag is equal to 1, it specifies that sps_bcw_enabled_flag shall be equal to 0. When no_bcw_constraint_flag is equal to 0, such a constraint is not imposed.
[0105] When no_ciip_constraint_flag is equal to 1, it specifies that sps_ciip_enabled_flag shall be equal to 0. When no_cipp_constraint_flag is equal to 0, such a constraint is not imposed.
[0106] When no_gpm_constraint_flag is equal to 1, it specifies that sps_gpm_enabled_flag shall be equal to 0. When no_gpm_constraint_flag is equal to 0, such a constraint is not imposed.
[0107] 3.1.4. Picture Parameter Set (PPS)
[0108] The syntax tables and semantics of multiple syntax elements in the latest VVC draft text (JVET-Q2001-vE / v15) are defined as follows:
[0109] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0110]
[0111]
[0112] num_ref_idx_default_active_minus1[i] plus 1 specifies the inferred value of NumRefIdxActive[0] for P slices or B slices where num_ref_idx_active_override_flag equals 0 when i equals 0, and the inferred value of NumRefIdxActive[1] for B slices where num_ref_idx_active_override_flag equals 0 when i equals 1. The value of num_ref_idx_default_active_minus1[i] shall be in the range from 0 to 14 (inclusive of 0 and 14).
[0113] pps_weighted_bipred_flag equal to 0 specifies that explicit weighted prediction is not applied to B slices of the reference PPS. pps_weighted_bipred_flag equal to 1 specifies that explicit weighted prediction is applied to B slices of the reference PPS. When sps_weighted_bipred_flag equals 0, the value of pps_weighted_bipred_flag shall be equal to 0.
[0114] 3.1.5 DPB Parameter Syntax
[0115] The syntax tables and semantics of multiple syntax elements in the latest VVC draft text (JVET-Q2001-vE / v15) are defined as follows:
[0116] 7.3.4 DPB Parameter Syntax
[0117]
[0118] 7.4.5 DPB Parameter Semantics
[0119] The dpb_parameters() syntax structure provides information on the DPB size, maximum picture reordering count, and maximum latency for one or more OLSs.
[0120] When the dpb_parameters() syntax structure is included in the VPS, the OLS to which the dpb_parameters() syntax structure applies is specified by the VPS. When the dpb_parameters() syntax structure is included in the SPS, it applies to the OLS of the lowest layer among the layers that include only the reference SPS, and the lowest layer is an independent layer.
[0121] max_dec_pic_buffering_minus1[i] plus 1 specifies the maximum required size of the DPB in picture storage buffer units when Htid is equal to i. The value of max_dec_pic_buffering_minus1[i] shall be in the range from 0 to MaxDpbSize - 1 (inclusive of 0 and MaxDpbSize - 1), where MaxDpbSize is specified in Clause A.4.2. When i is greater than 0, max_dec_pic_buffering_minus1[i] shall be greater than or equal to max_dec_pic_buffering_minus1[i - 1]. When there is no max_dec_pic_buffering_minus1[i] for i in the range from 0 to maxSubLayersMinus1 - 1 (inclusive of 0 and maxSubLayersMinus1 - 1), it is inferred to be equal to max_dec_pic_buffering_minus1[maxSubLayersMinus1] since subLayerInfoFlag is equal to 0.
[0122] max_num_reorder_pics[i] specifies the maximum allowed number of pictures that can be in the OLS before any picture in the OLS and after that picture in output order when Htid is equal to i. The value of max_num_reorder_pics[i] shall be in the range from 0 to max_dec_pic_buffering_minus1[i] (inclusive of 0 and max_dec_pic_buffering_minus1[i]). When i is greater than 0, max_num_reorder_pics[i] shall be greater than or equal to max_num_reorder_pics[i - 1]. When there is no max_num_reorder_pics[i] for i in the range from 0 to maxSubLayersMinus1 - 1 (inclusive of 0 and maxSubLayersMinus1 - 1), it is inferred to be equal to max_num_reorder_pics[maxSubLayersMinus1] since subLayerInfoFlag is equal to 0.
[0123] max_latency_increase_plus1[i] is not equal to 0 for calculating the value of MaxLatencyPictures[i], which specifies the maximum number of pictures that can be before any picture in the OLS and after that picture in decoding order when Htid is equal to i in the OLS, according to the output order.
[0124] When max_latency_increase_plus1[i] is not equal to 0, the value of MaxLatencyPictures[i] is specified as follows:
[0125] MaxLatencyPictures[i] = max_num_reorder_pics[i] + max_latency_increase_plus1[i] - 1 (7-110)
[0126] When max_latency_increase_plus1[i] is equal to 0, the corresponding restriction is not expressed.
[0127] The value of max_latency_increase_plus1[i] shall be in the range of 0 to 2 32 -2 (including 0 and 2 32 -2). When there is no max_latency_increase_plus1[i] for i in the range from 0 to maxSubLayersMinus1 - 1 (including 0 and maxSubLayersMinus1 - 1), since subLayerInfoFlag is equal to 0, it is inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1].
[0128] 3.2. Picture Header (PH) and Slice Header (SH) in VVC
[0129] Similar to HEVC, the slice header in VVC conveys information about a specific slice. This includes slice address, slice type, slice QP, least significant bit (LSB) of the picture order count (POC), RPS and RPL information, weighted prediction parameters, loop filter parameters, entry offsets for slices and WPP, etc.
[0130] VVC introduces a Picture Header (PH) which contains the header parameters for a specific picture. Each picture must have one and only one PH. The PH basically carries those parameters that would be in the slice header if the PH was not introduced, but each parameter has the same value for all slices of the picture. These include IRAP / GDR picture indication, inter / intra slice enable flag, POCLSB and optionally POC MSB, information on RPL, deblocking, SAO, ALF, QP delta and weighted prediction, coding block partition information, virtual boundary, collocated picture information, etc. It often happens that each picture in an entire picture sequence contains only one slice. To allow for each picture to not have at least two NAL units in such cases, the PH syntax structure is allowed to be included in the PH NAL unit or the slice header.
[0131] In VVC, information on collocated pictures for temporal motion vector prediction is signaled in the picture header or the slice header.
[0132] 3.2.1. Picture Header (PH)
[0133] The syntax tables and semantics of multiple syntax elements in the latest VVC working draft () are defined as follows:
[0134] 7.3.2.7 Picture Header Structure Syntax
[0135]
[0136]
[0137] 3.2.2. Slice Header (SH)
[0138] The syntax tables and semantics of multiple syntax elements in the latest VVC working draft () are defined as follows:
[0139] 7.3.7.1 General Slice Header Syntax
[0140]
[0141]
[0142]
[0143] The slice_type specifies the coding type of the slice according to Table 9.
[0144] Table 9 – Association with the name of slice_type
[0145] slice_type Name of slice_type 0 B (B stripe) 1 P (P stripe) 2 I (I stripe)
[0146] When it does not exist, the value of slice_type is inferred to be equal to 2.
[0147] When ph_intra_slice_allowed_flag is equal to 0, the value of slice_type shall be equal to 0 or 1. When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT), and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type shall be equal to 2.
[0148] 3.3. Latest progress in JVET-R0052
[0149] In JVET-R0052 method #2, it is proposed to add an allowed type index (i.e., ph_allowed_slice_types_idc), and whether to use B slices in a picture can be deduced from the newly added syntax element.
[0150]
[0151] In addition, another new syntax element ph_multiple_slice_types_in_pic_flag is added to the PH structure to specify whether more than one slice type can exist in the current picture. ph_multiple_slice_types_in_pic_flag being equal to 1 specifies that the coded slices of the picture can have different values of slice_type. ph_multiple_slice_types_in_pic_flag being equal to 0 specifies that all the coded slices of the picture have the same value of slice_type. When ph_multiple_slice_types_in_pic_flag is equal to 0, ph_slice_type is further signaled to specify the value of slice_type for all the slices of the picture, and the slice_type in the slice header is not decoded and is inferred to be equal to the value of ph_slice_type.
[0152] 7.3.2.7 Picture header structure syntax
[0153]
[0154]
[0155]
[0156] 7.3.7.1 General strip header syntax
[0157]
[0158]
[0159]
[0160] 7.4.3.7 Picture header structure semantics
[0161]
[0162]
[0163]
[0164]
[0165] [[The ph_inter_slice_allowed_flag being equal to 0 specifies that all coded slices of the picture have a slice_type equal to 2. The ph_inter_slice_allowed_flag being equal to 1 specifies that there may or may not be one or more coded slices in the picture with a slice_type equal to 0 or 1. [Editor (YK): Carefully check the need / correctness of the inference rules for those syntax elements conditioned on the flag being equal to 0.]]]
[0166] The ph_intra_slice_allowed_flag being equal to 0 specifies that all coded slices of the picture have a slice_type equal to 0 or 1. The ph_intra_slice_allowed_flag being equal to 1 specifies that there may or may not be one or more coded slices in the picture with a slice_type equal to 2. When not present, the value of the ph_intra_slice_allowed_flag is inferred to be equal to 1. [Editor (YK): Carefully check the need / correctness of the inference rules for those syntax elements conditioned on the flag being equal to 1.]]]
[0167] Note 2 – For bitstreams where sub-picture based bitstream merging should be performed without changing the PH NAL unit, it is expected that the encoder will [[ph_inter_slice_allowed_flag and ph_intra_slice_allowed_flag]] be set to equal 1.
[0168] 7.4.8.1 General strip header semantics
[0169] The slice_type specifies the coding / decoding type of the slice according to Table 9.
[0170] Table 9 – Association with the name of slice_type
[0171] slice_type Name of slice_type 0 B (B stripe) 1 P (P stripe) 2 I (I stripe)
[0172] When it is not present, the value of slice_types is [[inferred to be equal to 2]] derived as follows:
[0173] – If ph_multiple_slice_types_in_pic_flag is equal to 1, the value of slice_type is set to be equal to (slice_type_modified >= ph_allowed_slice_types_idc? slice_type_modified + 1 : slice_type_modified).
[0174] – Otherwise, the value of slice_type is set to be equal to the value of ph_slice_type.
[0175]
[0176] [[When ph_intra_slice_allowed_flag is equal to 0, the value of slice_type shall be equal to 0 or 1.]] When nal_unit_type is in the range from IDR_W_RADL to CRA_NUT (including IDR_W_RADL and CRA_NUT) and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, slice_type shall be equal to 2.
[0177] 7.4.8.2 Weighted Prediction Parameter Semantics
[0178] num_l1_weights specifies the number of weights signaled for the entries in reference picture list 1 when both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1. The value of num_l1_weights shall be in the range from 0 to Min(15, num_ref_entries[1][RplsIdx[1]]) (including 0 and Min(15, num_ref_entries[1][RplsIdx[1]])).
[0179] The variable NumWeightsL1 is derived as follows:
[0180]
[0181] NumWeightsL1 = NumRefIdxActive[1]
[0182] The new syntax element pps_multiple_slice_types_in_pic_flag can be further signaled in the PPS. When pps_multiple_slice_types_in_pic_flag is equal to 0, for all PHs that refer to the PPS, ph_multiple_slice_types_in_pic_flag is inferred to be equal to 0.
[0183] The relevant modifications to VVC Draft 8 are written in red and highlighted in yellow and are provided as follows:
[0184] 7.3.2.4 Picture Parameter Set RBSP Syntax
[0185]
[0186] PH of Method 1
[0187] 7.3.2.7 Picture Header Structure Syntax
[0188]
[0189] PH of Method 2
[0190]
[0191] 7.4.3.4 Picture Parameter Set RBSP Semantics
[0192]
[0193]
[0194] 3.4. Uncoded Sub - pictures and Potential Applications in JVET - R0151
[0195] In this document, it is shown how the mechanism of enabling uncoded sub - pictures can be used to extend VVC. When the sub - pictures do not completely fill the picture, by providing completely unused areas, uncoded sub - pictures can be used for efficient coding. Examples of OMAF use cases and 360° video coding of 4x3 cube maps are shown. In addition, uncoded sub - pictures can be used to reserve space that is not filled with coded data but with content generated from already coded content. Here, examples of high - level efficient geometric filling for 360° video are shown.
[0196] 4. Technical problems solved by the disclosed technical solutions
[0197] The current VVC text and the latest progress in JVET have the following problems:
[0198] 1. In the latest VVC draft text (in JVET-Q2001-vE / v15), there are two PH syntax elements related to the allowed slice types. For example, ph_inter_slice_allowed_flag and ph_intra_slice_allowed_flag, as listed in the picture header structure syntax. Using these two flags, only when ph_intra_slice_allowed_flag is true, the syntax elements related to intra-slice coding are signaled, and only when ph_inter_slice_allowed_flag is true, the syntax elements related to inter-slice coding are signaled. However, when ph_inter_slice_allowed_flag is equal to 1, the decoder does not know whether the picture contains B slices. Some applications (e.g., online games, video conferencing, video surveillance) typically only use P slices and I slices. Therefore, if there is an indication of whether B slices are allowed, the decoders of such applications will be able to select to request / use only bitstreams that do not include B slices. In addition, this indication can be used to avoid transmitting multiple unnecessary parameters.
[0199] 2. In JVET-R0052, the proposed changes are only applied to PH and SH. There is no higher-level control over whether it can only have the same slice type within a picture and / or which allowed slice types are enabled in the picture. In addition, when some syntax elements related only to bi-directional prediction do not exist, there is no description of how to infer the values.
[0200] 3. In item 1 of JVET-R0191, it is proposed to replace the constraint that the value of sps_ptl_dpb_hrd_params_present_flag should be equal to vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] with the following:
[0201] When there is an OLS that contains only one layer and the layer ID is equal to the nuh_layer_id of the SPS, the value of sps_ptl_dpb_hrd_params_present_flag should be equal to 1.
[0202] However, by the condition of "when there is an OLS with only one layer", this change introduces VPS dependency. Another problem is that for a single-layer bitstream, the value of sps_ptl_dpb_hrd_params_present_flag should be equal to 1, and this is not recorded by the changed constraint.
[0203] 5. List of Example Technical Solutions and Embodiments
[0204] To solve the above problems, the following summarized methods are disclosed. The following items should be regarded as examples explaining general concepts and should not be interpreted in a narrow way. In addition, these items can be applied individually or combined in any way.
[0205] One or more syntax elements can be added to the parameter set (e.g., SPS, VPS, PPS, APS, DCI) and / or the general constraint information syntax to indicate whether only X (e.g., I or B or P) slices are allowed within a picture; and / or to indicate the set of allowed slice types in the picture.
[0206] In parameter set and general constraint information syntax
[0207] 1. In a video unit such as SPS or PPS, add one or more syntax elements (e.g., sps_allowed_slice_idc) to specify whether the allowed slice types are in the picture of CLVS.
[0208] 1) In one example, add a first syntax element (e.g., sps_allowed_slice_idc), and its semantics can be defined as: sps_allowed_slice_idc being equal to X specifies that the picture only allows the following allowed slice types or any combination thereof:
[0209] i. {All I}, {All P}, {All B}, {I, P}, {I, B}, {P, B}, {I, B, P}
[0210] ii. In one instance, the first syntax element can be encoded and decoded using fixed length (e.g., u(1), u(2) or U(3)), unary coding, truncated unary coding, EG coding.
[0211] iii. Additionally, alternatively, the signaling and / or semantics and / or inference of one or more syntax elements signaled in the SPS or PPS can be modified such that they are only signaled when the first syntax element meets certain conditions.
[0212] a. In one example, one or more syntax elements are syntax elements for enabling coding tools that require more than one prediction signal, such as bidirectional prediction or hybrid intra and inter coding, or prediction using linear / nonlinear weighting from multiple prediction blocks.
[0213] b. In one example, one or more syntax elements may include but are not limited to:
[0214] a) sps_weighted_bipred_flag
[0215] b) sps_bdof_enabled_flag
[0216] c) sps_smvd_enabled_flag
[0217] d) sps_dmvr_enabled_flag
[0218] e) sps_bcw_enabled_flag
[0219] f) sps_ciip_enabled_flag
[0220] g) sps_gpm_enabled_flag
[0221] c. In one example, one or more syntax elements may be signaled only if the first syntax element specifies that the CLVS associated with the video unit may contain one or more B slices. Otherwise, the signaling is skipped and the value of the syntax element is inferred.
[0222] d. In one example, when sps_b_slice_allowed_flag is equal to 0, the syntax elements sps_weighted_bipred_flag, sps_bdof_enabled_flag, sps_smvd_enabled_flag, sps_dmvr_enabled_flag, sps_bcw_enabled_flag, sps_ciip_enabled_flag, and sps_gpm_enabled_flag are not signaled and their values are inferred.
[0223] a) In one example, when absent, they are all inferred to be 0.
[0224] iv. Additionally, alternatively, a second syntax element, such as no_b_slice_contraint_flag, may be signaled in the general constraint information syntax to indicate whether the first syntax element should be equal to 0.
[0225] a. In one example, the semantics of no_b_slice_contraint_flag are defined as follows:
[0226] Equal to 1 specifies that sps_allowed_slice_idc should be equal to X (e.g., the allowed slice types are represented as {I, B, P} or {B, P}, {all B}). no_b_slice_constraint_flag equal to 0 does not impose such a constraint.
[0227] v. Additionally, alternatively, if the first syntax element specifies that the CLVS does not contain B slices (e.g., sps_allowed_slice_idc equal to X representing only {I, P}, {all I}, {all P}), then one or more syntax elements signaled in the common constraint information syntax should be equal to 1.
[0228] a. In one example, one or more syntax elements may include but are not limited to:
[0229] a) no_bcw_constraint_flag
[0230] b) no_ciip_constraint_flag
[0231] c) no_gpm_constraint_flag
[0232] d) no_bdof_constraint_flag
[0233] e) no_dmvr_constraint_flag
[0234] vi. Additionally, alternatively, the signaling and semantics of one or more syntax elements signaled in dpb_parameters() can be modified such that they are only signaled when the first syntax element meets certain conditions.
[0235] a. In one example, one or more syntax elements may include but are not limited to:
[0236] a) max_num_reorder_pics
[0237] b. In one example, when the first syntax element indicates that B slices are not allowed, max_num_reorder_pics is not signaled and is inferred to be 0.
[0238] In PH / SH
[0239] 2. In PH / SH, the variable X is used to indicate whether B slices are allowed / used in the picture / strip, and this variable can be derived from the SPS syntax element and / or a new PH syntax element (e.g., ph_allowed_slice_idc) that specifies the allowed slice types and / or other syntax elements (e.g., BSliceAllowed used in JVET - R0052).
[0240] 1) In one example, a new PH syntax element is added, and how to signal this syntax element can depend on the allowed slice types in the SPS.
[0241] 2) Additionally, alternatively, one or more syntax elements signaled in PH can be modified in terms of signaling and / or semantics and / or inference according to the variable.
[0242] i. In one example, one or more syntax elements are those for enabling codec tools that require more than one prediction signal, such as bi - directional prediction or hybrid intra - and inter - frame coding, or prediction using linear / non - linear weighting from multiple prediction blocks.
[0243] ii. In one example, one or more syntax elements can include but are not limited to:
[0244] a) ph_collocated_from_l0_flag
[0245] b) mvd_l1_zero_flag
[0246] c) ph_disable_bdof_flag
[0247] d) ph_disable_dmvr_flag
[0248] e) num_l1_weights
[0249] iii. In one example, one or more syntax elements can be signaled only when the first syntax element specifies that the picture can contain one or more B slices. Otherwise, the signaling is skipped, and the value of the syntax element is inferred.
[0250] a) Additionally, alternatively, whether one or more syntax elements are signaled can depend on the first syntax element in bullets 1.1 and 2.1, such as (X is true or 1).
[0251] b) ph_disable_bdof_flag can be signaled only when (sps_bdof_pic_present_flag ) is true.
[0252] c) Only when (sps_dmvr_pic_present_flag ) is true, can ph_disable_dmvr_flag be signaled.
[0253] iv. In one example, when X is equal to 0 (or false), mvd_l1_zero_flag is not signaled, and its value is inferred to be 1.
[0254] v. In one example, the inference of one or more syntax elements depends on the value of the first syntax element.
[0255] a) In one example, for ph_disable_bdof_flag, the following applies:
[0256] – If sps_bdof_enabled_flag is equal to 1 then the value of ph_disable_bdof_flag is inferred to be equal to 0.
[0257] – Otherwise (sps_bdof_enabled_flag is equal to ), the value of ph_disable_bdof_flag is inferred to be equal to 1.
[0258] b) In one example, for ph_disable_dmvr_flag, the following applies:
[0259] – If sps_dmvr_enabled_flag is equal to then the value of ph_disable_dmvr_flag is inferred to be equal to 0.
[0260] – Otherwise (sps_dmvr_enabled_flag is equal to ), the value of ph_disable_dmvr_flag is inferred to be equal to 1.
[0261] c) In one example, when ph_temporal_mvp_enabled_flag and rpl_info_in_ph_flag are both equal to 1 and X is equal to 0 (or false), the value of ph_collocated_from_l0_flag is inferred to be equal to 1.
[0262] d) In one example, when X is equal to 0 (or false), the non - belief signal notifies num_l1_weights, and its value is inferred to be 0. Therefore, the weighted prediction parameters of reference picture list 1 are not signaled in the PH or SH of the picture.
[0263] Inference of syntax elements
[0264] 3. For the syntax elements related to coding / decoding tool X and / or a set of syntax elements that may exist in A (e.g., PH) or B (e.g., SH) but not in both, if A is included in B, then at least one of the indications of the existence of those syntax elements may not be signaled and may be inferred to be 0, i.e., existing in B.
[0265] 1) In one example, coding / decoding tool X may include one of the following:
[0266] i. Loop filtering techniques, such as de - blocking filter, ALF, SAO
[0267] ii. Weighted prediction
[0268] iii. QP delta information
[0269] iv. RPL information
[0270] 2) In one example, the condition “A is included in B” can be defined as “the slice header of the reference PPS contains the PH syntax structure” or “the current picture consists of only one slice”.
[0271] 3) In one example, “the indication of the existence of those syntax elements” can be defined as one or more of the following syntax elements:
[0272] i. qp_delta_info_in_ph_flag, rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, wp_info_in_ph_flag
[0273] 4) In one example, one or more of the following changes are proposed.
[0274] rpl_info_in_ph_flag being equal to 1 specifies that the reference picture list information exists in the PH syntax structure and does not exist in the slice header of the reference PPS that does not contain the PH syntax structure. rpl_info_in_ph_flag being equal to 0 specifies that the reference picture list information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure.
[0275] When dbf_info_in_ph_flag is equal to 1, it specifies that the deblocking filter information exists in the PH syntax structure and does not exist in the slice header of the reference PPS that does not contain the PH syntax structure. When dbf_info_in_ph_flag is equal to 0, it specifies that the deblocking filter information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure. When it does not exist, the value of dbf_info_in_ph_flag is inferred to be equal to 0.
[0276] When sao_info_in_ph_flag is equal to 1, it specifies that the SAO filter information exists in the PH syntax structure and does not exist in the slice header of the reference PPS that does not contain the PH syntax structure. When sao_info_in_ph_flag is equal to 0, it specifies that the SAO filter information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure.
[0277] When alf_info_in_ph_flag is equal to 1, it specifies that the ALF information exists in the PH syntax structure and does not exist in the slice header of the reference PP that does not contain the PH syntax structure. When alf_info_in_ph_flag is equal to 0, it specifies that the ALF information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure.
[0278] When wp_info_in_ph_flag is equal to 1, it specifies that the weighted prediction information may exist in the PH syntax structure and does not exist in the slice header of the reference PPS that does not contain the PH syntax structure. When wp_info_in_ph_flag is equal to 0, it specifies that the weighted prediction information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure. When it does not exist, the value of wp_info_in_ph_flag is inferred to be equal to 0.
[0279] The qp_delta_info_in_ph_flag being equal to 1 specifies that the QP delta information exists in the PH syntax structure and does not exist in the slice header of the reference PPS that does not contain the PH syntax structure. The qp_delta_info_in_ph_flag being equal to 0 specifies that the QP delta information does not exist in the PH syntax structure and may exist in the slice header of the reference PPS that does not contain the PH syntax structure.
[0280]
[0281] 4. The compliant bitstream shall follow the rule that when the POC value of the splicing point picture of the CLVS AU in the spliced bitstream is greater than the POC value of the previous picture, the NoOutputOfPriorPicsFlag shall be set to equal 1 for the splicing point picture.
[0282] 5. Whether signaling indicates the syntax elements allowing inter-layer slices / B-slices / P-slices in the picture and / or RPL / WP information, and / or the indication of the existence of RPL / WP information can depend on the picture type and whether layer independence is enabled.
[0283] 1) In one example, the syntax elements are not signaled for IRAP pictures and layer independence is enabled.
[0284] i. In one example, the ph_inter_slice_allowed_flag in VVC is not signaled for IRAP pictures and layer independence is enabled.
[0285] ii. In one example, the slice_type in VVC is not signaled for IRAP pictures and layer independence is enabled.
[0286] iii. In one example, the ph_slice_type in JVET-R0052 is not signaled for IRAP pictures and layer independence is enabled.
[0287] 2) In one example, the syntax elements are not signaled for IRAP pictures and layer independence is enabled, even if the existence of such information informs them in the PH.
[0288] i. When the gdr_or_irap_pic_flag is equal to 1 and the gdr_pic_flag is equal to 0, a new flag called idr_pic_flag is proposed to specify whether the picture associated with the picture header is an IDR picture. And the following can be applied:
[0289] a. When sps_idr_rpl_present_flag is equal to 0, layer independence is enabled, and idr_pic_flag is equal to 1, even when the value of rpl_info_in_ph_flag is equal to 1, RPL signaling does not exist in the PH.
[0290] b. When sps_idr_rpl_present_flag is equal to 0, layer independence is enabled, and idr_pic_flag is equal to 1, even when the value of wp_info_in_ph_flag is equal to 1, WP signaling does not exist in the PH.
[0291] 6. It is proposed that when sps_video_parameter_set_id is greater than 0 and there is an OLS with only one layer where nuh_layer_id is equal to the SPS's nuh_layer_id, or when sps_video_parameter_set_id is equal to 0, the value of sps_ptl_dpb_hrd_params_present_flag should be equal to 1.
[0292] Reference list related
[0293] 7. The signaling notification and / or generation of the reference picture list can depend on the allowed slice types in the CLVS pictures.
[0294] 1) For example, if B slices are not allowed in the CLVS, one or more syntax elements for constructing reference list 1 may not be signaled.
[0295] 2) For example, if B slices are not allowed in the CLVS, one or more processes for constructing reference list 1 may not be performed.
[0296] APS related
[0297] 8. It is required that two APSs should not have the same APS_id in the sequence, CLVS, or bitstream.
[0298] 1) Alternatively, it is required that two APSs of the same APS type (such as ALF APS or LMCS APS) should not have the same APS_id in the sequence, CLVS, or bitstream.
[0299] 2) Alternatively, two APSs of the same APS type (such as ALF APS or LMCS APS) are allowed to have the same APS_id, but they must have the same content in the sequence, CLVS, or bitstream.
[0300] 3) Alternatively, two APSs with the same APS type (such as ALF APS or LMCS APS) are allowed to have the same APS_id. And the APS notified by earlier signaling is replaced by the APS notified by later signaling.
[0301] 4) Alternatively, two APSs with the same APS type (such as ALF APS or LMCS APS) are allowed to have the same APS_id. And the APS notified by later signaling is ignored.
[0302] 9. Two different parameter sets (e.g., APS and SPS) can be dependent on each other, and the syntax elements or variables derived from the syntax elements in the first parameter set can be used to conditionally signal another syntax element in the second parameter set.
[0303] 1) Alternatively, the syntax elements or variables derived from the syntax elements in the first parameter set can be used to derive the value of another syntax element in the second parameter set.
[0304] Non-coded / decoded sub-picture related
[0305] 10. It is proposed that the boundaries of non-coded / decoded sub-pictures must be regarded as picture boundaries.
[0306] 11. It is proposed that loop filtering (such as ALF / deblocking / SAO) cannot cross the boundaries of non-coded / decoded sub-pictures.
[0307] 12. It is required that if there is only one sub-picture, it cannot be a non-coded / decoded sub-picture.
[0308] 13. It is required that non-coded / decoded sub-pictures cannot be extracted.
[0309] 14. It is proposed that information related to (multiple) non-coded / decoded sub-pictures can be signaled in the SEI message.
[0310] 15. It is required that non-coded / decoded sub-pictures can only have one strip.
[0311] 16. It is required that the top-left sub-picture cannot be a non-coded / decoded sub-picture.
[0312] 17. It is required that at least one of the sub-pictures is not a non-coded / decoded sub-picture.
[0313] 18. Whether and / or how to code / decoder the side information-related sub-pictures can depend on whether the sub-picture is a non-coded / decoded sub-picture.
[0314] 1) In one example, if it is a non-coded / decoded sub-picture, signaling of side information is not required.
[0315] 19. Additionally, alternatively, for the above requirements, they can be modified to be signaled conditionally according to the above situation.
[0316] Next, a list of preferred examples of some embodiments is provided.
[0317] The first set of clauses shows example embodiments of the techniques discussed in the previous section. The following clauses show example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0318] 1. A video processing method (e.g., Figure 3 the method 3000 shown in), including performing a conversion (3002) between a video including one or more layers each including one or more video regions and an encoded / decoded representation of the video according to format rules, where the format rules specify that one or more syntax elements are included in the encoded / decoded representation at one or more video region levels corresponding to the allowed slice types of the corresponding video regions.
[0319] 2. The method according to clause 1, wherein the format rules specify that one or more syntax elements include a first syntax element, and the value of the first syntax element indicates a combination of allowed slice types in the corresponding video region.
[0320] The following clauses show example embodiments of the techniques discussed in the previous section (e.g., item 2).
[0321] 3. The method according to any one of clauses 1 - 2, wherein the format rules specify that the syntax elements are included in a picture header or a slice header to indicate whether bidirectional prediction (B) slices are allowed for the corresponding picture or slice or whether the B slices are used for the corresponding picture or slice.
[0322] 4. The method according to clause 3, wherein the syntax elements in the sequence parameter set control the presence of the syntax elements included in the picture header or the slice header.
[0323] The following clauses show example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0324] 5. A video processing method, including: performing a conversion between a video including one or more layers and an encoded / decoded representation of the video according to format rules, where the one or more layers include one or more video pictures each including one or more video slices, and the format rules specify that syntax elements related to the enabling or use of the encoded / decoding mode at the slice level are included at most once between a picture header and a slice header according to a second rule.
[0325] 6. The method according to clause 5, wherein the encoded / decoding mode includes a loop filter or a weighted prediction mode, or a quantization parameter delta mode.
[0326] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., item 7).
[0327] 7. A video processing method, comprising: performing a conversion between a video and an encoded / decoded representation of the video, the video including one or more video pictures each including one or more video strips, according to format rules, where the format rules specify whether a control reference picture list for an allowed strip type in a video picture is signaled in the encoded / decoded representation or generated from the encoded / decoded representation.
[0328] 8. The method according to clause 7, wherein the format rules specify that, since the allowed strip type excludes bi-directional strips (B-strips), a syntax element corresponding to reference picture list 1 is omitted from the encoded / decoded representation.
[0329] 9. The method according to clause 7, wherein the format rules specify that, since the allowed strip type excludes bi-directional strips (B-strips), a process for generating reference picture list 1 is disabled for the video picture.
[0330] The following clauses illustrate example embodiments of the techniques discussed in the previous section (e.g., items 10 - 15).
[0331] 10. A video processing method, comprising: performing a conversion between videos, the videos including one or more video pictures each including one or more sub-pictures, where the encoded / decoded representation complies with format rules, and where the format rules specify the processing of non-encoded / decoded sub-pictures of the video picture.
[0332] 11. The method according to clause 10, wherein the format rules specify that, during the conversion, the boundaries of non-encoded / decoded sub-pictures are treated as picture boundaries.
[0333] 12. The method according to clause 10, wherein the format rules specify that loop filtering across the boundaries of non-encoded / decoded pictures is disabled.
[0334] 13. The method according to clause 10, wherein the format rules do not allow non-encoded / decoded sub-pictures to be merely sub-pictures of a video picture.
[0335] 14. The method according to any one of clauses 10 - 13, wherein the format rules specify that information for decoding assistance for non-encoded / decoded sub-pictures is included in an auxiliary enhancement information syntax element of the encoded / decoded representation.
[0336] 15. The method according to clause 10, wherein the format rules specify that non-encoded / decoded sub-pictures are allowed to have at most one strip.
[0337] 16. The method according to any one of the above clauses, wherein the video region includes a video picture or a video strip.
[0338] 17. The method according to any one of clauses 1 to 16, wherein the conversion includes encoding the video into a codec representation.
[0339] 18. The method according to any one of clauses 1 to 16, wherein the conversion includes decoding the codec representation to generate pixel values of the video.
[0340] 19. A video decoding device, comprising a processor configured to implement the method according to one or more of clauses 1 to 18.
[0341] 20. A video encoding device, comprising a processor configured to implement the method according to one or more of clauses 1 to 18.
[0342] 21. A computer program product storing computer code which, when executed by a processor, causes the processor to implement the method according to any one of clauses 1 to 18.
[0343] 22. A method, device or system described in this document.
[0344] The second set of clauses shows example embodiments of the techniques discussed in the previous section (e.g., items 1 - 19).
[0345] 1. A method for video processing (e.g., the method 700 as shown in Figure 7A ) includes: performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify that, in response to satisfying one or more conditions, a syntax element indicating whether a first syntax structure providing profile, layer, and level information and a second syntax structure providing decoded picture buffer information exist in a sequence parameter set is set to equal 1 to indicate that the first syntax structure and the second syntax structure exist in the sequence parameter set.
[0346] 2. The method according to clause 1, wherein the one or more conditions include 1) the video parameter set identifier referred to by the sequence parameter set is greater than 0, and there is an output layer set including only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value, or 2) the video parameter set identifier is equal to 0.
[0347] 3. The method according to clause 1 or 2, wherein the syntax element being equal to 1 further specifies that a third syntax structure providing general timing and hypothesized reference decoder parameter information and a fourth syntax structure providing output layer set timing and hypothesized reference decoder parameter information are allowed to exist in the sequence parameter set.
[0348] 4. The method according to clause 3, wherein the third syntax structure corresponds to the general_timing_hrd_parameters() syntax structure, and the fourth syntax structure corresponds to the ols_timing_hrd_parameters() syntax structure.
[0349] 5. The method according to any one of clauses 1 to 4, wherein the syntax element corresponds to the sps_ptl_dpb_hrd_params_present_flag, the first syntax structure corresponds to the profile_tier_level() syntax structure, and the second syntax structure corresponds to the dpb_parameters() syntax structure.
[0350] 6. A method for video processing (e.g., the method 710 as Figure 7B shown), comprising: performing a conversion between a video including one or more coding layers and a bitstream of the video according to format rules, wherein the format rules specify one or more parameter sets and / or a general constraint information syntax structure including one or more syntax elements indicating the allowed slice types in a picture of the coded layer video sequence.
[0351] 7. The method according to clause 6, wherein the format rules specify that the first syntax element is further included, and the value of the first syntax element indicates the allowed slice type or the combination of allowed slice types in the video region.
[0352] 8. The method according to clause 7, wherein the format rules specify that one or more syntax elements are signaled only when the first syntax element meets certain conditions.
[0353] 9. The method according to clause 7, wherein the format rules specify that the general constraint information syntax structure includes a second syntax element to indicate whether the first syntax element is equal to 0.
[0354] 10. The method according to clause 7, wherein the format rules specify that when the first syntax element specifies that no bi-predictive B slices are included in the coded layer video sequence, one or more syntax elements are equal to 1.
[0355] 11. A method for video processing (e.g., the method 720 as Figure 7C shown), comprising: performing a conversion between a video including one or more layers and a bitstream of the video according to format rules, wherein the one or more layers include one or more pictures 722 each containing one or more slices, and wherein the format rules specify that syntax elements are included in the picture header or slice header to indicate whether bi-predictive B slices are allowed for the corresponding picture or slice of the video or whether the bi-predictive B slices are used for the corresponding picture or slice of the video.
[0356] 12. The method according to clause 11, wherein the formatting rules specify that the syntax elements in the sequence parameter set control the presence of the syntax elements included in the picture header or slice header.
[0357] 13. The method according to clause 11, wherein the formatting rules specify how the signaling of the syntax elements in the picture header depends on the slice types allowed in the sequence parameter set.
[0358] 14. The method according to clause 11, wherein the formatting rules specify that the syntax elements control the signaling and / or semantics and / or inference of one or more syntax elements included in the picture header.
[0359] 15. A method for video processing (e.g., the method 730 as shown in Figure 7D ) includes: performing a conversion between a video including one or more layers and a bitstream of the video according to formatting rules, wherein the one or more layers include one or more pictures 732 each including one or more slices, and wherein the formatting rules specify that one or more syntax elements related to the enabling or use of the codec mode at the slice level are included at most once between the picture header and the slice header according to a second rule.
[0360] 16. The method according to clause 15, wherein the codec mode includes loop filtering or weighted prediction mode, or quantization parameter delta mode, or reference picture list information.
[0361] 17. The method according to clause 15, wherein the formatting rules specify that the slice header of the reference picture parameter set contains the picture header syntax structure, and the requirement for bitstream consistency is that the values of one or more syntax elements are equal to 0.
[0362] 18. A method for video processing (e.g., the method 740 as shown in Figure 7E ) includes: performing a conversion 742 between a video including one or more pictures and a bitstream of the video according to formatting rules, wherein the formatting rules specify setting the value of a variable indicating whether the pictures in the decoded picture buffer that are decoded before the current picture in decoding order in the bitstream are output before being removed from the decoded picture buffer based on the picture order count value of the current picture.
[0363] 19. The method according to clause 18, wherein the formatting rules specify that when the picture order count value of the current picture, which is a splicing point picture in the bitstream and a video sequence access unit of the codec layer, is greater than the picture order count value of the previous picture, the value of the variable is set to be equal to 1 for the current picture.
[0364] 20. A method for video processing (e.g., asFigure 7F The method shown (750) includes: performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify enabling control of picture type and layer independence i) whether a syntax element indicating an inter prediction strip or a B strip or a P strip is included in the picture and / or prediction information, and / or ii) an indication of the presence of prediction information.
[0365] 21. The method according to clause 20, wherein the format rules specify that the syntax element is not included when i) the picture type is an intra random access point picture and ii) layer independence is enabled.
[0366] 22. The method according to clause 21, wherein the format rules specify that the syntax element is not included when i) and ii) are satisfied, regardless of another syntax element indicating the presence of prediction information in the picture header.
[0367] 23. The method according to clause 21 or 22, wherein the format rules specify that when the picture is an intra random access point picture, it further includes a variable specifying whether the picture associated with the picture header is an instant decoding refresh IDR picture.
[0368] 24. The method according to any one of clauses 21 to 23, wherein the format rules specify that an indication of the presence of prediction information does not exist in the picture header.
[0369] 25. A method for video processing (e.g., the method shown as Figure 7G 760), includes: performing a conversion (762) between a video including one or more pictures and a bitstream of the video according to format rules, where the format rules specify that the use of the reference picture list during the conversion of the coded layer video sequence depends on the allowed strip type in the pictures of the video corresponding to the coded layer video sequence.
[0370] 26. The method according to clause 25, wherein the format rules specify that since the allowed strip type excludes bi - directional strips (B strips), the syntax elements corresponding to reference picture list 1 are omitted from the bitstream.
[0371] 27. The method according to clause 25, wherein the format rules specify that since the allowed strip type excludes bi - directional strips (B strips), the process for generating reference picture list 1 is disabled for video pictures.
[0372] 28. A method for video processing (e.g., as Figure 7HThe method 770) shown, includes: performing a conversion 772 between a video including one or more video sequences and a bitstream of the video according to format rules, where the format rules specify whether or under which conditions two sets of adaptive parameters in the video sequence or the bitstream are allowed to have the same adaptive parameter set identifier.
[0373] 29. The method according to clause 28, wherein the format rules specify that two sets of adaptive parameters do not have the same adaptive parameter set identifier.
[0374] 30. The method according to clause 28, wherein, in the case where two sets of adaptive parameters have the same adaptive parameter set type, the two sets of adaptive parameters do not have the same adaptive parameter set identifier.
[0375] 31. The method according to clause 28, wherein, in 1) the case where two sets of adaptive parameters have the same adaptive parameter set type and have the same content, or 2) the case where two sets of adaptive parameters have the same adaptive parameter set type, the two sets of adaptive parameters have the same adaptive parameter set identifier.
[0376] 32. A method for video processing (e.g., the method 780) as Figure 7I shown, includes: performing a conversion 782 between a video and a bitstream of the video according to format rules, where the format rules specify that a first parameter set and a second parameter set are dependent on each other such that whether or how a syntax element is included in the second parameter set is based on the first parameter set.
[0377] 33. The method according to clause 32, wherein the format rules specify that the syntax elements in the second parameter set are conditionally included or derived based on a syntax element or variable derived from another syntax element in the first parameter set.
[0378] 34. A method for video processing (e.g., the method 790) as Figure 7J shown, includes: performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, each picture including one or more sub - pictures 792, where the format rules specify the processing of non - coded / decoded sub - pictures of the pictures.
[0379] 35. The method according to clause 34, wherein the format rules specify that, during the conversion, the boundaries of non - coded / decoded sub - pictures are treated as picture boundaries.
[0380] 36. The method according to clause 34, wherein the format rules specify disabling loop filtering across the boundaries of non - coded / decoded sub - pictures.
[0381] 37. The method according to clause 34, wherein the formatting rules do not allow a non-coded / decoded sub-picture to be merely a sub-picture of a video picture.
[0382] 38. The method according to clause 34, wherein the formatting rules specify that non-coded / decoded sub-pictures are not extracted during the conversion.
[0383] 39. The method according to clause 34, wherein the formatting rules specify that the auxiliary enhancement information syntax elements of the bitstream include information for decoding assistance of non-coded / decoded sub-pictures.
[0384] 40. The method according to clause 34, wherein the formatting rules specify that a non-coded / decoded sub-picture is allowed to have at most one slice.
[0385] 41. The method according to clause 34, wherein the formatting rules specify that a non-coded / decoded sub-picture is not the top-left sub-picture of a picture.
[0386] 42. The method according to clause 34, wherein the formatting rules specify that at least one of the one or more sub-pictures is a coded / decoded sub-picture.
[0387] 43. The method according to clause 34, wherein the formatting rules specify that whether and / or how to code / decode side information related to one or more sub-pictures depends on whether the sub-picture is coded / decoded or non-coded / decoded.
[0388] 44. The method according to any one of clauses 1 to 43, wherein the conversion includes encoding video into a bitstream.
[0389] 45. The method according to any one of clauses 1 to 43, wherein the conversion includes decoding video from the bitstream.
[0390] 46. The method according to clauses 1 to 43, wherein the conversion includes generating a bitstream from video, and the method further includes: storing the bitstream in a non-transitory computer-readable recording medium.
[0391] 47. A video processing apparatus, including a processor configured to implement the method according to any one or more of clauses 1 to 46.
[0392] 48. A method of storing a bitstream of video, including the method according to any one of clauses 1 to 46, and further including storing the bitstream in a non-transitory computer-readable recording medium.
[0393] 49. A computer-readable medium storing program code, which when executed causes a processor to implement the method according to any one or more of clauses 1 to 46.
[0394] 50. A computer-readable medium stores a bitstream generated according to any one of the above methods.
[0395] 51. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of clauses 1 to 46.
[0396] As used herein, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block may correspond to a juxtaposed position in the bitstream defined by the syntax or bits propagated at different positions. For example, a macroblock may be encoded based on the transformed and encoded error residual values and also using bits in the headers and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above technical solutions. Similarly, the encoder may determine whether to include or exclude a particular syntax field and generate the encoded / decoded representation accordingly by including the syntax field or excluding the syntax field from the encoded / decoded representation.
[0397] The disclosed and other technical solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a combination of substances that affect a machine-readable propagated signal, or one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver device.
[0398] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed on one or more computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0399] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0400] For example, processors suitable for executing a computer program include general and special-purpose microprocessors, and any one or more of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, e.g., magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable hard disks; magneto-optical disks; and CD ROM and DVD ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0401] Although this patent document contains many details, it should not be construed as limiting any subject matter or the scope of the claims, but rather as a description of the features of particular embodiments of a particular technology. Certain features described in the context of separate embodiments of this patent document may also be implemented in combination in a single embodiment. Conversely, the various functions described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Additionally, although the above features may be described as acting in certain combinations, and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0402] Similarly, although the operations are described in a particular order in the figures, this should not be construed as requiring that such operations be performed in the particular order shown or in sequence to obtain the desired result, or that all of the illustrated operations be performed. Additionally, the separation of various system components in the embodiments of this patent document should not be construed as required in all embodiments.
[0403] Only some implementations and examples are described, and other implementations, enhancements, and variations may be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: Performing a conversion between a video including one or more pictures and a bitstream of the video according to format rules, wherein the format rules specify that, in response to satisfying one or more conditions, a syntax element indicating whether a first syntax structure providing level information and a second syntax structure providing decoded picture buffer information are present in a sequence parameter set is set to be equal to 1 to indicate that the first syntax structure and the second syntax structure are present in the sequence parameter set, wherein the syntax element being equal to 1 further specifies that a third syntax structure providing general timing information related to hypothetical reference decoder parameters and a fourth syntax structure providing output layer set timing information related to hypothetical reference decoder parameters are allowed to be present in the sequence parameter set.
2. The method according to claim 1, wherein, The one or more conditions are related to a video parameter set identifier referred to by the sequence parameter set.
3. The method according to claim 2, wherein, The one or more conditions are further related to whether there is an output layer set including only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value.
4. The method according to claim 1, wherein The one or more conditions include: 1) the video parameter set identifier referred to by the sequence parameter set is greater than 0, and there is an output layer set including only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value; or 2) the video parameter set identifier is equal to 0.
5. The method according to claim 1, wherein, The third syntax structure corresponds to the general_timing_hrd_parameters() syntax structure, and the fourth syntax structure corresponds to the ols_timing_hrd_parameters() syntax structure.
6. The method according to claim 1, wherein, The syntax element corresponds to the sps_ptl_dpb_hrd_params_present_flag, the first syntax structure corresponds to the profile_tier_level() syntax structure, and the second syntax structure corresponds to the dpb_parameters() syntax structure.
7. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
8. The method according to claim 1, wherein The conversion includes decoding the video from the bitstream.
9. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform a conversion between a video including one or more pictures and a bitstream of the video according to format rules, Among them, The format rules specify that, in response to satisfying one or more conditions, a syntax element indicating whether a first syntax structure providing level information and a second syntax structure providing decoded picture buffer information are present in a sequence parameter set is set to be equal to 1 to indicate that the first syntax structure and the second syntax structure are present in the sequence parameter set, wherein the syntax element being equal to 1 further specifies that a third syntax structure providing general timing information related to hypothetical reference decoder parameters and a fourth syntax structure providing output layer set timing information related to hypothetical reference decoder parameters are allowed to be present in the sequence parameter set.
10. The device according to claim 9, wherein, The one or more conditions are related to a video parameter set identifier referenced by the sequence parameter set.
11. The apparatus according to claim 10, wherein, The one or more conditions are further related to whether there is an output layer set that includes only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value.
12. The apparatus according to claim 9, wherein, The one or more conditions include: 1) the video parameter set identifier referenced by the sequence parameter set is greater than 0, and there is an output layer set that includes only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value; or 2) the video parameter set identifier is equal to 0.
13. The device according to claim 9, wherein The third syntax structure corresponds to the general_timing_hrd_parameters() syntax structure, and the fourth syntax structure corresponds to the ols_timing_hrd_parameters() syntax structure.
14. The apparatus according to claim 9, wherein, The syntax element corresponds to sps_ptl_dpb_hrd_params_present_flag, the first syntax structure corresponds to the profile_tier_level() syntax structure, and the second syntax structure corresponds to the dpb_parameters() syntax structure.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including one or more pictures and a bitstream of the video according to format rules, Among them, wherein the format rules specify that, in response to satisfying one or more conditions, a syntax element indicating whether a first syntax structure providing level information and a second syntax structure providing decoded picture buffer information are present in a sequence parameter set is set to be equal to 1 to indicate that the first syntax structure and the second syntax structure are present in the sequence parameter set, wherein the syntax element being equal to 1 further specifies that a third syntax structure providing general timing information related to hypothetical reference decoder parameters and a fourth syntax structure providing output layer set timing information related to hypothetical reference decoder parameters are allowed to be present in the sequence parameter set.
16. The medium according to claim 15, wherein, The one or more conditions include: 1) the video parameter set identifier referenced by the sequence parameter set is greater than 0, and there is an output layer set that includes only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value; or 2) the video parameter set identifier is equal to 0.
17. A method for storing a bitstream of a video, comprising: generating the bitstream of the video including one or more pictures according to format rules, storing the bitstream in a non-transitory computer-readable recording medium, wherein the format rules specify that, in response to satisfying one or more conditions, a syntax element indicating whether a first syntax structure providing level information and a second syntax structure providing decoded picture buffer information are present in a sequence parameter set is set to be equal to 1 to indicate that the first syntax structure and the second syntax structure are present in the sequence parameter set, Wherein, when the syntax element is equal to 1, it is also specified that a third syntax structure for providing general timing information related to the hypothesized reference decoder parameters and a fourth syntax structure for providing output layer set timing information related to the hypothesized reference decoder parameters are allowed to exist in the sequence parameter set.
18. The method according to claim 17, wherein The one or more conditions include: 1) the video parameter set identifier referred to by the sequence parameter set is greater than 0, and there is an output layer set that includes only one layer having a network abstraction layer (NAL) unit header layer identifier equal to a specific value; or 2) the video parameter set identifier is equal to 0.
19. A video processing apparatus, comprising a processor configured to implement the method according to any one of claims 7 to 8.
20. A method for storing a bitstream of video, comprising the method according to any one of claims 2 to 3, 5 to 8, and further comprising storing the bitstream into a non-transitory computer-readable recording medium.
21. A computer-readable medium storing program code, which when executed causes a processor to implement the method according to any one of claims 2 to 3, 5 to 8.
22. A video processing apparatus for storing a bitstream, wherein the video processing apparatus includes a processor configured to implement the method according to any one of claims 1 to 8 and to store the bitstream.
Citation Information
Patent Citations
Image decoding device, image decoding method, image coding device, and image coding method
US20160212437A1
Method and apparatus for decoded picture buffer management in video coding system using intra block copy
US20180295382A1