Strip type in picture

CN115362479BActive Publication Date: 2026-09-18DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180026190.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2021-03-29
Publication Date
2026-09-18
Estimated Expiration
2041-03-29

Smart Images

  • Figure CN115362479B_ABST
    Figure CN115362479B_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatuses for video processing are described. Video processing can include video encoding, video decoding, or video transcoding. An example method of video processing includes performing a conversion between a video and a bitstream of the video according to a rule. The rule specifies that one or more syntax elements in one or more video units are used to indicate whether a slice of a specified coding type is allowed for the conversion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] Pursuant to applicable patent law and / or in accordance with the rules of the Paris Convention, this application timely claims priority and benefit to U.S. Provisional Application No. 63 / 002,148, filed March 30, 2020. For all purposes required by law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process the codec representation of video using control information useful for decoding the codec representation.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a video bitstream according to rules. These rules specify that one or more syntax elements in one or more video units are used to indicate whether a stripe of a codec type is allowed for the conversion.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between video frames and a video bitstream according to rules. These rules specify one or more syntax elements within one or more video units to indicate whether different stripe types are permitted to be mixed within the video frames for the conversion.

[0008] In another example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more codec layer video sequences and a codec representation of the video, the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule specifying that a syntax structure is included at the sequence parameter set level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in a reference codec layer video sequence.

[0009] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more codec layer video sequences and a codec representation of the video, the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule specifying that a syntax structure is included at the picture parameter set level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in a reference picture.

[0010] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more codec layer video sequences and a codec representation of the video, the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule specifying that a syntax structure is included at the picture header level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in the picture.

[0011] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video and a codec representation of a video comprising one or more video images containing one or more stripes, wherein the conversion conforms to a rule specifying whether the stripe type of the stripe is included in the codec representation depends on a parameter set or the value of a syntax element in the image header of the image containing the stripe.

[0012] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video images containing one or more video stripes and a codec representation of the video, wherein the codec representation conforms to a format rule specifying whether the codec for an image allows for prediction of codec stripes (P-stripes) and bidirectional codec stripes (B-stripes).

[0013] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0014] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0015] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0016] These and other features will be described in this document. Attached Figure Description

[0017] Figure 1This is a block diagram of an example video processing system.

[0018] Figure 2 This is a block diagram of a video processing device.

[0019] Figure 3 This is a flowchart of an example method for video processing.

[0020] Figure 4 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0021] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0022] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0023] Figure 7 This is a flowchart representation of a video processing method based on this technology.

[0024] Figure 8 This is a flowchart representation of another video processing method according to the present technology. Detailed Implementation

[0025] Chapter headings are used in this document for ease of understanding, and the applicability of the technologies and embodiments disclosed in each chapter is not limited to that chapter alone. Furthermore, the use of H.266 technical terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0026] 1. Overview

[0027] This document relates to video codec technology. Specifically, it concerns improvements to signaling for permitted stripe types and related codec tools applicable only to bidirectional predictive stripes. These ideas can be applied, alone or in various combinations, to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Multi-Functional Video Codec (VVC) under development.

[0028] 2. Abbreviation

[0029] ALF (Adaptive Loop Filter)

[0030] APS (Adaptation Parameter Set)

[0031] AU (Access Unit)

[0032] AUD (Access Unit Delimiter)

[0033] AVC (Advanced Video Coding)

[0034] CLVS (Coded Layer Video Sequence)

[0035] CPB (Coded Picture Buffer) is a codec image buffer.

[0036] CRA (Clean Random Access)

[0037] CTU (Coding Tree Unit)

[0038] CVS (Coded Video Sequence) is a video sequence encoding and decoding mechanism.

[0039] DCI (Decoding Capability Information)

[0040] DPB (Decoded Picture Buffer)

[0041] DU (Decoding Unit)

[0042] EOB (End Of Bitstream) - End of Bitstream

[0043] EOS (End Of Sequence)

[0044] GDR (Gradual Decoding Refresh)

[0045] HEVC (High Efficiency Video Coding)

[0046] HRD (Hypothetical Reference Decoder)

[0047] IDR (Instantaneous Decoding Refresh)

[0048] JEM (Joint Exploration Model)

[0049] LMCS (Luma Mapping with Chroma Scaling)

[0050] MCTS (Motion-Constrained Tile Sets)

[0051] NAL (Network Abstraction Layer)

[0052] OLS (Output Layer Set)

[0053] PH (Picture Header)

[0054] PPS (Picture Parameter Set)

[0055] PTL (Profile, Tier and Level)

[0056] PU (Picture Unit)

[0057] RADL (Random Access Decodable Leading (Picture))

[0058] RAP (Random Access Point)

[0059] RASL (Random Access Skipped Leading (Picture))

[0060] RBSP (Raw Byte Sequence Payload)

[0061] RPL (Reference Picture List)

[0062] SAO (Sample Adaptive Offset)

[0063] SEI (Supplemental Enhancement Information)

[0064] SPS (Sequence Parameter Set)

[0065] STSA (Step-wise Temporal Sublayer Access)

[0066] SVC (Scalable Video Coding)

[0067] VCL (Video Coding Layer)

[0068] VPS (Video Parameter Set)

[0069] VTM (VVC Test Model)

[0070] VUI (Video Usability Information)

[0071] VVC (Versatile Video Coding) is a multi-functional video codec.

[0072] 3. Preliminary Discussion

[0073] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of new codec standards is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Thanks to continuous efforts in VVC standardization, the new codec technology has been adopted into the VVC standard at every JVET meeting.

[0074] 3.1. Parameter Set

[0075] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but it is included in the latest VVC draft text.

[0076] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.

[0077] A VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0078] APS is introduced to carry such image-level or stripe-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.

[0079] 3.1.1. Video Parameter Set (VPS)

[0080] Example syntax tables and semantics for multiple syntax elements are defined as follows:

[0081] 7.3.2.2 Video Parameter Set (RBSP) Syntax

[0082]

[0083] 3.1.2. Sequence Parameter Set (SPS)

[0084] Example syntax tables and semantics for multiple syntax elements are defined as follows:

[0085] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0086]

[0087]

[0088]

[0089] 3.1.3. General Constraint Marks

[0090] 7.3.3.2 General Constraint Information Syntax

[0091]

[0092]

[0093] A value of 1 for `no_bdof_constraint_flag` specifies that `sps_bdof_enabled_flag` should be 0. A value of 0 for `no_bdof_constraint_flag` does not impose such a constraint.

[0094] A value of 1 for `no_dmvr_constraint_flag` specifies that `sps_dmvr_enabled_flag` should be 0. A value of 0 for `no_dmvr_constraint_flag` does not impose this constraint.

[0095] A value of 1 for `no_bcw_constraint_flag` specifies that `sps_bcw_enabled_flag` should be 0. A value of 0 for `no_bcw_constraint_flag` does not impose this constraint.

[0096] A value of 1 for `no_ciip_constraint_flag` specifies that `sps_ciip_enabled_flag` should be 0. A value of 0 for `no_cipp_constraint_flag` does not impose such a constraint.

[0097] A value of 1 for no_gpm_constraint_flag indicates that sps_gpm_enabled_flag should be 0. A value of 0 for no_gpm_constraint_flag does not impose such a constraint.

[0098] 3.1.4. Image Parameter Set (PPS)

[0099] Example syntax tables and semantics for multiple syntax elements are defined as follows:

[0100] 7.3.2.4 Image Parameter Set (RBSP) Syntax

[0101]

[0102] The increment of 1 in num_ref_idx_default_active_minus1[i] specifies the inferred value of the variable NumRefIdxActive[0] for P-band or B-band where num_ref_idx_active_override_flag is equal to 0 when i equals 0, and specifies the inferred value of NumRefIdxActive[1] for B-band where num_ref_idx_active_override_flag is equal to 0 when i equals 1. The value of num_ref_idx_default_active_minus1[i] should be in the range of 0 to 14 (inclusive).

[0103] A value of 0 for pps_weighted_bipred_flag indicates that explicit weighted predictions are not applied to the B-strips of the reference PPS. A value of 1 for pps_weighted_bipred_flag indicates that explicit weighted predictions are applied to the B-strips of the reference PPS. When pps_weighted_bipred_flag is 0, the value of pps_weighted_bipred_flag should be 0.

[0104] 3.1.5. DPB Parameter Syntax

[0105] The syntax tables and semantics of multiple syntax elements are defined as follows:

[0106] 7.3.4 DPB Parameter Syntax

[0107]

[0108] 7.4.5 DPB Parameter Semantics

[0109] The dpb_parameters() syntax structure provides information on the DPB size, maximum number of image reorderings, and maximum latency for one or more OLS.

[0110] When the dpb_parameters() syntax structure is included in a VPS, the OLS to which the dpb_parameters() syntax structure applies is specified by the VPS. When the dpb_parameters() syntax structure is included in an SPS, it applies only to the OLS of the lowest layer among the layers that serve as the reference SPS, and that lowest layer is an independent layer.

[0111] The increment of 1 in `max_dec_pic_buffering_minus1[i]` specifies the maximum required size of the DPB in units of the image storage buffer when `Htid` equals `i`. The value of `max_dec_pic_buffering_minus1[i]` should be in the range of 0 to `MaxDpbSize - 1` (inclusive), where `MaxDpbSize` is as specified in Clause A.4.2. When `i` is greater than 0, `max_dec_pic_buffering_minus1[i]` should be greater than or equal to `max_dec_pic_buffering_minus1[i - 1]`. When there is no max_dec_pic_buffering_minus1[i] for i in the range from 0 to maxSubLayersMinus1 - 1 (inclusive), it is inferred to be equal to max_dec_pic_buffering_minus1[maxSubLayersMinus1] since subLayerInfoFlag is equal to 0.

[0112] `max_num_reorder_pics[i]` specifies the maximum allowed number of pictures in the OLS that can precede any picture in the OLS in decoding order and follow any picture in output order when `Htid` equals `i`. The value of `max_num_reorder_pics[i]` should be in the range of 0 to `max_dec_pic_buffering_minus1[i]` (inclusive). When `i` is greater than 0, `max_num_reorder_pics[i]` should be greater than or equal to `max_num_reorder_pics[i-1]`. When `i` is in the range of 0 to `maxSubLayersMinus1-1` (inclusive), if `max_num_reorder_pics[i]` does not exist, it is inferred to be equal to `max_num_reorder_pics[maxSubLayersMinus1]` since `subLayerInfoFlag` is equal to 0.

[0113] max_latency_increase_plus1[i] is not equal to 0 and is used to calculate the value of MaxLatencyPictures[i], which specifies the maximum number of pictures in the OLS that can precede any picture in the OLS in the output order and follow that picture in the decoding order when Htid equals i.

[0114] When max_latency_increase_plus1[i] is not equal to 0, the value of MaxLatencyPictures[i] is specified as follows:

[0115] MaxLatencyPictures[ i ] = max_num_reorder_pics[ i ] + max_latency_increase_plus1[ i ] - 1 (7-110)

[0116] When max_latency_increase_plus1[i] equals 0, the corresponding limit is not expressed.

[0117] The value of max_latency_increase_plus1[i] should be between 0 and 2. 32 - The range of 2 (inclusive of 0 and 2) 32 -2). When there is no max_latency_increase_plus1[i] in the range of 0 to maxSubLayersMinus1 - 1 (inclusive), it is inferred to be equal to max_latency_increase_plus1[maxSubLayersMinus1] since subLayerInfoFlag is equal to 0.

[0118] 3.2. Picture header (PH) and strip header (SH) in VVC

[0119] Similar to HEVC, the stripe header in VVC conveys information about a specific stripe. This includes the stripe address, stripe type, stripe QP, picture order count (POC) least significant bit (LSB), RPS and RPL information, weighted prediction parameters, loop filter parameters, slice and WPP entry offsets, etc.

[0120] VVC introduces a Picture Header (PH), which contains header parameters for a specific picture. Each picture must have one or only one PH. The PH essentially carries the parameters that would be present in the stripe header if no PH were introduced, but each parameter has the same value for all stripes of the picture. These include IRAP / GDR picture indication, inter-frame / intra-frame stripe allow flags, POCLSB and optionally POC MSB, information about RPL, deblocking, SAO, ALF, QP increment and weighted prediction, codec block segmentation information, virtual boundaries, juxtaposed picture information, etc. Often, each picture in the entire picture sequence contains only one stripe. To allow for situations where each picture does not have at least two NAL units, a PH syntax structure can be included in the PH NAL unit or the stripe header.

[0121] In VVC, signaling information about the juxtaposed images is provided in the image header or strip header for temporal motion vector prediction.

[0122] 3.2.1. Image header (PH)

[0123] The syntax tables and semantics of multiple syntax elements are defined as follows:

[0124] 7.3.2.7 Image Header Structure Syntax

[0125]

[0126]

[0127] 3.2.2. Strip header (SH)

[0128] The syntax tables and semantics of multiple syntax elements are defined as follows:

[0129] 7.3.7.1 General Strip Header Syntax

[0130]

[0131]

[0132]

[0133]

[0134] slice_type specifies the encoding / decoding type of the slice according to Table 9.

[0135] Table 9 – Name association with slice_type

[0136]

[0137] When it does not exist, the value of slice_type is inferred to be equal to 2.

[0138] When ph_intra_slice_allowed_flag equals 0, the value of slice_type should be 0 or 1. When nal_unit_type is within the range of IDR_W_RADL to CRA_NUT (inclusive of IDR_W_RADL and CRA_NUT), and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, slice_type should be 2.

[0139] 4. The technical problem solved by the disclosed technical solution

[0140] In some embodiments, two PH syntax elements related to the allowed slice types are included, such as ph_inter_slice_allowed_flag and ph_intra_slice_allowed_flag, as listed in the picture header structure syntax. Using these two flags, signaling is only sent to syntax elements related to I-slice encoding / decoding when ph_intra_slice_allowed_flag is true, and signaling is only sent to syntax elements related to inter-frame slice encoding / decoding when ph_inter_slice_allowed_flag is true. However, when ph_inter_slice_allowed_flag equals 1, the decoder is unaware whether the picture contains B-slices. Some applications (e.g., online games, video conferencing, video surveillance) typically use only P-slices and I-slices. Therefore, with an indication of whether B-slices are allowed, the decoder for such applications will be able to selectively request / use only bitstreams that do not include B-slices; furthermore, this indication can be used to avoid transmitting multiple unnecessary parameters.

[0141] 5. List of technical solutions

[0142] To address the above and other issues, the following summarized methods are presented. These items should be considered as examples for explaining general concepts, and not interpreted in a narrow way. Furthermore, these items can be applied individually or combined in any way.

[0143] One or more syntax elements can be added to the parameter set (e.g., SPS, VPS, PPS) and / or the general constraint information syntax and / or PH to indicate whether X (e.g., B or P) stripes are allowed.

[0144] In SPS and general constraint information syntax

[0145] 1. In SPS, add syntax elements (e.g., sps_X_slice_allowed_flag) to specify whether CLVS can contain one or more X slices; or to specify whether CLVS does not contain any X slices.

[0146] 1) In one example, add a first syntax element (e.g., sps_b_slice_allowed_flag), where sps_b_slice_allowed_flag equals 1 to specify that CLVS can contain one or more B slices, and sps_b_slice_allowed_flag equals 0 to specify that CLVS does not contain B slices.

[0147] i. Alternatively, the signaling notification and / or semantics and / or inference of one or more syntax elements in the SPS can be modified such that they are signaled only if the first syntax element satisfies certain conditions.

[0148] a. In one example, one or more syntax elements are syntax elements for enabling codec tools that require more than one prediction signaling, such as bidirectional prediction or mixed intra-frame and inter-frame coding and decoding, or prediction using linear / non-linear weighted prediction from multiple prediction blocks.

[0149] b. In one example, one or more syntax elements may include, but are not limited to:

[0150] a)sps_weighted_bipred_flag

[0151] b)sps_bdof_enabled_flag

[0152] c)sps_smvd_enabled_flag

[0153] d)sps_dmvr_enabled_flag

[0154] e)sps_bcw_enabled_flag

[0155] f)sps_ciip_enabled_flag

[0156] g)sps_gpm_enabled_flag

[0157] c. In one example, one or more syntax elements may be signaled only if the first syntax element specifies that the CLVS can contain one or more B stripes. Otherwise, the signaling is skipped, and the value of the syntax element is inferred.

[0158] d. In one example, when sps_b_slice_allowed_flag equals 0, no signaling is sent to the syntax elements sps_weighted_bipred_flag, sps_bdof_enabled_flag, sps_smvd_enabled_flag, sps_dmvr_enabled_flag, sps_bcw_enabled_flag, sps_ciip_enabled_flag, and sps_gpm_enabled_flag, and their values ​​are inferred.

[0159] a) In one example, they are all inferred to be 0 when they do not exist.

[0160] ii. Alternatively, the second syntax element, such as no_b_slice_contraint_flag, can be signaled in the general constraint information syntax to indicate whether the first syntax element should be equal to 0.

[0161] a. In one example, the semantics of no_b_slice_contraint_flag are defined as follows:

[0162] A value of 1 specifies that `sps_b_slice_allowed_flag` should be equal to 0. A value of 0 for `no_b_slice_constraint_flag` does not impose such a constraint.

[0163] iii. Alternatively, if the first syntax element specifies that CLVS does not contain B stripes, then one or more syntax elements of the signaling notification in the general constraint information syntax must be equal to 1.

[0164] a. In one example, one or more syntax elements may include, but are not limited to:

[0165] a)no_bcw_constraint_flag

[0166] b) no_ciip_constraint_flag

[0167] c) no_gpm_constraint_flag

[0168] d)no_bdof_constraint_flag

[0169] e)no_dmvr_constraint_flag

[0170] iv. Alternatively, the signaling notification and semantics of one or more syntax elements in dpb_parameters() can be modified so that they are signaled only if the first syntax element satisfies certain conditions.

[0171] a. In one example, one or more syntax elements may include, but are not limited to:

[0172] a)max_num_reorder_pics

[0173] b. In one example, when the first syntax element indicates that B stripes are not allowed, max_num_reorder_pics is not signaled and is inferred to be 0.

[0174] 2) In one example, a second syntax element is added (e.g., sps_p_slice_allowed_flag), where sps_p_slice_allowed_flag equals 1 to specify that CLVS can contain one or more P slices, and sps_p_slice_allowed_flag equals 0 to specify that CLVS does not contain P slices.

[0175] i. Alternatively, the sub-item symbol mentioned in section 1.1) can be applied by replacing sps_b_slice_allowed_flag with (!sps_p_slice_allowed_flag) or (!sps_p_slice_allowed_flag && ph_inter_slice_allowed_flag).

[0176] In PPS

[0177] 2. In PPS, add syntax elements (e.g., sps_X_slice_allowed_flag) to specify whether an image referencing the current PPS can contain one or more X slices; or to specify whether an image referencing the current PPS does not contain any X slices.

[0178] 1) In one example, add a first syntax element (e.g., pps_b_slice_allowed_flag), where pps_b_slice_allowed_flag equals 1 to specify that the image referenced by the current PPS can contain one or more B slices, and pps_b_slice_allowed_flag equals 0 to specify that the image referenced by the current PPS does not contain B slices.

[0179] 2) In one example, the first syntax element in bullet 1 (e.g., sps_b_slice_allowed_flag) and the first syntax element in bullet 2.1) (e.g., pps_b_slice_allowed_flag) should have the same constraints.

[0180] 3) Alternatively, the signaling notification and / or semantics and / or inference of one or more syntax elements of the signaling notification in the PPS can be modified based on the first syntax element.

[0181] i. In one example, one or more syntax elements are syntax elements for enabling codec tools that require more than one prediction signaling, such as bidirectional prediction or mixed intra-frame and inter-frame coding and decoding, or using linear / non-linear weighted prediction from multiple prediction blocks.

[0182] ii. In one example, whether signaling informs one or more syntax elements that a B-band is permitted under the condition that the inspection instruction of the first syntax element allows it.

[0183] a. Alternatively, if no signaling notification is received, the value can be inferred to be, for example, 0.

[0184] iii. In one example, one or more syntax elements may include, but are not limited to:

[0185] a.pps_weighted_bipred_flag

[0186] b.num_ref_idx_default_active_minus1[ 1 ]

[0187] In pH

[0188] 3. In PH, add syntax elements (e.g., ph_X_slice_allowed_flag) to specify whether an image can contain one or more X slices; or to specify whether an image does not contain any X slices.

[0189] 1) In one example, add a first syntax element (e.g., ph_b_slice_allowed_flag), where ph_b_slice_allowed_flag equals 1 to specify that the image can contain one or more B slices, and ph_b_slice_allowed_flag equals 0 to specify that the image does not contain B slices.

[0190] i. Alternatively, the first syntax element can be conditionally signaled (e.g., ph_b_slice_allowed_flag).

[0191] a. In one example, when sps_b_slice_allowed_flag and / or pps_b_slice_allowed_flag are true, ph_b_slice_allowed_flag can be signaled.

[0192] b. In one example, when sps_b_slice_allowed_flag and / or pps_b_slice_allowed_flag are false, ph_b_slice_allowed_flag may not be signaled and may be inferred as false.

[0193] 2) Alternatively, the signaling notification and / or semantics and / or inference of one or more syntax elements in the PH can be modified based on the first syntax element.

[0194] i. In one example, one or more syntax elements are syntax elements for enabling codec tools that require more than one prediction signaling, such as bidirectional prediction or mixed intra-frame and inter-frame coding and decoding, or using linear / non-linear weighted prediction from multiple prediction blocks.

[0195] ii. In one example, one or more syntax elements may include, but are not limited to:

[0196] a)ph_collocated_from_l0_flag

[0197] b)mvd_l1_zero_flag

[0198] c)ph_disable_bdof_flag

[0199] d)ph_disable_dmvr_flag

[0200] e)num_l1_weights

[0201] iii. In one example, one or more syntax elements may be signaled only if the first syntax element specifies that the picture can contain one or more B stripes. Otherwise, the signaling is skipped, and the value of the syntax element is inferred.

[0202] a) Alternatively, whether signaling notifies one or more syntax elements may depend on the first syntax element in bullet points 1.1 and 2.1, such as (sps_b_slice_allowed_flag && ph_b_slice_allowed_flag).

[0203] b) Only if (sps_bdof_pic_present_flag) Only when ) is true can ph_disable_bdof_flag be notified via signaling.

[0204] c) Only if (sps_dmvr_pic_present_flag) Only when ) is true can ph_disable_dmvr_flag be notified via signaling.

[0205] iv. In one example, when ph_b_slice_allowed_flag equals 0, no signaling is sent to mvd_l1_zero_flag, and its value is inferred to be 1.

[0206] v. In one example, the inference of one or more syntactic elements depends on the value of the first syntactic element.

[0207] a) In one example, for ph_disable_bdof_flag, the following applies:

[0208] - If sps_bdof_enabled_flag equals 1 If so, the value of ph_disable_bdof_flag is inferred to be equal to 0.

[0209] - Otherwise (sps_bdof_enabled_flag equals ...) The value of ph_disable_bdof_flag is inferred to be equal to 1.

[0210] b) In one example, for ph_disable_dmvr_flag, the following applies:

[0211] – If sps_dmvr_enabled_flag equals 1 If so, the value of ph_disable_dmvr_flag is inferred to be equal to 0.

[0212] - Otherwise (sps_dmvr_enabled_flag equals ...) The value of ph_disable_dmvr_flag is inferred to be equal to 1.

[0213] c) In one example, when both ph_temporal_mvp_enabled_flag and rpl_info_in_ph_flag are equal to 1 and ph_b_slice_allowed_flag is equal to 0, the value of ph_collocated_from_l0_flag is inferred to be equal to 1.

[0214] d) In one example, when ph_b_slice_allowed_flag equals 0, no signaling is sent to num_l1_weights, and its value is inferred to be 0. Therefore, no signaling is sent to the weighted prediction parameters of reference image list 1 in the PH or SH of the image.

[0215] 4. Whether signaling notification of stripe type and / or inference of stripe type can rely on syntax elements related to the allowed stripe types in the parameter set and / or picture header.

[0216] 1) In one example, based on "if (ph_inter_slice_allowed_flag) ")" to conditionally notify the stripe type of signaling.

[0217] 2) In one example, when there is no signaling notification for slice_type, the value of slice_type is inferred to be equal to (ph_inter_slice_allowed_flag ? 1: 2).

[0218] 3) When both ph_b_slice_allowed_flag and ph_intra_slice_allowed_flag are equal to 0, no signaling is sent to the syntax element slice_type, and the value is inferred to be equal to 1.

[0219] 5. Two flags can be added to indicate whether P and B are allowed respectively.

[0220] 1) In one example, the PH flag p_slices_allowed_flag (value 0 specifies that the image does not have P stripes) is added, and the PH flag b_slices_allowed_flag (value 0 specifies that the image does not have B stripes) may also be added.

[0221] 2) Alternatively, they may be conditionally signaled depending on whether inter-frame striping is applied.

[0222] One or more syntax elements can be added to the parameter set (e.g., SPS, VPS, PPS) and / or the general constraint information syntax and / or PH to indicate whether mixed stripe types are allowed, or to add constraints to disallow mixed inter-frame stripe types (e.g., B and P).

[0223] 6. It can be constrained that there should be no mixture of P-type and B-type bands within the image.

[0224] 1) In one example, it can be constrained that there should be no mixture of P-strip and B-strip types in the picture within any VVC bitstream (or any bitstream encoded and decoded using another video codec).

[0225] 2) In one example, a signaling syntax element (e.g., a flag) can be used in the bitstream (e.g., in a parameter set or DCI NAL unit), and the syntax element (e.g., a flag) equal to X (e.g., 1) specifies that there should not be a mixture of P-strip type and B-strip type within the picture in the bitstream.

[0226] 3) In one example, a signaling syntax element (e.g., a flag) can be used in SPS, and the syntax element (e.g., a flag) equal to X (e.g., 1) specifies that there should not be a mixture of P-strip type and B-strip type within the picture in CLVS.

[0227] 4) In one example, a signaling syntax element (e.g., a flag) can be used in PPS or PH, and the syntax element (e.g., a flag) equal to X (e.g., 1) specifies that there should not be a mixture of P-band type and B-band type within the picture.

[0228] 5) In one example, signaling notification syntax elements (e.g., flags, such as p_slices_allowed_flag) can be used in SPS / PPS / PH to indicate whether P-slices are allowed.

[0229] i. Alternatively, a syntax element (e.g., a flag) equal to X (e.g., 1) specifies that there should be no P stripes within the image.

[0230] ii. Additionally, in one example, if p_slices_allowed_flag is equal to 0, then the value of ph_collocated_from_l0_flag is not constrained; otherwise, the value of ph_collocated_from_l0_flag is required to be equal to 1.

[0231] 6) Alternatively, the syntax elements mentioned in the above sub-item symbols may be conditionally signaled.

[0232] iii. In one example, the syntax element for signaling notification in PH can be under the condition of "allow inter-slice" check (e.g., if (ph_inter_slice_allowed_flag)).

[0233] 7) Additionally, in one example, when a particular image contains only P-bands, the value of ph_collocated_from_l0_flag is required to be equal to 1.

[0234] iv. Alternatively, when a particular image contains only P-bands, ph_collocated_from_l0_flag may not be signaled and may be inferred to be equal to 1.

[0235] 6. Examples

[0236] The following are some example embodiments of aspects of the invention summarized in Section 5 above, which can be applied to the VVC specification. The most relevant parts that have been added or modified are... Underlined sections are used, and some deleted sections are indicated by [[]]. Note that the following examples can be combined.

[0237] 6.1. First Implementation Example of SPS-Related Changes

[0238] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0239]

[0240]

[0241] ...

[0243]

[0244] A value of 1 for `sps_weighted_bipred_flag` indicates that explicit weighted predictions can be applied to the B-bands of the reference SPS. A value of 0 for `sps_weighted_bipred_flag` indicates that explicit weighted predictions should not be applied to the B-bands of the reference SPS.

[0245] A value of 0 for sps_bdof_enabled_flag indicates that bidirectional optical flow inter-frame prediction is disabled. A value of 1 for sps_bdof_enabled_flag indicates that bidirectional optical flow inter-frame prediction is enabled.

[0246] A value of 1 for sps_smvd_enabled_flag indicates that symmetric motion vector difference can be used in motion vector decoding. A value of 0 for sps_smvd_enabled_flag indicates that symmetric motion vector difference is not used in motion vector encoding and decoding.

[0247] A value of 1 for sps_dmvr_enabled_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement is enabled. A value of 0 for sps_dmvr_enabled_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement is disabled.

[0248] `sps_bcw_enabled_flag` specifies whether bidirectional prediction using CU weights can be used for inter-frame prediction. If `sps_bcw_enabled_flag` equals 0, the syntax should be constrained so that bidirectional prediction using CU weights is not used in CLVS, and `bcw_idx` does not exist in the CLVS codec unit syntax. Otherwise (`sps_bcw_enabled_flag` equals 1), bidirectional prediction using CU weights can be used in CLVS.

[0249] `sps_ciip_enabled_flag` specifies that the `ciip_flag` can exist in the codec unit syntax of the inter-frame codec unit. `sps_ciip_enabled_flag` equal to 0 specifies that the `ciip_flag` does not exist in the codec unit syntax of the inter-frame codec unit. ...

[0251] 6.2. Second Implementation Example of PPS-Related Changes

[0252] 7.3.2.4 Image Parameter Set (RBSP) Syntax

[0253]

[0254]

[0255] Increment 1 to num_ref_idx_default_active_minus1[i]. When i equals 0, specify the inferred value of the variable NumRefIdxActive[0] for P-band or B-band where num_ref_idx_active_override_flag equals 0. When i equals 1, specify the inferred value of NumRefIdxActive[1] for B-band where num_ref_idx_active_override_flag equals 0. The value of num_ref_idx_default_active_minus1[i] should be in the range of 0 to 14 (inclusive).

[0256] A value of 0 for pps_weighted_bipred_flag indicates that explicit weighted predictions are not applied to the B-strips of the reference PPS. A value of 1 for pps_weighted_bipred_flag indicates that explicit weighted predictions are applied to the B-strips of the reference PPS. When pps_weighted_bipred_flag is 0, the value of pps_weighted_bipred_flag should be 0.

[0257] 6.3. Third embodiment of pH and SH related changes

[0258] 7.3.2.7 Image Header Structure Syntax

[0259]

[0260] ...

[0262] Alternatively, the following can be applied:

[0263]

[0264] Alternatively, the following can be applied:

[0265]

[0266] A value of 0 for `ph_intra_slice_allowed_flag` indicates that all codec slices in the image have a slice_type of either 0 or 1. A value of 1 for `ph_intra_slice_allowed_flag` indicates that the image may or may not have one or more codec slices with a slice_type of 2. When none of these slices exist, the value of `ph_intra_slice_allowed_flag` is inferred to be 1.

[0267] ...

[0269] A value of 1 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 0. A value of 0 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 1.

[0270] Alternatively, the following applies:

[0271] A value of 1 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 0. A value of 0 for ph_collocated_from_l0_flag indicates that the juxtaposed images used for temporal motion vector prediction are derived from reference image list 1.

[0272] ph_collocated_ref_idx specifies the reference index of the juxtaposed images used for temporal motion vector prediction.

[0273] When ph_collocated_from_l0_flag equals 1, ph_collocated_ref_idx refers to the entry in reference image list 0, and the value of ph_collocated_ref_idx should be in the range of 0 to num_ref_entries[0][RplsIdx[0]] - 1 (inclusive of 0 and num_ref_entries[0][RplsIdx[0]] - 1).

[0274] When ph_collocated_from_l0_flag equals 0, ph_collocated_ref_idx refers to the entry in reference image list 1, and the value of ph_collocated_ref_idx should be in the range of 0 to num_ref_entries[1][RplsIdx[1]] - 1 (inclusive of 0 and num_ref_entries[1][RplsIdx[1]] - 1).

[0275] When it does not exist, the value of ph_collocated_ref_idx is inferred to be equal to 0. ...

[0277] A `mvd_l1_zero_flag` of 1 indicates that the `mvd_coding(x0, y0, 1)` syntax structure is not parsed, and for `compIdx = 0..1` and `cpIdx = 0..2`, `MvdL1[x0][y0][compIdx]` and `MvdCpL1[x0][y0][cpIdx][compIdx]` are set to 0. A `mvd_l1_zero_flag` of 0 indicates that the `mvd_coding(x0, y0, 1)` syntax structure is parsed. ...

[0279] A value of 1 for ph_disable_bdof_flag indicates that inter-frame bidirectional prediction based on bidirectional optical flow is disabled in the stripe associated with PH. A value of 0 for ph_disable_bdof_flag indicates that inter-frame bidirectional prediction based on bidirectional optical flow can be enabled or disabled in the stripe associated with PH.

[0280] The following applies when ph_disable_bdof_flag is not present:

[0281] - If sps_bdof_enabled_flag equals 1 If so, the value of ph_disable_bdof_flag is inferred to be equal to 0.

[0282] - Otherwise (sps_bdof_enabled_flag equals 0) The value of ph_disable_bdof_flag is inferred to be equal to 1.

[0283] A value of 1 for ph_disable_dmvr_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement is disabled in the slice associated with the PH. A value of 0 for ph_disable_dmvr_flag indicates that inter-frame bidirectional prediction based on decoder motion vector refinement can be enabled or disabled in the slice associated with the PH.

[0284] The following applies when ph_disable_dmvr_flag is not present:

[0285] – If sps_dmvr_enabled_flag equals 1 If so, the value of ph_disable_dmvr_flag is inferred to be equal to 0.

[0286] - Otherwise (sps_dmvr_enabled_flag equals 0) The value of ph_disable_dmvr_flag is inferred to be equal to 1. ...

[0288] 7.3.7.1 General Strip Header Syntax

[0289] ...

[0291] The slice_type specifies the encoding / decoding type of the slice according to Table 9.

[0292] Table 9 – Name association with slice_type

[0293]

[0294] When it does not exist, the value of slice_type is inferred to be equal to [[2]].

[0295] when When ph_intra_slice_allowed_flag equals 0, the value of slice_type should be 0 or 1. When nal_unit_type is within the range of IDR_W_RADL to CRA_NUT (inclusive) and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, slice_type should be 2.

[0296] Alternatively, the following applies:

[0297] when When ph_intra_slice_allowed_flag equals 0, the value of slice_type should be 0 or 1. When nal_unit_type is within the range of IDR_W_RADL to CRA_NUT (inclusive) and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, slice_type should be 2. ...

[0299] 7.3.7.2 Weighted Prediction Parameter Syntax

[0300]

[0301] 7.4.8.2 Weighted Prediction Parameter Semantics ...

[0303] num_l1_weights specifies the number of weights for signaling notifications for entries in reference image list 1 when both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1. The value of num_l1_weights should be in the range of 0 to Min(15, num_ref_entries[1][RplsIdx[1]]) (inclusive).

[0304] The variable NumWeightsL1 is derived as follows:

[0305]

[0306] Figure 1 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0307] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection as indicated by component 1906. The bitstream (or codec) representation of the video received at input 1902, whether stored or communicated, can be used by component 1908 to generate pixel values ​​or transmit as displayable video to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.

[0308] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0309] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple) 3602 can be configured to implement one or more methods described in this document. The memories (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in a hardware circuit system.

[0310] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0311] like Figure 4As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, and this source device 110 may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110, and this target device 120 may be referred to as a video decoding device.

[0312] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0313] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0314] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0315] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with an external display device.

[0316] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or additional standards.

[0317] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 4 The video encoder 114 in the system 100 shown.

[0318] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0319] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0320] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0321] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes, in Figure 5 The example is represented separately.

[0322] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0323] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on inter-frame prediction signaling and intra-frame prediction signaling. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).

[0324] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0325] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip.

[0326] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0327] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0328] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.

[0329] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0330] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0331] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0332] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge pattern signaling notification.

[0333] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0334] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0335] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0336] Transform unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0337] After the transform unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0338] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample points of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in buffer 213.

[0339] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0340] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0341] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 4 The video decoder 124 in the system 100 shown.

[0342] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0343] exist Figure 6 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform functions typically associated with the video encoder 200. Figure 5 The encoding process described is the opposite of the decoding process.

[0344] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-coded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.

[0345] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0346] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0347] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information used to decode the encoded video sequence.

[0348] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 304 performs inverse quantization, for example, dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0349] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on the display device.

[0350] The following is a list of preferred solutions for some embodiments.

[0351] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0352] 1. A video processing method (e.g., Figure 3The method 3000 shown includes: performing a conversion between a video and a codec representation of a video comprising one or more codec layer video sequences (3002), the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule that specifies a syntax structure at the sequence parameter set level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in a reference codec layer video sequence.

[0353] 2. The method according to Solution 1, wherein the codec type includes a bidirectional (B) codec type of predictive (P) codec type.

[0354] 3. The method according to any one of solutions 1-2, wherein the syntax structure includes a first syntax element specifying whether one or more B codec strips are included in the reference codec layer video sequence.

[0355] 4. The method according to Solution 3, wherein the formatting rules further specify that, based on the value of the first syntax element, additional syntax elements are conditionally included at the sequence parameter set level.

[0356] 5. The method according to Solution 4, wherein the additional syntax elements include syntax elements indicating the use of multiple prediction blocks with non-linear weighting to represent video blocks in the codec layer video sequence.

[0357] 6. The method according to any one of solutions 1-5, wherein the syntax structure includes a second syntax element specifying whether one or more P codec stripes are included in the reference codec layer video sequence.

[0358] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0359] 7. A video processing method, comprising: performing a conversion between a video comprising one or more codec layer video sequences and a codec representation of the video, the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule specifying that a syntax structure is included at the picture parameter set level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in a reference picture.

[0360] 8. The method according to Solution 7, wherein the codec type includes a bidirectional (B) codec type of predictive (P) codec type.

[0361] 9. The method according to any one of solutions 7-8, wherein the syntax structure includes a first syntax element specifying whether one or more B codec stripes are included in the reference picture.

[0362] 10. The method according to Solution 9, wherein the formatting rule further specifies that, based on the value of the first syntax element, additional syntax elements are conditionally included at the image parameter set level.

[0363] 11. The method according to solution 10, wherein the additional syntax element includes a syntax element indicating the use of multiple prediction blocks with non-linear weighting to represent video blocks in the codec layer video sequence.

[0364] 12. The method according to any one of solutions 6-11, wherein the syntax structure includes a second syntax element specifying whether one or more P codec stripes are included in the reference codec layer video sequence.

[0365] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).

[0366] 13. A video processing method, comprising: performing a conversion between a video comprising one or more codec layer video sequences and a codec representation of the video, the one or more codec layer video sequences comprising one or more video pictures containing one or more video stripes; wherein the codec representation conforms to a format rule specifying that a syntax structure is included at the picture header level, wherein the syntax structure indicates whether one or more stripes of a codec type are included in the picture.

[0367] 14. The method according to solution 13, wherein the codec type includes a bidirectional (B) codec type of predictive (P) codec type.

[0368] 15. The method according to any one of solutions 13-14, wherein the syntax structure includes a first syntax element specifying whether one or more B codec stripes or P codec stripes are included in the picture.

[0369] 16. The method according to solution 15, wherein the formatting rules further specify that, based on the value of the first syntax element, additional syntax elements are conditionally included in the image header.

[0370] 17. The method according to solution 16, wherein the additional syntax element includes a syntax element indicating the use of multiple prediction blocks with non-linear weighting to represent video blocks in the codec layer video sequence.

[0371] 18. The method according to any one of solutions 13-17, wherein the syntax structure includes a second syntax element specifying whether one or more P codec stripes are included in the reference codec layer video sequence.

[0372] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 4).

[0373] 19. A video processing method comprising: performing a conversion between a video and a codec representation of a video including one or more video pictures containing one or more stripes, wherein the conversion conforms to a rule specifying whether the stripe type of the stripe is included in the codec representation depends on the value of a syntax element in a parameter set or a picture header of the picture containing the stripe.

[0374] 20. The method according to solution 19, wherein the stripe type is conditionally signaled based on the value of the expression “if( ph_inter_slice_allowed_flag && ( ph_b_slice_allowed_flag | | ph_intra_slice_allowed_flag) )”.

[0375] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 5 and 6).

[0376] 21. A video processing method comprising: performing a conversion between a video comprising one or more video pictures containing one or more video stripes and a video codec representation, wherein the codec representation conforms to a format rule specifying whether the codec of a picture allows for prediction of codec stripes (P-stripes) and bidirectional codec stripes (B-stripes).

[0377] 22. The method according to solution 21, wherein the formatting rules specify that a first syntax element indicating the enablement of P-bands is included in the group header of the picture, and a second syntax element indicating the enablement of B-bands is included in the group header.

[0378] 23. The method according to any one of solutions 21-22, wherein the format rules specify, conditionally based on inter-frame stripe signaling notification, that a first syntax element indicating the enablement of P stripe is included in the packet header of the picture, and a second syntax element indicating the enablement of B stripe is included in the packet header.

[0379] 24. The method according to solution 21, wherein the format rules specify that the syntax elements in the encoding / decoding representation instruct P-stripes and B-stripes to be used mutually exclusively for encoding / decoding all remaining pictures of the picture and the reference syntax element.

[0380] 25. The method according to solution 24, wherein the syntax element is included in the decoding capability indicator field.

[0381] 26. The method according to solution 24, wherein the syntax elements are included in the sequence parameter set.

[0382] 27. The method according to any one of solutions 1 to 26, wherein the conversion includes encoding the video into a codec representation.

[0383] 28. The method according to any one of solutions 1 to 26, wherein the conversion includes decoding the codec representation to generate pixel values ​​of the video.

[0384] 29. A video decoding apparatus comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0385] 30. A video encoding apparatus comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0386] 31. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of solutions 1 to 28.

[0387] 32. A method, apparatus, or system described in this document.

[0388] Figure 7 This is a flowchart representation of a video processing method according to the present technology. Method 700 includes, in operation 710, performing a conversion between video and video bitstream according to a rule. The rule specifies that one or more syntax elements in one or more video units are used to indicate whether a stripe of a codec type is allowed for the conversion.

[0389] In some embodiments, specifying a codec type includes a bidirectional (B) codec type or a predictive (P) codec type. In some embodiments, one or more video units include a sequence parameter set. In some embodiments, one or more syntax elements include a first syntax element in the sequence parameter set. A first syntax element in the sequence parameter set equal to 1 indicates that the codec layer video sequence (CLVS) corresponding to the sequence parameter set includes one or more stripes of type B codec, and a first syntax element equal to 0 indicates that the CLVS excludes stripes of type B codec. In some embodiments, a first syntax element in the sequence parameter set equal to 0 indicates that the codec layer video sequence (CLVS) corresponding to the sequence parameter set includes one or more stripes of type B codec, and a first syntax element equal to 1 indicates that the CLVS does not include stripes of type B codec. In some embodiments, the use of a first set of syntax elements in the sequence parameter set is modified based on the first syntax element. In some embodiments, a first set of syntax elements is indicated for a transition in response to a first syntax element indicating that stripes of a specified codec type are allowed. In some embodiments, a first set of syntax elements is inferred for a transition in response to a first syntax element indicating that stripes of a specified codec type are not allowed.

[0390] In some embodiments, the first set of syntax elements indicates the need for the use of more than one predictive signaling codec tool. In some embodiments, the first set of syntax elements includes at least one of the following: a syntax flag indicating whether explicit weighted prediction is applicable to B-slices; a syntax flag indicating whether bidirectional optical flow inter-frame prediction is enabled; a syntax flag indicating whether symmetric motion vector difference is enabled; a syntax flag indicating whether inter-frame bidirectional prediction based on decoder motion vector refinement is enabled; a syntax flag indicating whether bidirectional prediction using codec unit weights is enabled; a syntax flag indicating whether combined inter-frame merge and intra-frame prediction is enabled; or a syntax flag indicating whether motion compensation based on geometric segmentation is enabled. In some embodiments, syntax elements in the general constraint information are used to indicate the value of the first syntax element. In some embodiments, syntax elements in the general constraint information are represented as no_b_slice_contraint_flag. no_b_slice_contraint_flag equal to 1 indicates that the first syntax element is equal to 0, and no_b_slice_contraint_flag equal to 0 does not specify the value of the first syntax element.

[0391] In some embodiments, in response to the first syntax element indicating that CLVS excludes stripes of type B codec, the second set of syntax elements in the general constraint information is equal to 1. In some embodiments, the second set of syntax elements in the general constraint information includes at least one of the following: (1) a first flag indicating whether sps_bcw_enabled_flag is equal to 0, wherein sps_bcw_enabled_flag specifies whether bidirectional prediction using codec unit weights is applicable to inter-frame prediction; (2) a second flag indicating whether sps_ciip_enabled_flag is equal to 0, wherein sps_ciip_enabled_flag specifies whether ciip_flag exists in the codec unit syntax of the inter-frame codec unit; (3) an indication that sps_gpm_ The third flag indicates whether enabled_flag is equal to 0, where sps_gpm_enabled_flag specifies whether to enable motion compensation based on geometric segmentation, (4) the fourth flag indicates whether sps_bdof_enabled_flag is equal to 0, where sps_bdof_enabled_flag specifies whether to disable bidirectional optical flow inter-frame prediction, or (5) the fifth flag indicates whether sps_dmvr_enabled_flag is equal to 0, where sps_dmv_enabled_flag specifies whether to enable bidirectional inter-frame prediction based on decoder motion vector refinement.

[0392] In some embodiments, a third set of syntax elements in the Decode Picture Buffer (DPB) parameters is conditionally indicated based on a first syntax element. In some embodiments, the third set of syntax elements includes at least `max_num_reorder_pics[i]`, which specifies the maximum allowed number of pictures in the output layer set, such that when `Htid` equals `i`, the pictures in the output layer set can precede any picture in the output layer set in decoding order and follow that picture in output order. If the first syntax element indicates that B codec type striping is not allowed, `max_num_reorder_pics[i]` is omitted and inferred to be 0.

[0393] In some embodiments, one or more syntax elements include additional syntax elements from the sequence parameter set. An additional syntax element equal to 1 indicates that the codec layer video sequence (CLVS) includes one or more stripes of type P codec. An additional syntax element equal to 0 indicates that the CLVS does not include stripes of type P codec. In some embodiments, one or more syntax elements include additional syntax elements from the sequence parameter set. An additional syntax element equal to 0 indicates that the codec layer video sequence (CLVS) includes one or more stripes of type P codec, and an additional syntax element equal to 1 indicates that the CLVS does not include stripes of type P codec.

[0394] In some embodiments, one or more video units include a picture parameter set. In some embodiments, one or more syntax elements include a second syntax element in the picture parameter set. A second syntax element in the picture parameter set equal to 1 indicates that the video picture of the reference picture parameter set includes one or more stripes of type B codec, and a second syntax element equal to 0 indicates that the video picture excludes stripes of type B codec. In some embodiments, the first syntax element and the second syntax element have the same value. In some embodiments, the use of the first set of syntax elements in the picture parameter set is modified based on the second syntax element. In some embodiments, when the second syntax element indicates that stripes of a specified codec type are allowed for the conversion, the first set of syntax elements is indicated in the picture parameter set for the conversion. In some embodiments, when the second syntax element indicates that stripes of a specified codec type are not allowed for the conversion, the first set of syntax elements is inferred for the conversion.

[0395] In some embodiments, the first set of syntax elements indicates the need for the use of more than one codec tool for predictive signaling. In some embodiments, the first set of syntax elements includes at least one of the following: specifying whether explicit weighted prediction is applied to the pps_weighted_bipred_flag of the B-strip of the reference picture parameter set, or specifying the inferred value of the variable NumRefIdxActive[1] of the B-strip where num_ref_idx_active_override_flag is equal to 0, num_ref_idx_default_active_minus1[1].

[0396] In some embodiments, one or more video units include a picture header. In some embodiments, one or more syntax elements include a third syntax element in the picture header. A third syntax element equal to 1 indicates that the picture corresponding to the picture header includes one or more stripes of type B codec, and a third syntax element equal to 0 indicates that the picture does not include stripes of type B codec. In some embodiments, the third syntax element is conditionally indicated based on a first syntax element and / or a second syntax element. In some embodiments, the third syntax element is indicated when the first or second syntax element indicates that a stripe for which a codec type can be specified is allowed for the conversion. In some embodiments, in response to the first or second syntax element indicating that a stripe for which a codec type cannot be specified is not allowed for the conversion, the third syntax element is omitted and is inferred to be false.

[0397] In some embodiments, the use of the first set of syntax elements in the image header is modified for transformation based on a third syntax element. In some embodiments, the use of the first set of syntax elements in the image header is also modified based on the first syntax element and / or the second syntax element. In some embodiments, the first set of syntax elements is indicated in the image header for transformation in response to a stripe indicating that the first syntax element, the second syntax element, and / or the third syntax element allows the specification of a codec type. In some embodiments, a syntax element in the first set of syntax elements in the image header is indicated if the syntax element indicating the presence of one of the syntax elements in the sequence parameter set is equal to 1 and the third syntax element is equal to 1. In some embodiments, a syntax element in the first set of syntax elements in the image header includes a ph_disable_bdof_flag indicating whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the stripe associated with the image header, or a ph_disable_dmvr_flag indicating whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the stripe associated with the image header. In some embodiments, a first set of syntax elements is inferred for conversion in response to a first syntax element, a second syntax element, and / or a third syntax element indicating that a stripe is not allowed to specify a codec type. In some embodiments, if the corresponding syntax element in the sequence parameter set is equal to 1 and the third syntax element is equal to 1, one syntax element in the first set of syntax elements in the image header is inferred as 0. In some embodiments, if the corresponding syntax element in the sequence parameter set is equal to 0 and the third syntax element is equal to 0, one syntax element in the first set of syntax elements in the image header is inferred as 1. In some embodiments, one syntax element in the first set of syntax elements in the image header includes a ph_disable_bdof_flag specifying whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the stripe associated with the image header, or a ph_disable_dmvr_flag specifying whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the stripe associated with the image header.

[0398] In some embodiments, the first set of syntax elements in the image header includes one or more syntax elements indicating the use of an encoding / decoding tool that requires more than one prediction signaling. In some embodiments, the first set of syntax elements in the image header includes at least one of the following: whether to derive the ph_collocated_from_l0_flag specifying the juxtaposed image for temporal motion vector prediction from reference image list 0; and whether to parse mvd_coding(x0, y0, ...). 1) The `mvd_l1_zero_flag` of the syntax structure specifies whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the slice associated with the picture header. The `ph_disable_bdof_flag` specifies whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the slice associated with the picture header. Alternatively, it specifies `num_l1_weights` as the number of weights for signaling notifications of entries in reference picture list 1 when both `pps_weighted_bipred_flag` and `wp_info_in_ph_flag` are equal to 1. Here, `pps_weighted_bipred_flag` specifies whether explicit weighted prediction is applied to the B slice of the reference picture parameter set, and `wp_info_in_ph_flag` specifies whether weighted prediction information exists in the picture header syntax structure and not in the slice header of the reference picture parameter set that does not contain a picture header syntax structure.

[0399] In some embodiments, the slice type is indicated or inferred based on one or more syntax elements in one or more video units. In some embodiments, the slice type is indicated at least based on ph_inter_slice_allowed_flag, ph_intra_slice_allowed_flag, and one or more syntax elements in one or more video units. In some embodiments, when no slice type is indicated, the slice type is inferred to be equal to (ph_inter_slice_allowed_flag ? 1 : 2). In some embodiments, when ph_intra_slice_allowed_flag is equal to 0 and one or more syntax elements in one or more video units indicate that slices of codec type B are not allowed, the slice type is inferred to be 1. In some embodiments, one or more syntax elements are conditionally indicated based on whether inter-frame slices are applied to the transition.

[0400] Figure 8This is a flowchart representation of a video processing method according to the present technology. Method 800 includes, in operation 810, performing a conversion between video frames and video bitstreams according to rules. The rules specify that one or more syntax elements in one or more video units are used to indicate whether different stripe types are allowed to be mixed within the video frames for the conversion.

[0401] In some embodiments, the rule further specifies that mixing of bidirectional (B) and predictive (P) codec types is not permitted within a video picture of a video. In some embodiments, a first syntax element in one or more syntax elements indicates that mixing of B and P codec types is not permitted. In some embodiments, one or more video units include parameter sets or Decoding Capability Information (DCI) Network Abstraction Layer (NAL) units. In some embodiments, one or more video units include sequence parameter sets, picture parameter sets, or picture headers. In some embodiments, one or more syntax elements are conditionally indicated based on whether inter-frame striping is allowed. In some embodiments, when a video picture includes only stripes of the P codec type, a syntax flag in the picture header indicating whether the video picture is concatenated from reference list 0 is equal to 1. In some embodiments, the syntax flag is omitted and is inferred to be equal to 1.

[0402] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered in different places within the bitstream. For example, a macroblock can be encoded according to the error residual values ​​of the transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0403] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, such as one or more modules of computer program instructions encoded on a computer-readable medium for performing or controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0404] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communications network.

[0405] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic flows can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0406] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0407] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0408] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0409] Only some implementation methods and examples are described, and may be based on what is described and shown in this patent document.

Claims

1. A video processing method, comprising: Perform the conversion between the video and the video bitstream according to the rules. The rule specifies that one or more syntax elements in one or more video units are used to indicate whether a stripe specifying a codec type is allowed for the conversion. The specified codec type includes bidirectional B codec type or predictive P codec type. The rule further specifies that one or more syntax elements in the one or more video units are used to indicate whether the conversion allows the mixing of different stripes of the specified codec type within the video frames of the video.

2. The method according to claim 1, wherein, The one or more video units include a sequence parameter set.

3. The method according to claim 2, wherein, The one or more syntax elements include a first syntax element in the sequence parameter set, wherein the first syntax element in the sequence parameter set being equal to 1 indicates that the codec layer video sequence CLVS corresponding to the sequence parameter set includes one or more stripes of the bidirectional B codec type, and wherein the first syntax element being equal to 0 indicates that the CLVS excludes stripes of the bidirectional B codec type.

4. The method according to claim 2, wherein, The one or more syntax elements include a first syntax element in the sequence parameter set, wherein the first syntax element in the sequence parameter set being equal to 0 indicates that the codec layer video sequence CLVS corresponding to the sequence parameter set includes one or more stripes of the bidirectional B codec type, and wherein the first syntax element being equal to 1 indicates that the CLVS does not include stripes of the bidirectional B codec type.

5. The method according to claim 3 or 4, wherein, Modify the use of the first group of syntax elements in the sequence parameter set based on the first syntax element.

6. The method according to claim 5, wherein, In response to the first syntax element indicating that the specified codec type of stripe is allowed, the first set of syntax elements is indicated for the conversion.

7. The method according to claim 5, wherein, In response to the first syntax element indicating that stripes of the specified codec type are not allowed, the first set of syntax elements is inferred for the conversion.

8. The method according to claim 5, wherein, The first set of syntax elements indicates that more than one predictive signaling codec tool is required.

9. The method according to claim 5, wherein, The first set of syntax elements includes at least one of the following: a syntax flag indicating whether explicit weighted prediction is applicable to B-strips, a syntax flag indicating whether bidirectional optical flow inter-frame prediction is enabled, a syntax flag indicating whether symmetric motion vector difference is enabled, a syntax flag indicating whether inter-frame bidirectional prediction based on decoder motion vector refinement is enabled, a syntax flag indicating whether bidirectional prediction using codec unit weights is enabled, a syntax flag indicating whether combined inter-frame merge and intra-frame prediction is enabled, or a syntax flag indicating whether motion compensation based on geometric segmentation is enabled.

10. The method according to claim 3 or 4, wherein, The syntax elements in the general constraint information are used to indicate the value of the first syntax element.

11. The method according to claim 10, wherein, The syntax element in the general constraint information is represented as no_b_slice_contraint_flag, where no_b_slice_contraint_flag equal to 1 indicates that the first syntax element is equal to 0, and where no_b_slice_contraint_flag equal to 0 does not specify the value of the first syntax element.

12. The method according to claim 3 or 4, wherein, In response to the first syntax element indicating that the CLVS excludes the bidirectional B codec type stripe, the second set of syntax elements in the general constraint information is equal to 1.

13. The method according to claim 12, wherein, The second set of syntax elements in the general constraint information includes at least one of the following: (1) a first flag indicating whether sps_bcw_enabled_flag is equal to 0, wherein sps_bcw_enabled_flag specifies whether bidirectional prediction using codec unit weights is applicable to inter-frame prediction; (2) a second flag indicating whether sps_ciip_enabled_flag is equal to 0, wherein sps_ciip_enabled_flag specifies whether ciip_flag exists in the codec unit syntax of the inter-frame codec unit; (3) an indicator indicating whether sps_gpm_en The third flag indicates whether abled_flag is equal to 0, where sps_gpm_enabled_flag specifies whether to enable motion compensation based on geometric segmentation, (4) the fourth flag indicates whether sps_bdof_enabled_flag is equal to 0, where sps_bdof_enabled_flag specifies whether to disable bidirectional optical flow inter-frame prediction, or (5) the fifth flag indicates whether sps_dmvr_enabled_flag is equal to 0, where sps_dmv_enabled_flag specifies whether to enable bidirectional inter-frame prediction based on decoder motion vector refinement.

14. The method according to claim 3 or 4, wherein, The third set of syntax elements in the DPB parameters of the decoded image buffer is conditionally indicated based on the first syntax element.

15. The method according to claim 14, wherein, The third set of syntax elements includes at least max_num_reorder_pics[i], which specifies the maximum allowed number of pictures in the output layer set. When Htid equals i, the pictures in the output layer set can be decoded before any picture in the output layer set and output after the picture in the output order. In cases where the first syntax element indicates that stripes of the bidirectional B codec type are not allowed, max_num_reorder_pics[i] is omitted and inferred to be 0.

16. The method according to claim 3 or 4, wherein, The one or more syntax elements include additional syntax elements in the sequence parameter set, wherein the additional syntax element equal to 1 indicates that the codec layer video sequence CLVS includes one or more stripes of the predicted P codec type, and wherein the additional syntax element equal to 0 indicates that the CLVS does not include stripes of the predicted P codec type.

17. The method according to claim 3 or 4, wherein, The one or more syntax elements include additional syntax elements in the sequence parameter set, wherein the additional syntax element equal to 0 indicates that the codec layer video sequence CLVS includes one or more stripes of the predicted P codec type, and wherein the additional syntax element equal to 1 indicates that the CLVS does not include stripes of the predicted P codec type.

18. The method according to claim 1, wherein, The one or more video units include a set of image parameters.

19. The method according to claim 18, wherein, The one or more syntax elements include a second syntax element in the picture parameter set, wherein the second syntax element in the picture parameter set being equal to 1 indicates that the video picture referencing the picture parameter set includes one or more stripes of the bidirectional B codec type, and wherein the second syntax element being equal to 0 indicates that the video picture excludes stripes of the bidirectional B codec type.

20. The method according to claim 19, wherein, The first syntax element and the second syntax element have the same value.

21. The method according to claim 19 or 20, wherein, The use of the first set of syntax elements in the image parameter set is modified based on the second syntax element.

22. The method according to claim 21, wherein, In the case where the second syntax element indicates that the transformation allows stripes of the specified codec type, the first set of syntax elements is indicated for the transformation in the picture parameter set.

23. The method according to claim 21, wherein, If the second syntax element indicates that the transformation does not allow stripes of the specified codec type, the first set of syntax elements is inferred for the transformation.

24. The method according to claim 21, wherein, The first set of syntax elements indicates that more than one predictive signaling codec tool is required.

25. The method according to claim 24, wherein, The first set of syntax elements includes at least one of the following: specifying whether explicit weighted prediction is applied to the pps_weighted_bipred_flag of the B-band of the reference image parameter set, or specifying the inferred value of the variable NumRefIdxActive[1] of the B-band for which num_ref_idx_active_override_flag is equal to 0, num_ref_idx_default_active_minus1[1].

26. The method according to claim 1, wherein, The one or more video units include an image header.

27. The method according to claim 26, wherein, The one or more syntax elements include a third syntax element in the image header, wherein the third syntax element equal to 1 indicates that the image corresponding to the image header includes one or more stripes of the bidirectional B codec type, and wherein the third syntax element equal to 0 indicates that the image does not include stripes of the bidirectional B codec type.

28. The method according to claim 27, wherein, The third syntax element is conditionally indicated based on the first syntax element and / or the second syntax element.

29. The method according to claim 27 or 28, wherein, The third syntax element is indicated whereby the first or second syntax element indicates that the conversion allows stripes of the specified codec type.

30. The method according to claim 27 or 28, wherein, In response to the first or second syntax element indicating that the conversion does not allow stripes of the specified codec type, the third syntax element is omitted and is inferred to be false.

31. The method according to claim 27 or 28, wherein, Based on the third syntax element, the use of the first set of syntax elements in the image header is modified for the transformation.

32. The method according to claim 31, wherein, The use of the first set of syntax elements in the image header is also modified based on the first syntax element and / or the second syntax element.

33. The method according to claim 31, wherein, In response to the first syntax element, the second syntax element, and / or the third syntax element indicating that the stripe of the specified codec type is allowed, the first set of syntax elements indicates the conversion in the picture header.

34. The method according to claim 33, wherein, When the syntax element in the sequence parameter set that indicates the existence of one of the syntax elements in the first group of syntax elements is equal to 1 and the third syntax element is equal to 1, one of the syntax elements in the first group of syntax elements in the image header is indicated.

35. The method according to claim 33, wherein, One of the syntax elements in the first group of syntax elements in the image header includes a ph_disable_bdof_flag specifying whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the strip associated with the image header, or a ph_disable_dmvr_flag specifying whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the strip associated with the image header.

36. The method according to claim 31, wherein, In response to the first syntax element, the second syntax element, and / or the third syntax element indicating that the stripe of the specified codec type is not allowed, the first set of syntax elements is inferred for the conversion.

37. The method according to claim 36, wherein, When the corresponding syntax element in the sequence parameter set is equal to 1 and the third syntax element is equal to 1, one of the syntax elements in the first group of syntax elements in the image header is inferred to be 0.

38. The method according to claim 36, wherein, When the corresponding syntax element in the sequence parameter set is equal to 0 and the third syntax element is equal to 0, one of the syntax elements in the first group of syntax elements in the image header is inferred to be 1.

39. The method according to claim 37, wherein, One of the syntax elements in the first group of syntax elements in the image header includes a ph_disable_bdof_flag specifying whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the strip associated with the image header, or a ph_disable_dmvr_flag specifying whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the strip associated with the image header.

40. The method according to claim 31, wherein, The first set of syntax elements in the image header includes one or more syntax elements that indicate the use of an encoding / decoding tool that requires more than one predictive signaling.

41. The method according to claim 31, wherein, The first set of syntax elements in the image header includes at least one of the following: specifying whether to derive the ph_collocated_from_l0_flag for the juxtaposed image used for temporal motion vector prediction from reference image list 0; specifying whether to parse mvd_coding(x0, y0, ... 1) The `mvd_l1_zero_flag` of the syntax structure specifies whether to disable inter-frame bidirectional prediction based on bidirectional optical flow inter-frame prediction in the stripe associated with the picture header. The `ph_disable_bdof_flag` specifies whether to disable inter-frame bidirectional prediction based on decoder motion vector refinement in the stripe associated with the picture header. Alternatively, it specifies, when both `pps_weighted_bipred_flag` and `wp_info_in_ph_flag` are equal to 1, the number of weights of the signaling notifications for entries in the reference picture list 1, as `num_l1_weights`. Here, `pps_weighted_bipred_flag` specifies whether explicit weighted prediction is applied to the B-strip of the reference picture parameter set, and `wp_info_in_ph_flag` specifies whether weighted prediction information exists in the picture header syntax structure and not in the stripe header of the reference picture parameter set that does not contain a picture header syntax structure.

42. The method according to claim 1, wherein, Whether the strip type is indicated or inferred is based on one or more syntax elements in the one or more video units.

43. The method according to claim 42, wherein, The slice type is indicated at least by ph_inter_slice_allowed_flag, ph_intra_slice_allowed_flag, and the one or more syntax elements in the one or more video units.

44. The method according to claim 42 or 43, wherein, In the absence of an indication of the stripe type, the stripe type is inferred to be equal to (ph_inter_slice_allowed_flag ? 1 : 2).

45. The method according to claim 42, wherein, The slice type is inferred to be 1 if ph_intra_slice_allowed_flag equals 0 and the one or more syntax elements in the one or more video units indicate that slices of codec type B are not allowed.

46. ​​The method according to claim 1, wherein, The one or more syntax elements are conditionally indicated based on whether inter-frame stripes are applied to the transformation.

47. The method according to claim 1, wherein, The rule also stipulates that the bidirectional B codec type and the predictive P codec type are not allowed to be mixed within the video frame of the video.

48. The method according to claim 47, wherein, The second syntax element in one or more syntax elements indicates that mixing of the bidirectional B codec type and the predictive P codec type is not allowed.

49. The method according to claim 48, wherein, The one or more video units include parameter sets or decoding capability information DCI network abstraction layer NAL units.

50. The method according to claim 48, wherein, The one or more video units include a sequence parameter set, an image parameter set, or an image header.

51. The method according to claim 1, wherein, Based on whether or not inter-frame stripes are allowed, one or more syntax elements can be conditionally indicated.

52. The method according to claim 1, wherein, If the video image only includes stripes of the predicted P codec type, the image header indicates whether the syntax flags of the video image are set to 1 from reference list 0.

53. The method according to claim 52, wherein, The syntax marker is omitted and is inferred to be equal to 1.

54. A method for storing a bitstream of video, comprising: The video bitstream is generated according to the rules. The rule specifies that one or more syntax elements in one or more video units are used to indicate whether a stripe of a codec type is allowed for the bitstream that generates the video. The specified codec type includes bidirectional B codec type or predictive P codec type. The rule further specifies that one or more syntax elements in the one or more video units are used to indicate whether the bitstream that generates the video is allowed to mix different stripes of the specified codec type within the video frames of the video.

55. A video decoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 54.

56. A video encoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 54.

57. A computer program product having computer code stored thereon, said code causing the processor, when executed by a processor, to perform the method according to any one of claims 1 to 54.

58. A non-transitory computer-readable recording medium for storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: The video bitstream is generated according to the rules. The rule specifies that one or more syntax elements in one or more video units are used to indicate whether a stripe of a codec type is allowed for the bitstream that generates the video. The specified codec type includes bidirectional B codec type or predictive P codec type. The rule further specifies that one or more syntax elements in the one or more video units are used to indicate whether the bitstream that generates the video is allowed to mix different stripes of the specified codec type within the video frames of the video.

Citation Information

Patent Citations

  • Explicit way for signaling a collocated picture for high efficicency video coding (HEVC) using reference list0 and list1

    US20130128969A1

  • Constraints and unit types to simplify video random access

    US20200029094A1