Reference picture resampling

By optimizing the derivation of the reference image resampling flag variable and the signaling design of the access unit delimiter, the problems of overly strict semantics and inflexible signaling in video codec standards are solved, achieving more efficient multi-layer video codec processing.

CN115699731BActive Publication Date: 2026-08-04DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2021-06-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing video codec standards, the design of reference image resampling, access unit delimiters, and general constraint information suffers from problems such as overly strict semantic deduction, inflexible signaling design, byte alignment issues, and unclear relationships between inter-frame related syntax elements, which affect the efficiency and consistency of video codecs.

Method used

By adjusting the derivation of the reference image resampling flag variable, the signaling design of the access unit delimiter and general constraint information is optimized to ensure the flexibility and consistency of syntax elements. This includes adjusting the derivation of RprConstraintsActiveFlag, optimizing the constraint relationship of AUD and PH/SH syntax elements, and improving the signaling format of the GCI field.

Benefits of technology

It improves the flexibility and consistency of the video encoding and decoding process, optimizes encoding and decoding efficiency, reduces redundant signaling, and supports efficient processing of multi-layer video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699731B_ABST
    Figure CN115699731B_ABST
Patent Text Reader

Abstract

Examples of video encoding methods and apparatuses and video decoding methods and apparatuses are described. An example method of video processing includes performing a conversion between a current picture of a video and a bitstream of the video according to a rule. The rule specifies that a number of entries in a reference picture list of the current picture is greater than 0 in response to (1) one or more slices in the current picture being allowed to have a slice type other than an intra (I) slice type, and in response to (2) reference picture list (RPL) information being present in a picture header.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is an international patent application filed on June 2, 2021, under patent number PCT / CN2021 / 097844, which entered the Chinese national phase. It claims priority to PCT patent applications filed on June 4, 2020 (PCT / CN2020 / 094397), June 11, 2020 (PCT / CN2020 / 095688), and June 22, 2020 (PCT / CN2020 / 097390). The entire disclosure of the above applications is incorporated herein by reference and forms part of this disclosure. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to perform video encoding or decoding.

[0006] In one representative aspect, a method for processing video data is disclosed. The method includes: performing a conversion between a current frame of a video and a bitstream of the video according to rules. The rules specify that, in response to (1) one or more stripes in the current frame are allowed to have stripe types other than intra-frame (I) stripe types, and in response to (2) reference picture list (RPL) information present in the picture header, the number of entries in the reference picture list of the current frame is greater than 0.

[0007] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between a current frame of the video and the bitstream of the video according to rules. The rules specify that a variable indicating the number of sub-frames in each frame of the video is not used to derive a first syntax flag indicating whether a reference frame resampling constraint is satisfied.

[0008] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between a current frame of the video and a bitstream of the video according to rules. The rules specify a syntax flag indicating whether a reference frame resampling constraint is satisfied, determined based on whether a sub-frame of the current frame is treated as a frame in the conversion.

[0009] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between a current image of the video and the bitstream of the video according to rules. The rules specify syntax flags indicating whether a reference image resampling constraint is satisfied, determined based on whether the current image and its reference image are in the same layer and / or whether inter-layer prediction is allowed for the conversion.

[0010] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between video and a video bitstream according to rules. These rules stipulate that the value of a first syntax element in a picture header or stripe header is constrained by the value of a second syntax element present in an Access Unit Delimiter (AUD) in the bitstream.

[0011] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between video and a video bitstream according to rules. The rules specify whether a syntax element in a general constraint information syntax structure exists in the bitstream based on a constraint byte count greater than 0 or a reserved byte count.

[0012] In another representative aspect, a method for processing video data is disclosed. The method includes performing a conversion between a current block of video and the bitstream of the video according to rules. The rules specify the selective indication of the use of adaptive color transformation based on information about one or more prediction modes applicable to the current block.

[0013] In another representative aspect, a method for processing video data is disclosed. The method includes: performing a conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies the use of syntax fields, the syntax fields indicating the applicability of reference image resampling to corresponding segments of the video.

[0014] In another representative aspect, a method for processing video data is disclosed. The method includes: performing a conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that the value of a first syntax element in a picture header or stripe header is constrained based on the value of a second syntax element corresponding to an access unit delimiter.

[0015] In another representative aspect, a method for processing video data is disclosed. The method includes: performing a conversion between the video and its codec representation, wherein the codec representation conforms to format rules, wherein the format rules specify whether and how to include one or more syntax elements in a general constraint information field.

[0016] In yet another representative aspect, a video encoding apparatus is disclosed. This video encoding apparatus includes a processor configured to perform the methods described above.

[0017] In another representative aspect, a video decoding apparatus is disclosed. This video decoding apparatus includes a processor configured to perform the methods described above.

[0018] In another representative aspect, a computer-readable medium having code stored thereon is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0019] These and other features are described in this document. Attached Figure Description

[0020] Figure 1 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0021] Figure 2 This is a block diagram of an example hardware platform used for video processing.

[0022] Figure 3 This is a flowchart of an example method for video processing.

[0023] Figure 4 This is a block diagram illustrating an example video codec system.

[0024] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0025] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0026] Figure 7 This is a flowchart representation of a method for processing video data according to one or more embodiments of the present technology.

[0027] Figure 8 This is a flowchart representation of another method for processing video data according to one or more embodiments of the present technology.

[0028] Figure 9 This is a flowchart representation of another method for processing video data according to one or more embodiments of the present technology.

[0029] Figure 10 This is a flowchart representation of another method for processing video data according to one or more embodiments of the present technology.

[0030] Figure 11 This is a flowchart representation of another method for processing video data according to one or more embodiments of the present technology.

[0031] Figure 12 This is a flowchart representation of another method for processing video data according to one or more embodiments of the present technology.

[0032] Figure 13 This is a flowchart representation of yet another method for processing video data according to one or more embodiments of the present technology. Detailed Implementation

[0033] The chapter headings used in this document are for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not intended to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0034] 1. Overview

[0035] This disclosure relates to video codec technology. Specifically, this disclosure concerns the derivation of the Reference Picture Resampling (RPR) flag variable, the relationship between Access Unit Delimiters (AUDs) and syntax elements, and the signaling of General Constraint Information (GCI) in other NAL units of video codec. These concepts can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Universal Video Codec (VVC) currently under development.

[0036] 2. Abbreviations

[0037] APS Adaptive Parameter Set

[0038] AU Access Unit

[0039] AUD (Access Unit Delimiter)

[0040] AVC (Advanced Video Coding)

[0041] CLVS (Coded Layer Video Sequence)

[0042] CPB Coded Picture Buffer

[0043] CRA Clean Random Access

[0044] CTU (Coding Tree Unit)

[0045] CVS (Coded Video Sequence)

[0046] DPB Decoded Picture Buffer

[0047] DPS Decoding Parameter Set

[0048] End of Bitstream (EOB)

[0049] End of Sequence in EOS

[0050] GCI General Constraints Information

[0051] GDR Gradual Decoding Refresh

[0052] HEVC (High Efficiency Video Coding)

[0053] HRD (Hypothetical Reference Decoder)

[0054] Instantaneous Decoding Refresh (IDR)

[0055] JEM Joint Exploration Model

[0056] MCTS Motion-Constrained Tile Sets

[0057] NAL (Network Abstraction Layer)

[0058] PH Picture Header

[0059] PPS Image Parameter Set

[0060] PTL (Profile, Tier, Level)

[0061] PU Picture Unit

[0062] RRP Reference Picture Resampling

[0063] RBSP Raw Byte Sequence Payload

[0064] SE Syntax Element

[0065] SEI Supplemental Enhancement Information

[0066] SH Slice Header

[0067] SPS Sequence Parameter Set

[0068] Scalable Video Coding (SVC)

[0069] VCL (Video Coding Layer)

[0070] VPS Video Parameter Set

[0071] VTM VVC Test Model

[0072] VUI Video Usability Information

[0073] VVC (Versatile Video Coding)

[0074] 3. Introduction to Video Encoding and Decoding

[0075] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 video standards, while ISO / IEC developed the MPEG-1 and MPEG-4 video standards. These two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly at the same location, and the goal of the new codec standard is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The VVC project now aims to achieve Technical Finality (FDIS) at the July 2020 meeting.

[0076] 3.1. Image resolution changes within a sequence

[0077] In AVC and HEVC, the spatial resolution of an image cannot be changed unless a new sequence with a new SPS begins with an IRAP image. VVC allows the image resolution to be changed at some point within the sequence without encoding the IRAP image, which is always intra-frame encoded / decoded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling the reference image used for inter-frame prediction when the resolution of the reference image differs from the resolution of the current image being decoded.

[0078] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are suitable for scaling ratios from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as in motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process with scaling ratios ranging from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height and the left, right, top, and bottom scaling offsets specified for the reference and current images.

[0079] Other aspects of the VVC design that support this feature that differ from HEVC include: i) Picture resolution and the corresponding consistency window are signaled in the PPS instead of the SPS, while the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (the slot in the DPB used to store one decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.

[0080] 3.2. General Scalable Video Codec (SVC) in VVC

[0081] Scalable video codec (SVC, sometimes also referred to as scalability in video codec) refers to video codec using a base layer (BL), sometimes called a reference layer (RL), and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previous coding layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., neither the lowest nor the highest layer) can be an EL used for layers below the intermediate layer, such as a base layer or any intermediate enhancement layer, and simultaneously used as an RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0082] In SVC, the parameters used by the encoder or decoder are grouped into parameter sets based on the codec level they can use (e.g., video level, sequence level, picture level, stripe level, etc.). For example, parameters used by one or more codec video sequences from different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters used by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and additional parameters specific to a single stripe can be included in the stripe header. Likewise, an indication of which parameter set a particular layer uses at a given time can be provided at various codec levels.

[0083] Because VVC supports Reference Picture Resampling (RPR), it's possible to design support for bitstreams containing multiple layers, such as two layers with SD and HD resolutions in VVC, without requiring any additional signal processing level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires advanced syntax changes (compared to no scalability support). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standard, including extensions to AVC and HEVC, VVC scalability is designed to be as friendly as possible to single-layer decoder designs. The decoding capability of multi-layer bitstreams is specified in a way that only a single layer exists in the bitstream. For example, decoding capabilities, such as DPB size, are specified independently of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams don't require many changes to decode multi-layer bitstreams. Compared to the multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, IRAP AU requires a picture of every layer present in CVS.

[0084] 3.3. Parameter Set

[0085] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All AVC, HEVC, and VVC implementations support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.

[0086] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. With SPS and PPS, it is unnecessary to repeat infrequently changing information for each sequence or image, thus avoiding redundant signaling. Furthermore, using SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving fault tolerance.

[0087] The VPS was introduced to carry sequence-level header information that is common to all layers in a multi-layer bitstream.

[0088] The purpose of APS is to carry image-level or strip-level information, which requires a considerable number of bits to encode and decode. This information can be shared by multiple images and can have many different variations within a sequence.

[0089] 3.4. Semantics and variable markers of RPR-related syntactic elements

[0090] The semantics of RPR-related syntactic elements and the derivation of variable notation are as follows in the latest VVC draft text:

[0091] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics ...

[0093] When sps_ref_pic_resampling_enabled_flag equals 1, reference image resampling is enabled, and the current image of the reference SPS can have stripes of reference images from the active entry of the reference image list, which has one or more of the following seven parameters that are different from the current image: 1) pps_pic_width_in_luma_samples, 2) pps_pic_height_in_luma_samples, 3) pps_scaling_win_left_offset, 4) pps_scaling_win_right_offset, 5) pps_scaling_win_top_offset, 6) pps_scaling_win_bottom_offset, and 7) sps_num_subpics_minus1. A value of 0 for `sps_ref_pic_resampling_enabled_flag` indicates that reference image resampling is disabled, and the current image of the reference SPS cannot have stripes that reference images from the active entry of the reference image list, which has one or more of the above seven parameters that are different from the current image:

[0094] Note 3 - When sps_ref_pic_resampling_enabled_flag equals 1, for the current image, a reference image that has one or more of the above 7 parameters that are different from the current image can belong to the same layer as the layer containing the current image or a different layer.

[0095] A value of 1 for `sps_res_change_in_clvs_allowed_flag` indicates that the image spatial resolution can be changed within the CLVS of the reference SPS. A value of 0 for `sps_res_change_in_clvs_allowed_flag` indicates that the image spatial resolution will not be changed within any CLVS of the reference SPS. When it does not exist, the value of `sps_res_change_in_clvs_allowed_flag` is inferred to be 0. ...

[0097] 8.3.2 Decoding process based on the reference image list ...

[0099] fRefWidth is set to be equal to CurrPicScalWinWidthL of the reference image RefPicList[i][j]. refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, and refScalingWinBottomOffset are set to be equal to the values ​​of pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the reference image RefPicList[i][j], respectively. fRefNumSubpics is set to be equal to sps_num_subpics_minus1 of the reference image RefPicList[i][j].

[0100] 3.5. Access Module Delimiter (AUD)

[0101] The syntax and semantics of AUD in the latest VVC draft text are as follows:

[0102]

[0103] The AU delimiter is used to indicate the start of an AU, whether the AU is an IRAP or GDR AU, and the type of stripe present in the encoded / decoded picture within the AU containing the AU delimiter NAL unit. When the bitstream contains only one layer, there is no standard decoding procedure associated with the AU delimiter.

[0104] An aud_irap_or_gdr_au_flag value of 1 indicates that an AU containing an AU delimiter is an IRAP or GDR AU, while an irap_or_gdr_au_flag value of 0 indicates that an AU containing an AU delimiter is not an IRAP or GDR AU.

[0105] The `aud_pic_type` directive indicates that the `sh_slice_type` value for all slices of the encoded / decoded picture in the AU containing the AU delimiter NAL unit is a member of the set listed in Table 7 for a given value of `aud_pic_type`. In bitstreams conforming to this version of the specification, the value of `aud_pic_type` should be equal to 0, 1, or 2. Other values ​​of `aud_pic_type` are reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this version of the specification should ignore reserved values ​​of `aud_pic_type`.

[0106] Table 7 – Interpretation of aud_pic_type

[0107]

[0108] 3.6. GCI (General Constraint Information)

[0109] In the latest VVC draft text, the general level, hierarchy, and semantics are as follows:

[0110] 7.3.3 Grading, Layer, and Level Syntax

[0111] 7.3.3.1 General grade, layer, and level syntax

[0112]

[0113]

[0114] 7.3.3.2 General Constraint Information Syntax

[0115]

[0116]

[0117]

[0118] 3.7. Conditional Signaling for the GCI Field

[0119] In some embodiments, the GCI syntax structure has been modified. The GCI extended length indicator (gci_num_reserved_bytes) is moved from the last to the first (gci_num_constraint_bytes) in the GCI syntax structure to enable skip signaling for GCI fields. The value of gci_num_reserved_bytes should be equal to 0 or 9.

[0120] Added or modified parts are indicated by bold, italic, and underlined text, while deleted parts are indicated by the brackets [[]].

[0121]

[0122] 3.8. Luminance and Chromaticity QP Mapping Table

[0123] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0124]

[0125] The addition of 26 to `sps_qp_table_start_minus26[i]` specifies the starting luminance and chrominance QP used to describe the i-th chrominance QP mapping table. The value of `sps_qp_table_start_minus26[i]` should be in the range of −26 − QpBdOffset to 36 (inclusive). When it does not exist, the value of `sps_qp_table_start_minus26[i]` is inferred to be equal to 0.

[0126] The increment of 1 in sps_num_points_in_qp_table_minus1[i] specifies the number of points used to describe the i-th chromaticity QP map. The value of sps_num_points_in_qp_table_minus1[i] should be in the range of 0 to 63 + QpBdOffset (inclusive). When it does not exist, the value of sps_num_points_in_qp_table_minus1[0] is inferred to be equal to 0.

[0127] sps_delta_qp_in_val_minus1[i][j] specifies the increment value of the input coordinates used to derive the j-th pivot point of the i-th chromaticity QP map. When it does not exist, the value of sps_delta_qp_in_val_minus1[0][j] is inferred to be equal to 0.

[0128] sps_delta_qp_diff_val[i][j] specifies the increment value used to derive the output coordinates of the j-th pivot point of the i-th chromaticity QP map. When it does not exist, the value of sps_delta_qp_in_val_minus1[0][j] is inferred to be equal to 0.

[0129] The derivation of the i-th chroma QP mapping table ChromaQpTable[i] for i = 0..numQpTables − 1 is as follows:

[0130]

[0131]

[0132] When sps_same_qp_table_for_chroma_flag equals 1, for k in the range of −QpBdOffset to 63 (inclusive), ChromaQpTable[1][k] and ChromaQpTable[2][k] are set to equal ChromaQpTable[0][k].

[0133] The requirement for bitstream consistency is that for i in the range of 0 to numQpTables − 1 (inclusive) and j in the range of 0 to sps_num_points_in_qp_table_minus1[i] + 1 (inclusive), the values ​​of qpInVal[i][j] and qpOutVal[i][j] should be in the range of −QpBdOffset to 63 (inclusive).

[0134] 4. The technical problem solved by the disclosed technical solution

[0135] The existing designs of RPR, AU delimiters, and GCI have the following problems:

[0136] (1) The derivation of the semantics of sps_num_subpics_minus1 for sps_ref_pic_resampling_enabled_flag and the flag variable RprConstraintsActiveFlag[i][j] is too strict and unnecessary. When the reference picture has a different number of subpics, and sps_subpic_treated_as_pic_flag[] is equal to 0 for all subpics in the current picture, all disabled tools will be disabled when RprConstraintsActiveFlag[i][j] is equal to 1 for the current picture.

[0137] (2) Currently, in the AU delimiter (also known as AUD) syntax structure, two syntax elements (e.g., aud_irap_or_gdr_au_flag and aud_pic_type) are used to inform AU signaling whether it is an IRAP / GDR AU and the picture type of the AU. AUD SE is not used in any other part of the decoding process. However, in PH, there is an SE (e.g., ph_gdr_or_irap_pic_flag) that expresses a similar intent to aud_irap_or_gdr_au_flag. When AUD is present, the value of PH SE can be constrained by the value of AUD SE. Moreover, in SH, there is SE, sh_slice_type, which can also be constrained by the value of aud_pic_type.

[0138] (3) In the latest VVC draft text and in some embodiments, the design of GCI signaling has byte alignment between the GCI field and the extension byte. Therefore, in potential future extensions, when new GCI fields need to be added, they will be added after the current field via alignment bits, after which another byte alignment needs to be added.

[0139] (4) In some embodiments, the design of GCI signaling uses an 8-bit syntax element to specify the GCI field, byte alignment bits, and the number of bytes for the extended byte. However, for VVC version 1, the value of the syntax element is required to be equal to 0 or 9, i.e., 0 or a specific integer value greater than zero (depending on the number of bits required for all GCI fields in VVC version 1). In potential future expansions, after adding certain new GCI fields, the value of this 8-bit syntax element will need to be another specific value equal to 0 or greater than 9, for example, 11 if the new GCI field requires 9 to 16 bits. This means that in any version of VVC, the value of the 8-bit syntax element will be equal to 0 or a specific integer value greater than zero. Therefore, it is not necessary to signal the 8-bit value; it is sufficient to signal a 1-bit flag and derive the value from the value of the flag.

[0140] (5) The relationship between whether inter-frame related syntax elements are signaled in PH and whether non-empty RPL0 is signaled in PH is not well established. For example, when the Reference Picture List (RPL) is sent in PH and list0 is empty, the value of ph_inter_slice_allowed_flag can be restricted to 0. The reverse is also true.

[0141] (6) In the case of GDR images with zero recovery point of view (POC) distance, the GDR image itself is a recovery point image. This point should be considered / expressed in the specification.

[0142] (7) Considering the mixed NAL unit types and bitstream extraction and merging, a single GCI flag can be signaled to constrain both sps_idr_rpl_present_flag and pps_mixed_nalu_types_in_pic_flag.

[0143] 5. List of Implementation Examples

[0144] To address the aforementioned problems and some other unmentioned issues, the following summarized methods are disclosed. These inventions should be considered as examples for interpreting general concepts, and not interpreted narrowly. Furthermore, these inventions can be applied individually or in any combination.

[0145] 1) To address the first issue, the derivation of RprConstraintsActiveFlag[i][j] was modified to exclude sps_num_subpics_minus1, and comments were added to clarify that when the value of sps_num_subpics_minus1 differs for the current image and the reference image RefPicList[i][j], and for at least one k value in the range of 0 to sps_num_subpics_minus1 (inclusive), the current image sps_subpic_treated_as_pic_flag[k] equals 1, all tools that cannot be used when RprConstraintsActiveFlag[i][j] equals 1, such as PROF, need to be turned off by the encoder; otherwise, the bitstream will be a non-consistent bitstream because the extracted sub-images with sps_subpic_treated_as_pic_flag[k] equal to 1 will not be decoded correctly.

[0146] a. Additionally, in one example, the semantics of sps_ref_pic_resampling_enabled_flag are changed to not involve sps_num_subpics_minus1.

[0147] b. The above statement should only be applied if inter-layer prediction is permitted.

[0148] 2) To address the first problem, the derivation of RprConstraintsActiveFlag further depends on whether the sub-image is considered an image.

[0149] a. In one example, the derivation of RprConstraintsActiveFlag[i][j] is changed so that it can depend on at least one of the values ​​of sps_subpic_treated_as_pic_flag[k] in the range of 0 to sps_num_subpics_minus1 (inclusive) for the current picture k.

[0150] b. In one example, the derivation of RprConstraintsActiveFlag[i][j] is changed so that it can depend on whether at least one of sps_subpic_treated_as_pic_flag[k] in the range of 0 to sps_num_subpics_minus1 (inclusive) is equal to 1.

[0151] c. Alternatively, RPR can still be enabled when the number of subpics between the current image and its reference image differs and the current image is not considered a subpic of the image (e.g., all values ​​of sps_subpic_treated_as_pic_flag are false).

[0152] i. Alternatively, RprConstraintsActiveFlag[i][j] can be set based on other conditions, such as scaling the window size / offset.

[0153] d. Alternatively, RPR is always enabled when the number of subpicks between the current picture and its reference picture is different and the current picture has at least one subpick that is considered a picture (e.g., sps_subpic_treated_as_pic_flag is true), regardless of the values ​​of other syntax elements (e.g., scaling window).

[0154] i. Alternatively, in addition, RprConstraintsActiveFlag[i][j] is set to true for the above cases.

[0155] e. Alternatively, when RprConstraintsActiveFlag[i][j] is true, several tools (e.g., PROF / BDOF / DMVR) can be disabled accordingly.

[0156] i. When RprConstraintsActiveFlag[i][j] can be true after extraction, multiple tools (e.g., PROF / BDOF / DMVR) can be disabled accordingly.

[0157] f. Alternatively, the consistent bitstream should satisfy that, for the difference in the number of sub-pictures considered as pictures in the current picture and between the current picture and its reference picture, the encoding / decoding tools (e.g., PROF / BDOF / DMVR) that rely on the RprConstraintsActiveFlag[i][j] check can be disabled.

[0158] g. Alternatively, whether to invoke the decoding process of encoding / decoding tools (e.g., PROF / BDOF / DMVR) may depend on whether the current sub-image in the current image is considered an image and the number of sub-images between the current image and its reference image.

[0159] i. In one example, if the current sub-image in the current image is considered an image, and the number of sub-images between the current image and its reference image is different, then these codecs are disabled regardless of the SPS enable flag.

[0160] 3) To address the first problem, the derivation of RprConstraintsActiveFlag[i][j] was modified to allow it to depend on whether the current image and the reference image belong to the same layer and / or whether inter-layer prediction is allowed.

[0161] 4) To address the second problem, regarding the constraint of PH / SH SE values ​​via AUD SE, one or more of the following methods are disclosed:

[0162] a. In one example, the values ​​of SPS / PPS / APS / PH / SH syntax elements are constrained based on the value of the AUD syntax element (if it exists).

[0163] b. In one example, the value of the syntax element specifying whether a picture is a GDR / IRAP AU (if it exists) is constrained based on whether it is a GDR / IRAP AU (e.g., the value of the AUD syntax element aud_irap_or_gdr_au_flag).

[0164] c. For example, when the AUD syntax element specifies that the AU containing the AU delimiter is not an IRAP or GDR AU (e.g., aud_irap_or_gdr_au_flag equals 0), the value of the associated PH syntax element ph_gdr_or_irap_pic_flag should not be equal to a certain value (such as 1) that specifies the picture as an IRAP or GDR picture. For example, the following constraint can be added:

[0165] i. When aud_irap_or_gdr_au_flag exists and is equal to 0 (not IRAP or GDR), the value of ph_gdr_or_irap_pic_flag should not be equal to 1 (IRAP or GDR).

[0166] ii. Alternatively, when aud_irap_or_gdr_au_flag is equal to 0 (not IRAP or GDR), the value of ph_gdr_or_irap_pic_flag should not be equal to 1 (IRAP or GDR).

[0167] iii. Alternatively, when aud_irap_or_gdr_au_flag exists and is equal to 0 (not IRAP or GDR), the value of ph_gdr_or_irap_pic_flag should be equal to 0 (which may or may not be IRAP or GDR).

[0168] iv. Alternatively, when aud_irap_or_gdr_au_flag is equal to 0 (not IRAP or GDR), the value of ph_gdr_or_irap_pic_flag should be equal to 0 (which may or may not be IRAP or GDR).

[0169] d. In one example, the value of the syntax element specifying the slice type (e.g., the SH syntax element sh_slice_type) is constrained based on the value of the AUD syntax element (if present, e.g., aud_pic_type).

[0170] i. For example, when the AUD syntax element specifies that the sh_slice_type value that may exist in the AU is an intra-frame (I) slice (e.g., aud_pic_type equals 0), the value of the associated SH syntax element sh_slice_type should not be equal to a certain value (such as 0 / 1) specifying the prediction (P) / bidirectional prediction (B) slice. For example, the following constraint can be added:

[0171] 1. When aud_pic_type exists and is equal to 0 (I-slice), the value of sh_slice_type should be equal to 2 (I-slice).

[0172] 2. Alternatively, when aud_pic_type equals 0 (I-slice), the value of sh_slice_type should be equal to 2 (I-slice).

[0173] 3. Alternatively, when aud_pic_type exists and is equal to 0 (I-slice), the value of sh_slice_type should not be equal to 0 (B-slice) or 1 (P-slice).

[0174] 4. Alternatively, when aud_pic_type is equal to 0 (I-slice), the value of sh_slice_type should not be equal to 0 (B-slice) or 1 (P-slice).

[0175] e. For example, when the AUD syntax element specifies that the sh_slice_type value that may exist in the AU is a P or I slice (e.g., aud_pic_type equals 1), the value of the associated SH syntax element sh_slice_type should not be equal to a certain value that specifies a B slice (e.g., 0). For example, the following constraint can be added:

[0176] i. When aud_pic_type exists and is equal to 1 (P and I stripes may exist), the value of sh_slice_type should be equal to 1 (P stripe) or 2 (I stripe).

[0177] ii. Alternatively, when aud_pic_type equals 1 (P and I stripes may exist), the value of sh_slice_type should be equal to 1 (P stripes) or 2 (I stripes).

[0178] iii. Alternatively, when aud_pic_type exists and is equal to 1 (P and I stripes may exist), the value of sh_slice_type should not be equal to 0 (B stripes).

[0179] iv. Alternatively, when aud_pic_type equals 1 (P and I stripes may exist), the value of sh_slice_type should not be equal to 0 (B stripes).

[0180] f. In one example, the indicator values ​​of relevant syntax elements in the AUD syntax structure and the PH / SH syntax structure can be aligned.

[0181] i. For example, aud_pic_type equal to 0 indicates that the possible sh_slice_type values ​​in the AU are B, P, or I stripes.

[0182] ii. For example, aud_pic_type equal to 1 indicates that the possible sh_slice_type values ​​in the AU are P or I stripes.

[0183] iii. For example, an aud_pic_type of 2 indicates that the possible sh_slice_type value in the AU is an I-slice.

[0184] iv. Alternatively, sh_slice_type equal to 0 indicates that the codec type of the slice is I-slice.

[0185] v. Alternatively, sh_slice_type equal to 1 indicates that the codec type of the slice is P-slice.

[0186] vi. Alternatively, sh_slice_type equal to 2 indicates that the codec type of the slice is B slice.

[0187] g. In one example, the names of relevant syntax elements in the AUD syntax structure and the PH / SH syntax structure can be aligned.

[0188] i. For example, aud_irap_or_gdr_au_flag can be renamed to aud_gdr_or_irap_au_flag.

[0189] ii. Alternatively, ph_gdr_or_irap_pic_flag can be renamed to ph_irap_or_gdr_pic_flag.

[0190] h. In one example, the values ​​of syntax elements specifying whether an image is an IRAP or GDR image (e.g., ph_gdr_or_irap_pic_flag equals 0 or 1, and / or ph_gdr_pic_flag equals 0 or 1, and / or the PH SE named ph_irap_pic_flag equals 0 or 1) can be constrained based on whether the bitstream is a single-layer bitstream and whether the image belongs to an IRAP or GDR AU.

[0191] i. For example, when the bitstream is a single-layer bitstream (e.g., when the VPS syntax element vps_max_layers_minus1 equals 0, and / or the SPS syntax element sps_video_parameter_set_id equals 0), and the AUD syntax element specifies that the AU containing the AU delimiter is not an IRAP or GDR AU (e.g., AUD_rap_or_GDR_auflag equals 0), the following constraint can be added:

[0192] 1. The value of the associated PH syntax element ph_gdr_or_irap_pic_flag should not be equal to a certain value (such as 1) that specifies the image as an IRAP or GDR image.

[0193] 2. The value of the associated PH syntax element ph_gdr_pic_flag should not be equal to a certain value that specifies the image as a GDR image (such as 1).

[0194] 3. Suppose there is a PH SE named ph_irap_pic_flag and ph_irap_pic_flag is equal to a certain value (such as 1) that specifies that the image is definitely an IRAP image. Then, under the above conditions, the value of the associated PH syntax element ph_gdr_irap_flag should not be equal to 1.

[0195] ii. For example, when the bitstream is a single-layer bitstream (e.g., when the VPS syntax element vps_max_layers_minus1 equals 0, and / or the SPS syntax element sps_video_parameter_set_id equals 0), and the AUD syntax element specifies that the AU containing the AU delimiter is an IRAP or GDR AU (e.g., aud_irap_or_gdr_au_flag equals 1), the following constraint can be added:

[0196] 1. Suppose there is a PH SE named ph_irap_pic_flag and ph_irap_pic_flag is equal to a certain value (such as 1) that specifies that the image is definitely an IRAP image. Then, under the above conditions, ph_gdr_pic_flag or ph_irap_flag should both be equal to 1.

[0197] iii. For example, when vps_max_layers_minus1 (or sps_video_parameter_set_id) equals 0 (single layer), and aud_irap_or_gdr_au_flag exists and equals 0 (not IRAP or GDR AU), the following constraints will be defined:

[0198] 1. The value of ph_gdr_or_irap_pic_flag should not be equal to 1 (IRAP or GDR image).

[0199] a. Alternatively, the value of ph_gdr_or_irap_pic_flag should be equal to 0 (the image may or may not be an IRAP image, but it is definitely not a GDR image).

[0200] 2. Alternatively, the value of ph_gdr_pic_flag should not be equal to 1 (GDR image).

[0201] a. Alternatively, the value of ph_gdr_pic_flag should be equal to 0 (the image is definitely not a GDR image).

[0202] 3. Alternatively, the value of ph_irap_pic_flag (named, if any) should not be equal to 1 (it must be an IRAP image).

[0203] a. Alternatively, the value of ph_irap_pic_flag (named, if any) should be equal to 0 (the image is definitely not an IRAP image).

[0204] 5) To address the third issue, whether the GCI syntax elements in the signaling notification GCI syntax structure can depend on whether the number of signaling notification / derived constraint / reserved bytes (e.g., gci_num_constraint_bytes in JVET-S0092-v1) is not equal to 0 or greater than 0.

[0205] a. In one example, the GCI syntax in JVET-S0092-v1 is changed to: (1) the condition “if( gci_num_constraint_bytes > 8 )” is changed to “if( gci_num_constraint_bytes > 0 )”; (2) byte alignment, i.e., the syntax element gci_alignment_zero_bit and its condition “while( !byte_aligned() )” are removed; and (3) the signaling of reserved bytes (i.e. gci_reserved_byte[i]) is changed to the signaling of service bits (i.e. gci_reserved_bit[i]) such that the total number of bits under the condition “if( gci_num_constraint_bytes > 0 )” is equal to gci_num_constraint_bytes * 8.

[0206] b. Alternatively, the number of constrained / reserved bytes (e.g., gci_num_constraint_bytes in JVET-S0092) can be constrained to a given value that may depend on the version of the GCI syntax element / grade / fragment / standard information.

[0207] i. Alternatively, another non-zero value is set to Ceil( Log2( numGciBits ) ), where the variable numGciBits is deduced to be equal to the number of bits of all syntax elements under the condition "if( gci_num_constraint_bytes > 0 )", excluding the gci_reserved_bit[i] syntax element.

[0208] 6) To address the fourth issue, flags can be used to indicate the presence of GCI syntax elements / or GCI syntax structures, and when a GCI syntax element is present, zero or more reserved bits can be further signaled.

[0209] a. In one example, the GCI syntax in JVET-S0092-v1 is changed to: (1) the 8-bit gci_num_constraint_bytes is replaced by a single-bit flag, e.g., gci_present_flag; (2) the condition “if( gci_num_constraint_bytes > 8 )” is changed to “if( gci_present_flag )”; (3) byte alignment, i.e., the syntax element gci_alignment_zero_bit and its condition “while( !byte_aligned() )” are removed; and (4) the signaling for reserved bytes (i.e., gci_reserved_byte[i]) is changed to the signaling for reserved bits, such that the total number of bits under the condition “if( gci_present_flag )” is equal to gciNumConstraintBytes * 8, where gciNumConstraintBytes is derived from the value of gci_present_flag.

[0210] b. Alternatively, the number of constraint / reserved bytes (e.g., gci_num_constraint_bytes in JVET-S0092) (e.g., signaling notification or derivation) may depend on the version of the GCI syntax element / grade / fragment / standard information.

[0211] i. Alternatively, another non-zero value is set to Ceil( Log2( numGciBits) ) where the variable numGciBits is deduced to be equal to the number of bits of all syntax elements under the condition "if( gci_present_flag > 0 )", excluding the gci_reserved_bit[i] syntax element.

[0212] c. Alternatively, the 7-bit reserved bits can be further signaled when the flag indicates that the GCI syntax element is not present.

[0213] i. In one example, the 7 reserved bits are 7 zero bits.

[0214] d. Alternatively, the GCI syntax structure can be moved after general_sub_profile_idc or after ptl_sublayer_level_present_flag[i] or directly before the while loop (for byte alignment) in the PTL syntax structure.

[0215] i. Alternatively, in addition, when a flag indicates the presence of a GCI syntax element, the GCI syntax structure (e.g., general_constraint_info()) is further signaled.

[0216] ii. Alternatively, when a flag indicates that a GCI syntax element does not exist, the GCI syntax structure (e.g., general_constraint_info()) is not signaled, and the value of the GCI syntax element is set to its default value.

[0217] Related Adaptive Color Transformation (ACT)

[0218] 7) Signaling used by ACT (such as ACT on / off control flags) can be skipped based on predictive pattern information.

[0219] a. In one example, whether the signaling notification indicates the ACT on / off flag can depend on whether the prediction mode of the current block is non-intra-frame (e.g., not MODE_INTRA) and / or non-inter-frame (e.g., not MODE_INTER) and / or non-IBC (e.g., not MODE_IBC).

[0220] b. In one example, when all intra-frame (e.g., MODE_INTRA) and inter-frame (e.g., MODE_INTER) and IBC (e.g., MODE_IBC) modes are not applied to the video unit, the signaling used by ACT (e.g., ACT on / off control flags) can be skipped.

[0221] i. Alternatively, in addition, the use of ACT is presumed to be false when there is no signaling signal.

[0222] c. In one example, when ACT is used for a block, the X mode is not applied to that block.

[0223] i. For example, X could be a color palette.

[0224] ii. For example, X can be a mode different from MODE_INTRA, MODE_INTER, and MODE_IBC.

[0225] iii. Alternatively, the use of the X pattern is inferred to be false under the above conditions.

[0226] other

[0227] 8) To address the fifth problem, one or more of the following methods are disclosed:

[0228] a. In one example, it is required that when inter-slice is allowed in the picture (e.g., ph_inter_slice_allowed_flag is true) and RPL is signaled in PH instead of SH, then the reference picture list 0 (e.g., RefPicList[0]) should not be empty, i.e., contain at least one entry.

[0229] i. For example, a bitstream constraint can be specified such that when pps_rpl_info_in_ph_flag equals 1 and ph_inter_slice_allowed_flag equals 1, the value of num_ref_entries[0][RplsId[0]] should be greater than 0.

[0230] ii. Additionally, whether signaling notification is made and / or how signaling notification is made and / or the inference of the number of reference entries in list 0 (e.g., num_ref_entries[0][RplsIdx[0]]) may depend on whether inter-frame striping is allowed in the picture.

[0231] 1. In one example, when inter-slice is allowed in the picture (e.g., ph_inter_slice_allowed_flag is true) and RPL is signaled in PH (e.g., pps_rpl_info_in_ph_flag is true), the signaling notification can be changed to subtract 1 from the number of entries in reference picture list 0.

[0232] 2. In one example, when inter-slice is not allowed in the image (e.g., ph_inter_slice_allowed_flag is false) and RPL is signaled in PH (e.g., pps_rpl_info_in_ph_flag is true), the number of entries in the reference image list X (X equals 0 or 1) may not be signaled again.

[0233] b. In one example, it is required that when RPL is signaled in PH instead of SH, and the reference picture list 0 (e.g., RefPicList[0]) is empty, i.e. contains 0 entries, then only I stripes are allowed in the picture.

[0234] i. For example, a bitstream constraint can be specified such that when pps_rpl_info_in_ph_flag is equal to 1 and the value of num_ref_entries[0][RplsIdx[0]] is equal to 0, then the value of ph_inter_slice_allowed_flag should be equal to 0.

[0235] ii. Additionally, the indication of whether and / or how to signal the inter-frame allow flag (e.g., ph_inter_slice_allowed_flag) and / or intra-frame allow flag (e.g., ph_intra_slice_allowed_flag) may depend on the number of entries in reference picture list 0.

[0236] 1. In one example, when pps_rpl_info_in_ph_flag equals 1 and num_ref_entries[0][RplsIdx[0]] equals 0, the indication of whether inter-slice is allowed (e.g., ph_inter_slice_allowed_flag) can no longer be signaled.

[0237] a. In addition, the instruction is presumed to be false.

[0238] b. Additionally, another indication of whether intra-slice is allowed (e.g., ph_intra_slice_allowed_flag) may no longer be signaled.

[0239] c. In one example, the inference of how signaling is notified and / or whether signaling is notified of the strip type and / or strip type may depend on the number of entries in reference picture list 0 and / or 1.

[0240] i. In one example, when the reference image list 0 (e.g., RefPicList[0]) is empty, i.e. contains 0 entries, the strip type indication is no longer signaled.

[0241] 1. Alternatively, when the RPL is signaled in the PH instead of the SH, and the reference picture list 0 (e.g., RefPicList[0]) is empty, i.e. contains 0 entries, the strip type indication is no longer signaled.

[0242] 2. Alternatively, in addition, for the above cases, the band type is inferred to be I band.

[0243] ii. In one example, when the reference image list 0 (e.g., RefPicList[0]) is empty, i.e. contains 0 entries, the strip type should be equal to I stripe.

[0244] iii. In one example, when the reference image list 1 (e.g., RefPicList[1]) is empty, i.e. contains 0 entries, the strip type should not be equal to strip B.

[0245] iv. In one example, when the reference image list 1 (e.g., RefPicList[1]) is empty, i.e. contains 0 entries, the strip type should be equal to I or P strip.

[0246] d. In one example, it is required that if RPL is signaled in PH, it should be used for all stripes in the picture. Therefore, if the entire picture contains only I stripes, list 0 can only be empty. Otherwise, if list 0 is not empty, there must be at least one B or P stripe in this picture.

[0247] 9) To solve the sixth problem, one or more of the following methods are disclosed:

[0248] a. In one example, the recovery point image should follow the associated GDR image in decoding order only if the recovery POC count is greater than 0.

[0249] b. In one example, when the recovery point count is 0, the recovery point image is the GDR image itself.

[0250] c. In one example, the recovered images may or may not be recovered before the recovered point images in the decoding order.

[0251] d. In one example, when the recovery point count is 0, the recovery point image is the GDR image itself, and it may or may not have a recovery image.

[0252] e. In one example, the semantics of ph_recovery_poc_cnt in JVET-R2001-vA can be changed as follows. Most of the relevant parts that have been added or modified are indicated by bold, italics, and underline, and some deleted parts are indicated by [[]].

[0253] Ph_recovery_poc_cnt specifies the recovery points of the decoded images in the output order.

[0254] When the current image is a GDR image, the variable recoveryPointPocVal is derived as follows:

[0255] recoveryPointPocVal = PicOrderCntVal + ph_recovery_poc_cnt

[0256] The image picA that follows the current GDR image in decoding order within CLVS with PicOrderCntVal equal to recoveryPointPocVal is called the recovery point image. Otherwise, the first image in CLVS with PicOrderCntVal greater than recoveryPointPocVal in output order is called the recovery point image. The recovery point image should not precede the current GDR image in decoding order. The image associated with the current GDR image and whose PicOrderCntVal is less than recoveryPointPocVal is called the recovery image of the GDR image. The value of Ph_recovery_poc_cnt should be in the range of 0 to MaxPicOrderCntLsb − 1 (inclusive).

[0257] 10) To solve the seventh problem, one or more of the following methods are disclosed:

[0258] a. In one example, the first GCI syntax element (e.g., named no_idr_rpl_mixed_nalu_constraint_flag) can be signaled to restrict the use of both RPL and mixed NAL unit types sent with the IDR picture.

[0259] i. For example, when no_idr_rpl_mixed_nalu_constraint_flag equals 1, no signaling should be sent to the RPL for IDR images, and the VCL NAL units for each image should have the same nal_unit_type value. When no_idr_rpl_mixed_nalu_constraint_flag equals 0, this constraint is not imposed.

[0260] ii. For example, when no_idr_rpl_mixed_nalu_constraint_flag equals 1, reference picture list syntax elements should not exist in the strip header of IDR pictures (e.g., sps_idr_rpl_present_flag should equal 0), and pps_mixed_nalu_types_in_pic_flag should equal 0. When no_idr_rpl_mixed_nalu_constraint_flag equals 0, this constraint is not imposed.

[0261] iii. For example, when the first GCI syntax element does not exist (e.g., an indication that the GCI syntax element exists indicates that the GCI syntax element does not exist), the value of the first GCI syntax element can be inferred to be X (e.g., X is 0 or 1).

[0262] 11) It is recommended that the signaling and / or range and / or inference of the syntax elements of the point count in the QP table depend on other syntax elements.

[0263] a. It is recommended to set the maximum value of num_points_in_qp_table_minus1[i] to (K - the starting luminance and / or chrominance QP used to describe the i-th chrominance QP mapping table).

[0264] i. In one example, K depends on the maximum allowed QP value of the video.

[0265] 1. In one example, K is set to (maximum allowed QP value - 1), for example, 62 in VVC.

[0266] ii. In one example, the maximum value is set to (62 – (qp_table_start_minus26[i] + 26) where sps_qp_table_start_minus26[i] plus 26 specifies the starting luminance and chrominance QP used to describe the i-th chrominance QP mapping table.

[0267] b. Alternatively, the sum of the points in the QP table (e.g., num_points_in_qp_table_minus1[i] for the i-th QP table) and the starting luminance and / or chrominance QP used to describe the i-th chrominance QP mapping table (e.g., sps_qp_table_start_minus26[i] plus 26) should be less than the maximum permissible QP value (e.g., 63).

[0268] 12) The requirement for bitstream consistency is that for i in the range of 0 to numQpTables − 1 (inclusive) and j in the range of 0 to sps_num_points_in_qp_table_minus1[i] + K (e.g., K = 0 or 1) (inclusive) and sps_num_points_in_qp_table_minus1[i] + K), the value of qpInVal[i][j] should be in the range of −QpBdOffset to 62 (inclusive).

[0269] 13) The requirement for bitstream consistency is that for i in the range of 0 to numQpTables − 1 (inclusive) and j in the range of 0 to sps_num_points_in_qp_table_minus1[i] + K (e.g., K = 0 or 1) (inclusive) and sps_num_points_in_qp_table_minus1[i] + K), the value of qpOutVal[i][j] should be in the range of −QpBdOffset to 62 (inclusive).

[0270] 6. Examples

[0271] The following are some example embodiments of some aspects of the invention outlined in this section, which can be applied to the VVC specification. Most of the relevant parts that have been added or modified are indicated by bold, italics, or underline, and some deleted parts are indicated by [[]].

[0272] 6.1. Example 1

[0273] This embodiment applies to item 1 and its sub-items.

[0274] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics ...

[0276] When `sps_ref_pic_resampling_enabled_flag` equals 1, reference image resampling is enabled, and the current image of the reference SPS can have stripes of reference images from the active entry of the reference image list, which has the following characteristics different from the current image: One or more of the following parameters: 1) pps_pic_width_in_luma_samples, 2) pps_pic_height_in_luma_samples, 3) pps_scaling_win_left_offset, 4) pps_scaling_win_right_offset, 5) pps_scaling_win_top_offset 6) `pps_scaling_win_bottom_offset` [[and 7) `sps_num_subpics_minus1`], `sps_ref_pic_resampling_enabled_flag` equal to 0 indicates that reference picture resampling is disabled, and the current picture of the reference SPS cannot have stripes of reference pictures from the active entry of the reference picture list, which has stripes different from the current picture. 6 One or more of the following parameters: [[7]]

[0277] Note 3 – When sps_ref_pic_resampling_enabled_flag equals 1, for the current image, it has the above-mentioned features different from those of the current image. 6 [[7]] The reference image for one or more of the parameters can belong to the same layer as the layer containing the current image or a different layer.

[0278] A value of 1 for `sps_res_change_in_clvs_allowed_flag` indicates that the image spatial resolution can be changed within the CLVS of the reference SPS. A value of 0 for `sps_res_change_in_clvs_allowed_flag` indicates that the image spatial resolution will not be changed within any CLVS of the reference SPS. When it does not exist, the value of `sps_res_change_in_clvs_allowed_flag` is inferred to be 0. ...

[0280] ...

[0282] fRefWidth is set to be equal to CurrPicScalWinWidthL of the reference image RefPicList[i][j].

[0283] refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, and refScalingWinBottomOffset are set to the values ​​of pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the reference image RefPicList[i][j], respectively.

[0284]

[0285]

[0286] 6.2. Example 2

[0287] This embodiment is used for item 2.

[0288] 8.3.2 Decoding process based on the reference image list ...

[0290] fRefWidth is set to be equal to CurrPicScalWinWidthL of the reference image RefPicList[i][j]. refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, and refScalingWinBottomOffset are set to be equal to the values ​​of pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the reference image RefPicList[i][j], respectively.

[0291] fRefNumSubpics is set to be equal to sps_num_subpics_minus1 of the reference image RefPicList[i][j]. ...

[0293] 6.3. Example 3

[0294] This embodiment is used for item 3.

[0295] 8.3.2 Decoding process based on the reference image list

[0296] fRefWidth is set to equal to CurrPicScalWinWidthL of the reference image RefPicList[i][j]. fRefHeight is set to equal to CurrPicScalWinHeightL of the reference image RefPicList[i][j].

[0297] refPicWidth, refPicHeight, refScalingWinLeftOffset, refScalingWinRightOffset, refScalingWinTopOffset, and refScalingWinBottomOffset are respectively set to equal the values ​​of pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset of the reference image RefPicList[i][j].

[0298] RefNumSubpics is set to be equal to sps_num_subpics_minus1 of the reference image RefPicList[i][j]. ...

[0299] 6.4. Example 4

[0300] This example is used for item 5. The PTL syntax is changed as follows:

[0301]

[0302] The GCI syntax has been changed as follows:

[0303]

[0304] gci_num_constraint_bytes specifies the number of bytes in all syntax elements of the general_constraint_info() syntax structure, excluding the syntax element itself.

[0305] The variable numGciBits is deduced to be equal to the number of bits of all syntax elements under the condition "if( gci_num_constraint_bytes > 0 )", excluding the gci_reserved_bit[i] syntax element.

[0306] The number of syntax elements in gci_reserved_bit[i] should be less than or equal to 7.

[0307] Alternatively, the value of gci_num_constraint_bytes should be equal to 0 or Ceil(Log2(numGciBits)).

[0308] gci_reserved_bit[i] can have any value. The decoder should ignore the value of gci_reserved_bit[i] (if it exists).

[0309] 6.5. Example 5

[0310] This example is used for item 6. The PTL syntax is changed as follows:

[0311]

[0312] The GCI syntax has been changed as follows:

[0313]

[0314] A value of 1 for `gci_present_flag` indicates the presence of a GCI field. A value of 0 for `gci_present_flag` indicates the absence of a GCI field.

[0315] The variable numGciBits is deduced to be equal to the number of bits of all syntax elements under the condition "if( gci_present_flag )", excluding the gci_reserved_bit[i] syntax element.

[0316] When gci_present_flag equals 1, the variable gciNumConstraintBytes is set to equal Ceil(Log2(numGciBits)).

[0317] Note: The number of elements in the gci_reserved_bit[i] syntax is less than or equal to 7.

[0318] 6.6. Example 6

[0319] The PTL syntax has been changed as follows:

[0320]

[0321] The GCI syntax has been changed as follows:

[0322]

[0323] A `gci_present_flag` value of 1 indicates the existence of `general_constraint_info()`. A `gci_present_flag` value of 0 indicates the absence of `general_constraint_info()`.

[0324] The variable numGciBits is deduced to be equal to the number of bits of all syntax elements excluding the gci_reserved_bit[i].

[0325] The variable gciNumConstraintBytes is set to equal Ceil(Log2(numGciBits)).

[0326] Note: The number of elements in the gci_reserved_bit[i] syntax is less than or equal to 7.

[0327] Figure 1 This is a block diagram illustrating an example video processing system 1900 to which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet and Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0328] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described herein. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1906, the output of codec component 1904 can be stored or transmitted via connected communication. The stored or transmitted bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or displayable video that is sent to display interface 1810. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as "codec" operations or tools, it should be understood that encoding tools or operations are used at the encoder, and corresponding decoding tools or operations will be performed by the decoder to reverse the encoded results.

[0329] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, and so on. The technologies described herein can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0330] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple processors) 3602 can be configured to implement one or more methods described herein. The memories (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.

[0331] Figure 4 A block diagram of an example video codec system 100 that can utilize the techniques disclosed herein is shown.

[0332] like Figure 4 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, and the source device may be referred to as a video decoding device.

[0333] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0334] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage media / server 130b for access by destination device 120.

[0335] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0336] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.

[0337] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.

[0338] Figure 5 This is a block diagram illustrating an example of a video encoder 200. The video encoder can be... Figure 4 The video encoder 114 in the system 100 shown.

[0339] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0340] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (including a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0341] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0342] Furthermore, some components (e.g., motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for interpretative purposes... Figure 5 The example is shown separately.

[0343] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0344] The mode selection unit 203 may, for example, select one of multiple encoding / decoding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signaling and intra-frame prediction signaling. In the case of inter-frame prediction, the mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0345] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0346] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0347] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in list 0 or list 1, which includes the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0348] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference images in lists 0 and 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0349] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.

[0350] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may signal the motion information of the current video block to another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0351] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0352] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0353] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.

[0354] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0355] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0356] In other examples, the current video block may not have residual data, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.

[0357] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0358] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0359] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block, which is stored in buffer 213.

[0360] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0361] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0362] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 4 The video decoder 114 in the system 100 shown.

[0363] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0364] exist Figure 6 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The decoding process is the inverse of the encoding process described.

[0365] Entropy decoding unit 301 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and from the entropy-decoded video data, motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and merge modes.

[0366] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0367] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of video blocks, to calculate interpolated sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0368] The motion compensation unit 302 may use some syntax information to determine the size of the blocks of frames and / or stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0369] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0370] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0371] The following is a list of preferred solutions for some embodiments.

[0372] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., items 1 through 3).

[0373] 1. A method for video processing (e.g., Figure 3 Method 600 in the document includes execution (602).

[0374] The conversion between a video and its codec representation, wherein the codec representation conforms to a format rule, wherein the format rule specifies the use of a syntax field that indicates the applicability of reference image resampling to the corresponding segment of the video.

[0375] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 1).

[0376] 2. The method according to Solution 1, wherein the rule specifies that the value of the syntax field is derived independently of the values ​​of sub-images included in the sequence parameter set corresponding to the video segment.

[0377] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 2).

[0378] 3. The method according to any one of solutions 1 to 2, wherein the rule specifies that the value of the syntax field is derived based on whether the sub-image is considered as an image for the transformation.

[0379] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 3).

[0380] 4. The method according to any one of solutions 1 to 3, wherein the rule specifies that the value of the syntax field is derived based on whether the current image and the reference image of the current image belong to the same layer and / or whether inter-layer prediction is allowed.

[0381] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 4).

[0382] 5. A video processing method, comprising: performing a conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies that the value of a first syntax element in a picture header or a stripe header is constrained based on the value of a second syntax element corresponding to an access unit delimiter.

[0383] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., items 5 and 6).

[0384] 6. A video processing method, comprising: performing a conversion between a video and a codec representation of the video, wherein the codec representation conforms to a format rule, wherein the format rule specifies whether and how to include one or more syntax elements in a general constraint information field.

[0385] 7. The method according to Solution 6, wherein the rule specifying whether the one or more syntax elements are included in the codec representation is based on the number of bytes of the second field in the codec representation.

[0386] 8. The method according to any one of solutions 6 to 7, wherein the rule specifies that a plurality of reserved bits are present when the one or more syntax elements are included in the general constraint information syntax element.

[0387] 9. The method according to any one of solutions 1 to 8, wherein performing the conversion includes encoding the video to generate the codec representation.

[0388] 10. The method according to any one of solutions 1 to 8, wherein performing the conversion includes parsing and decoding the codec representation to generate the video.

[0389] 11. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions 1 to 10.

[0390] 12. A video encoding apparatus, comprising a processor configured to implement one or more of the methods described in solutions 1 to 10.

[0391] 13. A computer program product having computer code stored thereon, wherein when executed by a processor, the code causes the processor to implement the method of any one of solutions 1 to 10.

[0392] 14. A method, apparatus or system described herein.

[0393] Figure 7 This is a flowchart representation of a method 700 for processing video data according to one or more embodiments of the present technology. Method 700 includes, in operation 710, performing a conversion between a current frame of video and a bitstream of video according to a rule. The rule specifies that, in response to (1) one or more stripes in the current frame are allowed to have stripe types other than intra-frame (I) stripe types, and in response to (2) the presence of a reference picture list (RPL) information in the picture header, the number of entries in the reference picture list of the current frame being greater than 0.

[0394] In some embodiments, the reference image list is reference image list 0. In some embodiments, a first syntax flag in the image header indicates that one or more stripes in the current image are allowed to have stripe types other than intra (I) stripe types. In some embodiments, a second syntax flag in the image parameter set indicates whether the RPL information is present in the image header rather than in the stripe header. In some embodiments, in response to the presence of the RPL information in the image header, the number of entries in reference image list 0 is 0 only if the one or more stripes in the current image are all intra (I) stripes. In some embodiments, in response to the presence of the RPL information in the image header, the number of entries in reference image list 0 is greater than 0 if at least one of the one or more stripes in the current image is a prediction (P) stripe or a bidirectional prediction (B) stripe.

[0395] In some embodiments, the number of entries in the reference image list 0 of the current image and / or how they exist in the bitstream is based on whether one or more slices in the current image are allowed to have slice types other than intra-frame (I) slice types. In some embodiments, in response to one or more slices in the current image being allowed to have slice types other than intra-frame (I) slice types, the number of entries in the reference image list 0 of the current image is indicated in the bitstream as (number of entries in reference image list 0 - 1). In some embodiments, in response to the number of entries in the reference image list 0 of the current image being equal to 0, one or more slices of the current image are intra-frame slices. In some embodiments, in response to one or more slices in the image being only intra-frame (I) slices, the number of entries in the reference image list 0 of the current image is not present in the bitstream. In some embodiments, in response to the number of entries in the reference image list 0 of the current image being equal to 0, one or more slices of the current image are intra-frame slices. In some embodiments, a value of 0 for the first syntax flag indicates that one or more slices of the current image are intra-frame slices. In some embodiments, the bitstream indicates whether or how (1) one or more slices in the current picture are allowed to have a slice type other than intra (I) slice type, and (2) whether reference picture list (RPL) information is present in the picture header based on the number of entries in reference picture list 0 or reference picture list 1 of the current picture. In some embodiments, if the number of entries in reference picture list 0 is equal to 0, the first syntax flag in the picture header indicating whether one or more slices in the current picture are allowed to have a slice type other than intra (I) slice type is omitted in the bitstream. In some embodiments, if the number of entries in reference picture list 0 is equal to 0, one or more slices in the current picture are intra slices. In some embodiments, if the number of entries in reference picture list 1 is equal to 0, one or more slices in the current picture are not bidirectional prediction slices. In some embodiments, if the number of entries in reference picture list 1 is equal to 0, one or more slices in the current picture are either intra slices or prediction slices.

[0396] Figure 8 This is a flowchart representation of a method 800 for processing video data according to one or more embodiments of the present technology. The method 800 includes an operation 810 performing a conversion between a current frame of the video and a bitstream of the video according to a rule. The rule specifies that a variable indicating the number of sub-frames in each frame of the video is not used to derive a first syntax flag indicating whether a reference frame resampling constraint is satisfied.

[0397] In some embodiments, the rule further specifies that, if (1) the number of sub-images in the current image differs from the number of sub-images in the reference image of the current image and (2) at least one sub-image is associated with a flag indicating that a sub-image is considered an image, then encoding / decoding tools compatible with satisfying the reference image resampling constraint are disabled. In some embodiments, the rule specifies that variables are not used to derive a second syntax flag indicating whether reference image resampling is enabled. In some embodiments, the rule applies to transformations in response to inter-layer prediction being allowed.

[0398] Figure 9 This is a flowchart representation of a method 900 for processing video data according to one or more embodiments of the present technology. Method 900 includes, in operation 910, performing a conversion between a current frame of the video and a bitstream of the video according to a rule. The rule specifies a syntax flag indicating whether a reference frame resampling constraint is satisfied, determined based on whether a sub-frame of the current frame is treated as a frame during the conversion.

[0399] In some embodiments, sub-images of an image are associated with variables, each indicating whether a corresponding sub-image is considered an image, and syntax flags are determined based on at least one of the variables. In some embodiments, reference image resampling is allowed when the number of sub-images in the current image differs from the number of sub-images in the reference image of the current image, and none of the sub-images in the current image are considered images. In some embodiments, syntax flags indicating whether reference image resampling constraints are used for reference image resampling are determined based on the scaling size or scaling offset. In some embodiments, reference image resampling is always enabled when the number of sub-images in the current image differs from the number of sub-images in the reference image of the current image, and at least one sub-image in the current image is considered an image.

[0400] In some embodiments, in response to satisfying reference image resampling constraints for reference image resampling, the encoding / decoding tools associated with reference image resampling are disabled. In some embodiments, the encoding / decoding tools include optical flow prediction thinning (PROF), bidirectional optical flow (BDOF), or decoder-side motion vector thinning (DMVR). In some embodiments, in response to sub-images in the current image being considered images and the number of sub-images in the current image being different from the number of sub-images in the reference image of the current image, the encoding / decoding tools associated with reference image resampling are disabled.

[0401] Figure 10This is a flowchart representation of a method 1000 for processing video data according to one or more embodiments of the present technology. The method 1000 includes operation 1010 performing a conversion between a current frame of the video and a bitstream of the video according to rules. These rules specify a syntax flag indicating whether a reference frame resampling constraint is satisfied, determined based on whether the current frame and its reference frame are in the same layer and / or whether inter-layer prediction is allowed for the conversion. In some embodiments, this syntax flag is represented in the bitstream as RprConstraintsActiveFlag.

[0402] Figure 11 This is a flowchart representation of a method 1100 for processing video data according to one or more embodiments of the present technology. The method 1100 includes, in operation 1110, performing a conversion between video and a video bitstream according to a rule. The rule specifies that the value of a first syntax element in a picture header or stripe header is constrained by the value of a second syntax element present in an Access Unit Delimiter (AUD) in the bitstream.

[0403] In some embodiments, a first syntax element in the picture header specifies whether the current picture is a Progressive Decoding Refresh (GDR) picture or an Intra-Random Access Point (IRAP) picture, and wherein a second syntax element in the AUD specifies whether the access unit containing the AUD is a GDR access unit or an IRAP access unit. In some embodiments, in response to the second syntax element specifying that the access unit is neither a GDR access unit nor an IRAP access unit, the first syntax element is not equal to a specific value. In some embodiments, the specific value is 1. In some embodiments, the second syntax element may or may not be present in the bitstream. In some embodiments, the first syntax element is equal to 0. In some embodiments, a first syntax element in the slice header specifies the slice type of the slice in the current picture, and a second syntax element specifies the slice type associated with the access unit containing the AUD. In some embodiments, in response to specifying that the slice type associated with the access unit is an intra-slice, the first element is not equal to a specific value indicating that the slice type is a predictive slice or a bidirectional predictive slice. In some embodiments, the specific value is 0 or 1. In some embodiments, the second syntax element may or may not be present in the bitstream. In some embodiments, the first syntax element is equal to 2. In some embodiments, in response to the second syntax element specifying that the slice type present in the access unit is a predicted slice or an intra-frame slice, the first syntax element in the slice header is not equal to a specific value. In some embodiments, in response to the second syntax element being equal to 1, the first syntax element is equal to 1 or 2. In some embodiments, in response to the second syntax element being equal to 1, the first syntax element is not equal to 0.

[0404] In some embodiments, the values ​​of the first syntax element and the second syntax element are aligned. In some embodiments, the names of the first syntax element and the second syntax element are aligned. In some embodiments, the first syntax element in the picture header specifies whether the current picture is a Progressive Decoding Refresh (GDR) picture or an Intra-Random Access Point (IRAP) picture, and wherein the second syntax element specifies whether the bitstream is a single-layer bitstream or whether the current picture belongs to an IRAP access unit or a GDR access unit. In some embodiments, in response to the second syntax element specifying that the bitstream is a single-layer bitstream, the first syntax element specifies that the access unit is not an IRAP or GDR access unit. In some embodiments, the value of the first syntax element is not equal to a specific value. In some embodiments, in response to the second syntax element specifying that the bitstream is a single-layer bitstream, the first syntax element specifies that the access unit is an IRAP or GDR access unit.

[0405] Figure 12 This is a flowchart representation of a method 1200 for processing video data according to one or more embodiments of the present technology. Method 1200 includes, in operation 1210, performing a conversion between video and a video bitstream according to rules. The rules specify whether a syntax element in a general constraint information syntax structure exists in the bitstream based on a constraint byte number greater than 0 or a reserved byte number.

[0406] In some embodiments, the number of constrained or reserved bytes is limited to a specific value based on video characteristics. In some embodiments, characteristics include syntax elements of general constraint information, video grade, or version of standard information. In some embodiments, syntax flags are used to indicate the presence of syntax elements in the general constraint information syntax structure in the bitstream, and in response to the syntax flag indicating the presence of syntax elements, the total number of bits is aligned to 8×N using one or more reserved bits, where N is a positive integer. In some embodiments, in response to the syntax flag indicating the absence of syntax elements, seven reserved bits are included in the bitstream. In some embodiments, in response to the syntax flag indicating the presence of syntax elements, the general constraint information syntax structure is present in the bitstream. In some embodiments, in response to the syntax flag indicating the absence of syntax elements, the value of the general constraint information syntax structure is determined according to a default value. In some embodiments, the general constraint information syntax structure is positioned after the general grade syntax structure, or after another syntax structure specifying whether a sub-layer indicates the presence of level information.

[0407] Figure 13This is a flowchart representation of a method 1300 for processing video data according to one or more embodiments of the present technology. Method 1300 includes, at operation 1310, performing a conversion between a current block of video and a video bitstream according to a rule. The rule specifies the selective indication of the use of adaptive color transformation based on information about one or more prediction modes applicable to the current block.

[0408] In some embodiments, the syntax flag indicating whether adaptive color transformation is enabled is based on the prediction mode of the current block being non-intra-frame, non-inter-frame, and / or non-intra-frame block copy (IBC). In some embodiments, in response to the fact that intra-frame, inter-frame, or IBC prediction modes are not applicable to the current block, the use of adaptive color transformation is not indicated in the bitstream. In some embodiments, in response to the application of adaptive color transformation to the current block, different encoding / decoding modes are not used for the current block. In some embodiments, different encoding / decoding modes include palette coding mode.

[0409] In some embodiments, the conversion includes encoding the video into a bitstream. In some embodiments, the conversion includes decoding the video from the bitstream.

[0410] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that enables machine-readable propagation signals, or combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also contain code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0411] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer or multiple computers located in one location or distributed across multiple locations and interconnected via a communication network.

[0412] The processes and logic flows described in this specification can be executed by one or more programmable processors to execute one or more computer programs, thereby performing functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the device can be implemented as dedicated logic circuitry, such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).

[0413] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0414] Although this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as a description of features that may be specific to particular embodiments of a particular invention. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations of sub-combinations.

[0415] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or to perform all shown operations to achieve the desired effect. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0416] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A method for processing video data, comprising: Perform the conversion between the current image of the video and the bitstream of the video according to the first and second rules. The first rule includes a bitstream constraint, which stipulates that: in response to (1) one or more stripes in the current picture are allowed to have stripe types other than intra-frame I stripe types, and (2) reference picture list RPL information exists in the picture header, the number of entries in the reference picture list of the current picture is greater than 0, wherein the reference picture list is reference picture list 0; The current image of the video includes the current video block. According to the second rule, the range of values ​​for the fourth syntax element of the quantization parameter QP table associated with the transformation of the current video block depends on one or more other syntax elements, including the fifth syntax element. The second rule specifies that the range of values ​​for the fourth syntax element is based on the values ​​of the fifth syntax element, which specifies the starting lightness and chromaticity QP used to describe the QP table.

2. The method according to claim 1, wherein, The first syntax flag in the image header indicates whether one or more stripes in the current image are allowed to have stripe types other than the intra-frame I stripe type.

3. The method according to claim 1, wherein, The second syntax flag in the image parameter set indicates whether the RPL information exists in the image header rather than in the strip header.

4. The method according to claim 1, wherein, The number of entries in the reference image list of the current image is indicated by a third syntax element included in the bitstream.

5. The method according to any one of claims 1 to 4, wherein, The conversion includes encoding the video into the bitstream.

6. The method according to any one of claims 1 to 4, wherein, The conversion includes decoding the video from the bitstream.

7. A video decoding apparatus, comprising a processor configured to implement the method of any one of claims 1 to 4 and 6.

8. A video encoding apparatus, comprising a processor configured to implement the method of any one of claims 1 to 5.

9. A computer program product having computer code stored thereon, the code causing the processor to perform the method of any one of claims 1 to 6 when executed by a processor.

10. A non-transitory computer-readable recording medium storing a bit stream and a computer program, wherein, when executed by a processor, the computer program generates the bit stream by performing the following method, wherein, The method includes: The bitstream of the video is generated from the current image in the video according to the first and second rules. The first rule includes a bitstream constraint, which specifies that: in response to (1) one or more stripes in the current image are allowed to have stripe types other than intra-frame I stripe types, and (2) reference image list (RPL) information exists in the image header, the number of entries in the reference image list of the current image is greater than 0, wherein the reference image list is reference image list 0; The current image of the video includes the current video block. According to the second rule, the range of values ​​for the fourth syntax element of the quantization parameter QP table associated with the transformation of the current video block depends on one or more other syntax elements, including the fifth syntax element. The second rule specifies that the range of values ​​for the fourth syntax element is based on the values ​​of the fifth syntax element, which specifies the starting lightness and chromaticity QP used to describe the QP table.

11. The non-transitory computer-readable recording medium according to claim 10, wherein, The first syntax flag in the image header indicates whether one or more stripes in the current image are allowed to have stripe types other than the intra-frame I stripe type, and the second syntax flag in the image parameter set indicates whether the RPL information exists in the image header rather than in the stripe header.

12. The non-transitory computer-readable recording medium according to claim 10, wherein, The number of entries in the reference image list of the current image is indicated by a third syntax element included in the bitstream.

13. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform the conversion between the current image of the video and the bitstream of the video according to the first and second rules. The first rule includes a bitstream constraint, which stipulates that: in response to (1) one or more stripes in the current picture are allowed to have stripe types other than intra-frame I stripe types, and (2) reference picture list RPL information exists in the picture header, the number of entries in the reference picture list of the current picture is greater than 0, wherein the reference picture list is reference picture list 0; The current image of the video includes the current video block. According to the second rule, the range of values ​​for the fourth syntax element of the quantization parameter QP table associated with the transformation of the current video block depends on one or more other syntax elements, including the fifth syntax element. The second rule specifies that the range of values ​​for the fourth syntax element is based on the values ​​of the fifth syntax element, which specifies the starting lightness and chromaticity QP used to describe the QP table.

14. The apparatus according to claim 13, wherein, The first syntax flag in the image header indicates whether one or more stripes in the current image are allowed to have stripe types other than the intra-frame I stripe type.

15. The apparatus according to claim 13, wherein, The second syntax flag in the image parameter set indicates whether the RPL information exists in the image header rather than in the strip header.

16. The apparatus according to claim 13, wherein, The number of entries in the reference image list of the current image is indicated by a third syntax element included in the bitstream.

17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: Perform the conversion between the current image of the video and the bitstream of the video according to the first and second rules. in, The first rule includes a bitstream constraint, which specifies that in response to (1) one or more stripes in the current picture are allowed to have stripe types other than intra-frame I stripe types, and (2) reference picture list RPL information exists in the picture header, the number of entries in the reference picture list of the current picture is greater than 0, wherein the reference picture list is reference picture list 0; The current image of the video includes the current video block. According to the second rule, the range of values ​​for the fourth syntax element of the quantization parameter QP table associated with the transformation of the current video block depends on one or more other syntax elements, including the fifth syntax element. The second rule specifies that the range of values ​​for the fourth syntax element is based on the values ​​of the fifth syntax element, which specifies the starting lightness and chromaticity QP used to describe the QP table.

18. The non-transitory computer-readable storage medium according to claim 17, wherein, The first syntax flag in the image header indicates whether one or more stripes in the current image are allowed to have stripe types other than the intra-frame I stripe type, and the second syntax flag in the image parameter set indicates whether the RPL information exists in the image header rather than in the stripe header.

19. The non-transitory computer-readable storage medium according to claim 17, wherein, The number of entries in the reference image list of the current image is indicated by a third syntax element included in the bitstream.

20. A method for storing a video bitstream, comprising: Generate the bitstream of the video from the current image in the video according to the first and second rules, and The bitstream is stored in a non-transitory computer-readable recording medium. The first rule includes a bitstream constraint, which stipulates that: in response to (1) one or more stripes in the current picture are allowed to have stripe types other than intra-frame I stripe types, and (2) reference picture list RPL information exists in the picture header, the number of entries in the reference picture list of the current picture is greater than 0, wherein the reference picture list is reference picture list 0; The current image of the video includes the current video block. According to the second rule, the range of values ​​for the fourth syntax element of the quantization parameter QP table associated with the transformation of the current video block depends on one or more other syntax elements, including the fifth syntax element. The second rule specifies that the range of values ​​for the fourth syntax element is based on the values ​​of the fifth syntax element, which specifies the starting lightness and chromaticity QP used to describe the QP table.