Transmitting general constraints for video coding
By modifying the general constraint information syntax structure of VVC version 2, the incompatibility issues between video coding standards are resolved, the compatibility of VVC version 1 and version 2 decoders is ensured, and the stability and consistency of video coding are improved.
Patent Information
- Application Number
- CN202411250760.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-13
- Filing Date
- 2022-11-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-11-08
AI Technical Summary
The existing video coding standard VVC version 2 has unclear and inconsistent decoder implementation methods in the transmission of universal constraint information, resulting in incompatibility between different versions of the video coding standard and affecting decoding stability.
By introducing the transmission and initialization method of general constraint information for video coding, the general constraint information syntax structure of VVC version 2 is modified to ensure its compatibility with the VVC version 1 decoder and to be decoded without causing desynchronization, including adjusting the value processing method and inference rules of gci_num_additional_bits.
The high-level syntax of the VVC version 2 code stream is compatible with the VVC version 1 decoder, eliminating the ambiguity of decoder behavior and improving the stability and consistency of video encoding.
Smart Images

Figure CN118921493B_ABST
Abstract
Description
[0001] This application is a divisional application of the PCT international patent application PCT / US2022 / 079490 with an application date of November 8, 2022, which entered the Chinese national phase and has the Chinese patent application number 202280088199.0 and the invention name “General Constraint Information for Transmitting Video Coding”.
[0002] Cross-references
[0003] This application claims priority to U.S. Provisional Application No. 63 / 266,615, filed on January 10, 2022, entitled “Signaling Methods for General Constraints Information for Video Coding,” U.S. Provisional Application No. 63 / 266,616, filed on January 10, 2022, entitled “Initialization Method for General Constraint Information Flags for Video Coding,” and U.S. Provisional Application No. 63 / 266,765, filed on January 13, 2022, entitled “Signaling and Initialization Methods for General Constraints Information for Video Coding,” all of which are incorporated herein by reference in their entirety. Technical Field
[0004] The present disclosure relates generally to video processing and, more particularly, to signaling and initializing general constraint information for video coding. Background Art
[0005] The ubiquity of video cameras, such as smartphones, tablets, and computers, has made capturing videos or images easier than ever before. However, the amount of data required for even a short video can be very large. Video coding techniques (including video encoding and decoding) allow video data to be compressed to a smaller size, thereby enabling the storage and transmission of various videos. Video coding has been widely used in digital television broadcasting, Internet and mobile network video transmission, real-time applications (such as video chat, video conferencing), DVDs and Blu-ray discs, and other fields. In order to reduce the storage space used to store videos and / or the network bandwidth consumption used to transmit videos, it is desirable to improve the efficiency of video coding schemes. Summary of the Invention
[0006] Some embodiments relate to the transmission and initialization of general constraint information for video coding. In one example, a video decoding method includes: extracting an additional bit count M from a video bitstream, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits included in the video bitstream, the additional bits including flag bits, the flag bits indicating that respective additional coding tools are constrained for the video, and an expected value of the additional bit count is 0 or 6; if it is determined that the extracted additional bit count M is greater than 6, extracting M-6 bits following six flag bits from the bitstream; and decoding the remaining portion of the video bitstream into an image independently of the extracted M-6 bits and based at least in part on the constraints specified for the respective additional coding tools by the six flag bits.
[0007] In another example, a non-transitory computer-readable medium stores program code thereon, the program code being executable by one or more processing devices to perform operations. The operations include: extracting an additional bit count M from a bitstream of a video, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits included in the bitstream of the video, the additional bits including flag bits, the flag bits respectively indicating that each additional coding tool is constrained for the video, and an expected value of the additional bit count is 0 or 6; if it is determined that the extracted additional bit count M is greater than 6, extracting M-6 bits following the six flag bits in the bitstream; and independently of the extracted M-6 bits, decoding the remaining portion of the bitstream of the video into an image based at least in part on the constraints specified for each additional coding tool by the six flag bits.
[0008] In another example, a system includes a processing device and a non-transitory computer-readable medium communicatively coupled to the processing device. The processing device is configured to execute program code stored in the non-transitory computer-readable medium, thereby performing operations. The operations include: extracting an additional bit count M from a bitstream of a video, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits contained in the bitstream of the video, the additional bits including flag bits, the flag bits respectively indicating that each additional coding tool is constrained for the video, and an expected value of the additional bit count is 0 or 6; if it is determined that the extracted additional bit count M is greater than 6, extracting M-6 bits following the six flag bits in the bitstream; and independently of the extracted M-6 bits, and based at least in part on the constraints specified for each additional coding tool by the six flag bits, decoding the remaining portion of the bitstream of the video into an image.
[0009] These illustrative embodiments are not mentioned to limit or define the present disclosure, but to provide examples to aid understanding of the present disclosure. Other embodiments are discussed and further described in the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The features, embodiments, and advantages of the present disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings.
[0011] Figure 1 FIG2 is a block diagram illustrating an example of a video encoder for implementing the embodiments proposed herein.
[0012] Figure 2 FIG. 4 is a block diagram illustrating an example of a video decoder for implementing the embodiments proposed herein.
[0013] Figure 3 An example of dividing a picture in a video into coding tree units according to some embodiments of the present disclosure is described.
[0014] Figure 4 An example of dividing a coding tree unit into coding units according to some embodiments of the present disclosure is described.
[0015] Figure 5 An example of a video decoding process according to some embodiments of the present disclosure is depicted.
[0016] Figure 6 Another example of a video decoding process according to some embodiments of the present disclosure is depicted.
[0017] Figure 7 Another example of a video decoding process according to some embodiments of the present disclosure is depicted.
[0018] Figure 8 An example of a computing system that can be used to implement some embodiments of the present disclosure is depicted. DETAILED DESCRIPTION
[0019] Various embodiments provide for the transmission and initialization of general constraint information for video coding. As described above, more and more video data is generated, stored and transmitted. It is beneficial to improve both the efficiency of video coding technology and the stability of video coding so that the video signal can be successfully decoded at the decoder end. Problems related to video decoding stability include incompatibility and inconsistency. With the development of video coding technology, newer video coding standards have also emerged. One of the video coding standards is the Versatile Video Coding Standard Version 1, which was jointly issued by the International Organization for Standardization (ISO) under the name of "ISO / IEC 23090-3:2021 Information technology - Coding representation for immersive media - Part 3: Versatile video coding" and by the International Telecommunication Union (ITU) under the name of "ITU-T H.26 Recommendation (08 / 2020): Versatile Video Coding". In the present disclosure, Versatile Video Coding Standard Version 1 may be referred to as "VVC Version 1" or "VVCvl". VVC version 1 has been superseded by version 2 of the Versatile Video Coding standard, which will be jointly published by ISO as "ISO / IEC 23090-3:2022 Information technology — Coded representation for immersive media — Part 3: Versatile video coding" and by ITU as "ITU-TH.266 Recommendation (04 / 2022): Versatile video coding." In this disclosure, version 2 of the Versatile Video Coding standard may be referred to as "VVC version 2" or "VVCv2." To enable video signals encoded with older versions of the video coding standard to be successfully decoded by video decoders using newer versions of the standard, video coding schemes should be designed to be backwards compatible with older versions of the standard. However, the transmission of generic constraint information used in the current draft of VVC version 2 can cause video decoding to be out of sync, a serious incompatibility issue between different versions of the video coding standard. Furthermore, in the current draft of VVC version 2, the generic constraint flag associated with the generic constraint information may be undefined in some cases, leading to ambiguous and inconsistent decoder implementations. Various embodiments described herein address these issues by introducing a method for transmitting and initializing general constraint information for video coding, thereby improving the stability of video coding.
[0020] In the VVC standard, the general constraint information (GCI) syntax structure general_constraints_info() is used to indicate specific constraint properties of a codestream. GCI contains a list of constraint flags and non-flag syntax elements. The binary flag gci_present_flag is used to specify whether the GCI syntax element is present. In some embodiments, if the VVC version 2 codestream transmits general constraint information (i.e., the value of gci_present_flag is 1), and the VVC version 2 general constraint information consists of N additional coding tools that may be constrained, the value of the syntax element gci_num_additional_bits corresponding to these N additional coding tools can only be set to 0 or N. If gci_num_additional_bits is set to 0, the general constraint flags of these N additional coding tools will not be transmitted. If gci_num_additional_bits is set to N, the next N bits in the codestream are used to transmit the general constraint flags of these N additional coding tools. In one example, N is set to 6.
[0021] In some examples, a VVC version 2 codestream does not allow gci_num_additional_bits to be set to a value other than 0 or N. However, a VVC version 2 decoder can still process general constraint information that sets gci_num_additional_bits to a value other than 0 or N. For example, gci_num_additional_bits can be set to M, where M is greater than 0 and less than N, or M is greater than N. If M is greater than 0 and less than N, after decoding the gci_num_additional_bits syntax element, the decoder extracts M bits from the codestream and discards them. If M is greater than N, after decoding the gci_num_additional_bits syntax element, the decoder extracts N bits from the codestream and interprets them as a general constraint flag for N additional coding tools. The decoder then extracts MN bits from the codestream and discards them. In other examples, a VVC version 2 decoder does not need to process general constraint information that sets gci_num_additional_bits to a value greater than 0 but less than N. Legal VVC version 2 codestreams can only set the value of gci_num_reserved_bits to 0 or N. Future versions of VVC codestreams will not allow gci_num_additional_bits to be set to a value between 0 and N.
[0022] In some embodiments, the general constraint information flag is initialized to address the aforementioned issue of ambiguity and inconsistency in decoder implementations. In these embodiments, when gci_present_flag is equal to 1 and gci_num_additional_bits is equal to 0, general_constraints_info() does not impose constraints on the coding tools associated with the general constraint information flag. In examples where a flag value of 0 indicates no constraints, the value of the general constraint information flag is inferred to be 0 when the flag is not present.
[0023] The embodiments described in this disclosure provide methods for transmitting and inferring general constraint flags for additional coding tools in VVC version 2. Unlike the prior art, the high-level syntax codestreams generated by the methods described in this disclosure are compatible with VVC version 1 decoders and can be decoded without causing desynchronization between the behavior of VVC version 1 decoders and VVC version 2 decoders. The inference rules described in this disclosure eliminate ambiguity in the VVC version 2 decoding behavior of VVC version 2 GCI syntax elements. These techniques can become effective coding tools in various video coding standards.
[0024] Referring now to the drawings, FIG1 is a block diagram illustrating an example of a video encoder 100 that may be used to implement the embodiments presented herein. Figure 1 In the example shown, the video encoder 100 includes a segmentation module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, a loop filter module 120, an intra-frame prediction module 126, an inter-frame prediction module 124, a motion estimation module 122, a decoded picture buffer 130 and an entropy encoding module 116.
[0025] The input to the video encoder 100 is an input video 102, which consists of a series of pictures (also called frames or images). In a block-based video encoder, for each picture, the video encoder 100 uses a partitioning module 112 to partition the picture into blocks 104, each of which contains a number of pixels. These blocks can be macroblocks, coding tree units, coding units, prediction units, and / or prediction blocks. A picture may contain blocks of different sizes, and the block partitioning may vary between pictures in the video. Each block can be encoded using different prediction methods, such as intra prediction, inter prediction, or a mixture of intra and inter prediction.
[0026] Typically, the first picture of a video signal is an intra-frame coded picture, which is coded using only intra-frame prediction. In intra-frame prediction mode, only the coded data in the same picture is used to predict the blocks of the picture. Intra-frame coded pictures can be decoded without using information from other pictures. To perform intra-frame prediction, Figure 1The illustrated video encoder 100 may utilize an intra-frame prediction module 126. The intra-frame prediction module 126 is configured to generate an intra-frame prediction block (prediction block 134) using reconstructed samples from a reconstructed block 136 of a neighboring block in the same picture. Intra-frame prediction is performed based on the intra-frame prediction mode selected for the block. The video encoder 100 then calculates the difference between the block 104 and the intra-frame prediction block 134. This difference is referred to as a residual block 106.
[0027] To further eliminate redundancy in the block, transform module 114 transforms the residual block 106 into a transform domain by transforming the samples in the block. Examples of transforms may include, but are not limited to, discrete cosine transforms (DCT) or discrete sine transforms (DST). The transformed values may be referred to as transform coefficients, representing the residual block in the transform domain. In some examples, the residual block may be directly quantized without undergoing the transform by transform module 114. This mode is referred to as transform skip mode.
[0028] The video encoder 100 may further quantize the transform coefficients using the quantization module 115 to obtain quantized coefficients. Quantization involves dividing the samples by the quantization step size and then rounding them, while inverse quantization involves multiplying the quantized values by the quantization step size. This quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of video samples (transformed or untransformed), thereby using fewer bits to represent the video samples.
[0029] Quantization of coefficients / samples within a block can be performed independently, a quantization method used in some existing video compression standards (such as H.264 and HEVC). For an N*M block, the two-dimensional coefficients of the block can be converted into a one-dimensional array using a certain scan order for coefficient quantization and encoding. Quantization of coefficients within a block can utilize scan order information. For example, the quantization of a given coefficient in a block may depend on the state of the previous quantized value in the scan order. To further improve coding efficiency, multiple quantizers can be used. Which quantizer is used to quantize the current coefficient depends on the information before the current coefficient in the encoding / decoding scan order. This quantization method is called dependent quantization.
[0030] The degree of quantization can be adjusted using the quantization step size. For example, for scalar quantization, different quantization step sizes can be used to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The quantization step size can be indicated by a quantization parameter (QP). The quantization parameter is provided in the video encoding bitstream so that the video decoder can access and apply the quantization parameter for decoding.
[0031] Afterwards, the entropy coding module 116 encodes the quantized samples to further reduce the size of the video signal. The entropy coding module 116 is used to apply an entropy coding algorithm to the quantized samples. In some examples, the quantized samples are binarized into a binary container, and the coding algorithm further compresses the binary container into bits. Examples of binarization methods include, but are not limited to, truncated Rice (TR) and finite k-order exponential Golomb (EGk) combined binarization, and k-order exponential Golomb binarization. Examples of entropy coding algorithms include, but are not limited to, variable length coding (VLC) schemes, content adaptive VLC schemes (CAVLC), arithmetic coding schemes, binarization, content adaptive binary arithmetic coding (CABAC), grammar-based content adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The entropy-coded data is added to the bitstream of the output coded video 132.
[0032] As described above, reconstructed blocks 136 from neighboring blocks are used for intra-frame prediction of blocks in a picture. Generating reconstructed blocks 136 for a block involves calculating a reconstructed residual for the block. The reconstructed residual can be determined by applying inverse quantization and an inverse transform to the quantized residual of the block. The inverse quantization module 118 is used to inverse quantize the quantized samples to obtain dequantized coefficients. The inverse quantization module 118 applies a quantization scheme that is the opposite of the quantization scheme applied by the quantization module 115, using the same quantization step size as the quantization module 115. The inverse transform module 119 is used to apply the inverse transform of the transform applied by the transform module 114, such as an inverse DCT or inverse DST, to the dequantized samples. The output of the inverse transform module 119 is the reconstructed residual for the block in the pixel domain. The reconstructed residual can be added to the prediction block 134 for the block to obtain the reconstructed block 136 in the pixel domain. For blocks that skip transforms, the inverse transform module 119 is not applied. The dequantized samples are the reconstructed residual for the block.
[0033] Blocks in subsequent pictures after the first intra-predicted picture can be encoded using inter-prediction or intra-prediction. In inter-prediction, predictions for blocks in a picture are derived from one or more previously encoded video pictures. To perform inter-prediction, video encoder 100 uses inter-prediction module 124. Inter-prediction module 124 is used to perform motion compensation on blocks based on the motion estimate provided by motion estimation module 122.
[0034] The motion estimation module 122 compares the current block 104 of the current picture with the decoded reference pictures 108 to perform motion estimation. The decoded reference pictures 108 are stored in the decoded picture buffer 130. The motion estimation module 122 selects a reference block from the decoded reference pictures 108 that best matches the current block. The motion estimation module 122 further identifies the offset between the position (e.g., x, y coordinates) of the reference block and the position of the current block. This offset, called a motion vector (MV), is provided to the inter-frame prediction module 124 along with the selected reference block. In some cases, multiple reference blocks may be identified for the current block in multiple decoded reference pictures 108. Consequently, multiple motion vectors are generated and provided to the inter-frame prediction module 124 along with the corresponding reference blocks.
[0035] The inter-frame prediction module 124 uses the motion vector and other inter-frame prediction parameters to perform motion compensation to generate a prediction for the current block, namely, an inter-frame prediction block 134. For example, based on the motion vector, the inter-frame prediction module 124 can locate the prediction block indicated by the motion vector in the corresponding reference picture. If there is more than one prediction block, the prediction blocks are combined with certain weights to generate the prediction block 134 for the current block.
[0036] For an inter-frame prediction block, the video encoder 100 can subtract the inter-frame prediction block 134 from the block 104 to generate a residual block 106. The residual block 106 can be transformed, quantized, and entropy coded in the same manner as the residual of the intra-frame prediction block described above. Similarly, the reconstructed block 136 of the inter-frame prediction block can be obtained by inverse quantizing and inverse transforming the residual and then combining it with the corresponding prediction block 134.
[0037] To obtain the decoded picture 108 for motion estimation, the reconstructed block 136 is processed by the loop filter module 120. The loop filter module 120 is used to smooth pixel transitions, thereby improving video quality. The loop filter module 120 can be used to implement one or more loop filters, such as an unblocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), etc.
[0038] Figure 2 An example of a video decoder 200 is depicted, which is used to implement the embodiments presented herein. The video decoder 200 processes the encoded video 202 in the bitstream and generates a decoded picture 208. Figure 2 In the example shown, the video decoder 200 includes an entropy decoding module 216 , an inverse quantization module 218 , an inverse transform module 219 , a loop filter module 220 , an intra prediction module 226 , an inter prediction module 224 , and a decoded picture buffer 230 .
[0039] The entropy decoding module 216 is used to perform entropy decoding on the encoded video 202. The entropy decoding module 216 decodes the quantized coefficients, coding parameters (including intra-frame prediction parameters and inter-frame prediction parameters), and other information. In some examples, the entropy decoding module 216 decodes the bitstream of the encoded video 202 into a binary representation and then converts the binary representation into the quantized levels of the coefficients. The entropy-decoded coefficient levels are dequantized by the dequantization module 218 and then dequantized to the pixel domain by the inverse transform module 219. The functions of the dequantization module 218 and the inverse transform module 219 are similar to those described above with respect to Figure 1 The inverse quantization module 118 and the inverse transform module 119 are described. The inverse transformed residual block can be added to the corresponding prediction block 234 to generate a reconstructed block 236. For blocks that skip the transform, the inverse transform module 219 is not applied. The dequantized samples generated by the inverse quantization module 218 are used to generate the reconstructed block 236.
[0040] The prediction block 234 for a particular block is generated based on the prediction mode of the block. If the coding parameters of the block indicate that the block is intra-predicted, the reconstructed block 236 of the reference block in the same picture may be input to the intra-prediction module 226 to generate the prediction block 234 for the block. If the coding parameters of the block indicate that the block is inter-predicted, the prediction block 234 is generated by the inter-prediction module 224. The functions of the intra-prediction module 226 and the inter-prediction module 224 are similar to Figure 1 The intra-frame prediction module 126 and the inter-frame prediction module 124 in.
[0041] As above relative to Figure 1 As described above, inter-frame prediction involves one or more reference pictures. Video decoder 200 generates a decoded picture 208 of the reference picture by applying loop filter module 220 to the reconstructed block of the reference picture. Decoded picture 208 is stored in decoded picture buffer 230 for use by inter-frame prediction module 224 and output.
[0042] Now refer to Figure 3 , Figure 3 Describes an example of dividing a picture in a video into coding tree units according to some embodiments of the present disclosure. Figure 1 and Figure 2 As mentioned above, in order to encode the pictures of the video, the pictures are divided into blocks, such as Figure 3 The CTU (Coding Tree Unit) 302 in the VVC is shown. For example, the CTU 302 can be a block of 128x128 pixels. Figure 3 In some examples, such as Figure 4As shown, each CTU 302 in the picture can be divided into one or more CUs (coding units) 402, and the CU 402 can be further divided into prediction units or transform units (TUs) for prediction and transformation. According to different coding schemes, the CTU 302 can be divided into CUs 402 in different ways. For example, in VVC, the CU 402 can be rectangular or square and can be encoded without further division into prediction units or transform units. Each CU 402 can be as large as its root CTU 302, or it can be a subdivision of the root CTU 302 as small as a 4x4 block. Figure 4 As shown, the partitioning of CTU 302 to CU 402 in VVC can be quadtree partitioning, binary tree partitioning or ternary tree partitioning. Figure 4 In the figure, solid lines indicate quadtree partitioning, and dashed lines indicate binary or ternary tree partitioning.
[0043] General constraints in VVC version 1 (VVCv1)
[0044] In VVC version 1, the General Constraint Information (GCI) syntax structure general_constraints_info() is used to indicate specific constraint properties of a codestream. GCI contains a list of constraint flags and non-flag syntax elements. The binary flag gci_present_flag is used to specify whether a GCI syntax element is present. gci_present_flag equal to 1 specifies that a GCI syntax element is present in the general_constraints_info() syntax structure. gci_present_flag equal to 0 specifies that no GCI field is present in the general_constraints_info() syntax structure and that the general_constraints_info() syntax structure does not impose any constraints.
[0045] Generic constraint information can be transmitted using a high-level syntax in various contexts. For example, GCI can be transmitted in network packets that only contain decoding capability information, such as a Network Abstraction Layer (NAL) packet with nal_unit_type set to 13 (i.e., DCI_NUT as the name of nal_unit_type). Alternatively, GCI can be transmitted in a video parameter set or sequence parameter set.
[0046] The purpose of the GCI syntax structure is to enable the discovery of configuration information related to the functionality required to decode a codestream and to allow the transmission of interoperability points that impose constraints beyond the specifications of the configuration, tiers and levels (PTL) and at a finer granularity than allowed by previous video coding standards. Similar to sub-profiles, the use of the GCI syntax structure enables interoperability to be defined for decoder implementations that do not support all features of the VVC configuration but meet specific application requirements. Decoder implementations can examine the GCI syntax elements to check whether the codestream avoids the use of specific features, thereby determining how to configure the decoding process and identifying whether the codestream can be decoded by the decoder. Decoder implementations that support all features of the VVC configuration can ignore the GCI syntax element values, because such decoders can decode any codestream that conforms to the indicated PTL.
[0047] The general constraint information syntax structure specified in VVC version 1 is defined as follows.
[0048]
[0049]
[0050] The presence of the general constraint flag depends on the value of gci_present_flag. When the value of gci_present_flag is 1, the general constraint flag is present in the codestream. When the value of gci_present_flag is 0, the general constraint flag is not present in the codestream.
[0051] In addition to the general constraint flags defined by VVCvl, the syntax element gci_num_reserved_bits also specifies additional general constraint flags. gci_num_reserved_bits is an 8-bit unsigned integer whose value indicates the number of additional bits transmitted in the general constraint syntax structure. In the VVC specification, these additional bits (called syntax elements gci_reserved_zero_bit[i]) are extracted from the codestream and discarded. This provision enables the VVCvl decoder to be at least forward compatible with the high-level syntax part of the codestream generated by subsequent versions of VVC.
[0052] General constraints in VVC version 2 (VVCv2)
[0053] In the current draft of VVC version 2 ("Extensions to the VVC operating range (Draft 5)", ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 Joint Video Experts Group document, JVET-X2005), some additional coding tools are proposed that are constrained by the general constraint flag. It is proposed to rename the 8-bit field called gc_num_reserved_bits in VVCv1 to gci_num_additional_bits. The syntax of the VVCv2 proposal adjustment is as follows.
[0054]
[0055] There are six additional general constraint flags: gci_all_rap_pictures_constraint_flag, gci_no_extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_rrc_rice_extension_constraint_flag, gci_no_persistent_rice_adaptation_constraint_flag, and gci_no_reverse_last_sig_coeff_constraint_flag. The suggested interpretation ("semantics") of the gci_num_additional_bits syntax element is as follows.
[0056] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 1. Values of gci_num_additional_bits greater than 1 are reserved by ITU-T | ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 1, decoders conforming to this version of this document shall allow values of gci_num_additional_bits greater than 1 to appear in the syntax and shall ignore the value of all gci_reserved_zero_bit[i] syntax elements when gci_num_additional_bits is greater than 1.
[0057] It is recommended that if the VVCv2 codestream transmits general constraint information, the syntax element gci_num_additional_bits should be set to 0 if the 6 VVCv2 general constraint flags are not transmitted, and gci_num_additional_bits should be set to 1 if the 6 VVCv2 general constraint flags are transmitted.
[0058] However, the VVCv2 general constraint syntax specification proposed above will result in incompatibility with the VVCv1 decoder. Specifically, when transmitting an additional general constraint flag, the proposed VVCv2 syntax transmits this flag by setting gci_num_additional_bits to 1. In the VVCv2 decoder, when transmitting general constraint information, the gci_num_additional_bits syntax element is decoded from the bitstream as an 8-bit unsigned integer. If the value of gci_num_additional_bits is decoded to 1, 6 additional bits will be decoded from the bitstream. These 6 bits are interpreted as general constraint flags for additional decoding tools that may be constrained in VVCv2.
[0059] In a VVCv1 decoder, the same 8-bit field is interpreted as gci_num_reserved_bits. However, if the value of this syntax element is decoded as 1, only one additional bit is decoded from the bitstream. This bit is discarded and not used. Therefore, a VVCv1 decoder that is decoding a VVCv2 bitstream may experience desynchronization issues due to the presence of five additional bits in the bitstream that are not recognized by the VVCv1 specification.
[0060] As mentioned above, desynchronization is a serious incompatibility issue between different versions of the video coding standard. A VVCvl decoder cannot decode the entire content of a VVCv2 codestream because a VVCv2 codestream can use coding tools that are defined in the VVCv2 specification but are unknown to the VVCvl decoder. However, a VVCvl decoder should ideally be able to decode at least the high-level syntax portion of a VVCv2 codestream. By successfully decoding the high-level syntax, the video decoder can determine not only general constraint information but also configuration and level information. This information provides the decoder with hints about the capabilities required to decode the codestream. For example, general constraint information provides the decoder with an indication of which coding tools are constrained by the codestream. Configuration and level information provide the decoder with an indication of the uncompressed video data throughput (such as video data rate, frame rate, resolution, etc.) that needs to be supported.
[0061] If high-level syntax is successfully decoded, the video decoder can determine whether the current bitstream can be decoded. If not, the decoder can gracefully terminate the decoding process. Conversely, asynchrony during high-level syntax decoding means that the information provided in the high-level syntax cannot be correctly decoded. In the worst case, the decoder may decode completely incorrect syntax element values after asynchrony occurs, resulting in incorrect parameter settings and incorrect decoding of subsequent low-level syntax, causing decoding failure.
[0062] Furthermore, in the VVCv2 decoder, when transmitting generic constraints information, the gci_num_additional_bits syntax element is decoded from the codestream as an 8-bit unsigned integer. If the value of gci_num_additional_bits is decoded to 0, no further generic constraint flags are transmitted. In this case, no inferred value is specified for the additional generic constraint flag, and its value is undefined. Therefore, it is unclear whether the coding tools associated with the additional generic constraint flag should behave constrained or unconstrained, which may lead to inconsistent decoder implementations.
[0063] In the VVC specification, the name of the syntax element gci_reserved_zero_bit[i] misleadingly implies that the value of such syntax elements must be 0. Typically, when an encoder writes reserved syntax elements into a bitstream as placeholders, a default value is embedded in the name of the reserved syntax elements. However, the design of the general constraint syntax structure means that the encoder never writes to gci_reserved_zero_bit[i]. gci_reserved_zero_bit[i] is only used when a specific version of the VVC decoder reads a higher version of the VVC bitstream. In this case, the value of gci_reserved_zero_bit[i] cannot be guaranteed to be 0. The following are some solutions to address the above issues.
[0064] Transmitting general constraint information
[0065] In one embodiment, general constraint information is transmitted to solve the above-mentioned out-of-sync problem. If the VVCv2 bitstream transmits general constraint information (i.e., the value of gci_present_flag is 1), and the VVCv2 general constraint information consists of N additional coding tools that may be constrained, the value of the syntax element gci_num_additional_bits can only be set to 0 or N. If gci_num_additional_bits is set to 0, the general constraint flags of these N additional coding tools are not transmitted. If gci_num_additional_bits is set to N, the next N bits in the bitstream are used to transmit the general constraint flags of these N additional coding tools. In one example, N is set to 6.
[0066] The VVCv2 bitstream does not allow the value of gci_num_additional_bits to be set to a value other than 0 or N. However, the VVCv2 decoder can still process general constraint information with the value of gci_num_additional_bits set to a value other than 0 or N. For example, the value of gci_num_additional_bits can be set to a value greater than 0 and less than N, or greater than N. Let the decoded value of gci_num_additional_bits be M. Then, if M is greater than 0 but less than N (i.e., 0 < M < N), after decoding the gci_num_additional_bits syntax element, the decoder extracts M bits from the bitstream as the gci_reserved_zero_bit[i] syntax element and discards them. If M is greater than N (i.e., N < M), after decoding the gci_num_additional_bits syntax element, the decoder extracts N bits from the bitstream and interprets them as the general constraint flags of N additional coding tools. Then, the decoder extracts M - N bits from the bitstream as the gci_reserved_zero_bit[i] syntax element and discards them.
[0067] In other examples, the VVCv2 decoder does not need to process general constraint information with the value of gci_num_additional_bits set to a value greater than 0 but less than N. A legal VVCvl bitstream can only set the value of gci_num_reserved_bits to 0. A legal VVCv2 bitstream can only set the value of gci_num_reserved_bits to 0 or N. Future versions of the VVC bitstream will not allow the value of gci_num_additional_bits to be set to a value between 0 and N.
[0068] In one example of this embodiment, the general constraint information syntax of VVCv2 is modified to the 6 general constraint flags currently proposed for the VVCv2 coding tool, as shown in Table 1 below (the added parts are shown with underscores, and the deleted parts are shown with strikethroughs), where "if (gci_num_additional_bits>0)" is replaced with "if (gci_num_additional_bits>5)".
[0069] Table 1
[0070]
[0071]
[0072] In one example, if the coding tool in VVCv2 has 6 additional general constraint flags, the corresponding semantics of gci_num_additional_bits are (the additions are shown underlined):
[0073] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Other than 0 and 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow Other than 0 and 6 The gci_num_additional_bits value appears in the syntax and should be in gci_num_additional_bits Other than 0 and 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when
[0074] In the example of gci_num_additional_bits semantics above, in addition to the values of 0 or 6 mentioned above, gci_num_additional_bits is allowed to take values other than 0 or 6. In other words, the value of gci_num_additional_bits can be between 1 and 5. The value of gci_num_additional_bits can also be greater than 6. If the value of gci_num_additional_bits (denoted as M) is between 1 and 5, according to the syntax shown in Table 1 above, the decoder will skip the steps executed if the "if" condition is true and jump to the "else" step, setting the value of "numAdditionalBitsUsed" to 0. Then, in the "for" loop, M bits are read and discarded. This avoids desynchronization issues. If the value of gci_num_additional_bits M is greater than 6, the 6 additional general constraint flags are extracted, and the remaining M-6 bits are further extracted and discarded.
[0075] In another example, if the coding tool in VVCv2 has 6 additional general constraint flags, the corresponding semantics of gci_num_additional_bits are (the additions are shown underlined):
[0076] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Greater than 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow for greater than 6 The gci_num_additional_bits value appears in the syntax and should be used when gci_num_additional_bits is greater than 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when
[0077] In another embodiment of transmitting general constraint information, if the VVCv2 codestream transmits general constraint information (i.e., the value of gci_present_flag is 1), and the VVCv2 general constraint information consists of N additional coding tools that may be constrained, the syntax element gci_num_additional_bits may be set to the value M. The range of M is 0 to N (including 0 and N) (0≤M≤N). If gci_num_additional_bits is set to 0, the general constraint flags of these N additional coding tools are not transmitted. If gci_num_additional_bits is set to a non-zero value M, the next M bits in the codestream are used to transmit the general constraint flags of M of these N additional coding tools. Which M of the additional coding tools are constrained is determined by the order in which the general constraint flags appear in the general constraint information syntax table.
[0078] In an example of this embodiment, the general constraint information syntax of VVCv2 is modified to 6 general constraint flags currently proposed for the VVCv2 coding tool. The modification is shown in the following table.
[0079]
[0080] In another example of this embodiment, equivalent behavior can be achieved through a more compact syntax table.
[0081]
[0082]
[0083] An alternative arrangement of this embodiment can be expressed in the generic constraint information syntax by changing the order of the generic constraint flags of the VVCv2 coding tool.
[0084] If the coding tool in VVCv2 has 6 additional general constraint flags, the corresponding semantics of gci_num_additional_bits may be as follows.
[0085] gci_num_additional_bits specifies the additional bits in the general constraint information syntax structure except gci_alignment_ The number of additional GCI bits beyond the zero_bit syntax element (when present). In codestreams conforming to this version of this document, gci_ The value of num_additional_bits shall be from 0 to 6 (inclusive). Values of gci_num_additional_bits greater than 6 Reserved for future use by ITU-T|ISO / IEC. Although gci_num_additional_bits is required in this version of this document The value of gci_num_ can be 0 to 6 (inclusive), but decoders conforming to this version of this document shall allow gci_num_ to be greater than 6. The additional_bits value appears in the syntax and should be ignored if gci_num_additional_bits is greater than 6. The value of the gci_reserved_zero_bit[i] syntax element.
[0086] In this embodiment, depending on the value of gci_num_additional_bits, some or all of the additional general constraint flags may not be transmitted. In one arrangement of this embodiment, when the VVCv2 codestream transmits general constraint information (i.e., if the value of the gci_present_flag flag is 1), no constraints are imposed on the coding tools corresponding to the additional general constraint flags that are not transmitted. This behavior can be expressed by modifying the semantics of gci_num_additional_bits as follows.
[0087] gci_num_additional_bits specifies the additional bits in the general constraint information syntax structure except gci_alignment_ The number of additional GCI bits beyond the zero_bit syntax element (when present). In codestreams conforming to this version of this document, gci_ The value of num_additional_bits shall be from 0 to 6 (inclusive). Values of gci_num_additional_bits greater than 6 Reserved for future use by ITU-T|ISO / IEC. Although gci_num_additional_bits is required in this version of this document The value of gci_num_ can be 0 to 6 (inclusive), but decoders conforming to this version of this document shall allow gci_num_ to be greater than 6. The additional_bits value appears in the syntax and should be ignored if gci_num_additional_bits is greater than 6. gci_reserved_zero_bit[i] syntax element. In addition, when gci_present_flag is equal to 1, A coding tool imposes constraints for which there is no corresponding constraint flag in the syntax of general_constraints_info().
[0088] In another arrangement of this embodiment, when the VVCv2 codestream transmits general constraint information (i.e., if the value of gci_present_flag is 1) and the additional general constraint flag is not transmitted, the corresponding tool will be constrained. This behavior can be expressed as follows by modifying the semantics of the additional general constraint flag.
[0089] gci_num_additional_bits specifies the additional bits in the general constraint information syntax structure except gci_alignment_ The number of additional GCI bits beyond the zero_bit syntax element (when present). In codestreams conforming to this version of this document, gci_ The value of num_additional_bits shall be from 0 to 6 (inclusive). Values of gci_num_additional_bits greater than 6 Reserved for future use by ITU-T|ISO / IEC. Although gci_num_additional_bits is required in this version of this document The value of gci_num_ can be 0 to 6 (inclusive), but decoders conforming to this version of this document shall allow gci_num_ to be greater than 6. The additional_bits value appears in the syntax and should be ignored if gci_num_additional_bits is greater than 6. The value of the gci_reserved_zero_bit[i] syntax element.
[0090] gci_all_rap_pictures_constraint_flag equal to 1 specifies that all pictures in OlsInScope are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0. gci_all_rap_pictures_constraint_flag equal to 0 does not impose this constraint. When gci_present_flag is equal to 1 and gci_all_rap_ When pictures_constraint_flag does not exist, the value of gci_all_rap_pictures_constraint_flag is pushed It is considered equal to 1.
[0091] gci_no_extended_precision_processing_constraint_flag equal to 1 specifies that the sps_extended_precision_flag of all pictures in OlsInScope shall be equal to 0. gci_no_extended_precision_processing_constraint_flag equal to 0 does not impose this constraint. When gci_present_flag When gci_no_extended_precision_processing_constraint_flag is not present, gci_no_ The value of extended_precision_processing_constraint_flag is inferred to be equal to 1.
[0092] gci_no_ts_residual_coding_rice_constraint_flag equal to 1 specifies that the sps_ts_residual_coding_rice_present_in_sh_flag of all pictures in OlsInScope shall be equal to 0. gci_no_ts_residual_coding_rice_constraint_flag equal to 0 does not impose this constraint. When gci_present_flag When gci_no_ts_residual_coding_rice_constraint_flag is equal to 1 and gci_no_ts_residual_coding_rice_constraint_flag does not exist, gci_no_ts_ The value of residual_coding_rice_constraint_flag is inferred to be equal to 1.
[0093] gci_no_rrc_rice_extension_constraint_flag equal to 1 specifies that the sps_rrc_rice_extension_flag of all pictures in OlsInScope shall be equal to 0. gci_no_rrc_rice_extension_constraint_flag equal to 0 does not impose this constraint. When gci_present_flag is equal to 1 and gci_no_rrc_ When rice_extension_constraint_flag does not exist, gci_no_rrc_rice_extension_constraint_ The value of flag is inferred to be equal to 1.
[0094] gci_no_persistent_rice_adaptation_constraint_flag equal to 1 specifies that the sps_persistent_rice_adaptation_enabled_flag of all pictures in OlsInScope shall be equal to 0. gci_no_persistent_rice_adaptation_constraint_flag equal to 0 does not impose this constraint. When gci_ When present_flag is equal to 1 and gci_no_persistent_rice_adaptation_constraint_flag does not exist, The value of gci_no_persistent_rice_adaptation_constraint_flag is inferred to be equal to 1.
[0095] gci_no_reverse_last_sig_coeff_constraint_flag equal to 1 specifies that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope shall be equal to 0. gci_no_reverse_last_sig_coeff_constraint_flag equal to 0 does not impose this constraint. When gci_present_flag is equal to 1 and gci_ When no_reverse_last_sig_coeff_constraint_flag does not exist, gci_no_reverse_last_sig_ The value of coeff_constraint_flag is inferred to be equal to 1.
[0096] Initializes the general constraint information flag
[0097] This document describes an embodiment for initializing the general constraint information flags to address the aforementioned issue of ambiguity and inconsistency in decoder implementations. In this embodiment, when gci_present_flag is 1 and gci_num_additional_bits is 0, general_constraints_info() does not impose constraints on coding tools related to gci_all_rap_pictures_constraint_flag, gci_no_extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_rice_extension_constraint_flag, and gci_no_persistent_rice_adaptation_constraint_flag.
[0098] As an example, possible changes to the semantics of gci_num_additional_bits are shown below, which are based on the current version of the VVC version 2 specification (additions are shown underlined).
[0099] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 1. Values of gci_num_additional_bits greater than 1 are reserved by ITU-T | ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 1, decoders conforming to this version of this document shall allow values of gci_num_additional_bits greater than 1 to appear in the syntax and shall ignore the value of all gci_reserved_zero_bit[i] syntax elements when gci_num_additional_bits is greater than 1. When gci_present_flag is equal to 1 and gci_num_additional_bits is equal to 0, general_ constraints_info() does not match gci_all_rap_pictures_constraint_flag, gci_no_ extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_ constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_ rice_extension_constraint_flag and gci_no_persistent_rice_adaptation_ constraint_flag The associated coding tools impose constraints.
[0100] As another example, gci_num_ additional The possible changes to the semantics of _bits are shown below, based on the current version of the VVC version 2 specification.
[0101] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Other than 0 and 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow Other than 0 and 6 The gci_num_additional_bits value appears in the syntax and should be in gci_num_additional_bits Other than 0 and 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When gci_present_flag is equal to 1 and gci_num_additional_bits is equal to 0 When general_constraints_info() is not compatible with gci_all_rap_pictures_constraint_flag, gci_ no_extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_ rice_constraint_flag,gci_no_reverse_last_sig_coeff_constraint_flag,gci_no_ rrc_rice_extension_constraint_flag and gci_no_persistent_rice_adaptation_ constraint_flag The associated coding tools impose constraints.
[0102] In the current version of the VVC version 2 specification, the semantics of gci_num_additional_bits are only used when gci_present_flag is equal to 1. The scenario where gci_present_flag is equal to 0 is addressed in other sections of the VVC version 2 specification. Therefore, the "when gci_present_flag is equal to 1" in the above gci_num_additional_bits semantics is automatically satisfied. In addition, since the VVC version 2 specification does not allow the value of gci_num_additional_bits to be between 1 and 5, "gci_num_additional_bits equal to 0" is equivalent to "gci_num_additional_bits less than or equal to 5" or "gci_all_rap_pictures_constraint_flag, gci_no_extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_rice_extension_ constraint _flag and gci_no_persistent_rice_adaptation_constraint_flag do not exist". Therefore, the above changes to the semantics of gci_num_additional_bits are equivalent to the following.
[0103] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Other than 0 and 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow Other than 0 and 6 The gci_num_additional_bits value appears in the syntax and should be in gci_num_additional_bits Other than 0 and 6The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When gci_all_rap_pictures_constraint_flag, gci_no_ extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_ constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_ rice_extension_constraint_flag and gci_no_persistent_rice_adaptation_ When constraint_flag is not present, general_constraints_info() does not work with gci_all_rap_ pictures_constraint_flag, gci_no_extended_precision_processing_constraint_ flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_ coeff_constraint_flag, gci_no_rrc_rice_extension_constraint_flag and gci_no_ persistent_rice_adaptation_constraint_flag related coding tools impose constraints.
[0104] Below are a few examples where the inferred value of the additional general constraint flag can be set in a manner similar to the semantics described above. For example, the semantics of the additional general constraint flag can be modified as follows to include inferred value setting.
[0105] gci_all_rap_pictures_constraint_flag equal to 1 specifies that all pictures in OlsInScope are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0. gci_all_rap_pictures_constraint_flag equal to 0 does not impose this constraint. When not present, gci_all_rap_pictures_ The value of constraint_flag is inferred to be equal to 0.
[0106] gci_no_extended_precision_processing_constraint_flag equal to 1 specifies that the sps_extended_precision_flag of all pictures in OlsInScope shall be equal to 0. gci_no_extended_precision_processing_constraint_flag equal to 0 does not impose this constraint. When not present, gci_no_ The value of extended_precision_processing_constraint_flag is inferred to be equal to 0.
[0107] gci_no_ts_residual_coding_rice_constraint_flag is equal to 1 to specify that all pictures in OlsInScope sps _ts_residual_coding_rice_present_in_sh_flag shall all be equal to 0. gci_no_ts_residual_coding_rice_constraint_flag equal to 0 does not impose this constraint. When not present, gci_no_ The value of ts_residual_coding_rice_constraint_flag is inferred to be equal to 0.
[0108] gci_no_rrc_rice_extension_constraint_flag equal to 1 specifies that the sps_rrc_rice_extension_flag of all pictures in OlsInScope should be equal to 0. extension _constraint_flag equal to 0 does not impose this constraint. When not present, gci_no_rrc_rice_extension_ The value of constraint_flag is inferred to be equal to 0.
[0109] gci_no_persistent_rice_adaptation_constraint_flag is equal to 1 to specify the sps_ persistent _rice_adaptation_enabled_flag shall be equal to 0. gci_no_persistent_rice_adaptation_constraint_flag equal to 0 does not impose this constraint. When it does not exist, The value of gci_no_persistent_rice_adaptation_constraint_flag is inferred to be equal to 0.
[0110] gci_no_reverse_last_sig_coeff_constraint_flag equal to 1 specifies that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope shall be equal to 0. gci_no_reverse_last_sig_coeff_constraint_flag equal to 0 does not impose this constraint. When not present, gci_no_reverse_ The value of last_sig_coeff_constraint_flag is inferred to be equal to 0.
[0111] In another example, the semantics of the additional general constraint flag is modified as follows.
[0112] gci_all_rap_pictures_constraint_flag equal to 1 specifies that all pictures in OlsInScope are GDR pictures with ph_recovery_poc_cnt equal to 0, or IRAP pictures. pictures _constraint_flag equal to 0 does not impose this constraint. When gci_all_rap_pictures_constraint_flag is not When present, its value is inferred to be equal to 0.
[0113] gci_no_extended_precision_processing_constraint_flag equal to 1 specifies that the sps_extended_precision_flag of all pictures in OlsInScope shall be equal to 0. gci_no_extended_precision_processing_constraint_flag equal to 0 does not impose this constraint. When gci_no_extended_ When precision_processing_constraint_flag is not present, its value is inferred to be equal to 0.
[0114] gci_no_ts_residual_coding_rice_constraint_flag equal to 1 specifies that the sps_ts_residual_coding_rice_present_in_sh_flag of all pictures in OlsInScope shall be equal to 0. gci_no_ts_residual_coding_rice_constraint_flag equal to 0 does not impose this constraint. When gci_no_ts_ When residual_coding_rice_constraint_flag is not present, its value is inferred to be equal to 0.
[0115] gci_no_rrc_rice_extension_constraint_flag equal to 1 specifies that the sps_rrc_rice_extension_flag of all pictures in OlsInScope shall be equal to 0. gci_no_rrc_rice_extension_constraint_flag equal to 0 does not impose this constraint. When gci_no_rrc_rice_extension_constraint_ When flag is not present, its value is inferred to be equal to 0.
[0116] gci_no_persistent_rice_adaptation_constraint_flag equal to 1 specifies that the sps_persistent_rice_adaptation_enabled_flag of all pictures in OlsInScope shall be equal to 0. persistent If _rice_adaptation_constraint_flag is equal to 0, this constraint is not imposed. When gci_no_ When persistent_rice_adaptation_constraint_flag is not present, its value is inferred to be equal to 0.
[0117] gci_no_reverse_last_sig_coeff_constraint_flag equal to 1 specifies that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope shall be equal to 0. last _sig_coeff_constraint_flag is equal to 0 to not impose this constraint. When gci_no_reverse_last_sig_ When coeff_constraint_flag is not present, its value is inferred to be equal to 0.
[0118] In another example, the semantics of gci_num_additional_bits is modified as follows.
[0119] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Greater than 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow for greater than 6 The gci_num_additional_bits value appears in the syntax and should be used when gci_num_additional_bits is greater than 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When gci_num_additional_bits is equal to 0, all constraint flags specified by the additional GCI bits are deprecated. It is considered equal to 0.
[0120] In another example, the semantics of gci_num_additional_bits is modified as follows.
[0121] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). Version In the code stream, the value of gci_num_additional_bits should be equal to 0 or 6 . Greater than 6The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow for greater than 6 The gci_num_additional_bits value appears in the syntax and should be used when gci_num_additional_bits is greater than 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When not present, all constraint flags specified by the additional GCI bits are inferred to be equal to 0.
[0122] In another example, the syntax and semantics of the general constraint information syntax element are modified as follows.
[0123]
[0124]
[0125] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, gci_num_ additional The value of _bits should be equal to 0 or 6 . Greater than 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow for greater than 6 The gci_num_additional_bits value appears in the syntax and should be used when gci_num_additional_bits is greater than 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When gci_num_additional_bits is equal to 0, additional_general_constraints_info () The grammatical structure does not impose any constraints.
[0126] In another example, the semantics of the general constraint information syntax element is modified as follows.
[0127] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Greater than 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow for greater than 6 The gci_num_additional_bits value appears in the syntax and should be used when gci_num_additional_bits is greater than 6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When not present, the additional_general_constraints_info() syntax construct does not impose any constraint.
[0128] exist initialization In another embodiment of the general constraint information flags, when gci_present_flag is equal to 1 and gci_num_additional_bits is equal to 0, gci_all_rap_pictures_constraint_flag, gci_no_extended_precision_processing_constraint_flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_rice_extension_constraint_flag, and gci_no_persistent_rice_adaptation_constraint_flag may always impose the corresponding constraints specified by their respective semantics. In other words, when the GCI flags are present and gci_num_additional_bits is equal to 0, these six GCI flags are inferred to be equal to 1.
[0129] As an example, the possible changes in the semantics of gci_num_additional_bits are shown below, based on the current additions to the VVC version 2 specification.
[0130] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 1. Values of gci_num_additional_bits greater than 1 are reserved by ITU-T | ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 1, decoders conforming to this version of this document shall allow values of gci_num_additional_bits greater than 1 to appear in the syntax and shall ignore the value of all gci_reserved_zero_bit[i] syntax elements when gci_num_additional_bits is greater than 1. When gci_present_flag is equal to 1 and gci_num_additional_bits is equal to 0, gci_all_rap_ pictures_constraint_flag, gci_no_extended_precision_processing_constraint_ flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_ coeff_constraint_flag, gci_no_rrc_rice_extension_constraint_flag and gci_no_ persistent_rice_adaptation_constraint_flag is inferred to be equal to 1.
[0131] As another example, possible changes to the semantics of gci_num_additional_bits are shown below, based on the current additions to the VVC version 2 specification.
[0132] gci_num_additional_bits specifies the number of additional GCI bits in the general constraint information syntax structure, excluding the gci_alignment_zero_bit syntax element (when present). In codestreams conforming to this version of this document, the value of gci_num_additional_bits shall be equal to 0 or 6 . Other than 0 and 6 The value of gci_num_additional_bits is reserved by ITU-T|ISO / IEC for future use. Although this version of this document requires that the value of gci_num_additional_bits be equal to 0 or 6 , but decoders conforming to this version of this document shall allow Other than 0 and ∑6∑ The gci_num_additional_bits value appears in the syntax and should be in gci_num_additional_bits Other than ∑0∑ and ∑6 The values of all gci_reserved_zero_bit[i] syntax elements are ignored when When gci_present_flag is equal to 1 and gci_num_additional_ When bits are equal to 0, gci_all_rap_pictures_constraint_flag, gci_no_extended_precision_ processing_constraint_flag, gci_no_ts_residual_coding_rice_constraint_flag, gci_no_reverse_last_sig_coeff_constraint_flag, gci_no_rrc_rice_extension_ constraint_flag and gci_no_persistent_rice_adaptation_constraint_flag are inferred as Equal to 1.
[0133] As a supplement or alternative to the above embodiment, the misleading syntax element gci_reserved_zero_bit[i] may be renamed as gci_reserved_bit[i]. For example, the modified syntax and semantics may be as follows.
[0134]
[0135] gci_reserved_bit[i] MAY have any value. Its presence and value do not affect the decoding process specified by this version of this specification. Decoders conforming to this version of this specification SHOULD ignore all gci_reserved_bit[i] The value of the syntax element.
[0136] In another embodiment, the syntax element gci_reserved_zero_bit[i] is renamed gci_reserved_bit[i]. For example, the modified syntax and semantics may be as follows.
[0137]
[0138] gci_reserved_bit[i] Can have any value. Its existence and value do not affect this specification This version specifies the decoding process. Decoders conforming to this version of the specification should ignore all gci_reserved_bit[i] The value of the syntax element.
[0139] General restriction signs
[0140] gci_all_rap_pictures_constraint_flag
[0141] gci_all_rap_pictures_constraint_flag is used to indicate that pictures are restricted to IRAP or GDR pictures.
[0142] The Network Abstraction Layer (NAL) is a system interface that organizes VVC syntax elements into "NAL units." This structure allows VVC to be easily and efficiently customized to suit a wide range of use cases, from real-time communication applications to file formats for storage applications. The complete list of NAL unit types in the VVC standard is shown in the following table.
[0143]
[0144]
[0145] NAL units classified as Video Coding Layer (VCL) contain low-level syntax elements, while NAL units classified as non-VCL contain high-level syntax elements. Pictures of a video sequence are decoded from VCL NAL units. Different types of VCL categories can be used to indicate dependencies at a high level. For example, NAL units from TRAIL_NUT (0) to RSV_VCL_6 (6) can typically use inter-frame prediction tools, depending on whether a previously decoded (reference) picture can be accessed. Pictures encoded using inter-frame prediction tools can be compressed more efficiently than pictures compressed using only intra-frame prediction tools. However, this decoding dependency can cause problems when reference pictures are not available.
[0146] An intra random access point (IRAP) picture is a coded picture whose VCL NAL units all have the same nal_unit_type value, in the range of IDR_W_RADL to CRA_NUT (inclusive). An IRAP picture can be either a CRA picture or an IDR picture. During decoding, IRAP pictures do not use inter prediction from same-layer reference pictures. The first picture in the bitstream in decoding order is an IRAP picture or a gradual decoding refresh (GDR) picture. For single-layer bitstreams, the IRAP picture and all subsequent non-RASL pictures in decoding order in the CLVS can be correctly decoded, as long as the necessary parameter sets are available when the reference is needed, without having to decode any pictures that precede the IRAP picture in the decoding order. The pps_mixed_nalu_types_in_pic_flag value for an IRAP picture is 0. When pps_mixed_nalu_types_in_pic_flag of a picture is equal to 0 and the nal_unit_type of any slice of the picture is in the range of IDR_W_RADL to CRA_NUT (inclusive), the nal_unit_type value of all other slices of the picture is the same and the picture is known to be an IRAP picture after the first slice is received.
[0147] Therefore, IRAP pictures do not use inter-frame prediction across the same layer. This restriction allows IRAP pictures to serve as error recovery points for streaming video applications or as video seek locations for on-demand playback applications. However, IRAP pictures are generally less compressed than non-IRAP pictures.
[0148] The VVC standard introduced progressive decoding refresh (GDR) pictures as a trade-off between non-IRAP and IRAP pictures. GDR pictures have some "clean" parts that do not use inter-frame prediction, while the rest of the picture can freely use inter-frame prediction. By splitting the picture in this way, the "clean" parts can still be decoded correctly in the event of errors such as packet loss. The spatial position of the "clean" parts is rotated in consecutive GDR pictures, so that the entire picture can eventually be recovered from errors.
[0149] For video applications where streaming recovery capabilities for playback flexibility are important, it may be desirable to restrict all pictures to be either IRAP or GDR pictures. In VVCv2, the GCI flag gci_all_rap_pictures_constraint_flag was introduced to indicate this restriction at a high level. gci_all_rap_pictures_constraint_flag equal to 1 specifies that all pictures in OlsInScope are either GDR pictures with ph_recovery_poc_cnt equal to 0, or IRAP pictures. gci_all_rap_pictures_constraint_flag equal to 0 does not impose this constraint. When gci_all_rap_pictures_constraint_flag is not present, its value is inferred to be equal to 0.
[0150] When the profile_tier_level() syntax structure is included in a VPS, OlsInScope is one or more output layer sets (OLS) specified by the VPS. When the profile_tier_level() syntax structure is included in an SPS, OlsInScope is an OLS that includes only the lowest layer among the layers referenced by the SPS, and this lowest layer is an independent layer.
[0151] gci_no_extended_precision_processing_constraint_flag
[0152] The GCI flag gci_no_extended_precision_processing_constraint_flag communicates at a high level whether the VVCv2 tools for extended transform precision are constrained. In the VVC standard, pixels of a video signal are represented using integer values. All computations and processing described in the VVC standard are expressed using operations on integers. This constraint is important for reasons of complexity and interoperability. First, the computational cost of performing integer operations (addition, multiplication, and division) is generally lower than the equivalent floating-point operations. Second, floating-point operations are not deterministically standardized. Floating-point addition and multiplication are not necessarily commutative (for example, (a+b)+c is not necessarily equal to a+(b+c)), and the estimated values of floating-point operations on different platforms are not guaranteed to be exactly the same.
[0153] The bit depth of video signal samples is a property of the video source and is denoted as BitDepth in the VVC standard. In a hybrid video coding system, video samples are predicted using either inter-frame or intra-frame prediction tools. The difference between the original video sample and the predicted sample is called a residual. In the worst case, the bit depth of these residual coefficients may be increased (for example, BitDepth+1); however, in the VVC standard, these residual coefficients are clipped to maintain the bit depth BitDepth. In practice, this worst-case scenario does not occur because a real-world encoder will not select a prediction tool that produces a residual with a larger amplitude than the original video signal.
[0154] The residual coefficients are then typically analyzed by an integer discrete cosine transform (DCT) to produce transform coefficients. The discrete cosine transform (DCT) is a linear, reversible function that can be expressed as:
[0155]
[0156] in, Represents the N-dimensional space of real numbers. The VVC standard uses an integer approximation of DCT (i.e., Unlike prediction, some expansion of the bit depth is usually allowed during the transform phase, as the accuracy of the transform affects the coding gain. The bit depth of the approximated DCT coefficients and the resulting transform coefficients are design decisions that are a trade-off between hardware complexity and coding performance.
[0157] In VVCv1, the bit depth of the transform coefficient is 16. That is, the value range of each transform coefficient is [-2 15 ,2 15-1]. Multiplying the video samples with the integerized DCT coefficients usually produces intermediate transform coefficients with a bit depth greater than 16. To generate transform coefficients with the required bit depth, the intermediate transform coefficients are shifted to the right. This operation is inherently lossy.
[0158] In VVCv2, extended transform precision can be enabled by setting the SPS flag sps_extended_precision_flag to 1. If extended transform precision is enabled, the transform coefficient bit depth will be increased to (Log2TransformRange+1). The exact transform coefficient bit depth depends on BitDepth. sps_extended_precision_flag equal to 1 specifies that the extended dynamic range is used for transform coefficients during scaling and transformation, and for the binarization of the abs_remaining[] and dec_abs_level[] syntax elements. sps_extended_precision_flag equal to 0 specifies that the extended dynamic range is not used during scaling and transformation, and is not used for the binarization of the abs_remaining[] and dec_abs_level[] syntax elements. When not present, the value of sps_extended_precision_flag is inferred to be equal to 0. The derivation of the variable Log2TransformRange is as follows.
[0159] Log2TransformRange=sps_extended_precision_flag? Max(15,Min(20,BitDepth+6)):15
[0160] CoeffMin=-(1<<(sps_extended_precision_flag?Max(15,Min(20,BitDepth+6)):15))
[0161] CoeffMax=(1<<(sps_extended_precision_flag?Max(15,Min(20,BitDepth+6)):15))-1
[0162] gci_no_extended_precision_processing_constraint_flag equal to 1 specifies that the sps_extended_precision_flag of all pictures in OlsInScope shall be equal to 0. OlsInScope is also referred to herein as the "scoped output layer set." When the profile_tier_level() syntax structure is included in a VPS, OlsInScope is the output layer set (OLS) specified by the VPS. When the profile_tier_level() syntax structure is included in an SPS, OlsInScope is the OLS that includes only the lowest layer of the layers that reference the SPS, and this lowest layer is an independent layer. gci_no_extended_precision_processing_constraint_flag equal to 0 does not impose this constraint. When gci_no_extended_precision_processing_constraint_flag is not present, its value is inferred to be equal to 0.
[0163] gci_no_ts_residual_coding_rice_constraint_flag
[0164] The GCI flag gci_no_ts_residual_coding_rice_constraint_flag transmits at a high level whether the VVCv2 tool for explicit Rice parameter transmission is constrained. During entropy coding, the entropy coding process encodes each syntax element value into a bit sequence and inserts it into the codestream.
[0165] The two entropy coding processes used in VVC are Content Adaptive Binary Arithmetic Coding (CABAC) and Rice coding. The CABAC engine is adaptive and can compress syntax elements to a bit rate very close to the theoretical Shannon limit. However, arithmetic coding is very complex. In order to encode non-binary syntax elements (for example, syntax elements with values outside the range of 0 or 1) using CABAC, the syntax elements must first be binarized into a set of "containers". The following table gives two examples of binarization. The second example shows that for variable-length binarization, there may be no container for some values of the syntax element.
[0166] Table: Fixed-length binarization example
[0167] value Binarization 0 00 1 01 2 10 3 11
[0168] Table: Variable-length binarization example
[0169] value Binarization 0 0 1 10 2 110 3 1110 4 1111
[0170] Each container in the binarization can be encoded by a CABAC engine. However, to encode a container using CABAC, the engine must store and update the associated "content." This content models the probability distribution of the container. In the binarization of syntax elements, each container requires separate content.
[0171] Syntax elements with a wide range of values are not suitable for pure CABAC encoding because the binarization of these syntax elements is long, which would incur excessive overhead in terms of content storage and updates. For example, residual coefficient values are not suitable for pure CABAC encoding. Compared to CABAC, Rice coding can model the probability distribution of non-binary values with compact parameters.
[0172] Rice coding can also be called Golomb coding or Rice-Golomb coding. Golomb coding is an entropy encoder controlled by a single parameter M, which is restricted to positive integers. Rice coding is a subset of Golomb coding, in which the entropy encoder is controlled by a single Rice parameter R, which is a non-negative integer. Rice coding with Rice parameter R is equivalent to Golomb coding, where M = 2 R .
[0173] As follows, the non-negative integer value x is binarized into a Rice code with Rice parameter R. The calculation formula of the quotient q and the remainder r is:
[0174]
[0175] r=x modulo 2 R
[0176] The Rice encoding of x is a combination of prefix encoding and suffix encoding. The prefix code is determined by the unary encoding of q. For example, the prefix encoding used in VVC is a truncated unary encoding:
[0177]
[0178]
[0179] Since the unary encoding is truncated, the Rice encoding is also truncated, which means that the range of values of x that can be encoded is limited. In VVC, the truncated Rice encoding is applied to the range [0,6*2 R ]'s x value.
[0180] The suffix code is a binary representation of a fixed-length code of r with a length of R bits. For example, if R = 3, the suffix code is:
[0181] r Fixed-length encoding 0 000 1 001 2 010 3 011 4 100 5 101 6 110 7 111
[0182] In VVC, residual coefficients are encoded using a combination of CABAC, Rice coding, and Exponential Golomb coding. A small number of syntax element flags are defined that are sufficient to transmit small residual values. For example, sig_coeff_flag indicates whether the residual magnitude is zero. If sig_coeff_flag is 1, additional flags (typically named abs_level_gtx_flag) can be transmitted to indicate whether the residual magnitude is greater than 1, 2, 3, and so on. These flags are encoded by the CABAC engine. Since most residual coefficients have small magnitudes, CABAC can efficiently encode most residuals using relatively few containers.
[0183] Any remaining magnitude of the residual coefficients that cannot be transmitted via the residual coefficient flag is transmitted in the syntax element abs_remainder. R , then all the syntax elements are transmitted through the Rice coding of abs_remainder. If it is greater than 6*2 R , then the syntax element is passed through 6*2 R Rice code and (abs_remainder-6*2 R ) is associated with the exponential Golomb coding. The exponential Golomb coding process is not described here.
[0184] Although Rice coding is simpler than CABAC, it can also effectively compress to low bit rates under appropriate conditions. For residual coefficients with uniformly small values, Rice coding with smaller Rice parameters is more efficient. Conversely, when some residual coefficients have large values, larger Rice parameters may be more appropriate. To adapt to the statistics of the residual coefficients, the Rice parameter is adaptively determined by the value locSumAbs, which is calculated based on the amplitudes of adjacent residual coefficients.
[0185] locSumAbs 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 cRiceParam 0 0 0 0 0 0 0 1 1 1 1 1 1 1 2 2 locSumAbs 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 cRiceParam 2 2 2 2 2 2 2 2 2 2 2 2 3 3 3 3
[0186] Adaptive Rice parameter determination is designed for residual coefficients in "regular residual coding" (RRC), which are residual coefficients obtained by performing a discrete cosine transform (DCT). However, in VVCv1, it is also applicable to residual coefficients in "transform skip residual coding" (TSRC). In VVCv2, it was realized that an alternative mechanism for Rice parameter determination may be beneficial for transform-skip coefficients.
[0187] An alternative mechanism in VVCv2 allows the Rice parameter to be explicitly transmitted through the slice-level syntax element sh_ts_residual_coding_rise_idx_minusl. When this syntax element is transmitted, the Rice parameter will be set to R=sh_ts_residual_coding_rise_idx_minusl+1. This value of the Rice parameter is maintained for the duration of the slice.
[0188]
[0189] sh_ts_residual_coding_rice_idx_minusl plus 1 specifies the Rice parameter for the residual_ts_coding() syntax structure in the current slice. When not present, the value of sh_ts_residual_coding_rice_idx_minusl is inferred to be equal to 0.
[0190] Whether the alternative mechanism is enabled is controlled by the SPS level flag sps_ts_residual_coding_rice_present_in_sh_flag. If this flag is set to 0, the alternative Rice parameter transmission is not enabled. The GCI flag gci_no_ts_residual_coding_rice_constraint_flag signals at a high level whether the VVCv2 tools for explicit Rice parameter transmission are subject to constraints. gci_no_ts_residual_coding_rice_constraint_flag equal to 1 specifies that the sps_ts_residual_coding_rice_present_in_sh_flag of all pictures in OlsInScope should be equal to 0. gci_no_ts_residual_coding_rice_constraint_flag equal to 0 does not impose this constraint. When gci_no_ts_residual_coding_rice_constraint_flag is not present, its value is inferred to be equal to 0.
[0191] gci_no_rrc_rice_extension_constraint_flag
[0192] gci_no_rrc_rice_extension_constraint_flag specifies the sps_rrc_rice_extension_flag of all pictures in OlsInScope. gci_no_rrc_rice_extension_constraint_flag equal to 1 specifies that the sps_rrc_rice_extension_flag of all pictures in OlsInScope shall be equal to 0. gci_no_rrc_rice_extension_constraint_flag equal to 0 does not impose this constraint.
[0193] For high bit depth and high bit rate applications, there are many large quantization levels in many places of RRC. For such applications, larger Rice parameters will result in a reduction in the number of containers required to represent the remaining levels. The method of deriving Rice parameters in VVCvl may not be optimal for VVCv2 applications. Therefore, alternative Rice parameter derivation can be used, which can be transmitted through sps_rrc_rice_extension_flag. sps_rrc_rice_extension_flag equal to 1 specifies that alternative Rice parameter derivation is used for the binarization of abs_remaining[] and dec_abs_level[]. sps_rrc_rice_extension_flag equal to 0 specifies that alternative Rice parameter derivation is not used for the binarization of abs_remaining[] and dec_abs_level[]. When not present, the value of sps_rrc_rice_extension_flag is inferred to be equal to 0.
[0194] An example of using sps_rrc_rice_extension_flag to determine Rice parameters in VVCv2 is shown below. Given an array AbsLevel[x][y] of a transform block with a component index cIdx and an upper left luminance position (x0, y0), the variable locSumAbs is derived as specified in the following pseudo-code process (the underlined part is what VVCv2 adds relative to VVCvl).
[0195]
[0196] The lists Tx[] and Rx[] are defined as follows:
[0197] Tx[]={32,128,512,2048}(1523)
[0198] Rx[]={0,2,4,6,8}(1524)
[0199] The value of the variable shiftVal is derived as follows:
[0200]
[0201] The value of locSumAbs is updated as follows:
[0202] locSumAbs=Clip3(0,31,(locSumAbs>>shiftVal)-baseLevel*5) (1526X2)
[0203] Given the variable locSumAbs, first The Rice parameter cRiceParam is derived as specified in Table 128, Then more New as follows :
[0204] cRiceParam=cRiceParam+shiftVal (1526X3)
[0205] When baseLevel is equal to 0, the variable ZeroPos[n] is derived as follows:
[0206] ZeroPos[n]=(QState<2?1:2)< <cRiceParam (1518)
[0207] Table 128 - cRiceParam specifications based on locSumAbs
[0208] locSumAbs 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 cRiceParam 0 0 0 0 0 0 0 1 1 1 1 1 1 1 2 2 locSumAbs 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 cRiceParam 2 2 2 2 2 2 2 2 2 2 2 2 3 3 3 3
[0209] It can be seen from formula (1526X3) that compared with VVCv1, cRiceParam used to binarize the remaining part of the absolute level in VVCv2 can be larger.
[0210] gci_no_persistent_rice_adaptation_constraint_flag
[0211] gci_no_persistent_rice_adaptation_constraint_flag signals a constraint on the derivation of Rice parameters for binarization using the previous TU state. gci_no_persistent_rice_adaptation_constraint_flag equal to 1 specifies that the sps_persistent_rice_adaptation_enabled_flag of all pictures in OlsInScope shall be equal to 0. gci_no_persistent_rice_adaptation_constraint_flag equal to 0 does not impose this constraint. When gci_no_persistent_rice_adaptation_constraint_flag is not present, its value is inferred to be equal to 0.
[0212] sps_persistent_rice_adaptation_enabled_flag is equal to 1 and specifies that at the beginning of each TU, the Rice parameter derivation for binarizing abs_remainder[] and dec_abs_level[] is initialized using the statistical data accumulated by the previous TU. sps_persistent_rice_adaptation_enabled_flag is equal to 0 and specifies that the previous TU state is not used in the Rice parameter derivation. When not present, the value of sps_persistent_rice_adaptation_enabled_flag is inferred to be equal to 0. The following shows an example of using sps_persistent_rice_adaptation_enabled_flag to determine Rice parameters in VVCv2 (the underlined part is the content added by VVCv2 relative to VVCvl).
[0213] If the CTU is the first CTU in a slice or tile, the array PredictorPaletteSize[chType] (chType=0,1) is initialized to 0 according to the initialization process for calling content variables specified in subclause 9.3.2.2, and the array StatCoeff[i] (i=0...2) is initialized as follows:
[0214] StatCoeff[i]= sps_persistent_rice_adaptation_enabled_flag? 2*
[0215] Floor(Log2(BitDepth-10) :0 (1513X3)
[0216] StatCoeff[i] is used to calculate HisValue, which is used to calculate locSumAbs in equation (1517), and HisValue can be updated once per TU. With the help of HisValue, the derived values of Rice parameters at block boundaries can be made more accurate.
[0217] gci_no_reverse_last_sig_coeff_constraint_flag
[0218] gci_no_reverse_last_sig_coeff_constraint_flag equal to 1 specifies that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope shall be equal to 0. gci_no_reverse_last_sig_coeff_constraint_flag equal to 0 does not impose this constraint. When gci_no_reverse_last_sig_coeff_constraint_flag is not present, its value is inferred to be equal to 0.
[0219] In conventional residual coding (RRC), the position (x,y) of the last non-zero level in a TU is encoded by up to four syntax elements: last_sig_coeffx_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. This position is encoded using the difference between (x,y) and (0,0) of the current TU in VVCv1. This is reasonable for VVCv1 because there are many zero levels within a TU and most non-zero levels are located in the top-left corner of the TU. However, this may not be the case for VVCv2 applications, where many non-zero levels are distributed throughout the TU. Therefore, it may be beneficial to encode the position (x,y) relative to the bottom-right corner rather than the top-left corner (0,0). The sh_reverse_last_sig_coeff_flag provides tools to handle such applications.
[0220] sh_reverse_last_sig_coeff_flag equal to 1 specifies that for each transform block in the current slice, the coordinates of the last significant coefficient are encoded relative to ((Log2ZoTbWidth<<1)-1, (Log2ZoTbHeight<<1)-1). sh_reverse_last_sig_coeff_flag equal to 0 specifies that for each transform block in the current slice, the coordinates of the last significant coefficient are encoded relative to (0,0). When not present, the value of sh_reverse_last_sig_coeff_flag is inferred to be equal to 0.
[0221] In the slice header, sps_reverse_last_sig_coeff_enabled_flag is used to conditionally parse sh_reverse_last_sig_coeff_flag as follows.
[0222] -If last_sig_coeff_x_suffix does not exist, the following formula applies:
[0223] LastSignificantCoeffX = last_sig_coeff_x_prefix (193)
[0224] Otherwise (last_sig_coeff_x_suffix exists), the following formula applies:
[0225] LastSignificantCoeffX = ( 1 << ( (last_sig_coeff_x_prefix >> 1 ) - 1 ) )* (194)
[0226] (2+(last_sig_coeff_x_prefix&1))+last_sig_coeff_x_suffix
[0227] When sh_reverse_last_sig_coeff_flag is equal to 1, the value of LastSignificantCoeffX is modified as follows:
[0228] LastSignificantCoeffX = ( 1 << Log2ZoTbWidth ) - 1 -LastSignificantCoeffX (195)
[0229] -If last_sig_coeff_y_suffix does not exist, the following formula applies:
[0230] LastSignificantCoeffY = last_sig_coeff_y_prefix (196)
[0231] Otherwise (last_sig_coeff_y_suffix exists), the following formula applies:
[0232] LastSignificantCoeffY = ( 1 << ( ( last_sig_coeff_y_prefix >> 1 ) -1 ) ) * (197)
[0233] (2+(last_sig_coeff_y_prefix&1))+last_sig_coeff_y_suffix
[0234] When sh_reverse_last_sig_coeff_flag is equal to 1, the value of LastSignificantCoeffY is modified as follows:
[0235] LastSignificantCoeffY = ( 1 << Log2ZoTbHeight ) - 1 -LastSignificantCoeffY (198)
[0236] VVCv2 slice header
[0237]
[0238] Figure 5 An example of a process 500 for video decoding according to some embodiments of the present disclosure is depicted. One or more computing devices implement the process 500 by executing appropriate program code. Figure 5 For example, a computing device implementing the video decoder 200 may implement the operations described in Figure 5 The computing device includes, for example, an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, a loop filter module 220, an inter-frame prediction module 224, and an intra-frame prediction module 226. For illustrative purposes, process 500 is described with reference to some examples depicted in the figure. However, other implementations may also be used.
[0239] In step 502, process 500 involves accessing a bitstream of a video signal, such as encoded video 202. In step 504, process 500 involves extracting a general constraint information (GCI) flag from the bitstream of the video. As described above, a binary flag, gci_present_flag, is used to specify whether a GCI syntax element is present. A gci_present_flag equal to 1 specifies that a GCI syntax element is present in the general_constraints_info() syntax structure, and that the GCI syntax element is used to indicate constraints imposed on additional coding tools. A gci_present_flag equal to 0 specifies that no GCI syntax element is present and that no general constraints are imposed on the video. Depending on the encoder, the GCI flag may be extracted from a network packet of the video, a video parameter set of the video, or a sequence parameter set of the video.
[0240] In step 506, process 500 involves determining whether one or more general constraints are applied to the video based on the value of the GCI flag. If so (i.e., the GCI flag is 1), in step 508, process 500 involves extracting a value M representing the number of additional bits included in the video bitstream from the video bitstream. These additional bits include flag bits that indicate that additional coding tools are constrained for the video.
[0241] In step 510, process 500 involves determining whether M is greater than 5. If so, in step 512, process 500 involves extracting six flag bits from the codestream, the six flag bits representing flags indicating respective constraints of six additional coding tools. These six flags include: the flag gci_all_rap_pictures_constraint_flag, which indicates that the pictures of the video are restricted to intra random access point IRAP pictures or progressive decoding refresh GDR pictures; the flag gci_no_extended_precision_processing_constraint_flag, which indicates whether to constrain the extended transform precision; the flag gci_no_ts_residual_coding_rice_constraint_flag, which indicates whether to constrain explicit Rice parameter transmission; the flag gci_no_rrc_rice_extension_constraint_flag, which indicates an alternative Rice parameter derivation for binarization of the quantized residual of the video; the flag gci_no_persistent_rice_adaptation_constraint_flag, which indicates whether to initialize the Rice parameter derivation for binarization based on the previous transform unit; and the flag gci_no_reverse_last_sig_coeff_constraint_flag, which indicates whether to impose constraints on pictures in OlsInScope when decoding the position of the final non-zero level in the TU.
[0242] If M is greater than 6, process 500 involves extracting the remaining M-6 bits from the bitstream and discarding them in step 513. In other words, the video decoding will be performed independently of these M-6 bits.
[0243] At step 514, process 500 involves decoding the remainder of the video codestream into pictures according to the constraints indicated by the six flags for the six additional coding tools. For example, if the flag gci_all_rap_pictures_constraint_flag is 1, the decoder may determine that all pictures in one or more output layer sets are GDR pictures with ph_recovery_poc_cnt equal to 0, or IRAP pictures, and decode the GDR pictures or IRAP pictures in the one or more output layer sets. If the flag gci_no_extended_precision_processing_constraint_flag is 1, the decoder may determine that extended transform precision is constrained and decode the video by setting the sps_extended_precision_flag to 0 for the pictures in OlsInScope to not use extended dynamic range. If the flag gci_no_ts_residual_coding_rice_constraint_flag is 1, the decoder may determine that explicit Rice parameter transmission is constrained and decode the remainder of the video codestream by disabling alternative Rice parameter transmission for the pictures in OlsInScope. If the flag gci_no_rrc_rice_extension_constraint_flag is 1, the decoder can determine that the derivation of alternative Rice parameters for binarization of the quantized residual of the video is constrained and decode the rest of the video bitstream by disabling the transmission of alternative Rice parameters for the pictures in OlsInScope. If the flag gci_no_persistent_rice_adaptation_constraint_flag is 1, the decoder can determine that the initialization of the derivation of Rice parameters for binarization based on the previous transform unit state is constrained and decode the rest of the video bitstream without initializing the Rice parameters based on the previous transform unit state for the pictures in OlsInScope. If the flag gci_no_reverse_last_sig_coeff_constraint_flag is 1, the decoder can determine that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope is equal to 0, and for each transform block of the current slice, the coordinates of the final significant coefficient are encoded relative to the upper left corner (0,0); and decode the rest of the video bitstream by interpreting the decoded coordinates of the final significant coefficient as relative to the upper left corner (0,0) of each transform block of the current slice.
[0244] If it is determined in step 510 that M is not greater than 5, then in step 518, process 500 involves extracting M bits from the code stream and discarding them. In other words, decoding of the video will proceed independently of these M bits. In step 520, the decoder decodes the video without applying constraints to the six additional coding tools. If it is determined in step 506 that the GCI flag indicates that no general constraints are applied to the video (i.e., the GCI flag is 0), then process 500 involves decoding the video into pictures without general constraints. In some examples, based on the information relative to the GCI flag, the video is decoded into pictures without general constraints. Figure 2 The above process performs decoding. The decoded video can be output for display.
[0245] Figure 6 Another example of a process 600 for video decoding according to some embodiments of the present disclosure is depicted. One or more computing devices implement the process 600 by executing appropriate program code. Figure 6 For example, a computing device implementing the video decoder 200 may implement the operations described in Figure 6 The computing device includes, for example, an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, a loop filter module 220, an inter-frame prediction module 224, and an intra-frame prediction module 226. For illustrative purposes, process 600 is described with reference to some examples depicted in the figure. However, other implementations may also be used.
[0246] In step 602, process 600 involves accessing a bitstream of a video signal, such as encoded video 202. In step 604, process 600 involves extracting a general constraint information (GCI) flag from the bitstream of the video. As described above, a binary flag, gci_present_flag, is used to specify whether a GCI syntax element is present. A gci_present_flag equal to 1 specifies that a GCI syntax element is present in the general_constraints_info() syntax structure, and that the GCI syntax element is used to indicate constraints imposed on additional coding tools. A gci_present_flag equal to 0 specifies that no GCI syntax element is present and that no general constraints are imposed on the video. Depending on the encoder, the GCI flag may be extracted from a network packet of the video, a video parameter set of the video, or a sequence parameter set of the video.
[0247] In step 606, process 600 involves determining whether one or more general constraints are applied to the video based on the value of the GCI flag. If so (i.e., the GCI flag is 1), in step 608, process 600 involves extracting a value M representing the number of additional bits included in the video bitstream from the video bitstream. These additional bits include flag bits that indicate that additional coding tools are constrained for the video.
[0248] In step 610, process 600 involves determining whether M is greater than 6. If so, in step 612, process 600 involves extracting six flag bits from the codestream, the six flag bits representing flags indicating respective constraints of six additional coding tools. These six flags include: the flag gci_all_rap_pictures_constraint_flag, which indicates that the pictures of the video are restricted to intra random access point IRAP pictures or progressive decoding refresh GDR pictures; the flag gci_no_extended_precision_processing_constraint_flag, which indicates whether to constrain the extended transform precision; the flag gci_no_ts_residual_coding_rice_constraint_flag, which indicates whether to constrain explicit Rice parameter transmission; the flag gci_no_rrc_rice_extension_constraint_flag, which indicates an alternative Rice parameter derivation for binarization of the quantized residual of the video; the flag gci_no_persistent_rice_adaptation_constraint_flag, which indicates whether to initialize the Rice parameter derivation for binarization based on the previous transform unit; and the flag gci_no_reverse_last_sig_coeff_constraint_flag, which indicates whether to impose constraints on pictures in OlsInScope when decoding the position of the final non-zero level in the TU.
[0249] In step 614, process 600 involves extracting M-6 bits from the codestream following the six flag bits and discarding the extracted M-6 bits. In other words, decoding of the video will proceed independently of these M-6 bits. In step 616, process 600 involves decoding the remainder of the codestream of the video into pictures according to the constraints indicated by the six flags for the six additional coding tools. For example, if the flag gci_all_rap_pictures_constraint_flag is 1, the decoder may determine that all pictures in one or more output layer sets are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0, and decode the GDR pictures or IRAP pictures in the one or more output layer sets. If the flag gci_no_extended_precision_processing_constraint_flag is 1, the decoder may determine that extended transform precision is constrained and decode the video without using extended dynamic range by setting the sps_extended_precision_flag to 0 for the pictures in OlsInScope. If the flag gci_no_ts_residual_coding_rice_constraint_flag is 1, the decoder can determine that explicit Rice parameter transmission is constrained and decode the remaining part of the video bitstream by disabling alternative Rice parameter transmission for pictures in OlsInScope. If the flag gci_no_rrc_rice_extension_constraint_flag is 1, the decoder can determine that alternative Rice parameter derivation for binarization of quantized residual of video is constrained and decode the remaining part of the video bitstream by disabling alternative Rice parameter transmission for pictures in OlsInScope. If the flag gci_no_persistent_rice_adaptation_constraint_flag is 1, the decoder can determine that initialization of Rice parameter derivation for binarization based on previous transform unit state is constrained and decode the remaining part of the video bitstream without initializing Rice parameters based on previous transform unit state for pictures in OlsInScope.If the flag gci_no_reverse_last_sig_coeff_constraint_flag is 1, the decoder can determine that the sps_reverse_last_sig_coeff_enabled_flag of all pictures in OlsInScope is equal to 0, and for each transform block of the current slice, the coordinates of the final significant coefficient are encoded relative to the upper left corner (0,0); and decode the rest of the video bitstream by interpreting the decoded coordinates of the final significant coefficient as relative to the upper left corner (0,0) of each transform block of the current slice.
[0250] If it is determined in step 610 that M is not greater than 6, then M is 0 or 6. If M is 6, then in step 620, process 600 involves extracting the six flag bits described above. In step 622, the decoder decodes the video based on the extracted flag bits by imposing constraints on the six additional coding tools. If M is 0, then the flag bits are not extracted. In step 624, the decoder decodes the video by not imposing constraints on the six additional coding tools because the flag bits are not extracted. If it is determined in step 606 that the GCI flag indicates that no general constraints are imposed on the video (i.e., the GCI flag is 0), then in step 618, process 600 involves decoding the video into an image without the general constraints. In some examples, based on the information relative to the GCI flag, the video is decoded into an image. Figure 2 The above process performs decoding. The decoded video can be output for display.
[0251] Figure 7 Another example of a process 700 for video decoding according to some embodiments of the present disclosure is depicted. One or more computing devices implement the process 700 by executing appropriate program code. Figure 7 For example, a computing device implementing the video decoder 200 may implement the operations described in Figure 7 The computing device includes, for example, an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, a loop filter module 220, an inter-frame prediction module 224, and an intra-frame prediction module 226. For illustrative purposes, process 700 is described with reference to some examples depicted in the figure. However, other implementations may also be used.
[0252] In step 702, process 700 involves accessing a bitstream of a video signal, such as encoded video 202. In step 704, process 700 involves extracting a general constraint information (GCI) flag from the bitstream of the video. As described above, a binary flag gci_present_flag is used to specify whether a GCI syntax element is present. A gci_present_flag equal to 1 specifies that a GCI syntax element is present in the general_constraints_info() syntax structure, and the GCI syntax element is used to indicate constraints imposed on additional coding tools. A gci_present_flag equal to 0 specifies that no GCI syntax element is present and no general constraints are imposed on the video. Depending on the encoder, the GCI flag can be extracted from a network data packet of the video, a video parameter set of the video, or a sequence parameter set of the video.
[0253] In step 706, process 700 involves determining whether one or more general constraints are applied to the video based on the value of the GCI flag. If so (i.e., the GCI flag is 1), in step 708, process 700 involves extracting a value M representing the number of additional bits included in the video bitstream from the video bitstream. These additional bits include flag bits that indicate that additional coding tools are constrained for the video.
[0254] In step 710, process 700 involves determining whether M is greater than 5. If not, in step 712, process 700 involves extracting M bits from the bitstream and discarding these M bits. In step 714, process 700 involves decoding the remainder of the video bitstream into an image independently of these M bits. If, in step 710, it is determined that M is not greater than 5, in step 718, process 700 involves extracting the six flag bits described above. If M is greater than 6, in step 719, process 700 involves extracting the remaining M-6 bits from the bitstream and discarding them. In other words, decoding of the video will proceed independently of these M-6 bits. In step 720, the decoder decodes the video based on the extracted six flag bits by applying constraints to the six additional coding tools, as described above.
[0255] If it is determined in step 706 that the GCI flag indicates that no general constraints are imposed on the video (i.e., the GCI flag is 0), process 700 involves decoding the video into pictures without the general constraints. Figure 2 The above process performs decoding. The decoded video can be output for display.
[0256] Example of a computing system implementing the transmission of general constraint information
[0257] Any suitable computing system may be used to perform the operations described herein. For example, Figure 8Describes what can be achieved Figure 1 The video encoder 100 or Figure 2 800 . In some embodiments, the computing device 800 may include a processor 812 that is communicatively coupled to a memory 814 and executes computer-executable program code and / or accesses information stored in the memory 814. The processor 812 may include a microprocessor, an application-specific integrated circuit ("ASIC"), a state machine, or other processing device. The processor 812 may include any one or more of a variety of processing devices. Such a processor may include, or may be in communication with, a computer-readable medium storing instructions that, when executed by the processor 812, cause the processor to perform the steps described herein.
[0258] Memory 814 may include any suitable non-transitory computer-readable medium. Computer-readable media may include any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include disks, memory chips, ROM, RAM, ASICs, configured processors, optical storage, magnetic tape or other magnetic storage, or any other medium from which a computer processor can read instructions. These instructions may include processor-specific instructions generated by a compiler and / or interpreter from code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.
[0259] The computing device 800 may also include a bus 816. The bus 816 may be communicatively coupled to one or more components of the computing device 800. The computing device 800 may also include a plurality of external or internal devices, such as input or output devices. For example, the computing device 800 shown has an input / output ("I / O") interface 818 that is capable of receiving input from one or more input devices 820 or providing output to one or more output devices 822. The one or more input devices 820 and the one or more output devices 822 may be communicatively coupled to the I / O interface 818. This communicative coupling may be achieved in any suitable manner (e.g., via a printed circuit board connection, via a cable connection, via wireless transmission communication, etc.). Non-limiting examples of input devices 820 include a touch screen (e.g., one or more cameras for imaging a touch area, or a pressure sensor for detecting pressure changes caused by a touch), a mouse, a keyboard, or any other device that can be used to generate input events in response to physical manipulation by a user of the computing device. Non-limiting examples of output devices 822 include an LCD screen, an external display, a speaker, or any other device that can be used to display or otherwise present output generated by the computing device.
[0260] The computing device 800 can execute program code that configures the processor 812 to perform the above-mentioned Figure 1-7 The program code may include one or more steps of the video encoder 100 or the video decoder 200. The program code may reside in the memory 814 or any suitable computer-readable medium and may be executed by the processor 812 or any other suitable processor.
[0261] The computing device 800 may also include at least one network interface device 824. The network interface device 824 may include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks 828. Non-limiting examples of the network interface device 824 include an Ethernet network adapter, a modem, and / or similar devices. The computing device 800 may transmit messages in the form of electronic or optical signals through the network interface device 824.
[0262] General considerations
[0263] Numerous details are set forth herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will appreciate that the claimed subject matter may be practiced without these details. In other instances, methods, devices, or systems known to those skilled in the art have not been described in detail in order to avoid obscuring the claimed subject matter.
[0264] Unless otherwise specifically noted, it is understood that discussions in this specification using terms such as "process," "compute," "determine," "identify," or similar terms refer to the operations or processes of a computing device (e.g., one or more computers, or similar electronic computing devices) that operate on or transform data represented as physical electronic or magnetic quantities in a memory, register, or other information storage device, transmission device, or display device of a computing platform.
[0265] The systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a result conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device that implements one or more embodiments of the present subject matter. In the software used to program or configure the computing device, any suitable programming language, scripting language, or other type of language or combination of languages may be used to implement the teachings contained herein.
[0266] The method embodiments disclosed herein can be performed in the operation of such a computing device. The order of the steps presented in the above examples can be changed - for example, the steps can be reordered, combined and / or decomposed into sub-steps. Certain steps or processes can be performed in parallel.
[0267] The use of "suitable for" or "configured to" in this document is open and inclusive language and does not exclude devices adapted or configured to perform additional tasks or steps. Furthermore, the use of the term "based on" is also open and inclusive, as a process, step, calculation, or other operation "based on" one or more stated conditions or values may actually be based on conditions or values other than those stated. The headings, lists, and numbering included herein are for ease of explanation only and are not intended to be limiting.
[0268] Although the subject matter has been described in detail with respect to specific embodiments thereof, it is understood that those skilled in the art, after understanding the foregoing, can readily make changes, modifications, and equivalents to these embodiments. Therefore, it should be understood that this disclosure is provided for purposes of illustration and not limitation, and does not exclude such changes, modifications, and / or additions to the subject matter as would be apparent to one of ordinary skill in the art.
Claims
1. A video decoding method, comprising: Decoding an additional bit count M from a bitstream of a video, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits included in the bitstream of the video, the additional bits including flag bits indicating respective additional coding tools to be constrained for the video, and an expected value of the additional bit count M is 0, 6, or greater than 6; In response to determining that the decoded additional bit count M is greater than 6, decoding M-6 bits following the six flag bits in the code stream; and Independently of the decoded M-6 bits, and based at least in part on the constraints specified by the six flag bits for the respective additional coding tools, the remaining portion of the codestream of the video is decoded into images.
2. The method according to claim 1, further comprising: Before decoding the M-6 bits, decoding the six flag bits representing six flags from the bitstream of the video, the six flags respectively indicating six additional coding tools to be constrained for the video; as well as A remaining portion of the codestream of the video is decoded into pictures based at least in part on the constraints indicated by the six flags for the six additional coding tools.
3. The method according to claim 2, wherein: The six signs include: A first flag indicating that the pictures of the video are limited to intra random access point (IRAP) pictures or gradual decoding refresh (GDR) pictures; The second flag indicates whether to constrain the extended transform precision; The third flag indicates whether to constrain explicit Rice parameter transmission; a fourth flag indicating an alternative Rice parameter derivation for binarization of the quantized residual of the video; A fifth flag indicating whether to initialize Rice parameter derivation for binarization based on a previous transform unit; and The sixth flag indicates whether to impose constraints on pictures in the in-scope output layer set OlsInScope when decoding the position of the final non-zero level in the transform unit.
4. The method according to claim 3, wherein: The decoding of the remaining portion of the codestream of the video into images based at least in part on the constraints indicated by the six flags for the six additional coding tools comprises one or more of the following steps: Based on the first flag being 1, determining that all pictures in one or more output layer sets are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0, and decoding the GDR pictures or the IRAP pictures in the one or more output layer sets; Based on the second flag being 1, determining that the extended transform precision is constrained, and decoding the remaining portion of the codestream of the video by setting the sps_extended_precision_flag of the pictures in the output layer set OlsInScope within the scope to be equal to 0 so as not to use the extended dynamic range; Based on the third flag being 1, determining that explicit Rice parameter transmission is constrained, and decoding the remaining portion of the code stream of the video by disabling alternative Rice parameter transmission of pictures in the output layer set OlsInScope within the scope; Based on the fourth flag being 1, determining that the derivation of alternative Rice parameters for binarization of the quantized residual of the video is constrained, and decoding the remaining part of the code stream of the video by disabling the transmission of alternative Rice parameters for pictures in the output layer set OlsInScope within the range; Based on determining that the fifth flag is 1, determining that the initialization of the Rice parameter derivation for binarization based on the previous transform unit state is constrained, and decoding the remaining part of the code stream of the video without initializing the Rice parameters based on the previous transform unit state for the pictures in the output layer set OlsInScope within the range; or Based on the sixth flag being 1, determining that the coordinates of the final significant coefficient are encoded relative to the upper left corner of each transform block of the slice, and decoding the remaining portion of the video code stream by interpreting the decoded coordinates of the final significant coefficient as relative to the upper left corner of each transform block of the slice.
5. The method according to claim 3, further comprising one or more of the following steps: determining that the first flag does not exist in the codestream, and inferring a value of the first flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the second flag does not exist in the codestream, and inferring a value of the second flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the third flag does not exist in the codestream, and inferring a value of the third flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fourth flag does not exist in the codestream, and inferring a value of the fourth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fifth flag does not exist in the codestream, and inferring the value of the fifth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; or It is determined that the sixth flag does not exist in the codestream, and the value of the sixth flag is inferred to be 0, indicating that no constraint is imposed on the corresponding coding tool.
6. The method according to claim 1, wherein Before decoding the additional bit count M, the method further comprises: Decoding a general constraint information (GCI) flag from a bitstream of the video; and Based on the value of the GCI flag, it is determined that one or more general constraints are imposed on the video.
7. The method according to claim 6, wherein: The GCI flag is decoded from a network data packet of the video, a video parameter set of the video, or a sequence parameter set of the video.
8. A non-transitory computer-readable medium having program code stored thereon, the program code being executable by one or more processing devices to perform operations comprising: Decoding an additional bit count M from a bitstream of a video, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits included in the bitstream of the video, the additional bits including flag bits indicating respective additional coding tools to be constrained for the video, and an expected value of the additional bit count M is 0, 6, or greater than 6; In response to determining that the decoded additional bit count M is greater than 6, decoding M-6 bits following the six flag bits in the code stream; and Independently of the decoded M-6 bits, and based at least in part on the constraints specified by the six flag bits for the respective additional coding tools, the remaining portion of the codestream of the video is decoded into images.
9. The non-transitory computer-readable medium of claim 8, wherein: The operations further include: Before decoding the M-6 bits, decoding the six flag bits representing six flags respectively from the bitstream of the video, the six flags respectively indicating six additional coding tools to be constrained for the video; and A remaining portion of the codestream of the video is decoded into pictures based at least in part on the constraints indicated by the six flags for the six additional coding tools.
10. The non-transitory computer-readable medium of claim 9, wherein: The six signs include: A first flag indicating that the pictures of the video are limited to intra random access point (IRAP) pictures or gradual decoding refresh (GDR) pictures; The second flag indicates whether to constrain the extended transform precision; The third flag indicates whether to constrain explicit Rice parameter transmission; a fourth flag indicating an alternative Rice parameter derivation for binarization of the quantized residual of the video; A fifth flag indicating whether to initialize Rice parameter derivation for binarization based on a previous transform unit; and The sixth flag indicates whether to impose constraints on pictures in the in-scope output layer set OlsInScope when decoding the position of the final non-zero level in the transform unit.
11. The non-transitory computer-readable medium of claim 10, wherein: The decoding of the remaining portion of the codestream of the video into images based at least in part on the constraints indicated by the six flags for the six additional coding tools comprises one or more of the following steps: Based on the first flag being 1, determining that all pictures in one or more output layer sets are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0, and decoding the GDR pictures or the IRAP pictures in the one or more output layer sets; Based on the second flag being 1, determining that the extended transform precision is constrained, and decoding the remaining portion of the codestream of the video by setting the sps_extended_precision_flag of the pictures in the output layer set OlsInScope within the scope to be equal to 0 so as not to use the extended dynamic range; Based on the third flag being 1, determining that explicit Rice parameter transmission is constrained, and decoding the remaining portion of the code stream of the video by disabling alternative Rice parameter transmission of pictures in the output layer set OlsInScope within the scope; Based on the fourth flag being 1, determining that the derivation of alternative Rice parameters for binarization of the quantized residual of the video is constrained, and decoding the remaining part of the code stream of the video by disabling the transmission of alternative Rice parameters for pictures in the output layer set OlsInScope within the range; Based on determining that the fifth flag is 1, determining that the initialization of the Rice parameter derivation for binarization based on the previous transform unit state is constrained, and decoding the remaining part of the code stream of the video without initializing the Rice parameters based on the previous transform unit state for the pictures in the output layer set OlsInScope within the range; or Based on the sixth flag being 1, determining that the coordinates of the final significant coefficient are encoded relative to the upper left corner of each transform block of the slice, and decoding the remaining portion of the video code stream by interpreting the decoded coordinates of the final significant coefficient as relative to the upper left corner of each transform block of the slice.
12. The non-transitory computer-readable medium of claim 10, wherein: The operations may also include one or more of the following: determining that the first flag does not exist in the codestream, and inferring a value of the first flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the second flag does not exist in the codestream, and inferring a value of the second flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the third flag does not exist in the codestream, and inferring a value of the third flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fourth flag does not exist in the codestream, and inferring a value of the fourth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fifth flag does not exist in the codestream, and inferring the value of the fifth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; or It is determined that the sixth flag does not exist in the codestream, and the value of the sixth flag is inferred to be 0, indicating that no constraint is imposed on the corresponding coding tool.
13. The non-transitory computer-readable medium of claim 8, wherein: Before decoding the additional bit count M, the operation further includes: Decoding a general constraint information (GCI) flag from a bitstream of the video; and Based on the value of the GCI flag, it is determined that one or more general constraints are imposed on the video.
14. The non-transitory computer-readable medium of claim 13, wherein: The GCI flag is decoded from a network data packet of the video, a video parameter set of the video, or a sequence parameter set of the video.
15. A system comprising: processing equipment; A non-transitory computer-readable medium communicatively coupled to the processing device; wherein the processing device is configured to execute program code stored in the non-transitory computer-readable medium to perform operations, the operations comprising: Decoding an additional bit count M from a bitstream of a video, wherein the additional bit count M indicates the number of additional general constraint information (GCI) bits included in the bitstream of the video, the additional bits including flag bits indicating respective additional coding tools to be constrained for the video, and an expected value of the additional bit count M is 0, 6, or greater than 6; In response to determining that the decoded additional bit count M is greater than 6, decoding M-6 bits following the six flag bits in the code stream; and Independently of the decoded M-6 bits, and based at least in part on the constraints specified by the six flag bits for the respective additional coding tools, the remaining portion of the codestream of the video is decoded into images.
16. The system according to claim 15, wherein: The operations further include: Before decoding the M-6 bits, decoding the six flag bits representing six flags respectively from the bitstream of the video, the six flags respectively indicating six additional coding tools to be constrained for the video; and A remaining portion of the codestream of the video is decoded into pictures based at least in part on the constraints indicated by the six flags for the six additional coding tools.
17. The system according to claim 16, wherein: The six signs include: A first flag indicating that the pictures of the video are limited to intra random access point (IRAP) pictures or gradual decoding refresh (GDR) pictures; The second flag indicates whether to constrain the extended transform precision; The third flag indicates whether to constrain explicit Rice parameter transmission; a fourth flag indicating an alternative Rice parameter derivation for binarization of the quantized residual of the video; A fifth flag indicating whether to initialize Rice parameter derivation for binarization based on a previous transform unit; and The sixth flag indicates whether to impose constraints on pictures in the in-scope output layer set OlsInScope when decoding the position of the final non-zero level in the transform unit.
18. The system according to claim 17, wherein: The decoding of the remaining portion of the codestream of the video into images based at least in part on the constraints indicated by the six flags for the six additional coding tools comprises one or more of the following steps: Based on the first flag being 1, determining that all pictures in one or more output layer sets are GDR pictures or IRAP pictures with ph_recovery_poc_cnt equal to 0, and decoding the GDR pictures or the IRAP pictures in the one or more output layer sets; Based on the second flag being 1, determining that the extended transform precision is constrained, and decoding the remaining portion of the codestream of the video by setting the sps_extended_precision_flag of the pictures in the output layer set OlsInScope within the scope to be equal to 0 so as not to use the extended dynamic range; Based on the third flag being 1, determining that explicit Rice parameter transmission is constrained, and decoding the remaining portion of the code stream of the video by disabling alternative Rice parameter transmission of pictures in the output layer set OlsInScope within the scope; Based on the fourth flag being 1, determining that the derivation of alternative Rice parameters for binarization of the quantized residual of the video is constrained, and decoding the remaining part of the code stream of the video by disabling the transmission of alternative Rice parameters for pictures in the output layer set OlsInScope within the range; Based on determining that the fifth flag is 1, determining that the initialization of the Rice parameter derivation for binarization based on the previous transform unit state is constrained, and decoding the remaining part of the code stream of the video without initializing the Rice parameters based on the previous transform unit state for the pictures in the output layer set OlsInScope within the range; or Based on the sixth flag being 1, determining that the coordinates of the final significant coefficient are encoded relative to the upper left corner of each transform block of the slice, and decoding the remaining portion of the video code stream by interpreting the decoded coordinates of the final significant coefficient as relative to the upper left corner of each transform block of the slice.
19. The system according to claim 17, wherein: The operations may also include one or more of the following: determining that the first flag does not exist in the codestream, and inferring a value of the first flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the second flag does not exist in the codestream, and inferring a value of the second flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the third flag does not exist in the codestream, and inferring a value of the third flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fourth flag does not exist in the codestream, and inferring a value of the fourth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; determining that the fifth flag does not exist in the codestream, and inferring the value of the fifth flag to be 0, indicating that no constraint is imposed on the corresponding coding tool; or It is determined that the sixth flag does not exist in the codestream, and the value of the sixth flag is inferred to be 0, indicating that no constraint is imposed on the corresponding coding tool.
20. The system of claim 15, wherein: Before decoding the additional bit count M, the operation further includes: Decoding a general constraint information (GCI) flag from a network data packet of the video, a video parameter set of the video, or a sequence parameter set of the video, and decoding the GCI flag from a bitstream of the video; and Based on the value of the GCI flag, it is determined that one or more general constraints are imposed on the video.
Citation Information
Patent Citations
High level syntax for video coding and decoding
GB202007554D0
Method and appartus for video coding
US20210314587A1