Sub-block coding inference in video coding

By improving the sub-block flag inference rules, the problem of inaccurate sub-block flag inference in the VVC standard is solved, video coding efficiency is improved, and it becomes an effective coding tool in future video coding standards.

CN119545022BActive Publication Date: 2026-03-17GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411839700.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-13
Filing Date
2023-04-28
Publication Date
2026-03-17
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

In the existing general video coding (VVC) standard, the inference rules for sub-block markers are inaccurate, leading to errors in the initial context state estimation of the entropy coding model and reducing coding efficiency.

Method used

An improved sub-block flag inference rule is provided, which accurately infers the sub-block flag by determining the conditions of the transform block. This improves the consistency between the sub-block flag value and the transform coefficient level, thereby providing a more accurate initial context value for the entropy coding model.

Benefits of technology

It improves video coding efficiency by inferring more accurate sub-block markers, thereby enhancing the coding efficiency of the entropy coding model and becoming an effective coding tool in future video coding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545022B_ABST
    Figure CN119545022B_ABST
Patent Text Reader

Abstract

For a transform block of a frame of video encoded with regular residual coding, a decoder infers the sb coded flags of a sub-block of the transform block that is not present to be 1 if the inferred sub-block is a DC sub-block and / or the last sub-block in the transform block that contains a non-zero coefficient level. Otherwise, the sb coded flag is inferred to be 0. If the transform block is encoded with transform skip residual coding, the sb coded flag of a sub-block that is not present is inferred to be 1. The decoder determines an initial context value for the sub-block flags of an inferred sub-block based on the sub-block flags of the sub-block and uses the entropy coding model to determine the sub-block flags of the encoded sub-block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application number 202380033784.5, entitled "Sub-block Coding Inference in Video Coding", which entered the Chinese national phase of PCT international patent application PCT / US2023 / 066351, filed on April 28, 2023.

[0002] Cross-references to related applications

[0003] This application claims priority to U.S. Provisional Application No. 63 / 363,804, filed April 28, 2022, entitled "Inference Rules for Subblock Flags," and U.S. Provisional Application No. 63 / 364,713, filed May 13, 2022, entitled "Inference Rules for Subblock Flags," which are hereby incorporated in their entirety by reference. Technical Field

[0004] This disclosure generally relates to computer-implemented methods and systems for video processing. Specifically, this disclosure relates to inferring sub-block coding strategies in video coding. Background Technology

[0005] Ubiquitous camera-enabled devices, such as smartphones, tablets, and computers, have made capturing video or images easier than ever before. However, even short videos can have a considerable amount of data. Video encoding technologies (including video encoding and decoding) enable the compression of video data to a smaller size, thus allowing for the storage and transmission of a wide variety of videos. Video encoding has been used in a wide range of applications, such as digital TV broadcasting, video transmission over the internet and mobile networks, real-time applications (e.g., video chat, video conferencing), DVDs and Blu-ray discs, and more. To reduce the storage space used to store video and / or the network bandwidth consumed for transmitting video, it is necessary to improve the efficiency of video encoding schemes. Summary of the Invention

[0006] Some embodiments relate to sub-block coding strategies in inferring video coding. In one example, a method for decoding video encoded with Universal Video Coding (VVC) includes: accessing a binary string representing a frame of the video, the frame including a plurality of coding tree units (CTUs), each CTU including a plurality of transform blocks, and for each transform block of the video frame, determining a flag sb_coded_flag for each inferring sub-block of the transform block. Determining the flag sb_coded_flag includes: determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether a transform is applied to the transform block, and the second flag indicating whether the transform skipping residual coding process is disabled; in response to determining that the first flag is equal to 0 or the second flag is equal to 1, determining that a flag sb_coded_flag for a sub-block does not exist; in response to determining that one or more of the conditions are true, setting the flag sb_coded_flag for the inferring sub-block to a first value; and in response to determining that the conditions are not true, setting the flag sb_coded_flag for the inferring sub-block to a second value. These conditions include: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is the last sub-block in a transform block containing non-zero coefficient levels. A flag `sb_coded_flag` with a second value indicates that all values ​​at the transform coefficient levels of the sub-block can be inferred to be zero. Determining the flag `sb_coded_flag` further includes: in response to determining that the first flag is equal to 1 and the second flag is equal to 0, and in response to determining that the flag `sb_coded_flag` of the sub-block does not exist, inferring that the flag `sb_coded_flag` of the sub-block is the first value. The method further includes: for each transform block of a video frame, determining an initial context value for an entropy coding model used to encode the flag `sub_block_flag` of the coded sub-block based at least in part on the inferred flag `sub_block_flag` of the sub-block, determining the flag `sub_block_flag` of the coded sub-block according to the entropy coding model with the initial context value, and decoding the transform block by decoding at least a portion of a binary string based on the determined flag `sub_block_flag`. The first value of the flag indicates that at least one transform coefficient level in the corresponding sub-block has a non-zero value. The method also includes: reconstructing video frames at least partially based on the decoded transform blocks; and outputting the reconstructed frames of the video for display along with the other frames of the video.

[0007] In another example, a non-transitory computer-readable medium has program code stored thereon that can be executed by one or more processing devices to perform operations. These operations include: accessing a binary string representing a frame of video encoded using Universal Video Coding (VVC). The frame includes multiple Coding Tree Units (CTUs), and each CTU includes multiple transform blocks. These operations also include: for each transform block of the video frame, determining a flag `sb_coded_flag` for each inferred sub-block of the transform block. Determining the flag `sb_coded_flag` includes: determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether a transform is applied to the transform block, and the second flag indicating whether the transform skipping of the residual coding process is disabled; in response to determining that the first flag is equal to 0 or the second flag is equal to 1, determining that a flag `sb_coded_flag` for a sub-block does not exist; in response to determining that one or more of the conditions are true, setting the flag `sb_coded_flag` of the inferred sub-block to a first value; and in response to determining that the conditions are not true, setting the flag `sb_coded_flag` of the inferred sub-block to a second value. These conditions include: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is the last sub-block in a transform block containing non-zero coefficient levels. A flag `sb_coded_flag` with a second value indicates that all values ​​at the transform coefficient levels of the sub-block can be inferred to be zero. Determining the flag `sb_coded_flag` also includes: in response to determining that the first flag is equal to 1 and the second flag is equal to 0, and in response to determining that the flag `sb_coded_flag` of the sub-block does not exist, inferring that the flag `sb_coded_flag` of the sub-block is the first value. These operations also include: for each transform block of a video frame, determining an initial context value for an entropy coding model used to encode the flag `sub_block_flag` of the coded sub-block based at least in part on the inferred flag `sub_block_flag` of the sub-block, determining the flag `sub_block_flag` of the coded sub-block according to the entropy coding model with the initial context value, and decoding the transform block based on the determined flag `sub_block_flag` by decoding at least a portion of a binary string. The first value of the flag indicates that at least one of the transform coefficient levels of the corresponding sub-block has a non-zero value. These operations also include: reconstructing video frames based at least in part on the decoded transform blocks; and outputting the reconstructed frames of the video for display along with the other frames of the video.

[0008] In yet another example, a system includes a processing means and a non-transitory computer-readable medium communicatively coupled to the processing means. The processing means is configured to execute program code stored in the non-transitory computer-readable medium to perform operations. These operations include: accessing a binary string representing a frame of video encoded with Universal Video Coding (VVC), the frame including a plurality of coding tree units (CTUs), each CTU including a plurality of transform blocks, and for each transform block of the video frame, determining a flag sb_coded_flag for each inferred sub-block of the transform block. Determining the flag sb_coded_flag includes: determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether a transform is applied to the transform block, and the second flag indicating whether the transform skipping of the residual coding process is disabled; in response to determining that the first flag is equal to 0 or the second flag is equal to 1, determining that a flag sb_coded_flag for a sub-block does not exist; in response to determining that one or more of the conditions are true, setting the flag sb_coded_flag for the inferred sub-block to a first value, and in response to determining that the conditions are not true, setting the flag sb_coded_flag for the inferred sub-block to a second value. These conditions include: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is the last sub-block in a transform block containing non-zero coefficient levels. A flag `sb_coded_flag` with a second value indicates that all values ​​at the transform coefficient levels of the sub-block can be inferred to be zero. Determining the flag `sb_coded_flag` also includes: in response to determining that the first flag is equal to 1 and the second flag is equal to 0, and in response to determining that the flag `sb_coded_flag` of the sub-block does not exist, inferring that the flag `sb_coded_flag` of the sub-block is the first value. These operations also include: for each transform block of a video frame, determining an initial context value for an entropy coding model used to encode the flag `sub_block_flag` of the coded sub-block based at least in part on the inferred flag `sub_block_flag` of the sub-block, determining the flag `sub_block_flag` of the coded sub-block according to the entropy coding model with the initial context value, and decoding the transform block based on the determined flag `sub_block_flag` by decoding at least a portion of a binary string. The first value of the flag indicates that at least one of the transform coefficient levels of the corresponding sub-block has a non-zero value. These operations also include: reconstructing video frames based at least in part on the decoded transform blocks; and outputting the reconstructed frames of the video for display along with the other frames of the video.

[0009] These illustrative embodiments are mentioned not to limit or restrict this disclosure, but to provide examples to aid in understanding it. Further embodiments are discussed in the detailed description, and further description is provided. Attached Figure Description

[0010] The features, embodiments, and advantages of this disclosure can be better understood by reading the following detailed description with reference to the accompanying drawings.

[0011] Figure 1 This is a block diagram illustrating an example of a video encoder configured to implement the embodiments presented herein.

[0012] Figure 2 This is a block diagram illustrating an example of a video decoder configured to implement the embodiments presented herein.

[0013] Figure 3 Examples of coded tree unit partitioning of images in a video according to some embodiments of the present disclosure are depicted.

[0014] Figure 4 Examples of coding unit partitioning of coding tree units according to some embodiments of the present disclosure are described.

[0015] Figure 5 Examples of coded blocks according to some embodiments of the present disclosure, as well as predetermined scanning and encoding orders for the coded blocks, are described.

[0016] Figure 6 Examples of a process for decoding video frames according to some embodiments of the present disclosure are described.

[0017] Figure 7 Examples of a process for determining the value of a sub-block flag for each sub-block in a transform block, according to some embodiments of the present disclosure, are described.

[0018] Figure 8 Examples of computing systems that can be used to implement some embodiments of this disclosure are described. Detailed Implementation

[0019] Various embodiments provide mechanisms for inferring sub-block coding strategies in video coding. As discussed above, an increasing amount of video data is being generated, stored, and transmitted. Improving the efficiency of video encoding and decoding techniques, thereby representing video with less data without compromising the visual quality of the decoded video, is beneficial. One approach to improving coding efficiency is to compress video-associated data, including sub-block flags, into a binary bitstream using entropy coding with as few bits as possible. In context-based binary arithmetic entropy coding, the coding engine estimates the context probability indicating the likelihood that the next binary symbol will have a value of 1. This estimation requires an initial context probability estimate. The initial context probability estimate for the entropy coding model of sub-block flags can be derived based on the sub-block flags from neighboring sub-blocks of the current sub-block.

[0020] The sub-block flag `sb_coded_flag` indicates whether the corresponding sub-block in the transform block contains a non-zero transformed coefficient level. For example, if all the transformed coefficient levels in a sub-block are zero, then the sub-block does not need to be encoded, and the sub-block flag can be set to 0. In some examples, the sub-block flags of some sub-blocks are not represented by signals, and therefore need to be derived or inferred at the decoder side. However, because the values ​​of some sub-block flags are not inferred in accordance with the transformed coefficient levels contained in the corresponding sub-blocks, the inference rules in earlier versions of the Universal Video Coding (VVC) standard were inaccurate. This inconsistency leads to estimation errors in the initial context state of the entropy coding model for sub-block flags, thus reducing coding efficiency.

[0021] In some embodiments, the video decoder may determine the value of the sub-block flag of a sub-block within a transform block as follows: The decoder may determine that the first flag `transform_skip_flag[x0][y0][cIdx]` is 0 or the second flag `sh_ts_residual_coding_disabled_flag` is 1. If either of these conditions is true (indicating that the transform block is encoded using a regular residual coding process), then for a sub-block where `sb_coded_flag` is not present in the encoded bitstream, the decoder may determine whether one or more of these two conditions are true. These two conditions include: a first condition, which states that the sub-block is a DC sub-block; and a second condition, which states that the sub-block is the last sub-block in a transform block containing non-zero coefficient levels. If one or more of these two conditions are true, the decoder may infer that the sub-block flag of the sub-block is 1, indicating that the current sub-block has non-zero coefficients. Otherwise, the sub-block flag of the sub-block may be inferred to be 0, indicating that all transform coefficient levels in the sub-block may be inferred to be 0. If the first flag transform_skip_flag[x0][y0][cIdx] is 1 and the second flag sh_ts_residual_coding_disabled_flag is 0 (which indicates that the transform block is encoded using transform skip residual coding), then for sub-blocks where sb_coded_flag is not present in the encoded bitstream, the decoder can infer that the flag sb_coded_flag is 1.

[0022] As described herein, some embodiments improve video coding efficiency by providing improved inference rules for sub-block flags. Using the proposed inference rules, the value of the sub-block flag can be inferred consistent with the level of transform coefficients contained in the corresponding sub-block. The inferred sb_coded_flag value more accurately reflects the probability of sb_coded_flag, thus providing a more accurate estimate of the initial context value for entropy coding models. Therefore, coding efficiency can be improved. These techniques could become effective coding tools in future video coding standards.

[0023] Now refer to the attached diagram, Figure 1 This is a block diagram illustrating an example of a video encoder 100 configured to implement the embodiments presented herein. Figure 1 In the example shown, the video encoder 100 includes a partitioning module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, an in-loop filter module 120, an intra-frame prediction module 126, an inter-frame prediction module 124, a motion estimation module 122, a decoded image buffer 130, and an entropy coding module 116.

[0024] The input to the video encoder 100 is an input video 102 containing a sequence of images (also referred to as frames or pictures). In a block-based video encoder, for each image in the picture, the video encoder 100 uses a partitioning module 112 to divide the image into blocks 104, and each block contains multiple pixels. These blocks can be macroblocks, coding tree units, coding units, prediction units, and / or prediction blocks. An image can include blocks of different sizes, and the block partitioning of different images in the video can also be different. Each block can be encoded using different prediction methods, such as intra-frame prediction, inter-frame prediction, or a mixture of intra-frame and inter-frame prediction.

[0025] Typically, the first frame of a video signal is an intra-predicted frame, which is encoded using only intra-prediction. In intra-prediction mode, blocks of frames are predicted using only data from the same frame. Intra-predicted frames can be decoded without information from other frames. To perform intra-prediction, Figure 1 The video encoder 100 shown may employ an intra-prediction module 126. The intra-prediction module 126 is configured to generate an intra-prediction block (prediction block 134) using reconstructed samples from reconstructed blocks 136 of adjacent blocks of the same image. Intra-prediction is performed according to the intra-prediction mode selected for that block. The video encoder 100 then calculates the difference between block 104 and intra-prediction block 134. This difference is referred to as residual block 106.

[0026] To further remove redundancy from the block, the transform module 114 transforms the residual block 106 into the transform domain by applying a transform to the samples in the block. Examples of transforms may include, but are not limited to, the Discrete Cosine Transform (DCT) or the Discrete Sine Transform (DST). The transform values ​​may be referred to as the transform coefficients representing the residual block in the transform domain. In some examples, the residual block can be quantized directly without transforming it through the transform module 114. This is called the transform skip mode.

[0027] The video encoder 100 may further use the quantization module 115 to quantize the transform coefficients to obtain quantized coefficients. Quantization involves dividing the sample by the quantization step size and then rounding, while inverse quantization involves multiplying the quantized value by the quantization step size. This quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of video samples (transformed or untransformed), thereby using fewer bits to represent video samples.

[0028] Quantization of coefficients / samples within a block can be performed independently, and this quantization method is used in some existing video compression standards, such as H.264 and HEVC. For an N x M block, the two-dimensional (2D) coefficients of the block can be converted into a one-dimensional (1-D) array using a specific scan order for coefficient quantization and encoding. Quantization of coefficients within a block can utilize scan order information. For example, the quantization of a given coefficient in a block can depend on the state of previous quantization values ​​in the scan order. To further improve encoding efficiency, more than one quantizer can be used. Which quantizer is used to quantize the current coefficient depends on information prior to the current coefficient in the encoding / decoding scan order. This quantization method is called dependent quantization.

[0029] The degree of quantization can be adjusted using the quantization step size. For example, for scalar quantization, different quantization step sizes can be applied to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. The quantization step size can be indicated by the quantization parameter (QP). Providing the quantization parameter in the encoded bitstream of the video allows the video decoder to apply the same quantization parameter for decoding.

[0030] The quantized samples are then encoded by entropy coding module 116 to further reduce the size of the video signal. Entropy coding module 116 is configured to apply an entropy coding algorithm to the quantized samples. In some examples, the quantized samples are binary-coded into binary binaries, and the coding algorithm further compresses the binary binaries into bits. Examples of binary-coding methods include, but are not limited to, truncated Rice (TR) and finite k-order Exp-Golomb (EGk) binary-coding. To improve coding efficiency, a history-based Rice parameter derivation method is used, where the Rice parameters derived for a transform unit (TU) are based on variables obtained from or updated from previous TUs. Examples of entropy coding algorithms include, but are not limited to, variable-length coding (VLC) schemes, context-adaptive VLC schemes (CAVLC), arithmetic coding schemes, binary-coding, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The entropy-coded data is added to the bitstream of the output encoded video 132.

[0031] As discussed above, reconstructed blocks 136 from neighboring blocks are used for intra-frame prediction of blocks in the image. Generating reconstructed blocks 136 involves calculating the reconstruction residuals of the block. The reconstruction residuals can be determined by applying inverse quantization and inverse transform to the quantization residuals of the block. Inverse quantization module 118 is configured to apply inverse quantization to quantized samples to obtain dequantization coefficients. Inverse quantization module 118 applies the inverse of the quantization scheme applied by quantization module 115 using the same quantization step size as quantization module 115. Inverse transform module 119 is configured to apply the inverse transform of the transform applied by transform module 114 to the dequantized samples, such as inverse DCT or inverse DST. The output of inverse transform module 119 is the reconstruction residuals of the block in the pixel domain. The reconstruction residuals can be added to the predicted block 134 of the block to obtain reconstructed block 136 in the pixel domain. For blocks that skip the transform, inverse transform module 119 is not applied to these blocks. The dequantized samples are the reconstruction residuals of the block.

[0032] Inter-frame prediction or intra-frame prediction can be used to encode blocks in subsequent images following the first intra-frame predicted image. In inter-frame prediction, the prediction of blocks in an image comes from one or more previously encoded video images. To perform inter-frame prediction, the video encoder 100 uses an inter-frame prediction module 124. The inter-frame prediction module 124 is configured to perform motion compensation for blocks based on motion estimates provided by the motion estimation module 122.

[0033] The motion estimation module 122 compares the current block 104 of the current image with the decoded reference image 108 to perform motion estimation. The decoded reference image 108 is stored in the decoded image buffer 130. The motion estimation module 122 selects the reference block from the decoded reference image 108 that best matches the current block. The motion estimation module 122 further identifies the offset between the position of the reference block (e.g., x, y coordinates) and the position of the current block. This offset is called a motion vector (MV) and is provided to the inter-frame prediction module 124. In some cases, multiple reference blocks are identified for a block in multiple decoded reference images 108. Therefore, multiple motion vectors are generated and provided to the inter-frame prediction module 124.

[0034] Inter-frame prediction module 124 uses one or more motion vectors and other inter-frame prediction parameters to perform motion compensation to generate a prediction for the current block, i.e., inter-frame prediction block 134. For example, based on one or more motion vectors, inter-frame prediction module 124 can locate one or more prediction blocks pointed to by one or more motion vectors in one or more corresponding reference images. If there are more than one prediction block, these prediction blocks are combined with some weights to generate prediction block 134 for the current block.

[0035] For an inter-frame prediction block, the video encoder 100 can subtract the inter-frame prediction block 134 from block 104 to generate a residual block 106. The residual block 106 can be transformed, quantized, and entropy-coded in the same manner as the residual of the intra-frame prediction block discussed above. Similarly, the reconstructed block 136 of the inter-frame prediction block can be obtained by inverse quantization, inverse transforming the residual, and subsequently combining it with the corresponding prediction block 134.

[0036] To obtain the decoded image 108 for motion estimation, the reconstructed block 136 is processed by the in-loop filter module 120. The in-loop filter module 120 is configured to smooth pixel transitions, thereby improving video quality. The in-loop filter module 120 can be configured to implement one or more in-loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or an adaptive loop filter (ALF), etc.

[0037] Figure 2 A block diagram depicts an example of a video decoder 200 configured to implement the embodiments presented herein. The video decoder 200 processes encoded video 202 in a bitstream and generates decoded images 208. Figure 2 In the example shown, the video decoder 200 includes an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, an in-loop filter module 220, an intra-frame prediction module 226, an inter-frame prediction module 224, and a decoded image buffer 230.

[0038] Entropy decoding module 216 is configured to perform entropy decoding of encoded video 202. Entropy decoding module 216 decodes quantization coefficients, encoding parameters including intra-frame prediction parameters and inter-frame prediction parameters, and other information. In some examples, entropy decoding module 216 decodes the bitstream of encoded video 202 into a binary representation, and then converts the binary representation into quantization levels of coefficients. The entropy decoding coefficients are then inverse-quantized by inverse quantization module 218, and subsequently inverse-transformed to the pixel domain by inverse transform module 219. The functions of inverse quantization module 218 and inverse transform module 219 are similar to those described in the reference above. Figure 1 The inverse quantization module 118 and inverse transform module 119 are described above. Inverse transform residual blocks can be added to the corresponding prediction block 234 to generate a reconstruction block 236. For blocks that skip the transform, the inverse transform module 219 is not applied to these blocks. The dequantized samples generated by the inverse quantization module 218 are used to generate the reconstruction block 236.

[0039] The prediction block 234 for a specific block is generated based on the block's prediction mode. If the block's coding parameters indicate that the block is intra-frame predicted, the reconstructed block 236 of the reference block in the same image can be fed into the intra-frame prediction module 226 to generate the prediction block 234 for that block. If the block's coding parameters indicate that the block is inter-frame predicted, the prediction block 234 is generated by the inter-frame prediction module 224. The functions of the intra-frame prediction module 226 and the inter-frame prediction module 224 are respectively similar to... Figure 1 The intra-frame prediction module 126 and the inter-frame prediction module 124.

[0040] As referenced above Figure 1 As discussed, inter-frame prediction involves one or more reference images. The video decoder 200 generates a decoded image 208 of the reference images by applying an in-loop filter module 220 to a reconstructed block of the reference images. The decoded image 208 is stored in a decoded image buffer 230 for use by the inter-frame prediction module 224 and also for output.

[0041] Now for reference Figure 3 , Figure 3 Examples of coded tree unit partitioning of images in a video according to some embodiments of the present disclosure are described. (Refer to the above text.) Figure 1 and Figure 2 As discussed, in order to encode images for video, images are divided into blocks, for example, such as... Figure 3 The image shows a coding tree unit (CTU) 302 in a VVC. For example, a CTU 302 can be a 128x128 pixel block. According to the order, as... Figure 3 The CTUs are processed in the order shown. In some examples, each CTU 302 in the image can be divided into one or more coding units (CUs) 402, such as... Figure 4As shown, the CU can be further divided into prediction units or transform units (TUs) for prediction and transformation. Depending on the coding scheme, the CTU 302 can be divided into CUs 402 differently. For example, in VVC, the CU 402 can be rectangular or square and can be encoded without further division into prediction units or transform units. Each CU 402 can be as large as its root CTU 302, or a subdivision of the root CTU 302 as small as a 4x4 block. Figure 4 As shown, in VVC, partitioning CTU 302 into CU 402 can be done using a quadtree, a binary tree, or a ternary tree. Figure 4 In the diagram, solid lines indicate quadtree partitioning, while dashed lines indicate binary or ternary tree partitioning.

[0042] Residual coding process

[0043] In hybrid video coding systems, efficient compression performance can be achieved by selecting from various prediction tools. In VVC, prediction is performed at the CU (Cubic Unit) level. Each coding unit consists of one or more coding blocks (CBs) corresponding to the color components of the video signal. For example, if the chroma format of the video signal is YCbCr, then each coding unit consists of one luma coding block and two chroma coding blocks. A prediction unit (PU) with the same number of blocks and samples as the CU is derived by applying the selected prediction tool. Then, if the prediction is accurate, the difference between the current coding block and the predicted block (called the residual) consists mainly of small amplitude values ​​and is easier to encode than the original samples of the CB. Depending on hardware constraints, each residual block can be divided into one or more transform blocks (TBs). Encoding a single TB is most efficient for compressing residual data, but if the residual block is larger than the maximum transform size supported by VVC, it may be necessary to divide the residual block.

[0044] When a video signal contains camera-captured (“natural”) content, the residuals in each terabyte (TB) can be further compressed by applying an integer-level version of a transform, such as a discrete cosine transform. Lossy compression is typically achieved by quantizing the transform coefficients. The amplitude of the quantization coefficients (often referred to as the transform coefficient level) and the sign of the quantization coefficients are encoded into the bitstream through a residual coding process. For video signals containing screen-captured content, the residuals may not benefit from the application of a transform. For example, if the transform coefficients have high spatial frequency coefficients with relatively high amplitudes, the energy of the residuals will not be compressed into a small number of coefficients by the transform. In this case, the transform can be skipped, and the residual samples can be quantized directly.

[0045] The statistical distribution of transform coefficients typically differs from that of coefficients skipped from the transform. To efficiently encode both transform and skipped transform coefficients, VVC employs two residual coding procedures: regular residual coding (RRC) and transform skip residual coding (TSRC). When a transform is used, RRC is selected for the control unit (CU). When a skipped transform is used and TSRC is available, TSRC is selected for the CU. TSRC is unavailable if the slice header flag `sh_ts_residual_coding_disabled_flag` is set to 1. In this case, RRC is used for both transform and skipped transform CUs.

[0046] Both residual coding processes first collect coefficients into a smaller set of sub-blocks (e.g., 16 samples) called coded sub-blocks. As mentioned above, due to accurate prediction, the residuals are expected to consist primarily of small amplitude values. After quantization, the residuals are expected to consist primarily of zero-value coefficients. The coded sub-block structure achieves an efficient signal representation of a large number of zero-value coefficients. Each coded sub-block of coefficients is associated with a sub-block flag syntax element `sb_coded_flag`. If all coefficients in a sub-block have values ​​of 0, then `sb_coded_flag` is set to 0. For this type of sub-block, only the sub-block flag needs to be decoded from the bitstream, since the values ​​of all coefficients in the sub-block can be inferred to be 0.

[0047] The `sb_coded_flag` itself can be represented by a signal or inferred. In RRC, the position of the least significant coefficient in the TB is represented by a signal before any sub-block flag. The least significant coefficient is the last non-zero coefficient in a two-level hierarchical diagonal scan sequence, where the first level is a diagonal scan across sub-blocks of the CU, and the second level is a diagonal scan of coefficients traversing sub-blocks. Coefficient-level encoding is performed in reverse scan order, starting from the position of the least significant coefficient. Figure 5 Examples of coded blocks according to some embodiments of the present disclosure, and examples of predetermined scan and encoding orders for such coded blocks, are depicted. In this example, transform block 500 comprises 16 sub-blocks 502, and each sub-block may have 4×4 samples. Dashed lines indicate the scan order, and solid lines indicate the encoding order. The scan order is from top left to bottom right, and the encoding order is the reverse of the scan order, i.e., from bottom right to top left. In some examples, encoding begins at the sub-block containing the lowest effective coefficient of the coded block, such as... Figure 5 The subblock _L shown is shown in the image.

[0048] Since the subblock containing the lowest effective coefficient is guaranteed to contain at least one effective coefficient, its associated subblock flag is not represented as a signal but is inferred to be 1. The first subblock in the diagonal scan sequence, subblock (0,0), contains the transform coefficients corresponding to the lowest spatial frequency. Because the lowest spatial frequency is most likely to contain effective coefficients, the first subblock is not guaranteed to contain effective coefficients, but its associated subblock flag is also not represented as a signal and is inferred to be 1. The subblock flags associated with the subblocks between the first subblock and the subblock containing the lowest effective coefficient are represented as signals. Figure 5 In the example shown, the sub-block flags associated with the sub-block between sub-block (0,0) and sub-block _L are represented by signals. These sub-blocks are in Figure 5 The sub-block flags associated with the remaining sub-blocks of transform block 500 are not represented by signals.

[0049] In TSRC, the least significant coefficient position is not indicated by a signal. Coefficient level encoding is performed starting from position (0,0) in scan order. For each sub-block, except for the possible last sub-block, a sub-block flag is indicated by a signal. If the sub-block flag indicated by a signal for every other sub-block in the TB is 0, then the last sub-block flag is inferred to be 1. Otherwise, the last sub-block is also indicated by a signal.

[0050] The sub-block flag, represented by a signal, is encoded into a context-coded binary number using Context Adaptive Binary Arithmetic Coding (CABAC). Decoding the context-coded binary number depends on the context state, which is updated as the binary number is decoded to adapt to the statistics of the syntax elements. VVC maintains two states (multiple hypotheses) tracking each context-coded binary number. The context state of `sb_coded_flag` is initialized by deriving the `ctxInc` value, as shown below.

[0051] The derivation of ctxInc of the syntax element sb_coded_flag

[0052] The inputs to this process are the color component index cIdx, the brightness position (x0, y0) of the top-left sample of the current transform block relative to the top-left sample of the current image, the current sub-block scan position (xS, yS), the previously decoded binary number of the syntax element sb_coded_flag, and the binary logarithms log2TbWidth and log2TbHeight of the transform block width and height. The output of this process is the variable ctxInc.

[0053] The variable csbfCtx is deduced using the current position (xS, yS) and two previously decoded binary numbers log2TbWidth and log2TbHeight of the syntax element sb_coded_flag in the scan order, as follows: – The variables log2SbWidth and log2SbHeight are deduced as follows:

[0054] log2SbWidth = (Min(log2TbWidth, log2TbHeight) < 2 ? 1 : 2) (1)

[0055] log2SbHeight = log2SbWidth (2)

[0056] – Modify the variables log2SbWidth and log2SbHeight as follows:

[0057] – If log2TbWidth is less than 2 and cIdx is equal to 0, then the following applies.

[0058] log2SbWidth = log2TbWidth (3)

[0059] log2SbHeight = 4 - log2SbWidth (4)

[0060] Otherwise, if log2TbHeight is less than 2 and cIdx is equal to 0, the following applies.

[0061] log2SbHeight = log2TbHeight (5)

[0062] log2SbWidth = 4 - log2SbHeight (6)

[0063] – Initialize the variable csbfCtx with 0 and modify it as follows:

[0064] – If transform_skip_flag[x0][y0][cIdx] equals 1 and sh_ts_residual_coding_disabled_flag equals 0, then the following applies:

[0065] – When xS is greater than 0, modify csbfCtx as follows:

[0066] csbfCtx += sb_coded_flag[ xS - 1 ][ yS ] (7)

[0067] – When yS is greater than 0, modify csbfCtx as follows:

[0068] csbfCtx += sb_coded_flag[ xS ][ yS - 1 ] (8)

[0069] – Otherwise (transform_skip_flag[x0][y0][cIdx] equals 0 or sh_ts_residual_coding_disabled_flag equals 1), the following applies:

[0070] – When xS is less than (1<<(log2TbWidth-log2SbWidth))-1, modify csbfCtx as follows:

[0071] csbfCtx += sb_coded_flag[ xS + 1 ][ yS ] (9)

[0072] – When yS is less than (1<<(log2TbHeight-log2SbHeight))-1, modify csbfCtx as follows:

[0073] csbfCtx += sb_coded_flag[ xS ][ yS + 1 ] (10)

[0074] The context index increment ctxInc is derived using the color component indices cIdx and csbfCtx, as follows: – If transform_skip_flag[x0][y0][cIdx] equals 1 and sh_ts_residual_coding_disabled_flag equals 0, then ctxInc is derived as follows:

[0075] ctxInc = 4 + csbfCtx (11)

[0076] Otherwise (transform_skip_flag[x0][y0][cIdx] equals 0 or sh_ts_residual_coding_disabled_flag equals 1), ctxInc is derived as follows:

[0077] – If cIdx equals 0, then the following applies:

[0078] ctxInc = Min(csbfCtx, 1) (12)

[0079] Otherwise (cIdx is greater than 0), ctxInc is derived as follows:

[0080] ctxInc = 2 + Min(csbfCtx, 1) (13)

[0081] In VVC version 10 draft (JVET-T2001), a shared inference rule is used for sub-block flags in both RRC and TSRC. The semantics of sb_coded_flag are as follows, where the inference rule is shown in italics:

[0082] sb_coded_flag[xS][yS] indicates the following content of the sub-block at position (xS, yS) within the current transform block, where the sub-block is an array of transform coefficients:

[0083] When sb_coded_flag[xS][yS] equals 0, all transform coefficient levels of the sub-block at position (xS, yS) are inferred to be equal to 0.

[0084] When sb_coded_flag[xS][yS] does not exist, it is inferred to be equal to 1.

[0085] In RRC, this means that the sub-block flag of the sub-block following the sub-block containing the lowest valid coefficient is also inferred to be 1. Under this inference rule, Figure 5 In the example shown, the subblock flag of the unmarked "S" subblock is inferred to be 1. This means that each unmarked "S" subblock contains at least one non-zero transform coefficient level. However, the unmarked "S" subblock does not contain non-zero coefficients. Because the unmarked "S" subblock precedes Subblock_L in the encoding order, the transform coefficient levels contained in the unmarked "S" subblock are not represented by signals and are therefore inferred to have the correct value of 0, regardless of the inferred value of the subblock flag. However, according to equations (9), (10), (12), and (13), the inferred value of sb_coded_flag may affect the derivation of ctxInc. Because the inferred value of sb_coded_flag in RRC is inconsistent with the transform coefficient levels contained in the corresponding subblock, the context initialization may not be optimal, resulting in reduced encoding efficiency.

[0086] To address the aforementioned issues, the semantics of sb_coded_flag can be replaced with the following, where separate inference rules are defined for sb_coded_flag in both RRC and TSRC. New additions related to JVET-T2001 are shown underlined, and deletions are shown with strikethrough.

[0087]

[0088] In another example of the embodiment, the semantics of sb_coded_flag are replaced with the following. New additions related to JVET-T2001 are shown underlined, and deleted content is shown with strikethrough.

[0089]

[0090] In another example of the embodiment, the semantics of sb_coded_flag are replaced with the following. New additions related to JVET-T2001 are shown underlined, and deleted content is shown with strikethrough.

[0091]

[0092] In yet another example, the semantics of sb_coded_flag are replaced with the following. New additions related to JVET-T2001 are shown underlined, and deleted content is shown with strikethrough.

[0093]

[0094]

[0095] Using the proposed semantics, the sub-block flag associated with the first sub-block and the sub-block containing the lowest effective coefficients is still inferred as 1. However, the sub-block flag associated with sub-blocks in the scan order following the sub-block containing the lowest effective coefficients is inferred as 0. During the context initialization derivation, this may lead to the context state being initialized with the assumption that sb_coded_flag has a high probability of having a value of 0. In this case, if sb_coded_flag does indeed have a value of 0, its encoding efficiency may be high, and if it has a value of 1, its encoding efficiency may be low. Sub-block flags are encoded in a reverse diagonal scan order, meaning that the sub-block flag associated with sub-blocks containing transform coefficients with higher frequencies is encoded first. Such sub-blocks are unlikely to contain effective coefficients, so their sub-block flags are more likely to be 0. Therefore, the inference rule proposed for sub-block flags will make the encoding of sb_coded_flag syntax elements more efficient, thereby improving encoding efficiency.

[0096] Figure 6 An example of a process 600 for decoding video according to some embodiments of the present disclosure is described. One or more computing devices implement this by executing suitable program code. Figure 6 The operations described herein. For example, a computing device implementing the video decoder 200 can be implemented by executing the program code of the entropy decoding module 216, the inverse quantization module 218, and the inverse transform module 219. Figure 6The operations described herein. For illustrative purposes, process 600 is described with reference to some examples depicted in the accompanying drawings. However, other specific implementations are also possible.

[0097] At box 602, process 600 involves accessing a binary string or binary representation of a frame representing the video from the video bitstream of the video signal. When encoding is performed, the frame can be divided into stripes or tiles, or any type of division, processed as units by the video encoder. A frame may include a set of CTUs, such as... Figure 3 As shown. Each CTU includes one or more CUs, such as Figure 4 As shown in the example, each CU may contain one or more transform blocks for encoding.

[0098] At box 604, including 606-610, process 600 involves decoding each transformed block of a frame from a binary string to generate a decoded sample of the transformed block. At box 606, process 600 involves determining a sub-block flag (sb_coded_flag) for each inferred sub-block within the transformed block. Details regarding the determination of the sub-block flag can be found in [reference needed]. Figure 7 As shown.

[0099] At box 608, procedure 600 involves determining an initial context value for an entropy coding model used to encode the sub-block flags. As discussed in detail above, the context index increment ctxInc is determined based on the values ​​of the inferred sub-block flags to the right and below the first coded sub-block flag. The initial context value for the entropy coding model can then be determined by deriving an index into the context state table based on the context index increment ctxInc and retrieving the initial context value from the context state table. At box 609, procedure 600 involves decoding the sub-block flag sb_coded_flag for each coded flag in the transform block, wherein the first coded sub-block flag is decoded using the initial context value, and subsequent coded sub-block flags are decoded using context values ​​updated from the initial context value. At box 610, procedure 600 involves decoding the transform block by decoding a portion of the binary string corresponding to the transform block. Decoding may include decoding the transform coefficient levels of sub-blocks in the transform block where the inferred or decoded sb_coded_flag value is 1. Decoding may also include inferring transform coefficient levels of 0 for sub-blocks within transform blocks where the inferred or decoded sb_coded_flag value is 0. Decoding may also include, for example, using the references above. Figure 2 The discussed methods include inverse quantization, inverse transform (if needed), inter-frame and / or intra-frame prediction to reconstruct samples of sub-blocks.

[0100] At box 612, process 600 involves reconstructing video frames based on the decoded transform blocks. At box 614, process 600 involves outputting the decoded frames of the video, along with other decoded frames of the video, for display.

[0101] Figure 7 An example of a process 700 for determining the value of a sub-block flag for each sub-block in a transform block, according to some embodiments of the present disclosure, is described. One or more computing devices implement this process by executing suitable program code. Figure 7 The operations described herein. For example, a computing device implementing the video decoder 200 can achieve this by executing appropriate program code. Figure 7 The operations depicted herein are described with reference to some examples depicted in the accompanying drawings for illustrative purposes. However, other specific implementations are also possible.

[0102] At box 702, procedure 700 involves determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether a transform is applied to a transform block, and the second flag indicating whether the transform skipping residual coding process is disabled. In some examples, the first flag is `transform_skip_flag[x0][y0][cIdx]`, and the second flag is `sh_ts_residual_coding_disabled_flag`. `transform_skip_flag[x0][y0][cIdx]` indicates whether a transform is applied to the associated transform block. Array indices x0, y0 indicate the position (x0, y0) of the top-left luminance sample of the considered transform block relative to the top-left luminance sample of the image or frame. Array index cIdx indicates the indicator of the color components; equivalent to Y being 0, Cb being 1, and Cr being 2. `transform_skip_flag[x0][y0][cIdx]` equal to 1 indicates that the transform is not applied to the associated transform block. `transform_skip_flag[x0][y0][cIdx]` equal to 0 indicates that the decision of whether a transform is applied to the associated transform block depends on other syntax elements. `sh_ts_residual_coding_disabled_flag` equal to 1 indicates that the `residual_coding()` syntax structure is used to parse the residual samples of the transform skipped blocks of the current slice. `sh_ts_residual_coding_disabled_flag` equal to 0 indicates that the `residual_ts_coding()` syntax structure is used to parse the residual samples of the transform skipped blocks of the current slice. When `sh_ts_residual_coding_disabled_flag` does not exist, it is inferred to be equal to 0.

[0103] If the first flag equals 0 or the second flag equals 1 (indicating that the transform block is RRC encoded), then at box 704, procedure 700 involves determining that the sub-block flag sb_coded_flag of the current sub-block does not exist in the binary string of the frame. At box 706, procedure 700 involves determining whether one or more of two conditions are true. These two conditions include: a first condition, which states that the sub-block is a DC sub-block (e.g., (xS, yS) equals (0, 0)); and a second condition, which states that the sub-block is the last sub-block in a transform block containing non-zero coefficient levels. The second condition can be checked by determining whether (xS, yS) equals (LastSignificantCoeffX >> log2SbW, LastSignificantCoeffY >> log2SbH). Here, (xS, yS) is the current sub-block scan position, and LastSignificantCoeffX and LastSignificantCoeffY are the coordinates of the least significant coefficient (e.g., the last non-zero coefficient) of the transform block. log2TbWidth and log2TbHeight are the binary logarithms of the transform block width and transform block height, respectively.

[0104] If one or more of the two conditions are true, then at box 708, procedure 700 involves inferring the subblock flag of the current subblock (xS, yS) to a first value, for example, 1, to indicate that the current subblock has at least one non-zero transform coefficient level. Otherwise, at box 710, procedure 700 involves inferring the subblock flag of the current subblock (xS, yS) to a second value, for example, 0, to indicate that all transform coefficient levels in the current subblock can be inferred as 0.

[0105] If the first flag is equal to 1 and the second flag is equal to 0 (indicating that the transform block is TSRC encoded), then at box 714, procedure 700 involves determining that the sub-block flag sb_coded_flag of the current sub-block is not present in the binary string of the frame. At box 716, procedure 700 involves inferring that the sub-block flag sb_coded_flag is a first value (e.g., 1). A flag with a first value indicates that at least one transform coefficient level of the sub-block has a non-zero value.

[0106] Example of a computing system

[0107] Any suitable computing system can be used to perform the operations described in this paper. For example, Figure 8 Describes what can be achieved Figure 1 Video encoder 100 or Figure 2Examples of computing devices 800 for video decoder 200. In some embodiments, computing device 800 may include processor 812, which is communicatively coupled to memory 814 and executes computer-executable program code and / or accesses information stored in memory 814. Processor 812 may include a microprocessor, application-specific integrated circuit (“ASIC”), state machine, or other processing means. Processor 812 may include any number of processing means, including one processing means. Such a processor may include a computer-readable medium storing instructions, or be communicative to a computer-readable medium storing instructions that, when executed by processor 812, cause the processor to perform the operations described herein.

[0108] Memory 814 may include any suitable non-transitory computer-readable medium. Computer-readable medium may include any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include disks, memory chips, ROM, RAM, ASICs, configured processors, optical storage, magnetic tape or other magnetic storage, or any other medium from which a computer processor can read instructions. Instructions may include processor-specific instructions generated by a compiler and / or interpreter from code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.

[0109] The computing device 800 may also include a bus 816. The bus 816 may communicatively couple one or more components of the computing device 800. The computing device 800 may also include multiple external or internal devices, such as input or output devices. For example, the computing device 800 is shown having an input / output (“I / O”) interface 818 that can receive input from one or more input devices 820 or provide output to one or more output devices 822. One or more input devices 820 and one or more output devices 822 may be communicatively coupled to the I / O interface 818. The communication coupling may be implemented in any suitable manner (e.g., via a printed circuit board connection, via a cable connection, via wireless communication, etc.). Non-limiting examples of input devices 820 include touchscreens (e.g., one or more cameras for imaging a touch area or a pressure sensor for detecting pressure changes caused by a touch), mice, keyboards, or any other means that can be used to generate input events in response to physical actions of a user of the computing device. Non-limiting examples of output devices 822 include LCD screens, external monitors, speakers, or any other means that can be used to display or otherwise present output generated by the computing device.

[0110] The computing device 800 can execute program code that configures the processor 812 to execute the above reference. Figures 1 to 7 One or more of the operations described. The program code may include video encoder 100 or video decoder 200. The program code may reside in memory 814 or any suitable computer-readable medium and may be executed by processor 812 or any other suitable processor.

[0111] The computing device 800 may also include at least one network interface device 824. The network interface device 824 may include any device or group of devices adapted to establish a wired or wireless data connection to one or more data networks 828. Non-limiting examples of the network interface device 824 include Ethernet network adapters, modems, etc. The computing device 800 may transmit messages as electronic or optical signals via the network interface device 824.

[0112] Overall considerations

[0113] This document sets forth numerous specific details to provide a thorough understanding of the claimed subject matter. However, it will be understood by those skilled in the art that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to those of ordinary skill in the art have not been described in detail so as not to obscure the claimed subject matter.

[0114] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “operating,” “determining,” and “indicating” refer to the actions or processes of a computing device, such as one or more computers or similar electronic computing devices, which manipulate or transform data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0115] The systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a result conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access storage software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device for implementing one or more embodiments of this subject matter. Any suitable programming, scripting, or other type of language or combination of languages ​​may be used to implement the teachings contained in the software used in programming or configuring the computing device.

[0116] Embodiments of the methods disclosed herein can be executed in the operation of such a computing device. The order of the blocks presented in the above examples can be changed; for example, the boxes can be reordered, combined, and / or decomposed into sub-boxes. Some boxes or processes can be executed in parallel.

[0117] The use of “suitable for” or “configured to” in this document is intended to be open-ended and inclusive, and does not exclude means suitable for or configured to perform additional tasks or steps. Furthermore, the use of “based on” is also intended to be open-ended and inclusive, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values ​​may actually be based on additional conditions or values ​​beyond those stated conditions or values. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.

[0118] Although the subject matter has been described in detail with reference to specific embodiments thereof, it will be understood that those skilled in the art can readily make changes, modifications, and equivalents to these embodiments based on the foregoing. Therefore, it is understood that this disclosure has been presented for illustrative purposes rather than limiting, and such modifications, variations, and / or additions to the subject matter are not excluded and will be apparent to those skilled in the art.

Claims

1. A method for decoding a video bitstream, the method comprising: determining, for one sub-block of a current transform block, whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether transform skip is applied to the transform block, and the second flag indicating whether transform skip residual coding process is disabled, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that there is no flag sb_coded_flag for the sub-block, in response to determining that one or more of the conditions are true, inferring the flag sb_coded_flag for the sub-block to be a first value, the conditions including: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is a last sub-block in the transform block containing non-zero coefficient levels, or in response to determining that the conditions are not true, inferring the flag sb_coded_flag for the sub-block to be a second value, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that there is the flag sb_coded_flag for the sub-block, determining, based at least in part on flags sb_coded_flag of preceding sub-blocks, a context index of an arithmetic decoding process for decoding the flag sb_coded_flag of the sub-block, the preceding sub-blocks including decoded sub-blocks adjacent to the sub-block; and decoding the flag sb_coded_flag of the sub-block according to the arithmetic decoding process and the determined context index; and decoding the transform block by decoding at least a portion of the bitstream based on the determined flag sb_coded_flag; wherein the first value is 1 and the second value is 0.

2. The method of claim 1, wherein, the first flag is transform_skip_flag[ x0 ][y0 ][ cIdx ], and the second flag is sh_ts_residual_coding_disabled_flag, where (x0, y0) indicates the luma position of the top-left sample of the transform block relative to the top-left sample of a current picture, and cldx is a color component index.

3. The method of claim 1, wherein, determining that the first condition is true includes determining that an index of a scan position (xS, yS) of the sub-block is equal to (0, 0).

4. The method of claim 1, wherein, determining that the second condition is true comprises determining that a scan position (xS, yS) of the sub-block is equal to (LastSignificantCoeffX » log2SbW, LastSignificantCoeffY » log2SbH), where LastSignificantCoeffX denotes a position of a least significant coefficient in the transform block in horizontal direction, LastSignificantCoeffY denotes the position of the least significant coefficient in the transform block in vertical direction, log2SbW denotes a binary logarithm of a width of the sub-block, and log2SbH denotes a binary logarithm of a height of the sub-block.

5. The method of claim 1, wherein, determining the context index of the arithmetic decoding process comprises: determining a variable csbfCtx based on a position of the sub-block, a decoded sb_coded_flags of a previous sub-block in the scan order, a width of the transform block, and a height of the transform block; determining a context index increment ctxInc based on the variable csbfCtx; and determining the context index of the arithmetic decoding process using the context index increment ctxInc.

6. The method of claim 5, wherein, the previous sub-block comprises a first neighboring sub-block to the right of the sub-block and a second neighboring sub-block below the sub-block in the scan order.

7. A non-transitory computer readable medium storing program code and a bitstream, the program code executable by one or more processing apparatuses for performing operations to generate the bitstream: for a sub-block of a current transform block, determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether transform skip is applied to the transform block, and the second flag indicating whether a transform skip residual coding process is disabled, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that there is no flag sb_coded_flag for the sub-block, in response to determining that one or more of the conditions are true, inferring the flag sb_coded_flag for the sub-block to be a first value, the conditions including: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is a last sub-block in the transform block containing a non-zero coefficient level, or in response to determining that the condition is not true, inferring the flag sb_coded_flag for the sub-block to be a second value, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that there is the flag sb_coded_flag for the sub-block, determining a context index of an arithmetic coding process for encoding the flag sb_coded_flag for the sub-block based at least in part on a flag sb_coded_flag of a previous sub-block, the previous sub-block comprising a coded sub-block neighboring the sub-block; and encoding the flag sb_coded_flag for the sub-block according to the arithmetic coding process and the determined context index; and encoding the transform block based on the determined flag sb_coded_flag; wherein the first value is 1 and the second value is 0.

8. The non-transitory computer-readable medium of claim 7, wherein, The first flag is transform_skip_flag[ x0 ][ y0 ][ cldx ] and the second flag is sh ts residual coding disabled flag, where (x0, y0) indicates the luma position of the top-left sample of the transform block relative to the top-left sample of the current picture and cldx is a color component index.

9. The non-transitory computer-readable medium of claim 7, wherein, Determining that the first condition is true comprises determining that an index of the scan position (xS, yS) of the sub-block is equal to (0, 0).

10. The non-transitory computer-readable medium of claim 7, wherein, Determining that the second condition is true comprises determining that the scan position (xS, yS) of the sub-block is equal to (LastSignificantCoeffX » log2SbW, LastSignificantCoeffY » log2SbH), where LastSignificantCoeffX represents a position of a least significant coefficient in the transform block in a horizontal direction, LastSignificantCoeffY represents the least significant coefficient in the transform block in a vertical direction, log2SbW represents a binary logarithm of a width of the sub-block, and log2SbH represents a binary logarithm of a height of the sub-block.

11. The non-transitory computer-readable medium of claim 7, wherein, Determining the context index of the arithmetic coding process comprises: determining a variable csbfCtx based on a position of the sub-block, encoded sb_coded_flags of previous sub-blocks in a scan order, a width of the transform block, and a height of the transform block; determining a context index increment ctxInc based on the variable csbfCtx; and determining the context index of the arithmetic coding process using the context index increment ctxInc.

12. The non-transitory computer-readable medium of claim 11, wherein, The previous sub-blocks include a first neighboring sub-block to the right of the sub-block and a second neighboring sub-block below the sub-block in the scan order.

13. A method for encoding, comprising: for one sub-block of a current transform block, determining whether a first flag is 0 or a second flag is equal to 1, the first flag indicating whether transform skip is applied to the transform block and the second flag indicating whether a transform skip residual coding process is disabled, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that there is no flag sb_coded_flag for the sub-block, in response to determining that one or more conditions are true, the conditions including: a first condition that the sub-block is a DC sub-block; and a second condition that the sub-block is a last sub-block in the transform block containing a non-zero coefficient level, inferring the flag sb_coded_flag for the sub-block to be a first value, in response to determining that the conditions are not true, inferring the flag sb_coded_flag for the sub-block to be a second value, in response to determining that the first flag is equal to 0 or the second flag is equal to 1, and in response to determining that the first condition is true, determining a context index of an arithmetic coding process used for encoding the flag sb_coded_flag of the sub-block based at least in part on a flag sb_coded_flag of a previous sub-block, the previous sub-block comprising a coded sub-block adjacent to the sub-block in the scan order; and encoding the flag sb_coded_flag of the sub-block according to the arithmetic coding process and the determined context index; and encoding the transform block based on the determined flag sb_coded_flag. wherein the first value is 1 and the second value is 0.

14. The method of claim 13, wherein, The first flag is transform_skip_flag[ x0 ][ y0 ][ cldx ], and the second flag is sh ts residual coding disabled flag, where (x0, y0) indicates the luma position of the top-left sample of the transform block relative to the top-left sample of the current picture, and cldx is a color component index.

15. The method of claim 13, wherein, Determining that the first condition is true comprises determining that the index of the scan position (xS, yS) of the sub-block is equal to (0, 0).

16. The method of claim 13, wherein, Determining that the second condition is true comprises determining that the scan position (xS, yS) of the sub-block is equal to (LastSignificantCoeffX » log2SbW, LastSignificantCoeffY » log2SbH), where LastSignificantCoeffX represents the position of the least significant coefficient in the transform block in the horizontal direction, LastSignificantCoeffY represents the position of the least significant coefficient in the transform block in the vertical direction, log2SbW represents the binary logarithm of the width of the sub-block, and log2SbH represents the binary logarithm of the height of the sub-block.

17. The method of claim 13, wherein, Determining the context index of the arithmetic coding process comprises: determining a variable csbfCtx based on the position of the sub-block (xS, yS), the encoded sb_coded_flags of previous sub-blocks in the scan order, the width of the transform block, and the height of the transform block; determining a context index increment ctxInc based on the variable csbfCtx; and determining the context index of the arithmetic coding process using the context index increment ctxInc.

18. The method of claim 17, wherein, The previous sub-blocks comprise a first neighboring sub-block to the right of the sub-block and a second neighboring sub-block below the sub-block in the scan order.

Citation Information

Patent Citations

  • Residual and coefficients coding for video coding

    WO2021138353A1

  • Image decoding method related to signaling of flag indicating whether tsrc is available, and device therefor

    WO2021158048A1