Method and apparatus for encoding video data in transform skip mode

By relocating the log2_transform_skip_max_size_minus2 parameter to the SPS and optimizing residual coding techniques, the solution addresses inefficiencies in VVC/H.266 for larger transform blocks, achieving improved compression efficiency and decoder throughput.

JP2025078631AActive Publication Date: 2025-05-20ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025015158
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2025-01-31
Publication Date
2025-05-20
Estimated Expiration
2040-08-13

AI Technical Summary

Technical Problem

The existing video coding standards, such as VVC/H.266, face challenges in achieving lossless compression for transform blocks larger than 32x32 due to parse dependencies and inefficiencies in residual coding, particularly with the transform skip mode, which limits compression efficiency and decoder throughput.

Method used

The solution involves moving the log2_transform_skip_max_size_minus2 parameter from the Picture Parameter Set (PPS) to the Sequence Parameter Set (SPS) to eliminate parse dependencies, allowing transform skip mode for larger transform blocks up to the maximum allowed size, and implementing residual coding techniques that decouple inverse level mapping from CABAC parsing, enabling efficient lossless compression for larger transform blocks.

Benefits of technology

This approach enhances compression efficiency by allowing lossless compression for larger transform blocks, reduces parse dependency issues, and improves decoder throughput by optimizing residual coding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078631000001_ABST
    Figure 2025078631000001_ABST
Patent Text Reader

Abstract

To provide a method and apparatus for video processing.SOLUTION: A method and apparatus for video processing includes determining to skip a transform process for a prediction residual on the basis of a maximum transform size of a prediction block, and signaling the maximum transform size in a sequence parameter set (SPS).SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 899,738, filed September 12, 2019, and U.S. Provisional Patent Application No. 62 / 904,880, filed September 24, 2019, both of which are incorporated herein by reference in their entireties. [Background technology]

[0002] background

[0002] A video is a sequence of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video may be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC) / H.265 standard, the Versatile Video Coding (VVC) / H.266 standard, and the AVS standard, that specify specific video coding formats, are developed by standardization organizations. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of the new video coding standards becomes higher and higher. Summary of the Invention [Means for solving the problem]

[0003] Disclosure Summary

[0003] An embodiment of the present disclosure provides a method and apparatus for video processing. In an example embodiment, the method includes determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block, and signaling the maximum transform size in a sequence parameter set (SPS).

[0004]

[0004] In another embodiment, an apparatus includes a memory configured to store instructions and a processor, the processor configured to execute the instructions to cause the apparatus to determine to skip a transformation process for a prediction residual based on a maximum transform size of a prediction block, and to signal the maximum transform size in a sequence parameter set (SPS).

[0005] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method including: determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; and signaling the maximum transform size in a sequence parameter set (SPS).

[0006]

[0006] In another example embodiment, a method includes receiving a bitstream of a video sequence, determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determining to skip a transform process for a prediction residual of the prediction block based on the maximum transform size.

[0007]

[0007] In another embodiment, an apparatus includes a memory configured to store instructions and a processor, the processor being configured to execute the instructions to cause the apparatus to receive a bitstream of a video sequence, determine a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determine to skip a transform process for a prediction residual of the prediction block based on the maximum transform size.

[0008] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method including receiving a bitstream of a video sequence, determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determining to skip a transform process for a prediction residual of the prediction block based on the maximum transform size.

[0009] BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Embodiments and various aspects of the present disclosure are set forth in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief description of the drawings]

[0010] [Figure 1]

[0010] FIG. 2 is a schematic diagram illustrating the structure of an example video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0011] 1 illustrates a schematic diagram of an example encoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B]

[0012] 4 shows a schematic diagram of another example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3A]

[0013] 1 shows a schematic diagram of an example decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3B]

[0014] 4 shows a schematic diagram of another example decoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 4]

[0015] 1 shows a block diagram of an example apparatus for encoding or decoding video according to some embodiments of the present disclosure. [Diagram 5]

[0016] Table 1 illustrates an example syntax structure of a sequence parameter set (SPS) according to some embodiments of the present disclosure. [Figure 6]

[0017] Table 2 illustrates an example syntax structure of a picture parameter set (SPS) according to some embodiments of the present disclosure. [Figure 7]

[0018] 1 shows Table 3 illustrating an example syntax structure of a transform unit according to some embodiments of the present disclosure. [Figure 8]

[0019] 4 illustrates an example syntax structure related to signaling of a block differential pulse code modulation (BDPCM) mode, according to some embodiments of the present disclosure. [Figure 9]

[0020] Table 5 illustrates another example syntax structure of an SPS, according to some embodiments of the present disclosure. [Figure 10]

[0021] 11 shows Table 6 illustrating another example syntax structure of a transform unit according to some embodiments of the present disclosure. [Figure 11]

[0022] FIG. 2 is a schematic diagram illustrating an example of diagonal scanning of a 64×64 transform block (TB) according to some embodiments of the present disclosure. [Figure 12A]

[0023] 1 illustrates an example residual unit (RU) according to some embodiments of the present disclosure. [Figure 12B] 2 illustrates an example residual unit (RU) according to some embodiments of the present disclosure. [Figure 12C] 2 illustrates an example residual unit (RU) according to some embodiments of the present disclosure. [Figure 12D] 2 illustrates an example residual unit (RU) according to some embodiments of the present disclosure. [Figure 13]

[0024] FIG. 1 is a schematic diagram illustrating an example of diagonal scanning of a 64×64 TB where the TB is divided into four 32×32 RUs according to some embodiments of the present disclosure. [Figure 14A]

[0025] 7 shows Table 7 illustrating an example syntax structure for residual coding when a TB is split into RUs according to some embodiments of the present disclosure. [Figure 14B]

[0025] Table 7 illustrates an example syntax structure for residual coding when a TB is split into RUs, according to some embodiments of the present disclosure. [Figure 14C]

[0025] Table 7 illustrates an example syntax structure for residual coding when a TB is split into RUs, according to some embodiments of the present disclosure. [Figure 14D]

[0025] Table 7 illustrates an example syntax structure for residual coding when a TB is split into RUs, according to some embodiments of the present disclosure. [Figure 15A]

[0026] 11 shows Table 8 illustrating another example syntax structure for residual coding according to some embodiments of the present disclosure. [Figure 15B]

[0026] Table 8 illustrates another example syntax structure for residual coding according to some embodiments of the present disclosure. [Figure 15C]

[0026] Table 8 illustrates another example syntax structure for residual coding according to some embodiments of the present disclosure. [Figure 15D]

[0026] Table 8 illustrates another example syntax structure for residual coding according to some embodiments of the present disclosure. [Figure 16]

[0027] 13 shows Table 9 illustrating example parameter values ​​derived from a chroma format according to some embodiments of the present disclosure. [Figure 17]

[0028] 1 illustrates an example syntax structure of Versatile Video Coding Draft 6 for residual coding with inverse level mapping according to some embodiments of the present disclosure. [Figure 18]

[0029] 1 is a flowchart of an example decoding method according to some embodiments of the present disclosure. [Figure 19]

[0030] 11 illustrates an example syntax structure for residual coding where inverse level mapping is not performed, according to some embodiments of the present disclosure. [Figure 20]

[0031] 12 shows an example lookup table for selecting Rice parameters according to some embodiments of the present disclosure. [Figure 21]

[0032] 1 illustrates a flowchart of an example process for video processing according to some embodiments of the present disclosure. [Figure 22]

[0033] 1 shows a flowchart of another example process for video processing according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Detailed Description

[0034] Reference can now be made in detail to example embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations described in the following description of example embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present invention as described in the appended claims. Certain aspects of the present disclosure are described in more detail below. In case of conflict with incorporated terms and / or definitions, the terms and definitions provided herein shall control.

[0012]

[0035] The ITU-T Video Coding Expert Group (VCEG) and the ISO / IEC Moving Picture Expert Group (MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC) / H.266 standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC) / H.265 standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265, but with half the bandwidth.

[0013]

[0036] To achieve the same subjective quality as HEVC / H.265 at half the bandwidth, JVET has developed techniques that go beyond HEVC using the joint exploration model (JEM) reference software. As the coding techniques are incorporated into JEM, JEM has achieved significantly higher coding performance than HEVC.

[0014]

[0037] The VVC standard is a recent development and continues to add more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0015]

[0038] A video is a series of still pictures (or "frames") arranged in time sequence to preserve visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display capability) can be used to display such pictures in time sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for surveillance, conference hosting, or live broadcasting, etc.

[0016]

[0039] To reduce the storage space and transmission bandwidth required in such applications, video may be compressed before storage and transmission, and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The compression module is commonly called an "encoder" and the decompression module is commonly called a "decoder." The encoders and decoders may be collectively referred to as a "codec." The encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoders and decoders may include any suitable computer-implemented algorithm or process fixed in program code, computer-executable instructions, firmware, or a computer-readable medium. Video compression and decompression may be performed by a variety of algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x family, etc. In some applications, a codec can reconstruct video from a first encoding standard and recompress the reconstructed video using a second encoding standard, in which case the codec is sometimes called a "transcoder."

[0017]

[0040] A video encoding process can identify and retain useful information that can be used for reconstruction of a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such an encoding process may be called "lossy". Otherwise, it may be called "lossless". Most encoding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0018]

[0041] Useful information of a picture being encoded (called the "current picture") includes changes with respect to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes may include pixel position changes, luminance changes, or color changes, among which position changes are the most important. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.

[0019]

[0042] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture". A picture that is coded using a previous picture as a reference picture is called a "P-picture". A picture that is coded using both a previous picture and a future picture as reference pictures (i.e., the references are "bidirectional") is called a "B-picture".

[0020]

[0043] 1 illustrates the structure of an example video sequence 100 according to some embodiments of the present disclosure. Video sequence 100 may be live video or captured and archived video. Video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files saved on a storage device), or a video feed interface for receiving video from a video content provider (e.g., a video broadcast transceiver).

[0021]

[0044] As shown in FIG. 1, a video sequence 100 may include a series of pictures arranged in time along a timeline including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture whose reference picture is picture 102 itself. Picture 104 is a P-picture whose reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture whose reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following the picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the embodiment of the reference pictures to the example shown in FIG.

[0022]

[0045] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they may divide a picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, whose boundaries are indicated by dashed lines. In some embodiments, the elementary processing units may be referred to as "macroblocks" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" (CTUs) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The elementary processing units may have variable sizes of pictures, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any shape and size of pixels. The size and shape of the elementary processing units may be selected on a picture-by-picture basis based on a balance between coding efficiency and the level of detail to be maintained in the elementary processing units.

[0023]

[0046] A basic processing unit may be a logical unit that may include a collection of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic lightness information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements (wherein the luma and chroma components may have the same size of the basic processing unit). The luma and chroma components may be referred to as "coding tree blocks" (CTBs) in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be repeated on each of its luma and chroma components.

[0024]

[0047] Video coding has multiple stages of operation, examples of which are shown in FIGS. 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large to process, and therefore may be further divided into segments, referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size as the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that may include a collection of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated on each of its luma and chroma components. Note that such division can be done to further levels depending on the processing needs. Note also that different stages can use different schemes to divide the basic processing units.

[0025]

[0048] For example, in a mode decision stage (an example of which is shown in FIG. 2B), the encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may split the basic processing unit into multiple basic processing sub-units (e.g., CUs in the case of H.265 / HEVC or H.266 / VVC) and decide the prediction type for each individual basic processing sub-unit.

[0026]

[0049] As another example, in the prediction stage (one example of which is shown in FIG. 2A-2B), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and perform prediction operations at the level of the segments.

[0027]

[0050] As another example, in the transform stage (one example of which is shown in FIG. 2A-2B), the encoder can perform transform operations on residual elementary processing subunits (e.g., CUs). However, in some cases, the elementary processing subunits may still be too large to process. The encoder can further divide the elementary processing subunits into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and perform transform operations at the level of the segments. Note that the division scheme of the same elementary processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0028]

[0051] 1, the basic processing units 112 are further divided into 3×3 basic processing sub-units, the boundaries of which are indicated by dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different schemes.

[0029]

[0052] In some implementations, to provide parallel processing capabilities and error resilience for video encoding and decoding, a picture may be divided into multiple regions for processing, such that for each region of a picture, the encoding or decoding process can not depend on information from any other region of the picture. That is, each region of a picture can be processed independently. In this way, the codec can process different regions of a picture in parallel, thus improving coding efficiency. Also, if data of one region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video coding standards, a picture may be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two region types: "slice" and "tile". It should also be noted that different pictures of the video sequence 100 may have different partition schemes for dividing the picture into regions.

[0030]

[0053] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and this disclosure is not limited to those embodiments.

[0031]

[0054] FIG. 2A illustrates a schematic diagram of an example encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (called "original pictures") arranged in a temporal order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into elementary processing units, elementary processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of elementary processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode one elementary processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for a region of each original picture of video sequence 202 (eg, regions 114-118).

[0032]

[0055] In FIG. 2A , an encoder may send a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may generate a residual BPU 210 by subtracting the prediction BPU 208 from the original BPU. The encoder may send the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may send the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the “forward path.” During process 200A, the encoder may send quantized transform coefficients 216, after quantization stage 214, to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may generate a prediction reference 224 to be used in the prediction stage 204 for the next iteration of process 200A by adding the reconstructed residual BPU 222 to a prediction BPU 208. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0033]

[0056] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture of the video sequence 202.

[0034]

[0057] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.

[0035]

[0058] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, and may perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224, which can be used to reconstruct the original BPU as the predicted BPU 208.

[0036]

[0059] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating the predicted BPU 208, the encoder may generate the residual BPU 210 by subtracting it from the original BPU. For example, the encoder may subtract the values ​​(e.g., grayscale or RGB values) of pixels of the predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.

[0037]

[0060] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns" (each basis pattern is associated with a "transform coefficient"). The basis patterns may have the same size (e.g., the size of the residual BPU 210). Each basis pattern may represent a variation frequency (e.g., frequency of brightness variation) component of the residual BPU 210. No basis pattern can be reproduced from any combination (e.g., linear combination) of the other basis patterns. That is, this decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0038]

[0061] Different transform algorithms may use different basis patterns. For example, various transform algorithms such as discrete cosine transform or discrete sine transform may be used in transform stage 212. The transform in transform stage 212 is reversible. That is, the encoder may restore the residual BPU 210 by the inverse operation of the transform (called the "inverse transform"). For example, to restore the pixels of the residual BPU 210, the inverse transform may be to generate a weighted sum by multiplying the values ​​of the corresponding pixels of the basis pattern by the respective associated coefficients and adding the products. For the video coding standard, both the encoder and the decoder may use the same transform algorithm (and therefore the same basis pattern). Thus, the encoder may record only the transform coefficients, and the decoder may reconstruct the residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0039]

[0062] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns may represent different variation frequencies (e.g., brightness variation frequencies). Because the human eye is generally good at recognizing low-frequency variations, the encoder can ignore the information of high-frequency variations without causing significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns may be converted to zero, and the transform coefficients of the low-frequency basis patterns may be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients with the inverse operation of quantization (called "dequantization").

[0040]

[0063] Because the encoder ignores such division remainders in rounding operations, the quantization stage 214 may be lossy. In general, the quantization stage 214 may contribute the most information loss in the process 200A. The more information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different levels of information loss, the encoder may use different values ​​of the quantization parameter or other parameters of the quantization process.

[0041]

[0064] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless or lossy compression algorithms. In some embodiments, besides the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, a transform type in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), or encoder control parameters (e.g., bitrate control parameters). The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0042]

[0065] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may generate reconstructed transform coefficients by performing inverse quantization on the quantized transform coefficients 216. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may generate a prediction reference 224 to be used in the next iteration of process 200A by adding the reconstructed residual BPU 222 to the prediction BPU 208.

[0043]

[0066] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages of FIG. 2A.

[0044]

[0067] 2B shows a schematic diagram of another example encoding process 200B consistent with an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x family). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and splits prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.

[0045]

[0068] In general, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can predict a current BPU by using pixels from one or more already coded neighboring BPUs in the same picture. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can predict a current BPU by using regions from one or more already coded pictures. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0046]

[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra prediction. With respect to the original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, or polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at a pixel level, for example, by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard used. In the case of intra prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, or the orientation of the neighboring BPUs used relative to the original BPU.

[0047]

[0070] As another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the prediction BPU 208. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (called a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU in the current picture. For example, the search window may be centered on a location having the same coordinates in the reference picture as the original BPU of the current picture, and may extend outward by a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., using a pel recursion algorithm or a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the location of the original BPU as time progresses. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can search for a matching region and determine its associated motion vector for each reference picture. In some embodiments, the encoder can assign weights to pixel values ​​of the matching region in each matching reference picture.

[0048]

[0071] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, or zooming. In the case of inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, a number of reference pictures, or weights associated with the reference pictures.

[0049]

[0072] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder can move the matching region of the reference picture according to a motion vector that allows the encoder to predict the original BPU of the current picture. If multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder can move the matching region of the reference picture according to the respective motion vectors and average the pixel values ​​of the matching region. In some embodiments, if the encoder assigned weights to the pixel values ​​of the matching region of each matching reference picture, the encoder can add a weighted sum of the pixel values ​​of the moved matching region.

[0050]

[0073] In some embodiments, inter prediction may be unidirectional or bidirectional. Unidirectional inter prediction may use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter prediction may use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture in which reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0051]

[0074] Further referring to the forward path of process 200B, after spatial prediction stage 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique in which the encoder can select a prediction mode to minimize the value of a cost function depending on the bitrates of candidate prediction modes and the distortions of reconstructed reference pictures under the candidate prediction modes. Depending on the selected prediction mode, the encoder can generate a corresponding predicted BPU 208 and predicted data 206.

[0052]

[0075] In the reconstruction path of process 200B, if an intra prediction mode was selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU encoded and reconstructed in the current picture), the encoder can send the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter prediction mode was selected in the forward path, after generating the prediction reference 224 (e.g., the current picture encoded and reconstructed for all BPUs), the encoder can send the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as, for example, deblocking, sample adaptive offset, or adaptive loop filter. The loop filtered reference picture may be stored in a buffer 234 (or a “decoded picture buffer”) for later use (e.g., to be used as an inter-predicted reference picture for a future picture of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in a binary encoding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.

[0053]

[0076] FIG. 3A shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder can decode one fundamental processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0054]

[0077] In FIG. 3A, the decoder may send a portion of the video bitstream 228 associated with an encoded picture basic processing unit (referred to as an “encoding BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may send the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may send the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may generate a prediction reference 224 by adding the reconstructed residual BPU 222 to the prediction BPU 208. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may send the prediction reference 224 to a prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0055]

[0078] The decoder may iteratively perform process 300A to decode each encoded BPU of the encoded picture and generate a predicted reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all encoded BPUs of the encoded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0056]

[0079] In the binary decoding stage 302, the decoder may perform the inverse of the binary encoding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithms). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), or encoder control parameters (e.g., bitrate control parameters). In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder may depacketize the video bitstream 228 before sending it to the binary decoding stage 302.

[0057]

[0080] 3B shows a schematic diagram of another example decoding process 300B consistent with an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x family). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.

[0058]

[0081] In the process 300B, for an encoding basic processing unit (referred to as a "current BPU") of an encoded picture (referred to as a "current picture") being decoded, the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, or a parameter of the intra prediction operation, etc. The parameter of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPU, a parameter of extrapolation, or a direction of the neighboring BPU relative to the original BPU, etc. As another example, if inter prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, or a parameter of the inter prediction operation, etc. Parameters for the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions in each reference picture, or one or more motion vectors respectively associated with the matching regions.

[0059]

[0082] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may generate a prediction reference 224 by adding the prediction BPU 208 and the reconstructed residual BPU 222, as shown in FIG. 3A.

[0060]

[0083] In the process 300B, the decoder can send the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can send the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture to which all BPUs are decoded), the encoder can send the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner shown in FIG. 2B. The loop filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., to be used as inter-prediction reference pictures for future encoded pictures in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter prediction was used to encode the current BPU, the prediction data may further include parameters of a loop filter (e.g., a loop filter strength).

[0061]

[0084] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video, according to an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 can become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any combination of several central processing units (i.e., "CPU"), graphic processing units (i.e., "GPU"), neural processing units ("NPU"), microcontroller units ("MCU"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), or application specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped together as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0062]

[0085] The device 400 may also include a memory 404 configured to store data (e.g., an instruction set, computer code, or intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing stages of the process 200A, 200B, 300A, or 300B) and data for processing (e.g., the video sequence 202, the video bitstream 228, or the video stream 304). The processor 402 may access the program instructions and the data for processing (e.g., via a bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of several random access memories (RAM), read only memories (ROM), optical disks, magnetic disks, hard drives, solid state drives, flash drives, security digital (SD) cards, memory sticks, or compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped as a single logical component.

[0063]

[0086] Bus 410 may be a communications device that transfers data between components within apparatus 400, such as an internal bus (eg, a CPU memory bus) or an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port).

[0064]

[0087] For the sake of clarity and simplicity, in this disclosure, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be fully or partially integrated with any other components of the device 400.

[0065]

[0088] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, or a mobile communication network, etc.) In some embodiments, the network interface 406 may include any combination of a number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, or cellular network chips, etc.

[0066]

[0089] In some embodiments, optionally, apparatus 400 may further include a peripheral interface 408 to provide a connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), or a video input device (e.g., a camera, or an input interface communicatively coupled to a video archive), etc.

[0067]

[0090] It should be noted that a video codec (e.g., a codec performing processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within apparatus 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of apparatus 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, or NPUs).

[0068]

[0091] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIG. 2A or 2B, inverse quantization 218 in FIG. 3A or 3B), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used to code a picture or slice may be signaled at a high level, for example, using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slice_qp_delta syntax element in the slice header. Additionally, the QP value may be adapted at a local level per CU using a delta QP value sent at the granularity of a quantization group.

[0069]

[0092] In Versatile Video Coding Draft 6 (VVC6), the residual of a transform block (TB) of video data may be coded using a transform skip (TS) mode in which the transform stage is skipped. For example, a decoder may decode the video data using the TS mode by decoding the video data to obtain the residual, and then perform inverse quantization and reconstruction without performing an inverse transform. VVC6 limits the applicability of the TS mode by a maximum block size (here, the TS mode is applicable to a TB only if the width and height of the TB are at most 32 pixels). Such a maximum block size for applying the TS mode may be specified as a picture parameter set (PPS) level syntax log2_transform_skip_max_size_minus2, and may be in the range of 0 to 3. If not present, the value of log2_transform_skip_max_size_minus2 is inferred to be 0. The maximum value MaxTsSize of the maximum block width or height that limits the TS mode may be determined based on Equation (1). MaxTsSize = 1 << ( log2_transform_skip_max_size_minus2 + 2 ) Equation (1)

[0070]

[0093] That is, when log2_transform_skip_max_size_minus2 is 0, TS mode may be allowed if the TB width and height are at most 4. In the current design of VVC6, the maximum allowed value of log2_transform_skip_max_size_minus2 is 3, so the maximum allowed value of MaxTsSize is 32. If the TB width and height are at most MaxTsSize, a parameter transform_skip_flag may be signaled specifying whether TS mode is selected. If the TB width or height is greater than 32, TS mode is not allowed for that TB.

[0071]

[0094] In VVC6, the residual levels of TS mode are coded using non-overlapping coefficient groups (CGs) of size 4 × 4. The transform skip coefficient levels of the CGs are coded in three passes across multiple scan positions.

[0072]

[0095] The first pass can be represented by the following pseudocode: for(n = 0; n <= numSbCoeff - 1; n++ ) if (remainingCtxBin > 0), decode sig_coeff_flag (context) else, bypass decoding of sig_coeff_flag (bypass) if (remainingCtxBin > 0), decode coeff_sign_flag(context) else, bypass decoding of coeff_sign_flag (bypass) if (remainingCtxBin > 0), decode abs_level_gtx_flag[0](context) else, bypass decoding of abs_level_gtx_flag[0] (bypass) if (remainingCtxBin > 0), decode par_level_flag(context) else, bypass decoding of par_level_flag (bypass)

[0073]

[0096] The second pass may be represented by the following pseudocode: for(n = 0; n <= numSbCoeff - 1; n++ ) if (remainingCtxBin > 0), decode abs_level_gtx_flag[1](context) else, bypass decoding of abs_level_gtx_flag[1] (bypass) if (remainingCtxBin > 0), decode abs_level_gtx_flag[2](context) else, bypass decoding of abs_level_gtx_flag[2] (bypass) if (remainingCtxBin > 0), decode abs_level_gtx_flag[3](context) if (remainingCtxBin > 0), decode abs_level_gtx_flag[4](context) else, bypass decoding of abs_level_gtx_flag[4] (bypass)

[0074]

[0097] The third pass may be represented by the following pseudocode: for(n = 0; n <= numSbCoeff - 1; n++ ) rice = cctx.templateAbsSumTS(n, coeff); Decode abs_remainder_using_RG_Coding

[0075]

[0098] In the above description, syntax elements in TS mode for residual coding (referred to as “TS residual coding”) may be coded using either context coding (denoted as “context”) or bypass coding (denoted as “bypass”).

[0076]

[0099] In some embodiments, for TS residual coding, a coding tool called "level mapping" can be adopted. The absolute coefficient level parameter absCoeffLevel can be mapped to a modification level to be coded according to the values ​​of the quantized residual samples to the left and above the current residual sample. Let X0 denote the absolute coefficient level to the left of the current coefficient, and X1 denote the absolute coefficient level above the current coefficient. To represent the coefficients with the absolute coefficient level ("absCoeff"), the mapped parameter absCoeffMod can be coded. absCoeffMod can be derived in the manner represented by the following pseudocode: pred = max(X0, X1); if (absCoeff == pred) { absCoeffMod = 1; } else { absCoeffMod = (absCoeff < pred) ? absCoeff + 1 : absCoeff; }

[0077]

[0100] There are some problems with the current design of TS mode. In VVC6, TS mode is a coding tool that can achieve mathematical lossless compression for a block under both conditions that an appropriate quantization parameter value is selected and the loop filter stage is turned off. Because VVC6 does not allow TS mode for a TB with a width or height greater than 32, the current design of VVC6 cannot achieve mathematical lossless compression for the block when the TB width or height is greater than 32.

[0078]

[0101] In addition, the newly adopted level mapping process has a large impact on the throughput of Context Adaptive Binary Arithmetic Coding (CABAC), because the decoder needs to calculate prediction values ​​from above and from the left for each coefficient level. Because the derivation process of Rice parameters depends on the actual level, the calculation of the actual level with inverse mapping needs to be performed in the CABAC parsing loop. Such an interleaved manner of parsing and level decoding is undesirable because it may reduce the throughput of the decoder hardware implementation.

[0079]

[0102] In VVC6, in addition to log2_transform_skip_max_size_minus2 as above, another sequence parameter set (SPS) level flag, sps_max_luma_transform_size_64_flag, can specify the maximum TB size in luma samples. If sps_max_luma_transform_size_64_flag is equal to 1, the maximum TB size in luma samples is equal to 64. If sps_max_luma_transform_size_64_flag is equal to 0, the maximum TB size in luma samples is equal to 32. If the luma coding tree block size of a coding tree unit (CTU) is less than 64, the value of sps_max_luma_transform_size_64_flag is equal to 0. Based on sps_max_luma_transform_size_64_flag, the parameters MaxTbLog2SizeY and the maximum TB size MaxTbSizeY can be derived based on Equations (2) and (3). MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag ? 6 : 5 Formula (2) MaxTbSizeY = 1 << MaxTbLog2SizeY Equation (3)

[0080]

[0103] Based on equations (2)-(3), the maximum value of the PPS level syntax log2_transform_skip_max_size_minus2 may depend on the SPS level flag sps_max_luma_transform_size_64_flag. log2_transform_skip_max_size_minus2 specifies the maximum block size used for TS mode, and its value may be in the range of 0 to (3+sps_max_luma_transform_size_64_flag). The encoder may be configured to ensure that the value of log2_transform_skip_max_size_minus2 is within the allowed range. If not present, the value of log2_transform_skip_max_size_minus2 may be inferred to be 0. The maximum allowed MaxTsSize may be determined using equation (1). If the width and height of the TB are less than MaxTsSize, then TS mode may be allowed to encode the TB.

[0081]

[0104] As can be seen from the above description, in VVC6, log2_transform_skip_max_size_minus2 is signaled only when sps_transform_skip_enabled_flag is 1. sps_transform_skip_enabled_flag equal to 0 represents the absence of transform_skip_flag in the transform unit syntax. Therefore, when sps_transform_skip_enabled_flag is 0, it is not required to signal log2_transform_skip_max_size_minus2. This current signaling in VVC6 has a parse dependency problem between SPS and PPS. The above embodiment also has the same problem of parse dependency between PPS syntax log2_transform_skip_max_size_minus2 and SPS syntax sps_max_luma_transform_size_64_flag. Such parse dependency is generally undesirable.

[0082]

[0105] The embodiments of the present disclosure provide technical solutions to the above technical problems. To achieve lossless compression using TS mode for large TB, the present disclosure provides an embodiment in which TS mode can be extended to apply to TB sizes up to the maximum TB size allowed for the encoded video sequence. Different coefficient scanning methods are also provided for TS residual coding.

[0083]

[0106] Consistent with some embodiments of the present disclosure, log2_transform_skip_max_size_minus2 may be moved from the PPS to the SPS to remove the parse dependency between the SPS and the PPS. As an example, FIG. 5 shows Table 1 illustrating an example syntax structure of a sequence parameter set (SPS) according to some embodiments of the present disclosure. FIG. 6 shows Table 2 illustrating an example syntax structure of a picture parameter set (SPS) according to some embodiments of the present disclosure. Tables 1 and 2 show that log2_transform_skip_max_size_minus2 is moved from the PPS to the SPS, as shown by row 502 of Table 1 and rows 602-604 of Table 2.

[0084]

[0107] Consistent with some embodiments of the present disclosure, the maximum block size for applying TS mode blocks may be set as the maximum TB size (MaxTbSizeY), in which case log2_transform_skip_max_size_minus2 is not signaled. By doing so, TS mode may be allowed if the TB width and height are less than or equal to MaxTbSizeY. In some embodiments, MaxTbSizeY may be determined based on Equations (2)-(3).

[0085]

[0108] As an example, FIG. 7 shows Table 3 illustrating an example syntax structure of a transform unit according to some embodiments of the present disclosure. Table 3 shows that according to the example syntax structure of a transform unit, the width and height of a TB can be less than or equal to a maximum value MaxTbSizeY (i.e., 32), as shown by row 706. By doing so, since the maximum block size for applying the TS mode is the same as MaxTbSizeY, the TS mode can be allowed for all TBs, and no further check is needed to determine whether the width and height of the TB are less than or equal to MaxTbSizeY, as shown in rows 702-704. It should be noted that VVC6 also uses a Multiple Transform Selection (MTS) scheme for residual coding both inter-coded and intra-coded blocks. MTS uses multiple selection transforms from DCT8 / DST7. However, since MTS is allowed when both tbWidth and tbHeight are less than or equal to 32, further checks are needed during MTS coding.

[0086]

[0109] VVC6 offers another coding tool called Block Differential Pulse Code Modulation (BDPCM). In BDPCM mode, horizontal and vertical Differential Pulse Code Modulation (DPCM) is applied in the residual domain and the conversion stage is skipped. The maximum allowed block width or height for applying BDPCM mode is the same as that of TS mode.

[0087]

[0110] Consistent with some embodiments of the present disclosure, the maximum block size for applying the BDPCM mode may also be extended to be the maximum block size for applying the TS mode, such that the BDPCM mode may be allowed when the width and height of a coding unit (CU) are less than or equal to MaxTbSizeY. As an example, FIG. 8 shows Table 4 illustrating an example syntax structure related to signaling of a block differential pulse code modulation (BDPCM) mode, in accordance with some embodiments of the present disclosure. Table 4 shows that the maximum block size for applying the BDPCM mode may be extended to be the maximum block size for applying the TS mode, as shown by row 802.

[0088]

[0111] In some cases, the allowed values ​​of log2_transform_skip_max_size_minus2 may depend on the profile of the codec. For example, the main profile may specify that the value of log2_transform_skip_max_size_minus2 may be equal to the maximum TB size. Any bitstream signaling a log2_transform_skip_max_size_minus2 value that is not equal to the maximum TB size may be considered a non-compliant bitstream by the codec. For extended profiles beyond the main profile, the value of log2_transform_skip_max_size_minus2 may differ from the maximum TB size.

[0089]

[0112] Consistent with some embodiments of the present disclosure, methods and syntax structures are provided herein to ensure that the value of log2_transform_skip_max_size_minus2 is always equal to the maximum TB size, such as by not signaling log2_transform_skip_max_size_minus2 and inferring that it is equal to the maximum TB size, or by profile constraint configuration, etc. Doing so can reduce the burden on the decoder implementation, since there are fewer syntax element value combinations to test.

[0090]

[0113] In some embodiments, the SPS flag may be signaled to indicate that the maximum block size for applying the TS mode is 32 or 64. For example, the SPS flag may be signaled in the same manner as the maximum TB size is signaled. As an example, to specify that the maximum block size for applying the TS mode is 32, sps_max_transform_skip_size_64_flag may be set to 0. In another example, to specify that the maximum block size for applying the TS mode is 64, sps_max_transform_skip_size_64_flag may be set to 1. In some embodiments, if sps_max_transform_skip_size_64_flag is not signaled, its value may be inferred to be 0.

[0091]

[0114] In some embodiments, the maximum block size for applying the TS mode may be determined based on equation (4). MaxTsSize = sps_max_transform_skip_size_64_flag ? 64 : 32 Formula (4)

[0092]

[0115] In some embodiments, sps_max_transform_skip_size_64_flag may be signaled if sps_max_luma_transform_size_64_flag and sps_transform_skip_enabled_flag are both equal to 1.

[0093]

[0116] As an example, Figure 9 shows Table 5 illustrating an example syntax structure of an SPS for signaling sps_max_transform_skip_size_64_flag according to some embodiments of the present disclosure. Figure 10 shows Table 6 illustrating an example syntax structure of a transform unit for signaling sps_max_transform_skip_size_64_flag according to some embodiments of the present disclosure. Tables 5 and 6 show an implementation of signaling sps_max_transform_skip_size_64_flag as shown by row 902 of Table 5 and rows 1002-1006 of Table 6.

[0094]

[0117] Consistent with some embodiments of the present disclosure, since the maximum block size for applying TS mode or BDPCM mode may be extended to be the maximum TB size, the residual coding in TS mode or BDPCM mode may also be extended to allow encoding the maximum TB size in that respect. According to some disclosed embodiments, the residual coding may be directly extended to allow up to the maximum TB size without changing the scanning pattern.

[0095]

[0118] In some embodiments, similar to VVC Draft 6, a transform block can be divided into coefficient groups (CGs) and diagonal scanning can be performed. As an example, FIG. 11 is a schematic diagram illustrating a diagonal scanning example of a 64×64 transform block (TB) according to some embodiments of the present disclosure. FIG. 11 illustrates a diagonal scanning pattern (indicated by zigzag arrow lines) of a 64×64 TB (e.g., MaxTbSizeY=64). Each cell in FIG. 11 may represent a 4×4 CG. It should be noted that while FIG. 11 illustrates a 64×64 TB to illustrate the diagonal scanning process, the TB may be of any size or shape and is not limited to the example as shown herein. For example, if the TB is rectangular instead of square, only one of its dimensions is equal to 64.

[0096]

[0119] One challenge of scanning an entire TB (e.g., the 64×64 TB in FIG. 11) in residual coding is that the current residual coding in VVC only supports up to 32×32 block sizes, so the current VVC residual coding needs to be modified to support the above extension. In the current VVC design, even if a transform is applied to a 64×64 TB (e.g., in non-skip mode), the decoder may still need to apply residual coding only to a 32×32 block of coefficients that represents the top-left 32×32 block of the 64×64 TB. In such a case, all remaining high frequency coefficients are forced to zero (thus no coding of the remaining coefficients is required). For example, for an M×N TB (M is the block width and N is the block height), when M is equal to 64, only the left 32 columns of transform coefficients may be coded. Similarly, when N is equal to 64, only the top 32 rows of transform coefficients may be coded.

[0097]

[0120] Consistent with some embodiments of the present disclosure, to reuse existing VVC6 residual coding techniques, a large TB can be divided into smaller residual units (RUs). For example, if the width of the TB is greater than 32, the TB can be divided into two partitions horizontally. As another example, if the height of the TB is greater than 32, the TB can be divided into two partitions vertically. As yet another example, if both dimensions of the TB are greater than 32, the TB can be divided into four RUs horizontally and vertically. After division, 32×32 RUs can be encoded.

[0098]

[0121] As an example, Figures 12A-12D show example residual units (RUs) according to some embodiments of the present disclosure. In Figure 12A, a 64x64 TB is split into four 32x32 RUs (shown in dashed lines). In Figure 12B, a 64x16 TB is split horizontally into two 32x16 RUs (shown in dashed lines). In Figure 12C, a 32x64 TB is split vertically into two 32x32 RUs (shown in dashed lines). In Figure 12D, no splitting is done since neither the height nor the width exceeds 32, and the RU size is the same as the TB size. In some embodiments, the maximum allowed RU size is 32x32.

[0099]

[0122] As an example, FIG. 13 is a schematic diagram illustrating an example of diagonal scanning of a 64×64 TB, where the TB is divided into four 32×32 RUs, according to some embodiments of the present disclosure. In FIG. 13, the 64×64 TB is divided into four RUs (indicated by the thick solid lines in the TB), and the coefficients of each RU are scanned individually (e.g., independently) in the RU, following the same order as the scanning pattern for the 32×32 TB. In FIG. 13, the context model and Rice parameter derivation of one RU may be independent of another RU. In some embodiments, the maximum number of context coded bins may also be assigned independently for each RU. Such a scheme is different for VVC6, where the maximum number of context coded bins is defined at the TB level.

[0100]

[0123] As an example, FIGS. 14A-14D show Table 7 illustrating an example syntax structure for residual coding when a TB is split into RUs, according to some embodiments of the present disclosure.

[0101]

[0124] In VVC6, coded_sub_block_flag is signaled for each coefficient group (CG) of a TS mode block. coded_sub_block_flag=0 means that all of the coefficients of the CG are zero. coded_sub_block_flag=1 means that at least one coefficient in the CG is non-zero. However, if all coded_sub_block_flag of previously coded CGs (i.e. before the last CG) are zero, the coded_sub_block_flag of the last CG is not signaled and is inferred to be 1. This means that the parsing of the last CG of a TB depends on all previously decoded CGs. To remove dependencies between RUs, coded_sub_block_flag may be signaled for all of the CGs of the RU including the last CG.

[0102]

[0125] Consistent with some embodiments of the present disclosure, a further syntax coded_RU_flag may be introduced. In some embodiments, coded_RU_flag may be signaled if the number of RUs in the TB is greater than 1. In some embodiments, if coded_RU_flag is not present, it may be inferred to be 1. coded_RU_flag=0 may specify that all of the coefficients of the RU are zero. coded_RU_flag=1 may specify that at least one of the coefficients of the RU is non-zero. In some embodiments, if all coded_RU_flag except the last RU are zero, the coded_RU_flag of the last RU does not need to be signaled and may be inferred to be 1. As an example, the following pseudocode shows an example of signaling coded_RU_flag: inferRUCbf = 1; for( k =0; k < numofRUs; k++ ) { if( (k != lastRU | | !inferRUCbf ) signal coded_RU_flag; if( coded_RU_flag) inferRUCbf = 0; }

[0103]

[0126] 15A-15D show Table 8 illustrating another example syntax structure for residual coding when coded_RU_flag is signaled according to some embodiments of the present disclosure. In some embodiments, when coded_RU_flag is signaled, the last CG flag may be maintained in the same manner as in VVC6, i.e., if coded_sub_block_flag of all previous CGs in the same RU is zero, coded_sub_block_flag is not signaled and can be inferred to be 1.

[0104]

[0127] JVET (Joint Video Experts Team) AHG Lossless and Near-Lossless Encoding Tools (AHG18) releases lossless software based on VTM-6.0. The lossless software introduces a CU level flag called cu_transquant_bypass_flag. cu_transquant_bypass_flag=1 means that transform and quantization for that CU is skipped and the CU is coded in lossless mode. In the current version of the lossless software, sps_max_luma_transform_size_64_flag is set to 0, which means that the maximum TB size for luma samples is limited to 32x32. For chroma samples, the maximum TB size is adjusted based on the YUV color format (e.g., up to 16x16 for YUV420). In some embodiments, the luma transform block size may be increased up to 64x64 when cu_transquant_bypass_flag=1, and the residual coding techniques described above may be used when cu_transquant_bypass_flag=1.

[0105]

[0128] In some embodiments, the maximum TB size for a chroma component may be determined using equations (2) and (3). Based on equations (2) and (3), the maximum TB width, maxTbWidth, and the maximum TB height, maxTbHeight, may be determined based on equations (5) and (6). maxTbWidth = ( cIdx == 0 ) ? MaxTbSizeY : MaxTbSizeY / SubWidthC Equation (5) maxTbHeight = ( cIdx == 0 ) ? MaxTbSizeY : MaxTbSizeY / SubHeightC Formula (6)

[0106]

[0129] In Equation (5) and Equation (6), cIdx=0 means the luma component. cIdx=1 and cIdx=2 mean two chroma components. As an example, values ​​of SubWidthC and SubHeightC can be derived from a chroma format. Consistent with some embodiments of the present disclosure, FIG. 16 shows Table 9 illustrating example parameter values ​​derived from a chroma format according to some embodiments of the present disclosure.

[0107]

[0130] In VVC6, the inverse level mapping is embedded in the CABAC module. Figure 17 shows Table 10 illustrating an example syntax structure in VVC6 for residual coding with inverse level mapping according to some embodiments of the present disclosure.

[0108]

[0131] Consistent with some embodiments of the present disclosure, to improve the CABAC throughput of transform skip residual parsing, the Rice parameters may be derived based on the mapped level values ​​instead of based on the actual level values. In some embodiments, both the context model and the Rice parameters may depend on the mapped values, and no inverse mapping operation may be performed during the residual parsing process. By doing so, the inverse mapping may be decoupled from the residual parsing process. The inverse mapping may be performed after the completion of the parsing of the entire TB of the residual. In some embodiments, the inverse mapping and the residual parsing may be performed simultaneously in one pass, which allows the actual implementation to decide whether to interleave the parsing and mapping or to split them into two passes.

[0109]

[0132] As an example, Figure 18 is a flowchart of an example decoding method 1400 according to some embodiments of the present disclosure. The method 1800 may be performed when parsing and inverse mapping are separated. Figure 18 illustrates that inverse mapping is decoupled from residual parsing by being performed after completion of parsing of the residual across the TB and before inverse quantization.

[0110]

[0133] Consistent with some embodiments of the present disclosure, Figure 19 illustrates Table 10, which illustrates an example syntax structure for residual coding where inverse level mapping is not performed, in accordance with some embodiments of the present disclosure. In some embodiments, the inverse level mapping can be moved to the decoding process, which is described below.

[0111]

[0134] Consistent with some embodiments of the present disclosure, the following pseudocode illustrates an inverse level mapping process that can occur after residual parsing and before inverse quantization (as shown in FIG. 18): In the following pseudocode, TransCoeffLevel[xC][yC] represents the coefficient value at the (xC, yC) location after residual parsing, and TransCoeffLevelInvMapped[xC][yC] represents the coefficient value at the (xC, yC) location after inverse mapping. for (int yC = 0; yC < height; yC++) { for (int xC = 0; xC < width; xC++) { TransCoeffLevelInvMapped [xC][yC] = TransCoeffLevel [xC][yC]; if (TransCoeffLevel [xC][yC]) { topPos = abs (TransCoeffLevel [xC][yC-1]); leftPos = abs(TransCoeffLevel [xC - 1][yC]); if (topPos || leftPos) { int absMappedLevel = abs(TransCoeffLevel [xC][yC]); int sign = TransCoeffLevel [xC][yC] < 0; int pred1 = std::max(topPos, leftPos); if (absMappedLevel == 1) TransCoeffLevelInvMapped [xC][yC]= pred1; else TransCoeffLevelInvMapped [xC][yC] = absMappedLevel - (absMappedLevel <= pred1); TransCoeffLevelInvMapped [xC][yC] = sign ? -dst[xC][yC] : dst[xC][yC]; } } } }

[0112]

[0135] Consistent with some embodiments of the present disclosure, the Rice parameters may be derived based on the mapped values, which differs from VVC6 in that the Rice parameters are derived based on the actual level values. Assuming that the array TransCoeffLevel[xC][yC] is the mapped level value for the TB of a given color component at location (xC, yC), the variable locSumAbs may be derived based on the following pseudocode: locSumAbs = 0 AbsLevel [xC][yC] = abs(TransCoeffLevel[xC][yC]) if( xC > 0 ) locSumAbs += AbsLevel[ xC - 1 ][ yC ] if( yC > 0 ) locSumAbs += AbsLevel[ xC ][ yC - 1 ] locSumAbs = Clip3( 0, 31, locSumAbs )

[0113]

[0136] Consistent with some embodiments of the present disclosure, FIG. 20 illustrates Table 12, which shows an example lookup table for selecting Rice parameters, according to some embodiments of the present disclosure. In some disclosed embodiments, the value of locSumAbs can be adjusted based on a predefined offset value. In some embodiments, the offset value is calculated from offline training. The following pseudocode example shows that the offset value is 2: locSumAbs = 0 offset = 2; AbsLevel [xC][yC] = abs(TransCoeffLevel[xC][yC]) if( xC > 0 ) locSumAbs += AbsLevel[ xC - 1 ][ yC ] if( yC > 0 ) locSumAbs += AbsLevel[ xC ][ yC - 1 ] locSumAbs -= offset locSumAbs = Clip3( 0, 31, locSumAbs )

[0114]

[0137] Consistent with some embodiments of the present disclosure, Figures 21-22 show flowcharts of example processes 2100-2200 for video processing, according to some embodiments of the present disclosure. In some embodiments, the processes 2100-2200 may be performed by a codec (e.g., the encoder of Figures 2A-2B or the decoder of Figures 3A-3B). For example, the codec may be implemented as one or more software or hardware components of an apparatus for video processing (e.g., apparatus 400).

[0115]

[0138] As an example, FIG. 21 shows a flowchart of an example process 2100 for video processing according to some embodiments of the present disclosure. In step 2102, a codec (e.g., the encoder of FIGS. 2A-2B) may determine to skip a transform process for a prediction residual based on one of a maximum dimension of luma samples of a prediction block or a maximum dimension of a prediction block. The transform process may be the transform stage 212 of FIGS. 2A-2B. The prediction residual may be the residual BPU 210 of FIGS. 2A-2B. The transform block may be a block included in the prediction data 206 of FIGS. 2A-2B, such as a transform block (e.g., any of the transform blocks shown in FIGS. 11-13). The dimensions of the prediction block may include a height or a width.

[0116]

[0139] In some embodiments, the codec may decide to skip the transform process for the prediction residual by deciding to skip the transform process based on a determination that the dimension of the prediction block does not exceed a threshold. In some embodiments, the threshold may be MaxTbSizeY as shown and described in relation to equations (2)-(3). The threshold may have a maximum value equal to one of the maximum value of the dimension of the luma samples (e.g., 32, 64, or any number) or the maximum value of the dimension of the prediction block (e.g., 32, 64, or any number). In some embodiments, the maximum value of the dimension of the luma samples or the maximum value of the dimension of the prediction block may be a dynamic value (e.g., not constant).

[0117]

[0140] In some embodiments, the threshold is equal to the maximum value of a dimension of a luma sample that indicates luma information of the prediction block. In some embodiments, the maximum value of the threshold is 64. In some embodiments, the maximum value of the threshold is 32. In some embodiments, the minimum value of the threshold is 4. In some embodiments, the threshold may be equal to the maximum value of a dimension of a prediction block that is allowed to undergo the transformation process (e.g., MaxTsSize as shown and described in equation (1)).

[0118]

[0141] In some embodiments, the maximum value of the threshold is determined based on at least a first parameter of a first parameter set. For example, the first parameter set may be a sequence parameter set (SPS). In some embodiments, the value of the first parameter is 0 or 1. For example, the first parameter may be sps_max_luma_transform_size_64_flag as shown and described in connection with Table 5 of FIG. 9. In some embodiments, the threshold can be determined based on the value of the first parameter. For example, if the first parameter can be sps_max_luma_transform_size_64_flag and the threshold is MaxTbSizeY, when sps_max_luma_transform_size_64_flag is equal to 1, MaxTbSizeY may be equal to 64. When sps_max_luma_transform_size_64_flag is equal to 0, MaxTbSizeY is equal to 32.

[0119]

[0142] In some embodiments, the maximum value of the threshold may be determined based on at least a first parameter of a first parameter set. In some embodiments, the threshold may be determined based on a value of a second parameter of a second parameter set. In some embodiments, the second parameter set is a sequence parameter set (SPS). In some embodiments, the second parameter set is a picture parameter set (PPS). The second parameter may be log2_transform_skip_max_size_minus2 (e.g., as shown and described in connection with Table 1 of FIG. 5). The value of the second parameter may be determined based on the value of the first parameter. In some embodiments, the value of the second parameter (e.g., log2_transform_skip_max_size_minus2) has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter (e.g., sps_max_luma_transform_size_64_flag). For example, log2_transform_skip_max_size_minus2 may be in the range of 0 to (3+sps_max_luma_transform_size_64_flag). In some embodiments, the second parameter may have a first value in a first profile (e.g., a main profile) of the encoder and a second value in a second profile (e.g., an extended profile) of the encoder, where the first value and the second value are different.

[0120]

[0143] With further reference to FIG. 21, in step 2104, the codec may generate residual coefficients for the prediction residual by performing at least one of a lossless compression process or a quantization process on the prediction residual. The residual coefficients may be coefficients associated with a residual coding process as described herein. The quantization process may be the quantization stage 214 of FIGS. 2A-2B. The lossless compression process may include generating the residual coefficients using coefficient groups (CG). For example, the coefficient groups may be non-overlapping. In some embodiments, the coefficient groups have a size of 4×4.

[0121]

[0144] In some embodiments, the codec can generate the residual coefficients using a multiple transform selection (MTS) scheme. For example, the codec can determine whether the dimension of the prediction block does not exceed 32. If the dimension of the prediction block does not exceed 32, the codec can generate the residual coefficients using the MTS scheme.

[0122]

[0145] In some embodiments, the codec may further determine a transform skip coefficient level for the coefficient group using one of a context coding technique or a bypass coding technique. The codec may also determine a Rice parameter based on the transform skip coefficient level. The codec may further generate the bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter.

[0123]

[0146] In some embodiments, the codec may further map the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first predictive block to the left of the predictive block and a second value of a second residual coefficient of a second predictive block above the predictive block.

[0124]

[0147] In some embodiments, the codec can determine a transform skip coefficient level for a coefficient group using one of a context coding technique or a bypass coding technique, map the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first predictive block to the left of the predictive block and a second value of a second residual coefficient of a second predictive block above the predictive block, generate a context model for the context coding technique based on the modified transform skip coefficient level, determine a Rice parameter based on the modified transform skip coefficient level, generate residual coefficients using the coefficient group, and generate a bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter.

[0125]

[0148] 21, in step 2106, the codec may generate a bitstream by entropy encoding at least the residual coefficients. The bitstream may be the video bitstream 228 of FIGS. 2A-2B.

[0126]

[0149] 22 shows a flowchart of another example process 2200 for video processing according to some embodiments of the present disclosure. For example, the process 2200 may be performed by the decoder of FIGS.

[0127]

[0150] As shown in Figure 22, in step 2202, a decoder receives a bitstream containing coding information for a video sequence. The bitstream includes a sequence parameter set (SPS) for the video sequence.

[0128]

[0151] In step 2204, the decoder determines a maximum transform size of the prediction block based on parameters of a sequence parameter set (SPS) of the video sequence. The prediction block may be a block included in the prediction data 206 of FIGS. 2A-2B, such as a transform block (e.g., any of the transform blocks shown in FIGS. 11-13). In some embodiments, the maximum transform size may correspond to a maximum of dimensions of luma samples of the prediction block or a maximum of dimensions of the prediction block. The dimensions of the prediction block may include a height or a width. Detailed methods for determining the maximum transform size based on parameters of the SPS are described above in connection with FIGS. 5-10.

[0129]

[0152] In step 2206, the decoder determines to skip the transform process for the prediction residual of the prediction block based on the maximum transform size. The transform process may be the transform stage 212 of FIGS.

[0130]

[0153] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes instructions, which can be executed by a device (such as the disclosed encoder and decoder) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes or other magnetic data storage media, CD-ROMs, other optical data storage media, any physical media with a pattern of holes, RAM, PROMs, and EPROMs, FLASH-EPROMs or other flash memories, NVRAMs, caches, registers, other memory chips or cartridges, and networked versions of the above. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0131]

[0154] The embodiments can be further described using the following clauses. 1. determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling a maximum transform size in a sequence parameter set (SPS); A video processing method comprising: 2. Deciding to skip the transformation process on the prediction residuals and determining to skip the transformation process based on a determination that the dimension of the prediction block does not exceed a threshold, the threshold being: The maximum luma sample size of the prediction block, or Maximum prediction block size 2. The method of claim 1, wherein the maximum value is equal to one of 3. The method according to clause 2, wherein one of the maximum luma sample dimension or the maximum prediction block dimension is a dynamic value. 4. The method of any one of the preceding clauses, further comprising determining to skip the conversion process further based on a parameter indicating a conversion skip mode. 5. The method of claim 2, wherein the dimensions of the prediction block include a height or a width. 6. The method of claim 2, wherein the maximum value of the threshold is determined based on at least a first parameter of the first parameter set. 7. The method of claim 6, wherein the first parameter set is a sequence parameter set (SPS). 8. The method according to any one of clauses 6 to 7, wherein the value of the first parameter is 0 or 1. 9. The method according to any one of clauses 2 to 8, wherein the maximum value of the threshold is 64. 10. The method of any one of clauses 2 to 8, wherein the maximum value of the threshold is 32. 11. The method of any one of clauses 2 to 10, wherein the maximum value of the threshold is determined based on at least a first parameter of the first parameter set and a third parameter of the first parameter set. 12. The method of any one of clauses 2 to 11, wherein the minimum value of the threshold is 4. 13. The method according to any one of clauses 2 to 12, wherein the threshold is equal to the maximum value of the dimension of the luma samples indicating luminance information of the prediction block. 14. A method according to any one of clauses 6 to 13, wherein the maximum value of the threshold is determined based on the value of a second parameter of a second parameter set, the value of the second parameter being determined based on the value of the first parameter. 15. The method of claim 14, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter. 16. The method of claim 14, wherein the second parameter has a first value in a first profile of the encoder and a second value in a second profile of the encoder, the first value and the second value being different. 17. The method of any one of clauses 14 to 16, wherein the second parameter set is an SPS. 18. The method according to any one of clauses 14 to 16, wherein the second parameter set is a picture parameter set (PPS). 19. The method according to any one of clauses 2 to 12, wherein the threshold value is equal to the maximum value of the dimension of the prediction block that is allowed to undergo the transformation process. 20. The method of claim 19, wherein the threshold is determined based on a value of the first parameter. 21. The method of any one of the preceding clauses, further comprising generating residual coefficients for the prediction block using a multiple transform selection (MTS) scheme. 22. Determining whether a dimension of the prediction block does not exceed 32; generating residual coefficients using an MTS scheme based on a determination that a dimension of the prediction block does not exceed 32; 22. The method of claim 21, further comprising: 23. Determining whether a dimension of the prediction block does not exceed a threshold; based on a determination that a dimension of the prediction block does not exceed a threshold, performing block differential pulse code modulation (BDPCM) on the prediction residual prior to generating residual coefficients for the prediction block; 23. The method of any one of clauses 2 to 22, further comprising: 24. A method according to any one of the preceding clauses, further comprising generating residual coefficients for the prediction residual by performing a lossless compression process on the prediction residual, the lossless compression process comprising generating the residual coefficients using groups of coefficients, the groups of coefficients being non-overlapping. 25. The method of claim 24, wherein the coefficient groups have a size of 4x4. 26. Determining a transform skip coefficient level for a group of coefficients using one of a context coding technique or a bypass coding technique; determining a Rice parameter based on a transform skip coefficient level; generating a bitstream by entropy encoding at least one of a coefficient group, a transform skip coefficient level, or a Rice parameter; 26. The method according to any one of clauses 24 to 25, further comprising: 27. The method of any one of clauses 24 to 26, further comprising mapping a transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first predictive block to the left of the predictive block and a second value of a second residual coefficient of a second predictive block above the predictive block. 28. Determining a transform skip coefficient level for a group of coefficients using one of a context coding technique or a bypass coding technique; mapping a transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first predictive block to the left of the predictive block and a second value of a second residual coefficient of a second predictive block above the predictive block; generating a context model for a context encoding technique based on the modified transform skip coefficient level; determining a Rice parameter based on the modified transform skip factor level; generating residual coefficients using the coefficient groups; generating a bitstream by entropy encoding at least one of a coefficient group, a transform skip coefficient level, or a Rice parameter; 27. The method of any one of clauses 24 to 26, further comprising: 29. The method of claim 28, further comprising mapping transform skip coefficient levels to modified transform skip coefficient levels after the quantization process and during generation of the residual coefficients. 30. The method of claim 28, further comprising mapping the transform skip coefficient levels to modified transform skip coefficient levels after the quantization process and before generating the residual coefficients. 31. Determining the Rice parameters 31. The method of any one of clauses 28 to 30, comprising determining a Rice parameter based on modified transform skip factor levels of color components of the prediction block. 32. The method of claim 31, wherein the modified transform skip factor levels of the color components are offset by a predetermined offset value. 33. The method of claim 32, wherein the predetermined offset value is determined using a machine learning model in an offline training process. 34. Generating residual coefficients 34. The method of any one of clauses 23 to 33, comprising performing at least one of a lossless compression process or BDPCM on the prediction residual using diagonal scanning, wherein the maximum size of the prediction block for performing diagonal scanning is 64. 35. Generating residual coefficients based on a determination that a dimension of the prediction block exceeds 32, dividing the prediction block into a plurality of sub-blocks at the dimension; For each particular sub-block of the plurality of sub-blocks, performing at least one of a lossless compression process or a BDPCM on a prediction residual associated with the particular sub-block using diagonal scanning, wherein parameters and output results of each of the lossless compression processes or BDPCM associated with the plurality of sub-blocks are independent; 34. The method according to any one of clauses 23 to 33, comprising: 36. The method of clause 35, further comprising: dividing the prediction block into a plurality of sub-blocks in the two dimensions based on a determination that the two dimensions of the prediction block exceed 32. 37. A method according to any one of clauses 35 to 36, wherein the parameters and output results of each of the lossless compression processes or BDPCM associated with the plurality of sub-blocks include at least one of a context model associated with the context encoding technique, a Rice parameter, or a maximum number of context encoding bins associated with the context encoding technique. 38. A method according to any one of clauses 34 to 37, wherein the unit of diagonal scanning is a coefficient group. 39. The method of clause 38, further comprising setting, for each coefficient group of a particular sub-block, a first indicator parameter indicative of a value of the coefficient of the coefficient group. 40. The method of clause 38, further comprising setting, for each particular sub-block of the plurality of sub-blocks, a second indicator parameter indicative of values ​​of all coefficient groups of the particular sub-block. 41. For each coefficient group of a particular sub-block, setting a first indicator parameter indicating a value of a coefficient of the coefficient group; setting the first indicator parameter of the last coefficient group to one based on a determination that the first indicator parameters of all coefficient groups prior to the last coefficient group of the particular sub-block are zero; 41. The method of claim 40, further comprising: 42. Generating residual coefficients A method according to any one of clauses 24 to 41, comprising generating residual coefficients by performing a lossless compression process on the prediction residual based on a parameter indicating a lossless encoding mode, wherein the maximum value of the luma sample dimension is 64. 43. Receiving a video picture; Dividing a video picture into a plurality of blocks; performing one of intra prediction or inter prediction on the block to generate a predicted block; generating a prediction residual by subtracting the prediction block from the block; 20. The method of any one of the preceding clauses, further comprising: 44. A memory configured to store instructions; determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling a maximum transform size in a sequence parameter set (SPS); a processor configured to execute instructions to 13. An apparatus comprising: 45. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of an apparatus to cause the apparatus to perform a method, the method comprising: determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling a maximum transform size in a sequence parameter set (SPS); A non-transitory computer readable medium comprising: 46. ​​Receiving a bitstream of a video sequence; determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; determining to skip a transform process for a prediction residual of the prediction block based on a maximum transform size; A video processing method comprising: 47. Deciding to skip the transformation process on the prediction residuals determining to skip the transformation process in response to determining that a dimension of the prediction block does not exceed a threshold, the threshold being: The maximum luma sample size of the prediction block, or Maximum prediction block size 47. The method of claim 46, having a maximum value equal to one of 48. The method of clause 47, wherein the dimensions of the prediction block include height or width. 49. The method of claim 47, wherein the maximum value of the threshold is determined based on at least a first parameter of the SPS. 50. The method of claim 49, wherein the value of the first parameter is 0 or 1. 51. The method of any one of clauses 47 to 50, wherein the maximum value of the threshold is 64. 52. The method of any one of clauses 47 to 50, wherein the maximum value of the threshold is 32. 53. A method according to any one of clauses 47 to 52, wherein the maximum value of the threshold is determined based on at least a first parameter of the SPS and a third parameter of the SPS. 54. The method of any one of clauses 47 to 53, wherein the minimum value of the threshold is 4. 55. A method according to any one of clauses 47 to 54, wherein the threshold is equal to the maximum value of the dimension of the luma samples representing luminance information of the prediction block. 56. A method according to any one of clauses 49 to 54, wherein the maximum value of the threshold is determined based on the value of a second parameter of a second parameter set, the value of the second parameter being determined based on the value of the first parameter. 57. The method of claim 56, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter. 58. The method of clause 56, wherein the second parameter has a first value in a first profile of the encoder and a second value in a second profile of the encoder, the first value and the second value being different. 59. The method of any one of clauses 56 to 58, wherein the second parameter set is an SPS. 60. The method of any one of clauses 56 to 58, wherein the second parameter set is a picture parameter set (PPS). 61. A memory configured to store instructions; receiving a bitstream of a video sequence; determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; determining to skip a transform process for a prediction residual of the prediction block based on a maximum transform size; a processor configured to execute instructions to 13. An apparatus comprising: 62. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of an apparatus to cause the apparatus to perform a method, the method comprising: receiving a bitstream of a video sequence; determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence; determining to skip a transform process for a prediction residual of the prediction block based on a maximum transform size; A non-transitory computer readable medium comprising:

[0132]

[0155] It should be noted that relative terms herein, such as "first" and "second," are used only to distinguish one entity or operation from another, and do not require or imply an actual relationship or order between those entities or operations. Also, the words "comprising," "having," "containing," and "including," as well as other similar forms, are intended to be equivalent in meaning and to be open-ended in that the term or terms following any one of these terms is not an exhaustive enumeration of such term or terms, or limited to only the enumerated term or terms.

[0133]

[0156] As used herein, unless specifically stated otherwise, the term "or" encompasses all possible combinations unless impracticable. For example, if a component is described as including A or B, the component may include A, or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if a component is described as including A, B, or C, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C, unless specifically stated otherwise or impracticable.

[0134]

[0157] It is understood that the above embodiments may be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it may be stored in the above computer-readable medium. The software, when executed by a processor, may perform the disclosed method. The computing units and other functional units described in this disclosure may be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that more than one of the above modules / units may be integrated into one module / unit, and each of the above modules / units may be further divided into multiple sub-modules / sub-units.

[0135]

[0158] In the foregoing specification, the embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the above specification and examples be considered as examples only, with the true scope and spirit of the invention being indicated by the following claims. Additionally, the order of steps depicted in the figures is intended to be illustrative only, and is not intended to be limited to any particular order of steps. Thus, one skilled in the art will appreciate that these steps may be performed in different orders while performing the same method.

[0136]

[0159] In the drawings and specification, example embodiments have been disclosed. However, many variations and modifications to these embodiments may be made. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling the maximum transform size in a sequence parameter set (SPS); A video processing method comprising:

2. The method of claim 1 , further comprising: determining to skip the conversion process further based on a parameter indicating a conversion skip mode.

3. determining to skip the transformation process for the prediction residual; determining to skip the conversion process based on a determination that a dimension of the prediction block does not exceed a threshold, the threshold being: the maximum size of the luma samples of the prediction block; or The maximum size of the predicted block The method of claim 1 , wherein the first and second inputs have a maximum value equal to one of the first and second inputs.

4. The method of claim 3 , wherein one of the maximum value of the dimension of the luma samples or the maximum value of the dimension of the predictive block is a dynamic value.

5. The method of claim 3 , wherein the dimensions of the prediction block include a height or a width.

6. The method of claim 3 , wherein the maximum value of the threshold is 64.

7. The method of claim 3 , wherein the maximum value of the threshold is 32.

8. The method of claim 3 , wherein the minimum threshold value is four.

9. The method of claim 3 , wherein the threshold is equal to the maximum value of the dimension of the luma samples indicating luminance information of the prediction block.

10. The method of claim 3 , wherein the maximum value of the threshold is determined based on at least a first parameter of a first parameter set.

11. The method of claim 10 , wherein the first parameter set is a sequence parameter set (SPS).

12. The method of claim 10 , wherein the value of the first parameter is 0 or 1.

13. The method of claim 10 , wherein the maximum value of the threshold is determined based on at least the first parameter of the first parameter set and a third parameter of the first parameter set.

14. 11. The method of claim 10, wherein the maximum of the thresholds is determined based on a value of a second parameter of a second parameter set, the value of the second parameter being determined based on the value of the first parameter.

15. 15. The method of claim 14, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to 3 and the sum of the values ​​of the first parameter.

16. 15. The method of claim 14, wherein the second parameter has a first value in a first profile of an encoder and a second value in a second profile of the encoder, the first value and the second value being different.

17. The method of claim 14 , wherein the second set of parameters is the SPS.

18. The method of claim 14 , wherein the second parameter set is a Picture Parameter Set (PPS).

19. a memory configured to store instructions; a processor, the processor comprising: determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling the maximum transform size in a sequence parameter set (SPS); and configured to execute the instructions to cause the device to Device.

20. 1. A non-transitory computer readable medium storing a set of instructions, the set of instructions being executable by at least one processor of an apparatus to cause the apparatus to perform a method, the method comprising: determining to skip a transform process for a prediction residual based on a maximum transform size of a prediction block; signaling the maximum transform size in a sequence parameter set (SPS); A non-transitory computer readable medium comprising:

Citation Information

Patent Citations

  • Method and apparatus for encoding / decoding image

    US20160269730A1

  • Max transform size control

    US20200288131A1

  • Image coding method and device using transform skip flag

    WO2020149648A1

  • Usage of transquant bypass mode for multiple color components

    WO2020228716A1