Method and apparatus for encoding video data in transform skip mode
By moving the log2_transform_skip_max_size_minus2 parameter to the SPS and setting the maximum block size for the transform skip mode to MaxTbSizeY, the limitations of the VVC standard in compressing large transform blocks are overcome, enabling efficient lossless compression.
Patent Information
- Application Number
- JP2022513264
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-08-13
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2040-08-13
AI Technical Summary
Existing video coding standards face challenges in achieving efficient compression for large transform blocks, particularly in the Versatile Video Coding (VVC) standard, where the transform skip mode is limited to block sizes of 32x32 or smaller, preventing lossless compression for larger blocks.
The proposed solution involves moving the log2_transform_skip_max_size_minus2 parameter from the Picture Parameter Set (PPS) to the Sequence Parameter Set (SPS), allowing the maximum block size for the transform skip mode to be set as the maximum transform block size (MaxTbSizeY), which can be 32 or 64, thereby enabling the transform skip mode for larger blocks.
This approach allows for lossless compression of large transform blocks up to the maximum allowed size, improving coding efficiency and reducing the dependency on parsing between different parameter sets.
Smart Images

Figure 0007691413000001 
Figure 0007691413000002 
Figure 0007691413000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 899,738, filed on September 12, 2019, and U.S. Provisional Patent Application No. 62 / 904,880, filed on September 24, 2019, which are hereby incorporated by reference in their entirety.
Background Art
[0002] Background
[0002] Video is a series of still pictures (or "frames") that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually called encoding, and the restoration process is usually called decoding. Most commonly, there are various video coding formats that use standardized video coding techniques based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standard specifications such as the HEVC (High Efficiency Video Coding) / H.265 standard specification, the VVC (Versatile Video Coding) / H.266 standard specification, and the AVS standard specification, which specify a particular video coding format, have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standard specifications becomes higher and higher.
Summary of the Invention
Means for Solving the Problems
[0003] Summary of the Disclosure
[0003] Embodiments of this disclosure provide a method and an apparatus for video processing. In one example embodiment, the method includes determining to skip a conversion process for a prediction residual based on a maximum transform size of a prediction block, and signaling the maximum transform size in a Sequence Parameter Set (SPS).
[0004]
[0004] In another embodiment, the apparatus includes a memory configured to store instructions and a processor, and the processor is configured to execute instructions to cause the apparatus to determine to skip a transformation process for a prediction residual based on a maximum transformation size of a prediction block and to signal the maximum transformation size in a sequence parameter set (SPS).
[0005]
[0005] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of an apparatus to cause the apparatus to perform a method. The method includes determining to skip a transformation process for a prediction residual based on a maximum transformation size of a prediction block and signaling the maximum transformation size in a sequence parameter set (SPS).
[0006]
[0006] In another example embodiment, a method includes receiving a bitstream of a video sequence, determining a maximum transformation size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determining to skip a transformation process for a prediction residual of the prediction block based on the maximum transformation size.
[0007]
[0007] In another embodiment, the apparatus includes a memory configured to store instructions and a processor, and the processor is configured to execute instructions to cause the apparatus to receive a bitstream of a video sequence, determine a maximum transformation size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determine to skip a transformation process for a prediction residual of the prediction block based on the maximum transformation size.
[0008]
[0008] In another example embodiment, the non-transitory computer-readable medium stores a set of instructions that are executable by at least one processor of the apparatus to cause the apparatus to perform a method. The method includes receiving a bitstream of a video sequence, determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, and determining to skip a transform process for a prediction residual of the prediction block based on the maximum transform size.
[0009] Brief Description of the Drawings
[0009] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.
Brief Description of the Drawings
[0010]
Figure 1
[0010] It is a schematic diagram showing the structure of an example video sequence according to some embodiments of the present disclosure.
Figure 2A
[0011] A schematic diagram showing an example encoding process of a hybrid video coding system in accordance with an embodiment of the present disclosure is shown.
Figure 2B
[0012] A schematic diagram showing another example encoding process of a hybrid video coding system in accordance with an embodiment of the present disclosure is shown.
Figure 3A
[0013] A schematic diagram showing an example decoding process of a hybrid video coding system in accordance with an embodiment of the present disclosure is shown.
Figure 3B
[0014] A schematic diagram showing another example decoding process of a hybrid video coding system in accordance with an embodiment of the present disclosure is shown.
Figure 4
[0015] A block diagram of an example apparatus for encoding or decoding video according to some embodiments of the present disclosure is shown.
Figure 5
[0016] Table 1 showing an example of the syntax structure of a sequence parameter set (SPS) according to some embodiments of the present disclosure.
Figure 6
[0017] Table 2 showing an example of the syntax structure of a picture parameter set (SPS) according to some embodiments of the present disclosure.
Figure 7
[0018] Table 3 showing an example of the syntax structure of a conversion unit according to some embodiments of the present disclosure.
Figure 8
[0019] Table 4 showing an example of the syntax structure related to the signaling of the block differential pulse code modulation (BDPCM) mode according to some embodiments of the present disclosure.
Figure 9
[0020] Table 5 showing another example of the syntax structure of an SPS according to some embodiments of the present disclosure.
Figure 10
[0021] Table 6 showing another example of the syntax structure of a conversion unit according to some embodiments of the present disclosure.
Figure 11
[0022] FIG. is a schematic diagram showing an example of diagonal scanning of a 64×64 transform block (TB) according to some embodiments of the present disclosure.
Figure 12A
[0023] Examples of residual units (RUs) according to some embodiments of the present disclosure are shown.
Figure 12B
[0023] Examples of residual units (RUs) according to some embodiments of the present disclosure are shown.
Figure 12C
[0023] Examples of residual units (RUs) according to some embodiments of the present disclosure are shown.
Figure 12D
[0023] Examples of residual units (RUs) according to some embodiments of the present disclosure are shown.
Figure 13
[0024] FIG. is a schematic diagram showing an example of diagonal scanning of a 64×64 TB divided into four 32×32 RUs according to some embodiments of the present disclosure.
Figure 14A
[0025] Table 7 showing an example of a syntax structure for residual coding when a TB is divided into RUs according to some embodiments of the present disclosure is shown.
Figure 14B
[0025] Table 7 showing an example of a syntax structure for residual coding when a TB is divided into RUs according to some embodiments of the present disclosure is shown.
Figure 14C
[0025] Table 7 showing an example of a syntax structure for residual coding when a TB is divided into RUs according to some embodiments of the present disclosure is shown.
Figure 14D
[0025] Table 7 showing an example of a syntax structure for residual coding when a TB is divided into RUs according to some embodiments of the present disclosure is shown.
Figure 15A
[0026] Table 8 showing another example of a syntax structure for residual coding according to some embodiments of the present disclosure is shown.
Figure 15B
[0026] Table 8 showing another example of a syntax structure for residual coding according to some embodiments of the present disclosure is shown.
Figure 15C
[0026] Table 8 showing another example of a syntax structure for residual coding according to some embodiments of the present disclosure is shown.
Figure 15D
[0026] Table 8 showing another example of a syntax structure for residual coding according to some embodiments of the present disclosure is shown.
Figure 16
[0027] Table 9 showing example parameter values derived from a chroma format according to some embodiments of the present disclosure is shown.
Figure 17
[0028] Table 10 showing an example of a syntax structure of Versatile Video Coding Draft 6 for residual coding that performs inverse level mapping according to some embodiments of the present disclosure is shown.
Figure 18
[0029] A flowchart of an example decoding method according to some embodiments of the present disclosure is shown.
Figure 19
[0030] Table 11 shows an example of a syntax structure related to residual coding without inverse level mapping according to some embodiments of the present disclosure.
Figure 20
[0031] Table 12 shows an example of a lookup table for selecting Rice parameters according to some embodiments of the present disclosure.
Figure 21
[0032] A flowchart of an example of a process for video processing according to some embodiments of the present disclosure is shown.
Figure 22
[0033] A flowchart of another example of a process for video processing according to some embodiments of the present disclosure is shown.
Best Mode for Carrying Out the Invention
[0011] Detailed Description
[0034] Hereinafter, reference will be made to the accompanying drawings, in which like numerals in different drawings represent the same or similar elements, unless otherwise specified. The embodiments described in the following description of the exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present invention described in the appended claims. Specific aspects of the present disclosure will be described in more detail below. In case of conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided in this specification shall prevail.
[0012]
[0035] The Joint Video Experts Team (JVET) of ITU-T VCEG (ITU-T Video Coding Expert Group) and ISO / IEC MPEG (ISO / IEC Moving Picture Expert Group) is currently developing the VVC (Versatile Video Coding) / H.266 standard. The VVC standard aims to double the compression efficiency of its predecessor, the HEVC (High Efficiency Video Coding) / H.265 standard. That is, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 with half the bandwidth.
[0013]
[0036] To achieve the same subjective quality as HEVC / H.265 with half the bandwidth, JVET has been developing technologies beyond HEVC using the JEM (joint exploration model) reference software. Since the coding technologies were incorporated into JEM, JEM has achieved significantly higher coding performance than HEVC.
[0014]
[0037] The VVC standard has been recently developed and continues to add more coding technologies to provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0015]
[0038] A video is a series of still pictures (or "frames") arranged in a time series to store visual information. These pictures can be captured and stored in a time series using a video capture device (e.g., a camera), and such pictures can be displayed in a time series using a video playback device (e.g., a TV, computer, smartphone, tablet computer, video player, or any end-user terminal with a display function). Also, depending on the application, for purposes such as surveillance, holding a meeting, or live broadcast, the video capture device can transmit the captured video to a video playback device (e.g., a computer equipped with a monitor) in real time.
[0016]
[0039] To reduce the memory space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and restored before display. Compression and restoration can be performed by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for restoration is generally called a "decoder". The encoder and decoder may be collectively referred to as a "codec". The encoder and decoder can be implemented as any of various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and restoration can be performed according to various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, H.26x series. Depending on the application, the codec can restore the video from the first encoding standard and recompress the restored video using the second encoding standard. In this case, the codec may be called a "transcoder".
[0017]
[0040] The video encoding process can identify and retain the useful information that can be used for picture reconstruction and ignore the information that is not important for reconstruction. If the unimportant information that has been ignored cannot be fully reconstructed, such an encoding process may be called "irreversible". Otherwise, it may be called "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.
[0018]
[0041] The useful information of the encoded picture (referred to as the "current picture") includes changes with respect to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes may include changes in pixel position, luminance, or color, among which the position change is the most important. The position change of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0019]
[0042] A picture encoded without referring to another picture (i.e., it is its own reference picture) is called an "I picture". A picture encoded using a previous picture as a reference picture is called a "P picture". A picture encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".
[0020]
[0043] FIG. 1 shows the structure of video sequence example 100 according to some embodiments of the present disclosure. Video sequence 100 may be a live video, or a captured and archived video. Video 100 may be an actual video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., an actual video with augmented reality effects). Video sequence 100 may be input from a video capture device (e.g., a camera), a video archive including previously captured videos (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.
[0021]
[0044] As shown in FIG. 1, the video sequence 100 may include a series of pictures temporally arranged along a timeline including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, and there are more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a certain picture (e.g., picture 104) may not be present immediately before or after that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. The reference pictures of pictures 102-106 are merely examples, and it should be noted that the present disclosure does not limit the embodiments of the reference pictures as in the example shown in FIG. 1.
[0022]
[0045] Generally, a video codec does not perform the encoding or decoding of an entire picture all at once due to the computational complexity of such a task. More precisely, they can split a picture into basic segments and encode or decode the picture segment by segment. Such a basic segment is referred to as a basic processing unit (“BPU (basic processing unit)”) in the present disclosure. For example, the structure 110 of FIG. 1 shows an example of the structure of a certain picture (e.g., any one of pictures 102 to 108) of the video sequence 100. In the structure 110, the picture is split into 4×4 basic processing units, and their boundaries are indicated by dashed lines. In some embodiments, the basic processing unit may be called a “macroblock” in some video coding standard specifications (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or may be called a “coding tree unit” (“CTU (coding tree unit)”) in some other video coding standard specifications (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size of the picture, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any shape and size of pixels. The size and shape of the basic processing unit can be selected for each picture based on the balance between the coding efficiency and the level of detail to be maintained in the basic processing unit.
[0023]
[0046] The basic processing unit may be a logical unit that can include a group of a plurality of different types of video data stored in a computer memory (for example, in a video frame buffer). For example, the basic processing unit of a color picture may include a luma component (Y) representing achromatic lightness information, one or more chroma components representing color information (for example, Cb and Cr), and related syntax elements (wherein the luma component and the chroma components may have basic processing units of the same size). The luma component and the chroma components may be referred to as "coding tree blocks" ("CTB (coding tree block)") in some video coding standard specifications (for example, H.265 / HEVC or H.266 / VVC). Any operation performed on the basic processing unit can be repeatedly performed on each of its luma component and chroma components.
[0024]
[0047] Video coding has multiple operation stages, examples of which are shown in FIGS. 2A-2B and FIGS. 3A-3B. At each stage, the size of the basic processing unit may still be too large to process, and thus, in the present disclosure, it can be further divided into segments called "basic processing subunits". In some embodiments, the basic processing subunits may be called "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or may be called "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunits are also logical units that can include a group of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing subunit can be repeated for each of its luma and chroma components. Note that such a division can be performed at further levels according to the processing needs. Also note that different stages can divide the basic processing unit using different schemes.
[0025]
[0048] For example, in the mode decision stage (an example of which is shown in FIG. 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) should be used for a basic processing unit, and the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into a plurality of basic processing subunits (e.g., CUs in the case of H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.
[0026]
[0049] As another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform prediction operations at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (referred to as "prediction blocks" or "PBs" in, for example, H.265 / HEVC or H.266 / VVC), and perform prediction operations at the level of these segments.
[0027]
[0050] As another example, in the transformation stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform transformation operations on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (referred to as "transformation blocks" or "TBs" in, for example, H.265 / HEVC or H.266 / VVC), and perform transformation operations at the level of these segments. Note that the same basic processing subunit division scheme can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.
[0028]
[0051] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are indicated by dotted lines. Different basic processing units of the same picture may be divided into basic processing subunits with different schemes.
[0029]
[0052] In some embodiments, to provide parallel processing capabilities and error resilience for video encoding and decoding, the picture can be divided into multiple regions for processing such that for each region of the picture, the encoding or decoding process can be made independent of information from any other region of the picture. That is, each region of the picture can be processed independently. By doing so, the codec can process multiple different regions of the picture in parallel, thus improving the encoding efficiency. Also, if the data of a certain region is corrupted during processing or lost during network transmission, the codec can accurately encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video encoding standards, a picture can be divided into multiple different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two region types: "slice" and "tile". It should also be noted that different pictures of video sequence 100 can have different partitioning schemes for dividing the picture into regions.
[0030]
[0053] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, and their boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in FIG. 1 are merely examples and the present disclosure does not limit to those embodiments.
[0031]
[0054] FIG. 2A shows a schematic diagram of an encoding process example 200A consistent with an embodiment of the present disclosure. For example, the encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoder can encode the video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 can include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into a basic processing unit, a basic processing subunit, or a processing area. In some embodiments, the encoder can perform the process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder can perform the process 200A in an iterative manner where one basic processing unit can be encoded by one iteration of the process 200A. In some embodiments, the encoder can perform the process 200A in parallel for the areas (e.g., areas 114 to 118) of each original picture of the video sequence 202.
[0032]
[0055] In FIG. 2A, the encoder can send the basic processing unit of the original picture of the video sequence 202 (referred to as the "original BPU") to the prediction stage 204 in order to generate prediction data 206 and prediction BPU 208. The encoder can generate the residual BPU 210 by subtracting the prediction BPU 208 from the original BPU. The encoder can send the residual BPU 210 to the conversion stage 212 and the quantization stage 214 in order to generate the quantized transform coefficients 216. The encoder can send the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 in order to generate the video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, the encoder can send the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 in order to generate the reconstructed residual BPU 222 after the quantization stage 214. The encoder can generate the prediction reference 224 used in the prediction stage 204 for the next iteration of process 200A by adding the reconstructed residual BPU 222 to the prediction BPU 208. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path can be used to ensure that the encoder and the decoder use the same reference data for prediction together.
[0033]
[0056] The encoder can iteratively perform process 200A in order to encode each original BPU of the original picture (in the forward path) and generate the prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture of the video sequence 202.
[0034]
[0057] Referring to process 200A, the encoder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" can refer to receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or any action of any method for inputting data.
[0035]
[0058] In prediction stage 204, in the current iteration, the encoder can receive the original BPU and prediction reference 224, and can perform prediction operations to generate prediction data 206 and prediction BPU 208. Prediction reference 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce the redundancy of information by extracting prediction data 206 as prediction BPU 208 from prediction data 206 and prediction reference 224, which can be used to reconstruct the original BPU.
[0036]
[0059] Ideally, prediction BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, prediction BPU 208 generally differs slightly from the original BPU. To record such a difference, after generating prediction BPU 208, the encoder can generate residual BPU 210 by subtracting it from the original BPU. For example, the encoder can subtract the pixel value (e.g., grayscale value or RGB value) of prediction BPU 208 from the corresponding pixel value of the original BPU. Each pixel of residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and prediction BPU 208. Compared with the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.
[0037]
[0060] To further compress the residual BPU 210, in the transformation stage 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional “basis patterns” (each basis pattern is associated with a “transformation coefficient”). The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component of the residual BPU 210 (e.g., the frequency of brightness variation). No basis pattern can be reproduced from any combination (e.g., linear combination) of the other basis patterns. That is, this decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, where the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0038]
[0061] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as the discrete cosine transform or the discrete sine transform can be used in the transformation stage 212. The transformation in the transformation stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by means of the inverse operation of the transformation (referred to as “inverse transformation”). For example, to restore the pixels of the residual BPU 210, the inverse transformation may also generate a weighted sum by multiplying the values of the corresponding pixels of the basis patterns by their respective associated coefficients and adding those products. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis patterns from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.
[0039]
[0062] The encoder can further compress the conversion coefficients at the quantization stage 214. In the conversion process, different base patterns may represent different fluctuation frequencies (e.g., brightness fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore the information of high-frequency fluctuations without causing significant quality degradation in decoding. For example, at the quantization stage 214, the encoder can generate the quantized conversion coefficients 216 by dividing each conversion coefficient by an integer value (referred to as the "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some conversion coefficients of the high-frequency base pattern may be converted to zero, and the conversion coefficients of the low-frequency base pattern may be converted to smaller integers. The encoder can ignore the quantized conversion coefficients 216 with zero values, thereby further compressing the conversion coefficients. The quantization process is also reversible, where the quantized conversion coefficients 216 can be reconstructed into the conversion coefficients by the inverse operation of quantization (referred to as "inverse quantization").
[0040]
[0063] Since the encoder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be irreversible. Generally, the quantization stage 214 can contribute the most to information loss in the process 200A. The greater the information loss, the fewer bits the quantized conversion coefficients 216 may require. To obtain different levels of information loss, the encoder can use different values of the quantization parameter or other parameters of the quantization process.
[0041]
[0064] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques such as, for example, entropy encoding, variable-length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or other reversible or irreversible compression algorithms. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can also encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), or the encoder control parameters (e.g., bitrate control parameters). The encoder can use the output data of the binary encoding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0042]
[0065] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can generate the reconstructed transform coefficients by performing inverse quantization on the quantized transform coefficients 216. In the inverse transform stage 220, the encoder can generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can generate the prediction reference 224 used in the next iteration of process 200A by adding the reconstructed residual BPU 222 to the prediction BPU 208.
[0043]
[0066] Note that other variations of process 200A can be used to encode the video sequence 202. In some embodiments, the stages of process 200A can be performed by an encoder in a different order. In some embodiments, one or more stages of process 200A may be integrated into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be integrated into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages of FIG. 2A.
[0044]
[0067] FIG. 2B shows a schematic diagram of another encoding process example 200B that is consistent with an embodiment of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0045]
[0068] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., in-picture prediction or "intra prediction") can predict the current BPU by using pixels from one or more already-encoded adjacent BPUs within the same picture. That is, the prediction reference 224 in spatial prediction may include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can predict the current BPU by using regions from one or more already-encoded pictures. That is, the prediction reference 224 in temporal prediction may include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0046]
[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra prediction. With respect to the original BPU of the encoded picture, the prediction reference 224 may include one or more adjacent BPUs encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation techniques may include, for example, linear extrapolation or interpolation, or polynomial extrapolation or interpolation. In some embodiments, the encoder can perform extrapolation at the pixel level, for example, by extrapolating the values of the corresponding pixels for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation may be located relative to the original BPU in various directions, such as the vertical direction (e.g., above the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined in the video coding standard used. In the case of intra prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, or the direction of the adjacent BPUs used relative to the original BPU.
[0047]
[0070] As another example, in the time prediction stage 2044, the encoder can perform inter prediction. For the original BPU of the current picture, the prediction reference 224 can include one or more pictures (referred to as "reference pictures") that are encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the encoder can generate a reconstruction BPU by adding the reconstruction residual BPU 222 to the prediction BPU 208. When all the reconstruction BPUs of the same picture are generated, the encoder can generate a reconstructed picture as a reference picture. The encoder can perform an operation of "motion estimation" to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture can be determined based on the location of the original BPU in the current picture. For example, the search window may be centered at a location having coordinates within the same reference picture as the original BPU of the current picture and may extend outward by a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., using a per-recursive algorithm or a block matching algorithm), the encoder can determine such a region as the matching region. The matching region may have dimensions different from those of the original BPU (e.g., smaller, equal, larger, or different in shape). Since the reference picture and the current picture are temporally separated in the timeline (as shown, for example, in FIG. 1), as time elapses, the matching region can be regarded as "moving" to the location of the original BPU. The encoder can record the direction and distance of such motion as a "motion vector". When multiple reference pictures are used (as in picture 106 of FIG. 1, for example), the encoder can search for the matching region for each reference picture and determine the motion vector associated therewith. In some embodiments, the encoder can assign weights to the pixel values of the matching region of each matching reference picture.
[0048]
[0071] Motion estimation can be used to identify various types of motion, such as translational, rotational, or zooming. In the case of inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vectors associated with the matching region, the number of reference pictures, or the weights associated with the reference pictures.
[0049]
[0072] To generate the prediction BPU 208, the encoder can perform the operation of "motion compensation". Using motion compensation, the prediction BPU 208 can be reconstructed based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder can move the matching region of the reference picture according to the motion vector by which the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder can move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values of the matching regions. In some embodiments, when the encoder assigns weights to the pixel values of the matching regions of each matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching regions.
[0050]
[0073] In some embodiments, inter prediction may be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction with respect to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter prediction picture where the reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions with respect to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.
[0051]
[0074] Referring further to the forward path of Process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of Process 200B. For example, the encoder can perform a rate-distortion optimization technique in which the encoder can select a prediction mode to minimize the value of a cost function according to the bitrate of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0052]
[0075] In the reconstruction path of process 200B, when the intra prediction mode is selected in the forward path, after the generation of prediction reference 224 (e.g., the current BPU encoded and reconstructed within the current picture), the encoder can directly send the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). When the inter prediction mode is selected in the forward path, after the generation of prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the encoder can send the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, or adaptive loop filter. The reference picture on which loop filtering has been performed may be stored in buffer 234 (or "decode picture buffer") for later use (e.g., to be used as an inter prediction reference picture for future pictures of video sequence 202). The encoder may store one or more reference pictures used in the temporal prediction stage 2044 in buffer 234. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in the binary encoding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.
[0053]
[0076] FIG. 3A shows a schematic diagram of a decoding process example 300A consistent with an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to the compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), generally, the video stream 304 is not identical to the video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, the decoder can perform process 300A at the level of the basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner where the decoder can decode one basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.
[0054]
[0077] In FIG. 3A, the decoder can send the portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the “encoded BPU”) to the binary decoding stage 302. At the binary decoding stage 302, the decoder can decode the above portion into prediction data 206 and quantization transform coefficients 216. The decoder can send the quantization transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The decoder can send the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can generate the prediction reference 224 by adding the reconstructed residual BPU 222 to the prediction BPU 208. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., the decoded picture buffer of a computer memory). The decoder can send the prediction reference 224 to the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0055]
[0078] The decoder can repeatedly perform process 300A to decode each encoded BPU of the encoded picture and generate the prediction reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture of the video bitstream 228.
[0056]
[0079] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or other reversible compression algorithms) used by the encoder. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can also decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of transform, the parameters of the quantization process (e.g., quantization parameters), or the encoder control parameters (e.g., bit rate control parameters). In some embodiments, when the video bitstream 228 is packet transmitted over the network, the decoder can depacketize the video bitstream 228 before sending it to the binary decoding stage 302.
[0057]
[0080] FIG. 3B shows a schematic diagram of another decoding process example 300B consistent with an embodiment of the present disclosure. Process 300B can be changed from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.
[0058]
[0081] In Process 300B, with respect to the encoded basic processing unit (referred to as the "current BPU") of the encoded picture (referred to as the "current picture") being decoded, the prediction data 206 decoded from the binary decoding stage 302 by the decoder can include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, or parameters of the intra prediction operation. The parameters of the intra prediction operation can include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as a reference, the size of the adjacent BPUs, extrapolation parameters, or the direction of the adjacent BPUs with respect to the original BPU. As another example, if inter prediction was used by the encoder to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, or parameters of the inter prediction operation. The parameters of the inter prediction operation can include, for example, the number of reference pictures associated with the current BPU, the weights respectively associated with the reference pictures, the location (e.g., coordinates) of one or more matching regions in each reference picture, or one or more motion vectors respectively associated with the matching regions.
[0059]
[0082] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are shown in Figure 2B and will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208. As shown in Figure 3A, the decoder can generate a prediction reference 224 by adding the predicted BPU 208 and the reconstructed residual BPU 222.
[0060]
[0083] In process 300B, the decoder can send the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly send the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can send the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply the loop filter to the prediction reference 224 in the manner shown in FIG. 2B. The reference picture subjected to loop filtering may be stored in the buffer 234 (e.g., the decoded picture buffer of the computer memory) for later use (e.g., as an inter prediction reference picture for a picture to be encoded in the future of the video bitstream 228). The decoder may store one or more reference pictures used in the temporal prediction stage 2044 in the buffer 234. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter prediction has been used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).
[0061]
[0084] FIG. 4 is a block diagram of an example apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in FIG. 4, the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of operating on or processing information. For example, the processor 402 may include any combination of several central processing units (i.e., “CPUs”), graphics processing units (i.e., “GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP (intellectual property) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system on chips (SoCs), or application specific integrated circuits (ASICs). In some embodiments, the processor 402 may be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 may include a plurality of processors including processor 402a, processor 402b, and processor 402n.
[0062]
[0085] Device 400 may also include a memory 404 configured to store data (such as, for example, an instruction set, computer code, or intermediate data). For example, as shown in FIG. 4, the stored data may include program instructions (such as, for example, program instructions for implementing stages of processes 200A, 200B, 300A, or 300B) and processing data (such as, for example, video sequence 202, video bitstream 228, or video stream 304). The processor 402 can execute the program instructions to access the program instructions and the processing data (for example, via bus 410) and perform operations or manipulations on the processing data. The memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 may include any combination of several random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid state drives, flash drives, SD (security digital) cards, memory sticks, or compact flash (registered trademark) (CF) cards. The memory 404 may also be a group of memories grouped as a single logical component (not shown in FIG. 4).
[0063]
[0086] The bus 410 may be a communication device that transfers data between components within the device 400, such as an internal bus (such as, for example, a CPU memory bus) or an external bus (such as, for example, a universal serial bus port, a peripheral component interconnect express port).
[0064]
[0087] For simplicity of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuit may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Further, the data processing circuit may be a single independent module or may be fully or partially integrated with any other component of the device 400.
[0065]
[0088] Device 400 may further include a network interface 406 to provide wired or wireless communication with a network (such as, for example, the Internet, an intranet, a local area network, or a mobile communication network). In some embodiments, network interface 406 may include any combination of some network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth® adapters, infrared adapters, near field communication (“NFC”) adapters, or cellular network chips.
[0066]
[0089] In some embodiments, optionally, device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As shown in FIG. 4, peripheral devices may include, but are not limited to, a cursor control device (such as, for example, a mouse, touchpad, or touch screen), a keyboard, a display (such as, for example, a cathode ray tube display, a liquid crystal display, or a light emitting diode display), or a video input device (such as, for example, a camera, or an input interface communicatively coupled to a video archive).
[0067]
[0090] Note that the video codec (e.g., the codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, or NPU, etc.).
[0068]
[0091] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIGS. 2A or 2B, inverse quantization 218 in FIGS. 3A or 3B), quantization parameters (QP) are used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for encoding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element of the picture parameter set (PPS) and the slice_qp_delta syntax element of the slice header. Further, the QP value can be adapted at a local level for each CU using the delta QP values sent at the granularity of the quantization group.
[0069]
[0092] In VVC6 (Versatile Video Coding Draft 6), the residual of a transform block (TB) of video data can be encoded using a transform skip (TS) mode in which the transform stage is skipped. For example, the decoder can decode the video data using the TS mode by decoding the video data to obtain the residual, and then perform inverse quantization and reconstruction without performing inverse transformation. VVC6 limits the applicability of the TS mode by the maximum block size (here, the TS mode is applicable to the TB only when the width and height of the TB are at most 32 pixels). Such a maximum block size for applying the TS mode can be specified as the picture parameter set (PPS) level syntax log2_transform_skip_max_size_minus2, and can exist within the range of 0 to 3. If it does not exist, the value of log2_transform_skip_max_size_minus2 is inferred to be 0. The maximum value MaxTsSize of the width or height of the maximum block that limits the TS mode can be determined based on Equation (1). MaxTsSize = 1 << ( log2_transform_skip_max_size_minus2 + 2 ) Equation (1)
[0070]
[0093] That is, when log2_transform_skip_max_size_minus2 is 0, the TS mode can be allowed if the TB width and height are at most 4. In the current design of VVC6, since the maximum allowable value of log2_transform_skip_max_size_minus2 is 3, the maximum allowable value of MaxTsSize is 32. If the width and height of the TB are at most MaxTsSize, a parameter transform_skip_flag that specifies whether the TS mode is selected can be signaled. If the width or height of the TB is greater than 32, the TS mode is not allowed for that TB.
[0071]
[0094] In VVC6, the residual levels in the TS mode are coded using non - overlapping coefficient groups (CGs) of size 4×4. The transform skip coefficient levels of the CGs are coded in 3 passes over multiple scan positions.
[0072]
[0095] The first pass can be represented by the following pseudo - code. for(n = 0; n <= numSbCoeff - 1; n++ ) if (remainingCtxBin > 0), decode sig_coeff_flag (context) else, bypass - decode sig_coeff_flag (bypass) if (remainingCtxBin > 0), decode coeff_sign_flag (context) else, bypass - decode coeff_sign_flag (bypass) if (remainingCtxBin > 0), decode abs_level_gtx_flag[0] (context) else, bypass - decode abs_level_gtx_flag[0] (bypass) if (remainingCtxBin > 0), decode par_level_flag (context) else, bypass - decode par_level_flag (bypass)
[0073]
[0096] The second pass can be represented by the following pseudo - code. for(n = 0; n <= numSbCoeff - 1; n++ ) if (remainingCtxBin > 0), decode abs_level_gtx_flag[1] (context) else, bypass - decode abs_level_gtx_flag[1] (bypass) If (remainingCtxBin > 0), decode abs_level_gtx_flag[2] (context) Else, bypass decode abs_level_gtx_flag[2] (bypass) If (remainingCtxBin > 0), decode abs_level_gtx_flag[3] (context) If (remainingCtxBin > 0), decode abs_level_gtx_flag[4] (context) Else, bypass decode abs_level_gtx_flag[4] (bypass)
[0074]
[0097] The third pass can be represented by the following pseudo-code. for(n = 0; n <= numSbCoeff - 1; n++ ) rice = cctx.templateAbsSumTS(n, coeff); Decode abs_remainder_using_RG_Coding
[0075]
[0098] In the above description, the syntax elements in the TS mode for residual coding (referred to as "TS residual coding") can be coded using either context coding (displayed as "context") or bypass coding (displayed as "bypass").
[0076]
[0099] In some embodiments, for TS residual coding, an encoding tool called "level mapping" can be employed. The absolute coefficient level parameter absCoeffLevel can be mapped to a transform level such that it is encoded according to the values of the quantized residual samples to the left and above the current residual sample. Let X0 denote the absolute coefficient level to the left of the current coefficient and X1 denote the absolute coefficient level above the current coefficient. To represent the coefficient using the absolute coefficient level ("absCoeff"), the mapped parameter absCoeffMod can be encoded. absCoeffMod can be derived in a manner represented by the following pseudo-code. pred = max(X0, X1); if (absCoeff == pred) { absCoeffMod = 1; } else { absCoeffMod = (absCoeff < pred)? absCoeff + 1 : absCoeff; }
[0077]
[0100] The current design of the TS mode has several issues. In VVC6, the TS mode is an encoding tool that can achieve mathematically lossless compression for a block under the two conditions that an appropriate quantization parameter value is selected and the loop filter stage is turned off. Since VVC6 does not allow the TS mode for TBs with a width or height greater than 32, the current design of VVC6 cannot achieve mathematically lossless compression for a block when the TB width or height is greater than 32.
[0078]
[0101] In addition, since the newly adopted level mapping process requires the decoder to calculate prediction values from above and to the left for each coefficient level, it has a significant impact on the throughput of context adaptive binary arithmetic coding (CABAC). Since the process of deriving the Rice parameter depends on the actual level, the calculation of the actual level with inverse mapping needs to be performed within the CABAC parse loop. Such an interleaved way of parsing and level decoding is undesirable because it can reduce the throughput of the decoder hardware implementation.
[0079]
[0102] In VVC6, in addition to log2_transform_skip_max_size_minus2 as described above, another sequence parameter set (SPS) level flag, sps_max_luma_transform_size_64_flag, can specify the maximum TB size for luma samples. When sps_max_luma_transform_size_64_flag is equal to 1, the maximum TB size for luma samples is equal to 64. When sps_max_luma_transform_size_64_flag is equal to 0, the maximum TB size for luma samples is equal to 32. When the luma coding tree block size of the coding tree unit (CTU) is less than 64, the value of sps_max_luma_transform_size_64_flag is equal to 0. Based on sps_max_luma_transform_size_64_flag, the parameters MaxTbLog2SizeY and the maximum TB size MaxTbSizeY can be derived based on Equations (2) and (3). MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6: 5 Equation (2) MaxTbSizeY = 1 << MaxTbLog2SizeY Equation (3)
[0080]
[0103] Based on formulas (2) to (3), the maximum value of PPS level syntax log2_transform_skip_max_size_minus2 may depend on the SPS level flag sps_max_luma_transform_size_64_flag. log2_transform_skip_max_size_minus2 specifies the maximum block size used in the TS mode, and its value may exist within the range of 0 to (3 + sps_max_luma_transform_size_64_flag). The encoder may be configured to ensure that the value of log2_transform_skip_max_size_minus2 is within the allowable range. If it does not exist, the value of log2_transform_skip_max_size_minus2 may be inferred to be 0. The maximum allowable MaxTsSize can be determined using formula (1). When the width and height of the TB are less than MaxTsSize, the TS mode may be allowed to encode the TB.
[0081]
[0104] As can be seen from the above description, in VVC6, log2_transform_skip_max_size_minus2 is signaled only when sps_transform_skip_enabled_flag is 1. That sps_transform_skip_enabled_flag is equal to 0 means that there is no transform_skip_flag in the transform unit syntax. Therefore, when sps_transform_skip_enabled_flag is 0, it is not necessary to signal log2_transform_skip_max_size_minus2. This current signaling in VVC6 has a problem of parsing dependency between the SPS and the PPS. The above embodiments also have the same problem of parsing dependency between the PPS syntax log2_transform_skip_max_size_minus2 and the SPS syntax sps_max_luma_transform_size_64_flag. Such parsing dependencies are generally not desirable.
[0082]
[0105] Embodiments of the present disclosure provide a technical solution to the above technical problem. To achieve lossless compression for large TBs using the TS mode, the present disclosure provides embodiments in which the TS mode can be extended to be applicable to TB sizes up to the maximum TB size allowed for an encoded video sequence. Different coefficient scanning methods are also provided for TS residual encoding.
[0083]
[0106] In accordance with some embodiments of the present disclosure, log2_transform_skip_max_size_minus2 can be moved from the PPS to the SPS to remove the parsing dependency between the SPS and the PPS. As an example, FIG. 5 shows Table 1 which is an example of the syntax structure of a sequence parameter set (SPS) according to some embodiments of the present disclosure. FIG. 6 shows Table 2 which is an example of the syntax structure of a picture parameter set (SPS) according to some embodiments of the present disclosure. Tables 1 and 2 show that log2_transform_skip_max_size_minus2 is moved from the PPS to the SPS as shown by row 502 of Table 1 and rows 602 - 604 of Table 2.
[0084]
[0107] In accordance with some embodiments of the present disclosure, the maximum block size for applying the TS mode block can be set as the maximum TB size (MaxTbSizeY), in which case log2_transform_skip_max_size_minus2 is not signaled. By doing so, the TS mode can be allowed when the width and height of the TB are less than or equal to MaxTbSizeY. In some embodiments, MaxTbSizeY can be determined based on equations (2) - (3).
[0085]
[0108] As an example, FIG. 7 shows Table 3 which shows an example of the syntax structure of a conversion unit according to some embodiments of the present disclosure. Table 3 shows that according to the example of the syntax structure of the conversion unit, as indicated by row 706, the width and height of the TB can be at most the maximum value MaxTbSizeY (i.e., 32). By doing so, since the maximum block size for applying the TS mode is the same as MaxTbSizeY, the TS mode can be allowed for all TBs, and as shown in rows 702 to 704, no further check is required to determine whether the width and height of the TB are less than or equal to MaxTbSizeY. It should be noted that VVC6 also uses a multiple transform selection (MTS (Multiple Transform Selection)) scheme for residual encoding both of inter-encoded blocks and intra-encoded blocks. MTS uses multiple selected transforms from DCT8 / DST7. However, since MTS is allowed when both tbWidth and tbHeight are 32 or less, a further check is required during MTS encoding.
[0086]
[0109] VVC6 provides another encoding tool called block differential pulse code modulation (BDPCM). In the BDPCM mode, horizontal and vertical differential pulse code modulation (DPCM) is applied in the residual region and the conversion stage is skipped. The maximum allowable block width or height for applying the BDPCM mode is the same as that of the TS mode.
[0087]
[0110] In accordance with some embodiments of the present disclosure, the maximum block size for applying the BDPCM mode can also be extended to be the same as the maximum block size for applying the TS mode. By doing so, the BDPCM mode can be allowed when the width and height of the coding unit (CU) are less than or equal to MaxTbSizeY. As an example, FIG. 8 shows Table 4, which is an example of a syntax structure related to the signaling of the block differential pulse code modulation (BDPCM) mode according to some embodiments of the present disclosure. As shown by row 802 in Table 4, it shows that the maximum block size for applying the BDPCM mode can be extended to be the same as the maximum block size for applying the TS mode.
[0088]
[0111] In some cases, the allowable value of log2_transform_skip_max_size_minus2 may depend on the codec profile. For example, the main profile can specify that the value of log2_transform_skip_max_size_minus2 can be the same as the maximum TB size. Any bitstream that signals a log2_transform_skip_max_size_minus2 value different from the maximum TB size can be regarded as a non-compliant bitstream by the codec. In the case of an extended profile beyond the main profile, the value of log2_transform_skip_max_size_minus2 can be different from the maximum TB size.
[0089]
[0112] In accordance with some embodiments of the present disclosure, methods and syntax structures are provided herein for inferring that log2_transform_skip_max_size_minus2 is signaled without signaling it and is the same as the maximum TB size, or for ensuring that the value of log2_transform_skip_max_size_minus2 is always the same as the maximum TB size, such as through profile constraint configurations. By doing so, the number of combinations of syntax element values to be tested is reduced, thus reducing the burden on decoder implementation.
[0090]
[0113] In some embodiments, the SPS flag may be signaled to indicate that the maximum block size for applying the TS mode is 32 or 64. For example, the SPS flag can be signaled in the same way as the signaling of the maximum TB size. As an example, to specify that the maximum block size for applying the TS mode is 32, sps_max_transform_skip_size_64_flag can be set to 0. In another example, to specify that the maximum block size for applying the TS mode is 64, sps_max_transform_skip_size_64_flag can be set to 1. In some embodiments, if the signaling of sps_max_transform_skip_size_64_flag is not performed, its value can be inferred to be 0.
[0091]
[0114] In some embodiments, the maximum block size for applying the TS mode can be determined based on Equation (4). MaxTsSize = sps_max_transform_skip_size_64_flag? 64: 32 Equation (4)
[0092]
[0115] In some embodiments, sps_max_transform_skip_size_64_flag may be signaled when both sps_max_luma_transform_size_64_flag and sps_transform_skip_enabled_flag are equal to 1.
[0093]
[0116] As an example, FIG. 9 shows Table 5 which is an example of the syntax structure of the SPS for signaling sps_max_transform_skip_size_64_flag according to some embodiments of the present disclosure. FIG. 10 shows Table 6 which is an example of the syntax structure of the transform unit for signaling sps_max_transform_skip_size_64_flag according to some embodiments of the present disclosure. Tables 5 and 6 show the implementation of the signaling of sps_max_transform_skip_size_64_flag as indicated by row 902 of Table 5 and rows 1002 - 1006 of Table 6.
[0094]
[0117] In accordance with some embodiments of the present disclosure, since the maximum block size for applying the TS mode or the BDPCM mode can be extended such that it is the maximum TB size, the residual coding in the TS mode or the BDPCM mode can also be extended in that regard to allow encoding the maximum TB size. According to some disclosed embodiments, the residual coding can be directly extended to allow up to the maximum TB size without changing the scanning pattern.
[0095]
[0118] In some embodiments, similar to VVC draft 6, the transform block can be divided into coefficient groups (CGs) and diagonal scanning can be performed. As an example, FIG. 11 is a schematic diagram showing an example of diagonal scanning of a 64×64 transform block (TB) according to some embodiments of the present disclosure. FIG. 11 shows the diagonal scanning pattern (indicated by the zigzag arrow line) of a 64×64 TB (e.g., MaxTbSizeY = 64). Each cell in FIG. 11 can represent a 4×4 CG. Although FIG. 11 shows a 64×64 TB to illustrate the diagonal scanning process, it should be noted that the TB can be of any size or any shape and is not limited to the examples shown herein. For example, if the TB is rectangular instead of square, only one of its dimensions is equal to 64.
[0096]
[0119] One of the challenges in scanning the entire TB (e.g., the 64×64 TB in FIG. 11) in residual signification is that the current residual signification in VVC only supports up to a block size of 32×32, so the current VVC residual signification needs to be changed to support the above extension. In the current VVC design, even if a transformation is applied to a 64×64 TB (e.g., in non - skip mode), the decoder may still need to apply the residual signification only to a 32×32 block of coefficients representing the upper - left 32×32 block of the 64×64 TB. In such a case, all the remaining high - frequency coefficients are forced to be zero (therefore, encoding of the remaining coefficients is not necessary). For example, in the case of an M×N TB (where M is the block width and N is the block height), when M is equal to 64, only the left 32 columns of the transform coefficients can be encoded. Similarly, when N is equal to 64, only the upper 32 rows of the transform coefficients can be encoded.
[0097]
[0120] In accordance with some embodiments of the present disclosure, to reuse the existing VVC6 residual signification technology, a large TB can be divided into small residual units (RUs). For example, if the width of the TB is greater than 32, the TB can be divided into two partitions horizontally. As another example, if the height of the TB is greater than 32, the TB can be divided into two partitions vertically. In yet another example, if both dimensions of the TB are greater than 32, the TB can be divided into four RUs both horizontally and vertically. After division, a 32×32 RU can be encoded.
[0098]
[0121] As an example, FIGS. 12A to 12D show examples of residual units (RUs) according to some embodiments of the present disclosure. In FIG. 12A, a 64×64 TB is divided into four 32×32 RUs (shown by the dashed line). In FIG. 12B, a 64×16 TB is horizontally divided into two 32×16 RUs (shown by the dashed line). In FIG. 12C, a 32×64 TB is vertically divided into two 32×32 RUs (shown by the dashed line). In FIG. 12D, since neither the height nor the width exceeds 32, no division is performed, and the RU size is the same as the TB size. In some embodiments, the maximum allowable RU size is 32×32.
[0099]
[0122] As an example, FIG. 13 is a schematic diagram showing an example of diagonal scanning of a 64×64 TB in which the TB is divided into four 32×32 RUs according to some embodiments of the present disclosure. In FIG. 13, the 64×64 TB is divided into four RUs (shown by the thick solid line within the TB), and the coefficients of each RU are scanned individually (e.g., independently) within the RU in the same order as the scanning pattern for the 32×32 TB. In FIG. 13, the context model and Rice parameter derivation of a certain RU can be independent of other RUs. In some embodiments, the maximum number of context coded bins can also be assigned independently for each RU. Such a scheme is different from the case of VVC6 where the maximum number of context coded bins is defined at the TB level.
[0100]
[0123] As an example, FIGS. 14A to 14D show Table 7 showing examples of syntax structures related to residual coding when the TB is divided into RUs according to some embodiments of the present disclosure.
[0101]
[0124] In VVC6, for each coefficient group (CG) of a TS mode block, coded_sub_block_flag is signaled. coded_sub_block_flag = 0 means that all coefficients of the CG are zero. coded_sub_block_flag = 1 means that at least one coefficient in the CG is non - zero. However, if all coded_sub_block_flag of the previously coded CG (i.e., before the last CG) are zero, the coded_sub_block_flag of the last CG is not signaled and is inferred to be 1. This means that the parsing of the last CG of a TB depends on all previously decoded CGs. To remove the dependency between RUs, coded_sub_block_flag may be signaled for all CGs of the RUs including the last CG.
[0102]
[0125] In accordance with some embodiments of the present disclosure, an additional syntax coded_RU_flag may be introduced. In some embodiments, coded_RU_flag may be signaled when the number of RUs in a TB is greater than 1. In some embodiments, if coded_RU_flag does not exist, it may be inferred to be 1. coded_RU_flag = 0 may specify that all coefficients of the RU are zero. coded_RU_flag = 1 may specify that at least one coefficient of the RU is non - zero. In some embodiments, if all coded_RU_flag except for the last RU are zero, the coded_RU_flag of the last RU need not be signaled and can be inferred to be 1. As an example, the following pseudo - code shows an example of signaling coded_RU_flag. inferRUCbf = 1; for( k =0; k < numofRUs; k++ ) { if( (k != lastRU | |!inferRUCbf ) signal coded_RU_flag; if( coded_RU_flag) inferRUCbf = 0; }
[0103]
[0126] For example, FIGS. 15A to 15D show Table 8, which is another syntax structure example related to residual coding when coded_RU_flag is signaled according to some embodiments of the present disclosure. In some embodiments, when coded_RU_flag is signaled, the last CG flag can be maintained in the same way as in the case of VVC6. That is, when the coded_sub_block_flag of all previous CGs within the same RU is zero, the coded_sub_block_flag is not signaled and can be inferred to be 1.
[0104]
[0127] The JVET (Joint Video Experts Team) AHG reversible and nearly reversible coding tool (AHG18) releases reversible software based on VTM-6.0. The reversible software introduces a CU-level flag called cu_transquant_bypass_flag. cu_transquant_bypass_flag = 1 means that the transformation and quantization of that CU are skipped, and that CU is coded in reversible mode. In the current version of the reversible software, sps_max_luma_transform_size_64_flag is set to 0, which means that the maximum TB size in luma samples is limited to 32×32. For chroma samples, the maximum TB size is adjusted based on the YUV color format (for example, 16×16 at most for YUV420). In some embodiments, the luma transform block size can be increased up to 64×64 when cu_transquant_bypass_flag = 1, and the above-described residual coding technique can be used when cu_transquant_bypass_flag = 1.
[0105]
[0128] In some embodiments, the maximum TB size for chroma components can be determined using equations (2) and (3). Based on equations (2) and (3), the maximum TB width maxTbWidth and the maximum TB height maxTbHeight can be determined based on equations (5) and (6). maxTbWidth = ( cIdx == 0 ) ? MaxTbSizeY : MaxTbSizeY / SubWidthC Equation (5) maxTbHeight = ( cIdx == 0 ) ? MaxTbSizeY : MaxTbSizeY / SubHeightC Equation (6)
[0106]
[0129] In equations (5) and (6), cIdx = 0 means the luma component. cIdx = 1 and cIdx = 2 mean the two chroma components. As an example, the values of SubWidthC and SubHeightC can be derived from the chroma format. Consistent with some embodiments of the present disclosure, FIG. 16 shows Table 9 showing examples of parameter values derived from the chroma format according to some embodiments of the present disclosure.
[0107]
[0130] In VVC6, inverse level mapping is embedded in the CABAC module. FIG. 17 shows Table 10 showing an example of the syntax structure in VVC6 for residual coding that performs inverse level mapping according to some embodiments of the present disclosure.
[0108]
[0131] In accordance with some embodiments of the present disclosure, in order to improve the CABAC throughput of the conversion skip residual parse, instead of being based on the actual level value, the Rice parameter can be derived based on the mapped level value. In some embodiments, both the context model and the Rice parameter can depend on the mapped value, and it is possible that the inverse mapping operation is not performed during the residual parse process. By doing so, the inverse mapping can be separated from the residual parse process. The inverse mapping can be executed after the completion of the parse of the residuals of the entire TB. In some embodiments, the inverse mapping and the residual parse may be performed simultaneously within one pass, which allows the actual implementation to decide whether to interleave the parse and the mapping or divide them into two passes.
[0109]
[0132] As an example, FIG. 18 is a flowchart of a decoding method example 1400 according to some embodiments of the present disclosure. Method 1800 can be performed when the parse and the inverse mapping are separated. FIG. 18 shows that the inverse mapping is separated from the residual parse by being executed after the completion of the parse of the residuals of the entire TB and before inverse quantization.
[0110]
[0133] In accordance with some embodiments of the present disclosure, FIG. 19 shows Table 10 which shows an example of a syntax structure related to residual coding in which the inverse level mapping is not performed according to some embodiments of the present disclosure. In some embodiments, the inverse level mapping can be moved to the decoding process, which will be described below.
[0111]
[0134] In accordance with some embodiments of the present disclosure, the following pseudocode shows an inverse level mapping process that can be performed after residual parsing (as shown in FIG. 18) and before inverse quantization. In the following pseudocode, TransCoeffLevel [xC][yC] represents the coefficient value at the (xC, yC) position after residual parsing, and TransCoeffLevelInvMapped [xC][yC] represents the coefficient value at the (xC, yC) position after inverse mapping. for (int yC = 0; yC < height; yC++) { for (int xC = 0; xC < width; xC++) { TransCoeffLevelInvMapped [xC][yC] = TransCoeffLevel [xC][yC]; if (TransCoeffLevel [xC][yC]) { topPos = abs (TransCoeffLevel [xC][yC-1]); leftPos = abs(TransCoeffLevel [xC - 1][yC]); if (topPos || leftPos) { int absMappedLevel = abs(TransCoeffLevel [xC][yC]); int sign = TransCoeffLevel [xC][yC] < 0; int pred1 = std::max(topPos, leftPos); if (absMappedLevel == 1) TransCoeffLevelInvMapped [xC][yC]= pred1; else TransCoeffLevelInvMapped [xC][yC] = absMappedLevel - (absMappedLevel <= pred1); TransCoeffLevelInvMapped [xC][yC] = sign ? -dst[xC][yC] : dst[xC][yC]; } } } }
[0112]
[0135] Consistent with some embodiments of the present disclosure, based on the mapped value, a Rice parameter can be derived, which is different from VVC6 where the Rice parameter is derived based on the actual level value. Assuming that the array TransCoeffLevel [xC][yC] is the mapped level value for the TB of a given color component at location (xC,yC), the variable locSumAbs can be derived based on the following pseudo-code. locSumAbs = 0 AbsLevel [xC][yC] = abs(TransCoeffLevel[xC][yC]) if( xC > 0 ) locSumAbs += AbsLevel[ xC - 1 ][ yC ] if( yC > 0 ) locSumAbs += AbsLevel[ xC ][ yC - 1 ] locSumAbs = Clip3( 0, 31, locSumAbs )
[0113]
[0136] In accordance with some embodiments of the present disclosure, FIG. 20 shows Table 12, which is an example of a look-up table for selecting Rice parameters according to some embodiments of the present disclosure. In some disclosed embodiments, the value of locSumAbs can be adjusted based on a predefined offset value. In some embodiments, the offset value is calculated from offline training. The following pseudo-code example shows that the offset value is 2. locSumAbs = 0 offset = 2; AbsLevel [xC][yC] = abs(TransCoeffLevel[xC][yC]) if( xC > 0 ) locSumAbs += AbsLevel[ xC - 1 ][ yC ] if( yC > 0 ) locSumAbs += AbsLevel[ xC ][ yC - 1 ] locSumAbs -= offset locSumAbs = Clip3( 0, 31, locSumAbs )
[0114]
[0137] In accordance with some embodiments of the present disclosure, FIGS. 21-22 show flowcharts of process examples 2100-2200 for video processing according to some embodiments of the present disclosure. In some embodiments, processes 2100-2200 can be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device for video processing (e.g., device 400).
[0115]
[0138] As an example, FIG. 21 shows a flowchart of a process example 2100 for video processing according to some embodiments of the present disclosure. In step 2102, a codec (e.g., the encoder of FIGS. 2A - 2B) can determine to skip a conversion process for a prediction residual based on either the maximum value of the dimensions of the luma samples of the prediction block or the maximum value of the dimensions of the prediction block. The conversion process may be the conversion stage 212 of FIGS. 2A - 2B. The prediction residual may be the residual BPU 210 of FIGS. 2A - 2B. The conversion block may be a block included in the prediction data 206 of FIGS. 2A - 2B such as a conversion block (e.g., any of the conversion blocks shown in FIGS. 11 - 13). The dimensions of the prediction block may include the height or the width.
[0116]
[0139] In some embodiments, the codec may determine to skip the conversion process for the prediction residual by determining to skip the conversion process based on a determination that the dimensions of the prediction block do not exceed a threshold. In some embodiments, the threshold may be MaxTbSizeY as shown and described in relation to equations (2) - (3). The threshold may have a maximum value equal to either the maximum value of the dimensions of the luma samples (e.g., 32, 64, or any number) or the maximum value of the dimensions of the prediction block (e.g., 32, 64, or any number). In some embodiments, the maximum value of the dimensions of the luma samples or the maximum value of the dimensions of the prediction block may be a dynamic value (e.g., not constant).
[0117]
[0140] In some embodiments, the threshold is equal to the maximum value of the dimensions of the luma samples indicating the luminance information of the prediction block. In some embodiments, the maximum value of the threshold is 64. In some embodiments, the maximum value of the threshold is 32. In some embodiments, the minimum value of the threshold is 4. In some embodiments, the threshold may be equal to the maximum value of the dimensions of the prediction block (e.g., MaxTsSize as shown and described in equation (1)) for which the conversion process is allowed to be performed.
[0118]
[0141] In some embodiments, the maximum value of the threshold is determined based on at least a first parameter of the first parameter set. For example, the first parameter set may be a sequence parameter set (SPS). In some embodiments, the value of the first parameter is 0 or 1. For example, the first parameter may be sps_max_luma_transform_size_64_flag as shown and described in Table 5 of FIG. 9. In some embodiments, the threshold can be determined based on the value of the first parameter. For example, when the first parameter can be sps_max_luma_transform_size_64_flag and the threshold is MaxTbSizeY, MaxTbSizeY may be equal to 64 when sps_max_luma_transform_size_64_flag is equal to 1. When sps_max_luma_transform_size_64_flag is equal to 0, MaxTbSizeY is equal to 32.
[0119]
[0142] In some embodiments, the maximum value of the threshold can be determined based on at least a first parameter of the first parameter set. In some embodiments, the threshold can be determined based on the value of a second parameter of the second parameter set. In some embodiments, the second parameter set is a sequence parameter set (SPS). In some embodiments, the second parameter set is a picture parameter set (PPS). The second parameter may be log2_transform_skip_max_size_minus2 (as shown and described in connection with Table 1 of FIG. 5, for example). The value of the second parameter can be determined based on the value of the first parameter. In some embodiments, the value of the second parameter (e.g., log2_transform_skip_max_size_minus2) has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter (e.g., sps_max_luma_transform_size_64_flag). For example, log2_transform_skip_max_size_minus2 may be in the range of 0 to (3 + sps_max_luma_transform_size_64_flag). In some embodiments, the second parameter may have a first value in a first profile of the encoder (e.g., the main profile) and a second value in a second profile of the encoder (e.g., the extended profile), and the first value and the second value are different.
[0120]
[0143] Referring further to FIG. 21, at step 2104, the codec can generate residual coefficients for the prediction residual by performing at least one of a reversible compression process or a quantization process on the prediction residual. As described herein, the residual coefficients may be coefficients associated with a residual coding process. The quantization process may be the quantization stage 214 of FIGS. 2A - 2B. The reversible compression process may include generating residual coefficients using coefficient groups (CG). For example, the coefficient groups may be non - overlapping. In some embodiments, the coefficient groups have a size of 4×4.
[0121]
[0144] In some embodiments, the codec can generate residual coefficients using a Multiple Transform Selection (MTS) scheme. For example, the codec can determine whether the size of the prediction block exceeds 32. If the size of the prediction block does not exceed 32, the codec can generate residual coefficients using the MTS scheme.
[0122]
[0145] In some embodiments, the codec can further determine a transform skip coefficient level for a coefficient group using either context encoding technology or bypass encoding technology. The codec can also determine a Rice parameter based on the transform skip coefficient level. The codec can further generate a bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter.
[0123]
[0146] In some embodiments, the codec can further map the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block to the left of the prediction block and a second value of a second residual coefficient of a second prediction block above the prediction block.
[0124]
[0147] In some embodiments, the codec uses either context encoding technology or bypass encoding technology to determine a transform skip coefficient level for a coefficient group, and based on a first value of a first residual coefficient of a first prediction block to the left of the prediction block and a second value of a second residual coefficient of a second prediction block above the prediction block, maps the transform skip coefficient level to a modified transform skip coefficient level, generates a context model for the context encoding technology based on the modified transform skip coefficient level, determines a Rice parameter based on the modified transform skip coefficient level, generates residual coefficients using the coefficient group, and generates a bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter.
[0125]
[0148] Referring further to FIG. 21, at step 2106, the codec can generate a bitstream by entropy encoding at least the residual coefficients. The bitstream may be the video bitstream 228 of FIGS. 2A-2B.
[0126]
[0149] FIG. 22 shows a flowchart of another process example 2200 for video processing according to some embodiments of the present disclosure. For example, process 2200 may be performed by the decoder of FIGS. 3A-3B.
[0127]
[0150] As shown in FIG. 22, at step 2202, the decoder receives a bitstream including encoding information of a video sequence. The bitstream includes a sequence parameter set (SPS) of the video sequence.
[0128]
[0151] In step 2204, the decoder determines the maximum transform size of the prediction block based on the parameters of the sequence parameter set (SPS) of the video sequence. The prediction block may be a block included in the prediction data 206 of FIGS. 2A-2B, such as a transform block (e.g., any of the transform blocks shown in FIGS. 11-13). In some embodiments, the maximum transform size may correspond to the maximum value of the dimensions of the luma samples of the prediction block or the maximum value of the dimensions of the prediction block. The dimensions of the prediction block may include height or width. A detailed method for determining the maximum transform size based on the parameters of the SPS is described above in connection with FIGS. 5-10.
[0129]
[0152] In step 2206, the decoder determines to skip the transform process for the prediction residual of the prediction block based on the maximum transform size. The transform process may be the transform stage 212 of FIGS. 2A-2B.
[0130]
[0153] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. General forms of the non-transitory medium include, for example, floppy (registered trademark) disks, flexible disks, hard disks, solid state drives, magnetic tapes, or other magnetic data storage media, CD-ROMs, other optical data storage media, any physical medium having a pattern of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM or other flash memories, NVRAM, caches, registers, other memory chips or cartridges, and the networked versions described above. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0131]
[0154] Embodiments can be further described using the following clauses. 1. Determining to skip a conversion process for a prediction residual based on a maximum conversion size of a prediction block; signaling the maximum conversion size in a sequence parameter set (SPS); and a video processing method including the above. 2. The determination to skip the conversion process for the prediction residual includes determining to skip the conversion process based on a determination that the dimensions of the prediction block do not exceed a threshold value, where the threshold value is either the maximum value of the dimensions of the luma samples of the prediction block, or the maximum value of the dimensions of the prediction block The method according to clause 1, having a maximum value equal to one of the above. 3. The method according to clause 2, wherein either the maximum value of the dimensions of the luma samples or the maximum value of the dimensions of the prediction block is a dynamic value. 4. The method according to any one of the preceding clauses, further including determining to skip the conversion process further based on a parameter indicating a conversion skip mode. 5. The method according to clause 2, wherein the dimensions of the prediction block include height or width. 6. The method according to clause 2, wherein the maximum value of the threshold is determined based on at least a first parameter of a first parameter set. 7. The method according to clause 6, wherein the first parameter set is a sequence parameter set (SPS). 8. The method according to any one of clauses 6 - 7, wherein the value of the first parameter is 0 or 1. 9. The method according to any one of clauses 2 - 8, wherein the maximum value of the threshold is 64. 10. The method according to any one of clauses 2 - 8, wherein the maximum value of the threshold is 32. 11. The method according to any one of clauses 2 - 10, wherein the maximum value of the threshold is determined based on at least a first parameter of a first parameter set and a third parameter of the first parameter set. 12. The method according to any one of clauses 2 - 11, wherein the minimum value of the threshold is 4. 13. The method according to any one of clauses 2 - 12, further including determining to skip the conversion process based on a parameter indicating a conversion skip mode. 13. The method according to any one of clauses 2 to 12, wherein the threshold value is equal to the maximum value of the dimensions of the luma samples indicating the luminance information of the prediction block. 14. The method according to any one of clauses 6 to 13, wherein the maximum value of the threshold value is determined based on the value of the second parameter of the second parameter set, and the value of the second parameter is determined based on the value of the first parameter. 15. The method according to clause 14, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter. 16. The method according to clause 14, wherein the second parameter has a first value in the first profile of the encoder and a second value in the second profile of the encoder, and the first value and the second value are different. 17. The method according to any one of clauses 14 to 16, wherein the second parameter set is the SPS. 18. The method according to any one of clauses 14 to 16, wherein the second parameter set is the picture parameter set (PPS). 19. The method according to any one of clauses 2 to 12, wherein the threshold value is equal to the maximum value of the dimensions of the prediction block allowed to perform the conversion process. 20. The method according to clause 19, wherein the threshold value is determined based on the value of the first parameter. 21. The method according to any one of the preceding clauses, further comprising generating residual coefficients for the prediction block using a multiple transform selection (MTS) scheme. 22. Determining whether the dimension of the prediction block exceeds 32, Generating residual coefficients using the MTS scheme based on the determination that the dimension of the prediction block does not exceed 32, The method according to clause 21, further comprising. 23. Determining whether the dimension of the prediction block exceeds the threshold value, Performing block differential pulse code modulation (BDPCM) on the prediction residual before generating the residual coefficients for the prediction block based on the determination that the dimension of the prediction block does not exceed the threshold value. The method according to any one of clauses 2 to 22, further comprising 24. Further comprising generating a residual coefficient regarding the prediction residual by performing a reversible compression process on the prediction residual, the reversible compression process including generating the residual coefficient using a coefficient group, the coefficient group being non-overlapping, the method according to any one of the preceding clauses. 25. The method according to clause 24, wherein the coefficient group has a size of 4×4. 26. Determining a transform skip coefficient level regarding the coefficient group using one of context encoding technology or bypass encoding technology; Determining a Rice parameter based on the transform skip coefficient level; Generating a bitstream by entropy encoding at least one of the coefficient group, the transform skip coefficient level, or the Rice parameter; The method according to any one of clauses 24 to 25, further comprising 27. Further comprising mapping the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block to the left of the prediction block and a second value of a second residual coefficient of a second prediction block above the prediction block, the method according to any one of clauses 24 to 26. 28. Determining a transform skip coefficient level regarding the coefficient group using one of context encoding technology or bypass encoding technology; Mapping the transform skip coefficient level to a modified transform skip coefficient level based on a first value of a first residual coefficient of a first prediction block to the left of the prediction block and a second value of a second residual coefficient of a second prediction block above the prediction block; Generating a context model for the context encoding technology based on the modified transform skip coefficient level; Determining a Rice parameter based on the modified transform skip coefficient level; Generating a residual coefficient using the coefficient group; generating a bitstream by entropy encoding at least one of a coefficient group, a transform skip coefficient level, or a Rice parameter; The method according to any one of clauses 24 to 26, further comprising . 29. The method according to clause 28, further comprising mapping a transform skip coefficient level to a modified transform skip coefficient level after performing a quantization process and during generation of residual coefficients. 30. The method according to clause 28, further comprising mapping a transform skip coefficient level to a modified transform skip coefficient level after performing a quantization process and before generation of residual coefficients. 31. Determining a Rice parameter The method according to any one of clauses 28 to 30, comprising determining a Rice parameter based on a modified transform skip coefficient level of a color component of a prediction block. 32. The method according to clause 31, wherein a modified transform skip coefficient level of a color component is offset by a predetermined offset value. 33. The method according to clause 32, wherein the predetermined offset value is determined using a machine learning model in an offline training process. 34. Generating residual coefficients The method according to any one of clauses 23 to 33, comprising performing at least one of a reversible compression process or BDPCM on a prediction residual using diagonal scanning, wherein a maximum size of a prediction block for performing diagonal scanning is 64. 35. Generating residual coefficients Based on a determination that a dimension of a prediction block exceeds 32, dividing the prediction block into a plurality of sub-blocks in that dimension; For each specific sub-block of a plurality of sub-blocks, performing at least one of a reversible compression process or BDPCM on the prediction residual associated with the specific sub-block using diagonal scanning, wherein the parameters and output results of each of the reversible compression process or BDPCM associated with the plurality of sub-blocks are independent, and The method according to any one of clauses 23 to 33, including 36. The method according to clause 35, further comprising dividing the prediction block into a plurality of sub-blocks in two dimensions based on a determination that two dimensions of the prediction block exceed 32. 37. The method according to any one of clauses 35 to 36, wherein the parameters and output results of each of the reversible compression process or BDPCM associated with the plurality of sub-blocks include at least one of a context model associated with context encoding technology, a Rice parameter, or a maximum number of context encoding bins associated with context encoding technology. 38. The method according to any one of clauses 34 to 37, wherein the unit of diagonal scanning is a coefficient group. 39. The method according to clause 38, further comprising setting a first indicator parameter indicating the value of the coefficients of the coefficient group for each coefficient group of a specific sub-block. 40. The method according to clause 38, further comprising setting a second indicator parameter indicating the values of all coefficient groups of a specific sub-block for each specific sub-block of the plurality of sub-blocks. 41. Setting a first indicator parameter indicating the value of the coefficients of the coefficient group for each coefficient group of a specific sub-block; and Based on a determination that the first indicator parameters of all coefficient groups before the last coefficient group of the specific sub-block are zero, setting the first indicator parameter of the last coefficient group to 1; and The method according to clause 40, further comprising 42. Generating residual coefficients Generating residual coefficients by performing a reversible compression process on prediction residuals based on a parameter indicating a reversible symbolization mode, including generating, where the maximum value of the dimension of the luma samples is 64, the method according to any one of clauses 24 to 41. 43. Receiving a video picture, Dividing the video picture into a plurality of blocks, Generating a prediction block by performing either intra prediction or inter prediction on the block, Generating a prediction residual by subtracting the prediction block from the block, The method according to any one of the preceding clauses, further comprising. 44. A memory configured to store instructions, Determining to skip a transform process on a prediction residual based on a maximum transform size of the prediction block, Signaling the maximum transform size in a sequence parameter set (SPS), A processor configured to execute instructions to perform, An apparatus, comprising. 45. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of the apparatus to cause the apparatus to perform a method, the method comprising Determining to skip a transform process on a prediction residual based on a maximum transform size of the prediction block, Signaling the maximum transform size in a sequence parameter set (SPS), A non-transitory computer-readable medium, comprising. 46. Receiving a bitstream of a video sequence, Determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, Determining to skip a transform process on a prediction residual of the prediction block based on the maximum transform size, A video processing method, comprising. Determining to skip a conversion process for a prediction residual, includes determining to skip the conversion process in response to a determination that a dimension of a prediction block does not exceed a threshold value, the threshold value being a maximum value of dimensions of luma samples of the prediction block, or a maximum value of dimensions of the prediction block having a maximum value equal to one of them, the method according to clause 46. 48. The method according to clause 47, wherein a dimension of the prediction block includes height or width. 49. The method according to clause 47, wherein a maximum value of the threshold value is determined based on at least a first parameter of the SPS. 50. The method according to clause 49, wherein a value of the first parameter is 0 or 1. 51. The method according to any one of clauses 47 to 50, wherein a maximum value of the threshold value is 64. 52. The method according to any one of clauses 47 to 50, wherein a maximum value of the threshold value is 32. 53. The method according to any one of clauses 47 to 52, wherein a maximum value of the threshold value is determined based on at least a first parameter of the SPS and a third parameter of the SPS. 54. The method according to any one of clauses 47 to 53, wherein a minimum value of the threshold value is 4. 55. The method according to any one of clauses 47 to 54, wherein the threshold value is equal to a maximum value of dimensions of luma samples indicating luminance information of the prediction block. 56. The method according to any one of clauses 49 to 54, wherein a maximum value of the threshold value is determined based on a value of a second parameter of a second parameter set, and the value of the second parameter is determined based on the value of the first parameter. 57. The method according to clause 56, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to a sum of 3 and the value of the first parameter. 58. The method according to clause 56, wherein the second parameter has a first value in a first profile of the encoder and a second value in a second profile of the encoder, and the first value and the second value are different. 59. The method according to any one of clauses 56 to 58, wherein the second parameter set is an SPS. 60. The method according to any one of clauses 56 to 58, wherein the second parameter set is a picture parameter set (PPS). 61. A memory configured to store instructions, receiving a bitstream of a video sequence, determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, determining to skip a transform process for a prediction residual of the prediction block based on the maximum transform size, and a processor configured to execute instructions to perform An apparatus comprising. 62. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of an apparatus to cause the apparatus to perform a method, the method comprising receiving a bitstream of a video sequence, determining a maximum transform size of a prediction block based on a sequence parameter set (SPS) of the video sequence, determining to skip a transform process for a prediction residual of the prediction block based on the maximum transform size, A non-transitory computer-readable medium comprising.
[0132]
[0155] Relational terms such as "first" and "second" in this specification are used only to distinguish one entity or action from another entity or action, and it should be noted that they do not require or imply an actual relationship or order between these entities or actions. Also, the words "comprising", "having", "containing", and "including", and other similar forms, are equivalent in meaning, and one or more terms following any one of these words are intended to be in an open-ended form in that they are not an exhaustive listing of such one or more terms, or are not limited to only the one or more terms listed.
[0133]
[0156] In this specification, unless otherwise specified, the term "or" encompasses all possible combinations as long as they are not infeasible. For example, if a component is described as being able to include A or B, then unless otherwise specifically described or infeasible, the component can include A, or B, or A and B. As a second example, if a component is described as being able to include A, B, or C, then unless otherwise specified or infeasible, the component can include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0134]
[0157] It is understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. When implemented by software, it can be stored in the above computer-readable medium. The software can perform the disclosed method when executed by a processor. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that a plurality of the above modules / units can be integrated as one module / unit, and each of the above modules / units can be further divided into a plurality of sub-modules / sub-units.
[0135]
[0158] In the foregoing specification, embodiments have been described with respect to numerous specific details that may vary according to embodiments. Specific adaptations and modifications of the described embodiments can be made. Considering the specification and implementation of the invention disclosed herein, other embodiments may be apparent to those skilled in the art. The above specification and examples are intended to be regarded as merely illustrative, and the true scope and spirit of the present invention are indicated by the following claims. Also, the order of steps shown in the drawings is merely for the purpose of explanation and is not intended to be limited to any particular order of steps. Therefore, those skilled in the art can understand that these steps can be performed in a different order while implementing the same method.
[0136]
[0159] Embodiment examples have been disclosed in the drawings and the specification. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are used only in a general and illustrative sense and are not intended to be limiting.
Claims
Claim 1 Determining to skip a transformation process for a prediction residual based on a maximum transformation size of a prediction block, wherein determining to skip the transformation process for the prediction residual includes determining to skip the transformation process based on a determination that dimensions of the prediction block do not exceed a threshold value, and the threshold value has a maximum value equal to a maximum value of dimensions of luma samples of the prediction block or a maximum value of the dimensions of the prediction block, and Encoding a syntax element indicating the maximum transformation size in a sequence parameter set (SPS), and A video processing method comprising. Claim 2 The method according to claim 1, further comprising determining to skip the transformation process further based on a parameter indicating a transformation skip mode. Claim 3 The method according to claim 1, wherein one of the maximum value of the dimensions of the luma samples or the maximum value of the dimensions of the prediction block is a dynamic value. Claim 4 The method according to claim 1, wherein the dimensions of the prediction block include height or width. Claim 5 The method according to claim 1, wherein the maximum value of the threshold is 64. Claim 6 The method according to claim 1, wherein the maximum value of the threshold is 32. Claim 7 The method according to claim 1, wherein the minimum value of the threshold is 4. Claim 8 The method according to claim 1, wherein the threshold is equal to the maximum value of the dimensions of the luma samples indicating luminance information of the prediction block. Claim 9 The method according to claim 1, wherein the maximum value of the threshold is determined based on at least a first parameter of a first parameter set. Claim 10 The method according to claim 9, wherein the first parameter set is a sequence parameter set (SPS). Claim 11 The method according to claim 9, wherein the value of the first parameter is 0 or 1. Claim 12 The method according to claim 9, wherein the maximum value of the threshold is determined based on at least the first parameter of the first parameter set and a third parameter of the first parameter set. Claim 13 The method according to claim 9, wherein the maximum value of the threshold is determined based on a value of a second parameter of a second parameter set, and the value of the second parameter is determined based on the value of the first parameter. Claim 14 The method according to claim 13, wherein the value of the second parameter has a minimum value of 0 and a maximum value equal to the sum of 3 and the value of the first parameter.
15. The method according to claim 13, wherein the second parameter has a first value in a first profile of an encoder and a second value in a second profile of the encoder, and the first value and the second value are different.
16. The method according to claim 13, wherein the second parameter set is the SPS.
17. The method according to claim 13, wherein the second parameter set is a picture parameter set (PPS).
18. A memory configured to store instructions, and a processor, wherein the processor is configured to determine to skip a conversion process for a prediction residual based on a maximum transform size of a prediction block, encode a syntax element indicating the maximum transform size in a sequence parameter set (SPS), and execute the instructions to cause the device to perform, determining to skip the conversion process for the prediction residual includes determining to skip the conversion process based on a determination that the dimensions of the prediction block do not exceed a threshold, the threshold having a maximum value equal to one of a maximum value of dimensions of luma samples of the prediction block or a maximum value of the dimensions of the prediction block, device.
19. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method including determining to skip a conversion process for a prediction residual based on a maximum transform size of a prediction block, encoding a syntax element indicating the maximum transform size in a sequence parameter set (SPS), wherein determining to skip the conversion process for the prediction residual includes determining to skip the conversion process based on a determination that the dimensions of the prediction block do not exceed a threshold, the threshold having a maximum value equal to one of a maximum value of dimensions of luma samples of the prediction block or a maximum value of the dimensions of the prediction block, non-transitory computer-readable medium.
Citation Information
Patent Citations
Video coding method and apparatus using transform skip flag
JP2021513755A
Method and apparatus for encoding / decoding image
US20160269730A1