Method for signaling maximum transform size and residual coding

By using signal notification maximum transform size and residual coding methods, the video coding process is optimized, solving the problem of insufficient coding efficiency in existing technologies. This achieves more efficient video compression and decompression, and improves the performance of the encoder and decoder.

CN115552900BActive Publication Date: 2025-12-23ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180015506.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-01-29
Publication Date
2025-12-23
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Existing video coding standards have not fully utilized the maximum transform size of signal notification and residual coding methods in efficient video coding techniques, resulting in coding efficiency that needs to be improved.

Method used

The value of the coded tree block size is determined by receiving the bit stream, and based on this value, a flag indicating the maximum transform size and a first flag indicating that the transform skip mode is enabled are signaled. This determines whether to notify the second flag of the residual coding method and optimizes the coding process.

Benefits of technology

It improves the compression efficiency of video encoding, reduces the need for storage space and transmission bandwidth, and enhances the parallel processing capability and fault tolerance of encoders and decoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115552900B_ABST
    Figure CN115552900B_ABST
Patent Text Reader

Abstract

SUMMARY The present invention provides systems and methods for signaling a maximum transform size and residual coding methods. According to certain disclosed embodiments, the method comprises: receiving a bitstream comprising a set of pictures; determining, from the received bitstream, a value of a coding tree block size; and based on the value of the coding tree block size, determining whether to signal a flag indicating a maximum transform size of luma samples.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 980,117, filed on February 21, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to video processing, and more specifically, to methods and apparatus for signaling maximum transform size and residual coding methods. Background Technology

[0004] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common being prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Universal Video Coding (VVC / H.266) standard, and the AVS standard, specifying particular video coding formats. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is becoming increasingly higher. Summary of the Invention

[0005] Embodiments of this disclosure provide a method and apparatus for signaling a maximum transform size. In some exemplary embodiments, the method includes: receiving a bitstream comprising a set of images; determining a value of a coding tree block size based on the received bitstream; and determining, based on the value of the coding tree block size, whether to signal a flag indicating the maximum transform size of a luminance sample.

[0006] The apparatus may include a memory storing an instruction set; and one or more processors configured to execute the instruction set to cause the apparatus to: receive a bitstream comprising a set of images; determine a value for a coded tree block size based on the received bitstream; and determine, based on the value of the coded tree block size, whether to signal a flag for a maximum transform size for luminance samples.

[0007] Embodiments of the present disclosure also provide a non-transitory computer readable medium storing a set of instructions, which can be executed by at least one processor of a computer to cause the computer to perform a method for signaling a maximum transform size, the method comprising: receiving a bitstream comprising a set of pictures, determining a value of a coding tree block size from the received bitstream, and determining whether to signal a flag indicating a maximum transform size for luma samples based on the value of the coding tree block size.

[0008] Embodiments of the present disclosure also provide a method and apparatus for signaling a residual coding method. In some example embodiments, the method comprises: receiving a bitstream comprising a set of pictures; determining a value of a first flag from the received bitstream, the value of the first flag indicating whether a transform skip mode is enabled; and determining whether to signal a second flag indicating a residual coding method based on the value of the first flag.

[0009] The apparatus can comprise a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the apparatus to perform: receiving a bitstream comprising a set of pictures; determining a value of a first flag from the received bitstream, the value of the first flag indicating whether a transform skip mode is enabled; and determining whether to signal a second flag indicating a residual coding method based on the value of the first flag.

[0010] Embodiments of the present disclosure also provide a non-transitory computer readable medium storing a set of instructions, which can be executed by at least one processor of a computer to cause the computer to perform a method for signaling a residual coding method, the method comprising: receiving a bitstream comprising a set of pictures; determining a value of a first flag from the received bitstream, the value of the first flag indicating whether a transform skip mode is enabled; and determining whether to signal a second flag indicating a residual coding method based on the value of the first flag. BRIEF DESCRIPTION OF DRAWINGS

[0011] Embodiments of the present disclosure and various aspects thereof are illustrated by way of example in the following detailed description and the accompanying drawings, in which the various features are not drawn to scale.

[0012] Figure 1 is a structural diagram of an example video sequence consistent with some embodiments of the present disclosure.

[0013] Figure 2A is a schematic diagram illustrating an example encoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0014] Figure 2Bis a schematic diagram illustrating another example encoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0015] Figure 3A is a schematic diagram illustrating an example decoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0016] Figure 3B is a schematic diagram illustrating another example decoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0017] Figure 4 is a block diagram of an example apparatus for encoding or decoding video consistent with some embodiments of the present disclosure.

[0018] Figure 5A is an example method for signaling a maximum transform size consistent with some embodiments of the present disclosure.

[0019] Figure 5B is an example method for signaling a maximum transform size consistent with some embodiments of the present disclosure.

[0020] Figure 6A is an example method for signaling a residual coding method consistent with some embodiments of the present disclosure.

[0021] Figure 6B is an example method for signaling a residual coding method consistent with some embodiments of the present disclosure. DETAILED DESCRIPTION

[0022] Reference will now be made in detail to the example embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers represent the same or similar elements between the several figures. The implementations set forth in the following description of example embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present disclosure as described in the appended claims. Specific aspects of the present disclosure are described in more detail below. To the extent not explicitly recited in the claims, terms and / or definitions provided herein control.

[0023] The Joint Video Expert Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Motion Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality using half the bandwidth of HEVC / H.265.

[0024] To achieve the same subjective quality as HEVC / H.265 using half of the bandwidth, JVET has been developing techniques beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM achieves higher coding performance than HEVC.

[0025] The VVC standard is recently developed and continues to include more coding techniques that provide better compression performance. VVC is based on the hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0026] A video is a set of still images (or “frames”) arranged in a temporal order to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in a temporal order, and video playback devices (e.g., televisions, computers, smartphones, tablet computers, video players, or any end-user terminal with display functionality) can be used to display such images in a temporal sequence. Moreover, in some applications, video capture devices can transmit captured videos to video playback devices (e.g., computers with monitors) in real time, for example, for surveillance, conferencing, or live broadcasting.

[0027] To reduce the storage space and transmission bandwidth required for such applications, videos can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or special-purpose hardware. A module for compression is often referred to as an “encoder,” and a module for decompression is often referred to as a “decoder.” An encoder and a decoder can be collectively referred to as a “codec.” An encoder and a decoder can be implemented in any of a variety of suitable hardware, software, or a combination thereof. For example, a hardware implementation of an encoder and a decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic or any combination thereof. A software implementation of an encoder and a decoder can include program code, computer-executable instructions, firmware, or any suitable computer- implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress a video from a first encoding standard and recompress the decompressed video using a second encoding standard, in which case the codec can be referred to as a “transcoder.”

[0028] A video coding process can identify and preserve useful information for reconstructing an image and ignore unimportant reconstruction information. If the ignored unimportant information cannot be completely reconstructed, such a coding process can be referred to as "lossy." Otherwise, it can be referred to as "lossless." Most coding processes are lossy, which is a tradeoff for reducing the required storage space and transmission bandwidth.

[0029] Useful information of an encoded image (referred to as "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include position changes, intensity changes, or color changes of pixels, with position changes being the most concerned. Position changes of a group of pixels representing an object can reflect the motion of the object between the reference image and the current image.

[0030] An image that is encoded without reference to another image (i.e., it is its own reference image) is referred to as an "I-image." An image that is encoded using a previous image as a reference image is referred to as a "P-image," and an image that is encoded using a previous image and a future image as reference images is referred to as a "B-image" (the reference is "bidirectional").

[0031] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination of both (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capturing device (e.g., a camera), a repository containing previously captured video archives (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.

[0032] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a time line, including images 102, 104, 106, and 108. The images 102-106 are consecutive, with more images in between image 106 and 108. In Figure 1 particular, image 102 is an I-image, whose reference image is image 102 itself. Image 104 is a P-image, whose reference image is image 102, as shown by the arrow. Image 106 is a B-image, whose reference images are images 104 and 108, as shown by the arrows. In some embodiments, a reference image of an image (e.g., image 104) can not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It is noted that the reference images of images 102-106 are merely examples, and the present disclosure is not limited to the embodiments of reference images as shown. Figure 1 particular, image 102 is an I-image, whose reference image is image 102 itself. Image 104 is a P-image, whose reference image is image 102, as shown by the arrow. Image 106 is a B-image, whose reference images are images 104 and 108, as shown by the arrows. In some embodiments, a reference image of an image (e.g., image 104) can not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It is noted that the reference images of images 102-106 are merely examples, and the present disclosure is not limited to the embodiments of reference images as shown.

[0033] Typically, due to the computational complexity of the coding task, a video codec does not encode or decode an entire image at once. Instead, they can partition the image into basic segments and encode or decode the image segments one by one. In this disclosure, such basic segments are referred to as basic processing units (“BPUs”). For example, Figure 1 The structure 110 in FIG. 1 illustrates an example structure of an image (e.g., any of the images 102-108) of the video sequence 100. In the structure 110, the image is partitioned into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, a basic processing unit can be referred to as a “macroblock” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have a variable size in pixels, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size. The size and shape of the basic processing units can be selected for an image based on a balance of coding efficiency and level of detail to be maintained in the basic processing units. A CTU is the largest unit of blocks and can include up to 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). A CTU can be further partitioned into coding units (CUs) using quad-tree, binary-tree, ternary-tree, or a combination thereof.

[0034] A basic processing unit can be a logical unit that can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color image can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as “coding tree blocks” (“CTBs”). Any operation performed on a basic processing unit can be repeated on each of its luma and chroma components.

[0035] Video coding has multiple stages of operations, examples of which are shown in Figures 2A-2B and Figures 3A-3BAs shown. For each stage, the size of the basic processing unit can still be too large for processing, and thus can be further divided into segments, referred to in this disclosure as“basic processing subunits.” In some embodiments, the basic processing subunits can be referred to as“blocks” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as“coding units” (“CUs”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits can have the same size as the basic processing unit or have a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which can include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operations performed on the basic processing subunit can be repeated on each of its luma and chroma components. It should be noted that such division can be performed to further levels as needed for processing. It should also be noted that different stages can use different schemes to divide the basic processing unit.

[0036] For example, at the mode decision stage (examples of which are shown in Figure 2B , the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for a basic processing unit that can be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., as CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.

[0037] For another example, at the prediction stage (examples of which are shown in Figures 2A-2B , the encoder can perform the prediction operation at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits can still be too large for processing. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as“prediction blocks” or“PBs” in H.265 / HEVC or H.266 / VVC), at which level the prediction operation can be performed.

[0038] For another example, at the transform stage (examples of which are shown in Figures 2A-2BAs shown in FIG. 1, the encoder can perform a transform operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit can still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., referred to as “transform blocks” or “TBs” in H.265 / HEVC or H.266 / VVC), on which the transform operation can be performed. It is noted that the division scheme of the same basic processing subunit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and the transform blocks of the same CU can have different sizes and numbers.

[0039] In Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3x3 basic processing subunits, the boundaries of which are shown in dashed lines. Different basic processing units of the same picture can be divided into basic processing subunits in different schemes.

[0040] In some embodiments, to provide the capability of parallel processing and the fault tolerance capability for video encoding and decoding, a picture can be divided into regions for processing such that, for a region of the picture, the encoding or decoding process can not depend on information from any other region of the picture. In other words, each region of the picture can be processed individually. By doing so, the codec can process different regions of the picture in parallel, thereby improving the encoding efficiency. Furthermore, when the data of a region is corrupted in processing or lost in network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing the fault tolerance capability. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: “slices” and “tiles”. It is also noted that different pictures of the video sequence 100 can have different division schemes for dividing the pictures into regions.

[0041] For example, in Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3x3 basic processing subunits, the boundaries of which are shown in dashed lines. Different basic processing units of the same picture can be divided into basic processing subunits in different schemes. Figure 1 The basic processing units, basic processing subunits, and structure regions of the structure 110 are merely examples, and the disclosure does not limit embodiments thereof.

[0042] Figure 2A A schematic diagram of an exemplary encoding process 200A according to an embodiment of the disclosure is shown. For example, the encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoding process 200A can include the following steps. Figure 2AAs shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 in Figure 1 , the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in a temporal order. Similar to the structure 110 in Figure 1 , each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform the process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder can perform the process 200A in an iterative manner, where the encoder can encode a basic processing unit in one iteration of the process 200A. In some embodiments, the encoder can perform the process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0043] Referring to Figure 2A , the encoder can feed a basic processing unit of an original picture of the video sequence 202 (referred to as an “original BPU”) to the prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder can feed the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder can feed the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 to generate the video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the “forward path.” During the process 200A, after the quantization stage 214, the encoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A can be referred to as the “reconstruction path.” The reconstruction path can be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0044] The encoder can iteratively perform the process 200A to encode each original BPU of an original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of an original picture, the encoder can proceed to encode the next picture in the video sequence 202.

[0045] Referring to process 200A, an encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term “receive” can refer to any action that gets, obtains, retrieves, acquires, reads, accesses, or any action that is used to input data in any manner.

[0046] At a prediction stage 204, at a current iteration, the encoder can receive an original BPU and a prediction reference 224, and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting, from the prediction data 206 and the prediction reference 224, prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208.

[0047] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To account for these differences, upon generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the values (e.g., grayscale values or RGB values) of the corresponding pixels of the predicted BPU 208 from the values of the pixels of the original BPU. Each pixel of the residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. The prediction data 206 and the residual BPU 210 can have a smaller number of bits compared to the original BPU, but they can be used to reconstruct the original BPU without a noticeable quality degradation. Thus, the original BPU is compressed.

[0048] To further compress the residual BPU 210, at a transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional “base patterns.” Each base pattern is associated with a “transform coefficient.” The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent a component of the residual BPU 210 of a varying frequency (e.g., of luminance). None of the base patterns can be reproduced from any combination (e.g., linear combination) of any of the other base patterns. In other words, the decomposition can decompose the variations of the residual BPU 210 into the frequency domain. This decomposition is similar to a discrete Fourier transform of a function, where the base images are similar to the base functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the base functions.

[0049] Different transform algorithms can use different basis patterns. Various transform algorithms can be used at transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at transform stage 212 is invertible. That is, the encoder can recover the residual BPU 210 through an inverse operation of the transform, called “inverse transform.” For example, to recover a pixel of the residual BPU 210, the inverse transform can be multiplying the values of the corresponding pixels of the basis pattern by the respective correlation coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transform algorithm (and thus have the same basis pattern). Thus, the encoder can record only the transform coefficients from which the decoder can reconstruct the residual BPU 210 without receiving the basis pattern from the encoder. The transform coefficients can have fewer bits compared to the residual BPU 210, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.

[0050] The encoder can further compress the transform coefficients at quantization stage 214. During the transform process, different basis patterns can represent different frequencies of variation (e.g., frequencies of luminance variation). Because the human eye is generally better at recognizing low-frequency variations, the encoder can ignore information of high-frequency variations without causing noticeable quality degradation in decoding. For example, at quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value, called “quantization parameter,” and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the quantized transform coefficients 216 of zero value, whereby the transform coefficients are further compressed. This quantization process is also invertible, where the quantized transform coefficients 216 can be reconstructed to transform coefficients in an inverse operation of quantization, called “inverse quantization.”

[0051] Because the encoder ignores the remainder of the division in the rounding operation, quantization stage 214 can be lossy. Generally, quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer the number of bits required for the quantized transform coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.

[0052] At the binarization stage 226, the encoder can binarize the prediction data 206 and the quantized transform coefficients 216 using a binarization technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can binarize other information at the binarization stage 226, such as the prediction modes used at the prediction stage 204, parameters of the prediction operations, the type of transform at the transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. The encoder can use the output data of the binarization stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0053] Referring to the reconstruction path of the process 200A, at the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of the process 200A.

[0054] It should be noted that other variants of the process 200A can be used to encode the video sequence 202. In some embodiments, the stages of the process 200A can be performed by the encoder in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be split into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages of the process 200A. Figure 2A

[0055] Figure 2B A schematic diagram illustrating another example encoding process 200B is shown, in accordance with an embodiment of the disclosure. The process 200B can be a modification of the process 200A. For example, the process 200B can be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). In comparison to the process 200A, the forward path of the process 200B further includes a mode decision stage 230, and splits the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and the reconstruction path of the process 200B further additionally includes a loop filtering stage 232 and a buffer 234.

[0056] ​In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or “intra-prediction”) can use pixels from one or more already encoded neighboring BPU in the same image to predict a current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPU. Spatial prediction can reduce spatial redundancy inherent in images. Temporal prediction (e.g., inter-image prediction or “inter-prediction”) can use regions from one or more already encoded images to predict a current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce temporal redundancy inherent in images.

[0057] Referring to the process 200B, in the forward path, the encoder performs prediction operations at a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra-prediction. For an original BPU of an image being encoded, the prediction reference 224 can include one or more neighboring BPU in the same image that have been encoded (in the forward path) and reconstructed (in the reconstruction path). The encoder can generate a predicted BPU 208 by interpolating the neighboring BPU. Interpolation techniques can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at a pixel level, e.g., by interpolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPU used for interpolation can be located in various directions relative to the original BPU, e.g., in a vertical direction (e.g., at the top of the original BPU), a horizontal direction (e.g., at the left of the original BPU), a diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the neighboring BPU used, the size of the neighboring BPU used, parameters for interpolation, the direction of the neighboring BPU used relative to the original BPU, etc.

[0058] For another example, at the temporal prediction stage 2044, the encoder can perform inter prediction. For the original BPU of the current image, the prediction reference 224 can include one or more images (referred to as “reference images”) that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images can be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, the encoder can generate a reconstructed image as a reference image. The encoder can perform an operation of “motion estimation” to search for a matching region in a range (referred to as a “search window”) of the reference image. The location of the search window in the reference image can be determined based on the location of the original BPU in the current image. For example, the search window can be centered at a location in the reference image that has the same coordinates as the original BPU in the current image, and can extend outward by a predetermined distance. When the encoder identifies (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder can determine such a region as a matching region. The matching region can have a different size (e.g., smaller, equal, larger, or have a different shape) than the original BPU. Because the reference image and the current image are separated in time on a timeline (e.g., as shown in FIG. 1), the matching region can be considered to “move” to the location of the original BPU over time. The encoder can record the direction and distance of such motion as a “motion vector.” When multiple reference images are used (e.g., as in FIG. 1), the encoder can search for matching regions and determine their associated motion vectors for each reference image. In some embodiments, the encoder can assign weights to the pixel values of the matching regions of the respective matching reference images. Figure 1 Figure 1

[0059] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0060] To generate the predicted BPU 208, the encoder can perform an operation of “motion compensation.” Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference image according to the motion vector, where the encoder can predict the original BPU of the current image. When multiple reference images are used (e.g., as in FIG. 1), the encoder can move the matching region of each reference image according to its associated motion vector, where the encoder can predict the original BPU of the current image. Figure 1 ​​In some embodiments, the encoder can move the matching region of the reference image according to the individual motion vectors and average pixel values of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of the individual matching reference images, the encoder can add the weighted sum of the pixel values of the moved matching region.

[0061] In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in FIG. 1 is a uni-directional inter prediction image, where the reference image (i.e., image 102) precedes image 104. Bi-directional inter prediction can use one or more reference images in both temporal directions relative to the current image. For example, Figure 1 Image 106 in FIG. 1 is a bi-directional inter prediction image, where the reference images (i.e., images 104 and 108) precede image 104 in both temporal directions.

[0062] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stage 2044, at a mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, where the encoder can select the prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the corresponding predicted BPU 208 and prediction data 206.

[0063] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all the BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the in-loop filter stage 232. At this stage, the encoder can apply in-loop filters to the prediction reference 224 to reduce or eliminate the distortion (e.g., blockiness artifacts) introduced by inter prediction. The encoder can apply various in-loop filter techniques at the in-loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The in-loop filtered reference picture can be stored in the buffer 234 (or “decoded picture buffer”) for later use (e.g., as an inter prediction reference picture for future pictures of the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode parameters of the in-loop filters (e.g., in-loop filter strength) at the binary encoding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.

[0064] Figure 3A A schematic diagram illustrating an example decoding process 300A in accordance with embodiments of the present disclosure is shown. The process 300A can be a decompression process corresponding to the compression process 200A in Figure 2A some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in Figures 2A-2B some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in the process 200A in

[0065] As described above with respect to the process 200A in Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") to a binarization decoding stage 302, where the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of the process 300A.

[0066] The decoder can iteratively perform the process 300A to decode each encoded BPU of an encoded image and generate a prediction reference 224 for a next encoded BPU of the encoded image. After decoding all encoded BPUs of an encoded image, the decoder can output the image to a video stream 304 for display and continue decoding a next encoded image in the video bitstream 228.

[0067] At the binarization decoding stage 302, the decoder can perform an inverse operation of the binarization encoding technique used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information at the binarization decoding stage 302, such as prediction modes, parameters of prediction operations, transform types, parameters of quantization processes (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can depacketize the video bitstream 228 before feeding it to the binarization decoding stage 302.

[0068] Figure 3B A schematic diagram illustrating another example decoding process 300B according to embodiments of the disclosure is shown. The process 300B can be a modification of the process 300A. For example, the process 300B can be used by a decoder conforming to a hybrid video coding standard (e.g., the H.26x family of standards). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.

[0069] In process 300B, for a coded base processing unit (referred to as "current BPU") of a decoded coded picture (referred to as "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on what prediction mode is used by the encoder to code the current BPU. For example, if the current BPU is coded using intra prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, locations (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of interpolation, directions of the neighboring BPUs relative to the original BPU, and the like. For another example, if the current BPU is coded using inter prediction by the encoder, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating the inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, a number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions in the respective reference pictures, one or more motion vectors respectively associated with the matching regions, and the like.

[0070] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) at the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) at the temporal prediction stage 2044, details of performing such spatial or temporal prediction are described in Figure 2B , which will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208, to which the decoder can add the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A .

[0071] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction at the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If the current BPU is decoded using inter prediction at the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can feed the prediction reference 224 to the loop filter stage 232 as described in Figure 2BThe illustrated manner applies the loop filter to the prediction reference 224. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength).

[0072] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding a video according to embodiments of the present disclosure. As Figure 4 illustrated, the apparatus 400 can include a processor 402. When the processor 402 executes instructions as described herein, the apparatus 400 can become a special purpose machine for video encoding or decoding. The processor 402 can be any type of circuitry capable of manipulating or processing information. For example, the processor 402 can include any combination of central processing units (or “CPUs”), graphics processing units (or “GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), a field-programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), and the like. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as Figure 4 illustrated, the processor 402 can include multiple processors, including a processor 402a, a processor 402b, and a processor 402n.

[0073] The apparatus 400 can also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, and the like). For example, as Figure 4As shown, the stored data can include program instructions (e.g., for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 can include a high-speed random access memory or a nonvolatile memory. In some embodiments, memory 404 can include any combination of any quantity of random access memory (RAM), read only memory (ROM), optical disk, magnetic disk, hard drive, solid-state drive, flash drive, secure digital (SD) card, memory stick, compact flash (CF) card, and the like. Memory 404 can also be a group of memories grouped as a single logical component (not shown in FIG. 4). Figure 4

[0074] Bus 410 can be a communication device that transfers data between components within device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.

[0075] For ease of explanation and without causing ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as “data processing circuitry.” The data processing circuitry can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Moreover, the data processing circuitry can be a single separate module, or can be combined in whole or in part into any other component of device 400.

[0076] Device 400 can also include network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, and the like). In some embodiments, network interface 406 can include any combination of any quantity of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, and the like.

[0077] In some embodiments, optionally, device 400 can further include peripheral interface 408 to provide connection to one or more peripheral devices. As Figure 4 shown, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like. ​

[0078] It should be noted that a video codec (e.g., a codec that performs process 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in apparatus 400. For example, some or all stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of apparatus 400, such as a program instance that can be loaded into memory 404. For another example, some or all stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of apparatus 400, such as a special-purpose data processing circuit (e.g., FPGA, ASIC, NPU, etc.).

[0079] In a first reference embodiment, a sequence parameter set (SPS) syntax element sps_max_luma_transform_size_64_flag is signaled to specify the maximum transform size.

[0080] sps_max_luma_transform_size_64_flag equal to 1 specifies that the maximum transform size in luma samples is equal to 64. sps_max_luma_transform_size_64_flag equal to 0 specifies that the maximum transform size in luma samples is equal to 32. In the first reference embodiment, sps_max_luma_transform_size_64_flag is always signaled regardless of the value of the coding tree block size (CtbSizeY). However, if CtbSizeY is smaller than 64, sps_max_luma_transform_size_64_flag does not need to be signaled and can be inferred to be 0.

[0081] sps_max_luma_transform_size_64_flag, and can be inferred to be 0.

[0082] Further, in a second reference embodiment, a slice-level residual coding selection method can be used. The following are exemplary semantics for slice_ts_residual_coding_disabled_flag:

[0083]

[0084] In the second reference embodiment, slice_ts_residual_coding_disabled_flag is always signaled. However, since slice_ts_residual_coding_disabled_flag specifies the residual coding method of transform skip mode, if transform skip mode is disabled at the SPS level, slice_ts_residual_coding_disabled_flag does not need to be signaled in the slice header and can be inferred to be 0.

[0085] The present disclosure provides methods and apparatuses that reduce the above-mentioned coding redundancy.

[0086] Figure 5A is an exemplary method for signaling a maximum transform size consistent with some embodiments of the present disclosure. In some embodiments, method 500A can be performed by one or more software or hardware components of an encoder, a decoder, and an apparatus, such as apparatus 400 of Figure 4 , for example. For example, a processor, such as processor 402 of Figure 4 , can perform method 500A. In some embodiments, method 500A can be implemented by a computer program product, comprising computer executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of Figure 4 ). The method can comprise the following steps.

[0087] In step 501, a bitstream comprising a set of pictures is received. As described, a primary processing unit of a color picture can comprise a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size of the primary processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as “coding tree blocks” (“CTBs”). Any operations performed on a primary processing unit can be repeatedly performed on each of its luma and chroma components.

[0088] In step 503, a value of a coding tree block size is determined from the received bitstream. The value of the coding tree block size is signaled in the received bitstream. For example, the value of CtbSizeY in the exemplary semantics is determined.

[0089] In step 505, it is determined whether a flag indicating a maximum transform size of luma samples is signaled based on the value of the coding tree block size. For example, a sequence parameter set (SPS) syntax element sps max luma transform size 64 flag can be used to specify the maximum transform size. sps max luma transform size 64 flag equal to 1 specifies that the maximum transform size in luma samples is equal to 64. sps max luma transform size 64 flag equal to 0 specifies that the maximum transform size in luma samples is equal to 32.

[0090] In some embodiments, step 505 can include steps 505-1, 505-3, and 505-5 as shown in FIG. 5. Figure 5B

[0091] In step 505-1, it is determined from the received bitstream whether the value of the coding tree block size satisfies a first condition or satisfies a second condition.

[0092] In some embodiments, the first condition can be a value greater than 32, and the second condition can be a value less than or equal to 32. Satisfaction of the first condition can result in the flag indicating the maximum transform size being signaled. Otherwise, satisfaction of the second condition can result in a determination that the flag is not signaled.

[0093] In step 505-3, the flag is signaled in response to the value of the coding tree block size being greater than 32.

[0094] An example SPS syntax is described below in Table 1. In some example embodiments for signaling the maximum transform size, sps max luma transform size 64 flag is signaled only when CtbSizeY is greater than 32. Table 1 shows a portion of an example SPS syntax table for signaling the maximum transform size according to some disclosed embodiments. In Table 1, the row with “if (CtbSizeY > 32)” shows a change to the syntax for the first reference embodiment. As described, in the first reference embodiment, sps max luma transform size 64 flag is always signaled regardless of the value of the coding tree block size (CtbSizeY). In contrast, according to some embodiments of the disclosure, sps max luma transform size 64 flag is not always signaled. As shown in Table 1, the signaling of sps max luma transform size 64 flag is conditioned on CtbSizeY.

[0095] ​Table 1: Exemplary SPS syntax table for disclosed method of signaling maximum transform size

[0096]

[0097] In step 505-5, responsive to the value of the coding tree block size being less than or equal to 32, it is determined that the flag is not signaled.

[0098] In the signaled method consistent with Table 1, if CtbSizeY is not greater than 32, the value of sps_max_luma_transform_size_64_flag is inferred to be 0. The semantics of CtbSizeY is derived as shown below.

[0099]

[0100] In some embodiments, the first condition and the second condition can be different from the above cases. In the following example, the first condition can be a value not equal to 32, and the second condition can be a value equal to 32.

[0101] In an alternative step to step 505-3, responsive to the value of the coding tree block size not being equal to 32, the flag is signaled.

[0102] In an alternative step to step 505-5, responsive to the value of the coding tree block size being equal to 32, it is determined that the flag is not signaled.

[0103] An exemplary syntax description is as follows. The proposed syntax change shown in Table 1 can also be implemented using the condition "CtbSizeY!= 32" instead of "CtbSizeY > 32". In this case: if CtbSizeY is not equal to 32, sps_max_luma_transform_size_64_flag is signaled; if CtbSizeY is equal to 32, sps_max_luma_transform_size_64_flag is inferred to be 0.

[0104] An example syntax description is as follows. The proposed syntax change shown in Table 1 can also be implemented using the condition “CtbSizeY >= 64” instead of “CtbSizeY > 32”. In this case: if CtbSizeY is greater than or equal to 64, sps_max_luma_transform_size_64_flag is signaled; if CtbSizeY is less than 64, sps_max_luma_transform_size_64_flag is inferred to be 0. In other words, sps_max_luma_transform_size_64_flag needs to be signaled only when CtbSizeY is greater than 32. If CtbSizeY is not greater than 32, sps_max_luma_transform_size_64_flag is not signaled and inferred to be equal to 0. Alternatively, sps_max_luma_transform_size_64_flag is signaled only when CtbSizeY is greater than or equal to 64. If CtbSizeY is less than 64, sps_max_luma_transform_size_64_flag is not signaled and inferred to be equal to 0.

[0105] In some example embodiments for signaling the maximum transform size, the signaling of sps_max_luma_transform_size_64_flag is conditioned on sps_log2_ctu_size_minus5 instead of using CtbSizeY. As described above, sps_log2_ctu_size_minus5 plus 5 specifies the luma coding tree block size of each coding tree unit (CTU). A CTU can be 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). Table 2 below shows a portion of an example SPS syntax table for signaling the maximum transform size according to some disclosed embodiments. In Table 2, the row with "if (sps_log2_ctu_size_minus5 > 0)" shows the change to the syntax for the first reference embodiment. As described, in the first reference embodiment, sps_max_luma_transform_size_64_flag is always signaled regardless of the value of the coding tree block size (e.g., CtbSizeY). In contrast, according to some embodiments of the disclosure, sps_max_luma_transform_size_64_flag is not always signaled. As shown in Table 2, in some embodiments, the signaling of sps_max_luma_transform_size_64_flag is conditioned on sps_log2_ctu_size_minus5. As shown in Table 2, if sps_log2_ctu_size_minus5 is greater than 0, then sps_max_luma_transform_size_64_flag is signaled; if sps_log2_ctu_size_minus5 is 0, then the value of sps_max_luma_transform_size_64_flag is inferred to be 0.

[0106] In some embodiments, the value of the coding tree block size in step 501 can also include a value associated with the luma coding tree block size of each coding tree unit. This value (e.g., sps_log2_ctu_size_minus5) can be determined from the received bitstream.

[0107] In some embodiments, in an alternative step to step 505-1, it can be determined whether the value associated with the luma coding tree block size of each coding tree unit satisfies a first condition (e.g., greater than 0) or a second condition (e.g., equal to 0).

[0108] In some embodiments, in an alternative step to step 505-3, the flag (e.g., sps_max_luma_transform_size_64_flag) is signaled in response to a value associated with a size of a luma coding tree block of each coding tree unit satisfying a first condition (e.g., being greater than 0).

[0109] In some embodiments, in an alternative step to step 505-5, the flag (e.g., sps_max_luma_transform_size_64_flag) is determined not to be signaled in response to a value associated with a size of a luma coding tree block of each coding tree unit satisfying a second condition (e.g., being equal to 0).

[0110] Table 2: An exemplary SPS syntax table for the disclosed method for signaling a maximum transform size

[0111]

[0112] The above syntax changes shown in Table 2 can also be implemented using the condition “sps_log2_ctu_size_minus5!= 0”. In this case: if sps_log2_ctu_size_minus5 is not equal to 0, sps_max_luma_transform_size_64_flag is signaled; if sps_log2_ctu_size_minus5 is equal to 0, sps_max_luma_transform_size_64_flag is inferred to be 0.

[0113] In some embodiments, the first condition can also be a value (e.g., sps_log2_ctu_size_minus5) that is not equal to 0, as shown in the above example. In an alternative step to step 505-3, the flag (e.g., sps_max_luma_transform_size_64_flag) is signaled in response to a value associated with a size of a luma coding tree block of each coding tree unit satisfying the first condition (e.g., sps_log2_ctu_size_minus5!= 0).

[0114] FIG. 6 is an exemplary method for signaling a residual coding method, consistent with some embodiments of the present disclosure. In some embodiments, method 600A can be performed by an encoder, a decoder, and one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). Figure 4 In some embodiments, method 600A can be performed by an encoder, a decoder, and one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). Figure 4The method 600A can be performed by the processor 402 of the apparatus 400. In some embodiments, the method 600A can be implemented by a computer program product, comprising computer executable instructions, such as program code, executable by Figure 4 computing devices (e.g., a computer, a processor, or other programmable processing device) to create means for performing the functions of the method 600A. In some embodiments, the method 600A can be implemented by the apparatus 400 executing such instructions. The method can comprise the following steps.

[0115] In step 601, a bitstream comprising a set of pictures is received. As described, a base processing unit of a color picture can comprise a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the base processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as “coding tree blocks” (“CTBs”). Any operation performed on a base processing unit can be performed repeatedly on each of its luma and chroma components.

[0116] In step 603, a value of a first flag indicating whether a transform skip mode is enabled is determined from the received bitstream. The value of the first flag can be signaled in the received bitstream. For example, a value of sps_transform_skip_enabled_flag can be determined.

[0117] In step 605, it is determined whether a second flag indicating a residual coding method is signaled based on the value of the first flag. For example, the second flag can be a slice level residual coding flag. The slice level residual coding flag slice ts residual coding disabled flag specifies the residual coding method of the transform skip mode. If the transform skip mode is disabled at the SPS level, there is no need to signal in the slice header and it can be inferred to be 0.

[0118] In some embodiments, step 605 can comprise steps 605-1, 605-3, and 605-5, as shown in Figure 6B

[0119] In step 605-1, it is determined from the received bitstream whether the value of the first flag satisfies a first condition or a second condition.

[0120] In some embodiments, the first condition can be that the value of the first flag is 1, and the second condition can be that the value of the first flag is 0. Satisfying the first condition can result in the second flag being signaled. Otherwise, satisfying the second condition can result in a determination that the second flag is not signaled.

[0121] ​In step 605-3, in response to the value of the first flag (e.g., sps_transform_skip_enabled_flag) being 1, a second flag (e.g., slice_ts_residual_coding_disabled_flag) is signaled.

[0122] In step 605-5, in response to the value of the first flag (e.g., sps_transform_skip_enabled_flag) being 0, it is determined that the second flag (e.g., slice_ts_residual_coding_disabled_flag) is not signaled.

[0123] An example SPS syntax is described below in Table 3. In some example embodiments for signaling the residual coding method, if sps_transform_skip_enabled_flag is equal to 1, a slice level residual coding flag is signaled. If sps_transform_skip_enabled_flag is equal to 0, the value of slice_ts_residual_coding_disabled_flag is inferred to be 0. The semantics of sps_transform_skip_enabled_flag are shown below.

[0124]

[0125] Table 3 shows a portion of an example slice header syntax table for signaling the residual coding method according to some disclosed embodiments. In Table 3, the row with "if (sps_transform_skip_enabled_flag)" shows the changes to the syntax for the second reference embodiment. As described above, in the second reference embodiment, slice_ts_residual_coding_disabled_flag is always signaled. In contrast, according to some embodiments of the present disclosure, slice_ts_residual_coding_disabled_flag is not always signaled. Slice_ts_residual_coding_disabled_flag is only signaled when sps_transform_skip_enabled_flag is equal to 1. If sps_transform_skip_enabled_flag is equal to 0, slice_ts_residual_coding_disabled_flag is not signaled and is inferred to be equal to 0. The proposed changes can reduce the signaling overhead at the slice header level.

[0126] Table 3: Exemplary slice header syntax table for signaling residual coding method

[0127]

[0128] In some embodiments, non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by a device, such as the disclosed encoders and decoders, for performing the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or cartridge, and a networked version of any of the foregoing. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0129] The disclosed embodiments can be further described using the following clauses:

[0130] 1. A method of signaling of video data, comprising:

[0131] receiving a bitstream comprising a set of pictures;

[0132] determining, from the received bitstream, a value of a coding tree block size; and

[0133] based on the value of the coding tree block size, determining whether to signal a flag, the flag indicating a maximum transform size for a plurality of luma samples.

[0134] 2. The method of clause 1, wherein determining whether to signal the flag comprises:

[0135] in response to the value of the coding tree block size being greater than 32, signaling the flag; or

[0136] in response to the value of the coding tree block size being less than or equal to 32, determining to not signal the flag.

[0137] 3. The method of clause 1, wherein determining whether to signal the flag comprises:

[0138] in response to the value of the coding tree block size not being equal to 32, signaling the flag; or

[0139] in response to the value of the coding tree block size being equal to 32, determining to not signal the flag.

[0140] 4. The method of clause 1, wherein the values of the coding tree block size further include a value associated with a luma coding tree block size of a coding tree unit.

[0141] 5. The method of clause 4, wherein determining whether to signal the flag comprises:

[0142] in response to the value associated with the luma coding tree block size being greater than 0, signaling the flag; or

[0143] in response to the value associated with the luma coding tree block size being equal to 0, determining not to signal the flag.

[0144] 6. The method of clause 4, wherein determining whether to signal the flag comprises:

[0145] in response to the value associated with the luma coding tree block size not being equal to 0, signaling the flag; or

[0146] in response to the value associated with the luma coding tree block size being equal to 0, determining not to signal the flag.

[0147] 7. The method of any of clauses 1-6, wherein the flag indicates that the maximum transform size is 64.

[0148] 8. A method of signaling of video data, comprising:

[0149] receiving a bitstream comprising a set of pictures;

[0150] determining, from the received bitstream, a value of a first flag, the value of the first flag indicating whether a transform skip mode is enabled, and

[0151] based on the value of the first flag, determining whether to signal a second flag, the second flag indicating a residual coding method.

[0152] 9. The method of clause 8, wherein determining whether to signal the second flag comprises:

[0153] in response to the value of the first flag being 1, signaling the second flag; or

[0154] in response to the value of the first flag being 0, determining that the second flag is not signaled.

[0155] 10. The method of any of clauses 8 and 9, wherein:

[0156] the first flag is signaled in a sequence parameter set, and

[0157] The second flag is signaled in a slice header.

[0158] 11. An apparatus for signaling of video data, comprising:

[0159] a memory storing a set of instructions; and

[0160] one or more processors configured to execute the set of instructions to cause the apparatus to perform:

[0161] receiving a bitstream comprising a set of pictures;

[0162] determining a value of a coding tree block size from the received bitstream; and

[0163] based on the value of the coding tree block size, determining whether to signal a flag, the flag indicating a maximum transform size for luma samples.

[0164] 12. The apparatus of clause 11, wherein in determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further perform:

[0165] in response to the value of the coding tree block size being greater than 32, signaling the flag; or

[0166] in response to the value of the coding tree block size being less than or equal to 32, determining not to signal the flag.

[0167] 13. The apparatus of clause 11, wherein in determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further perform:

[0168] in response to the value of the coding tree block size not being equal to 32, signaling the flag; or

[0169] in response to the value of the coding tree block size being equal to 32, determining not to signal the flag.

[0170] 14. The apparatus of clause 11, wherein the value of the coding tree block size further comprises a value associated with a luma coding tree block size of a coding tree unit.

[0171] 15. The apparatus of clause 14, wherein in determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further perform:

[0172] in response to the value associated with the luma coding tree block size being greater than 0, signaling the flag; or

[0173] in response to a value associated with the luma coding tree block size being equal to 0, determining not to signal the flag.

[0174] 16. The apparatus of clause 14, wherein in determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further perform:

[0175] in response to a value associated with the luma coding tree block size being not equal to 0, signaling the flag; or

[0176] in response to a value associated with the luma coding tree block size being equal to 0, determining not to signal the flag.

[0177] 17. The apparatus of any of clauses 11-16, wherein the flag indicates that the maximum transform size is 64.

[0178] 18. An apparatus for signaling of video data, comprising:

[0179] a memory storing a set of instructions; and

[0180] one or more processors configured to execute the set of instructions to cause the apparatus to perform:

[0181] receiving a bitstream comprising a set of pictures;

[0182] from the received bitstream, determining a value of a first flag, the value of the first flag indicating whether a transform skip mode is enabled, and

[0183] based on the value of the first flag, determining whether to signal a second flag, the second flag indicating a residual coding method.

[0184] 19. The apparatus of clause 18, wherein in determining whether to signal the second flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further perform:

[0185] in response to the value of the first flag being 1, signaling the second flag; or

[0186] in response to the value of the first flag being 0, determining not to signal the second flag.

[0187] 20. The apparatus of any of clauses 18 and 19, wherein:

[0188] the first flag is signaled in a sequence parameter set, and

[0189] the second flag is signaled in a slice header.

[0190] 21. A non-transitory computer-readable medium storing a set of instructions capable of being executed by at least one processor of a computer to cause the computer to perform a method of signaling of video data, the method comprising:

[0191] receiving a bitstream comprising a group of pictures;

[0192] determining a value of a coding tree block size from the received bitstream; and

[0193] based on the value of the coding tree block size, determining whether to signal a flag, the flag indicating a maximum transform size of luma samples.

[0194] 22. The non-transitory computer-readable medium of clause 21, wherein determining whether to signal the flag comprises:

[0195] in response to the value of the coding tree block size being greater than 32, signaling the flag; or

[0196] in response to the value of the coding tree block size being less than or equal to 32, determining not to signal the flag.

[0197] 23. The non-transitory computer-readable medium of clause 21, wherein determining whether to signal the flag comprises:

[0198] in response to the value of the coding tree block size not being equal to 32, signaling the flag; or

[0199] in response to the value of the coding tree block size being equal to 32, determining not to signal the flag.

[0200] 24. The non-transitory computer-readable medium of clause 21, the value of the coding tree block size further comprising a value associated with a luma coding tree block size of a coding tree unit.

[0201] 25. The non-transitory computer-readable medium of clause 24, wherein determining whether to signal the flag comprises:

[0202] in response to the value associated with the luma coding tree block size being greater than 0, signaling the flag; or

[0203] in response to the value associated with the luma coding tree block size being equal to 0, determining not to signal the flag.

[0204] 26. The non-transitory computer-readable medium of clause 24, wherein determining whether to signal the flag comprises:

[0205] In response to a value not equal to 0 associated with the size of the luminance coding tree block, the flag is signaled; or

[0206] In response to a value of 0 associated with the size of the luminance-coded tree block, it is determined that no signal is needed to notify the flag.

[0207] 27. A non-transitory computer-readable medium according to any one of clauses 21-26, wherein the mark indicates that the maximum transformation size is 64.

[0208] 28. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer to cause the computer to perform a signal notification method for video data, the method comprising:

[0209] Receive a bitstream containing a set of images;

[0210] Based on the received bit stream, the value of a first flag is determined. The value of the first flag indicates whether transition skip mode is enabled.

[0211] Based on the value of the first flag, it is determined whether to signal the second flag, which indicates the residual coding method.

[0212] 29. The non-transitory computer-readable medium as described in Clause 28, wherein determining whether to signal the second flag includes:

[0213] In response to the value of the first flag being 1, the second flag is signaled; or

[0214] In response to the value of the first flag being 0, it is determined that no signal should be sent to the second flag.

[0215] 30. A non-transitory computer-readable medium according to any one of clauses 28 and 29, wherein:

[0216] The first flag is signaled in the sequence parameter set, and

[0217] The second mark is indicated by a signal in the strip header.

[0218] It should be noted that relational terms such as “first” and “second” in this document are used only to distinguish an entity or operation from another entity or operation, without requiring or implying any actual relationship or order between these entities or operations. Furthermore, the words “including,” “having,” “containing,” and “including,” and other similar forms, are semantically equivalent and open-ended, because one or more items following any of these words do not imply an exhaustive list of such one or more items, or that the list is limited to one or more items.

[0219] As used herein, the term "or" includes all possible combinations, unless otherwise specifically stated, unless otherwise apparent from context or impossible. For example, if a database is said to include A or B, then the database can include A, or B, or A and B, unless otherwise expressly stated, or impossible. As a second example, if a database is said to include A, B, or C, then the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise expressly stated, or impossible.

[0220] It should be understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. The software, when executed by the processor, can perform the disclosed method. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.

[0221] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary from embodiment to embodiment. Certain modifications and changes can be made to the described embodiments, by those of ordinary skill in the art. Other embodiments will be apparent from consideration of the specification and practice of the disclosure disclosed herein. The specification and examples given are considered to be exemplary only, with the true scope and spirit of the application being indicated by the following claims. The sequence of steps shown in the drawings is for illustrative purposes only and is not intended to be limiting to any particular sequence of steps. It will be understood by those skilled in the art that these steps can be performed in different orders while implementing the same method.

[0222] In the drawings and specification, there have been disclosed exemplary embodiments. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation, as the scope of the inventive subject matter is intended to be limited only by the claims.

Claims

1. A method of encoding video data, comprising: receiving a video frame to be encoded; signaling, in a sequence parameter set (SPS), a first flag for a basic processing unit of the video frame to be encoded, the first flag indicating whether a transform skip mode is enabled; in response to the first flag indicating that the transform skip mode is not enabled, disabling signaling of a second flag, the second flag indicating whether transform skip residual coding is used for resolving residual samples of a transform skip block of a current slice; determining a value of a coding tree block size associated with the SPS based on a parameter signaled in the sequence parameter set (SPS); and determining whether to signal a third flag associated with a maximum transform size of a plurality of luma samples based on the value of the coding tree block size; wherein determining whether to signal the third flag comprises, in response to the value of the coding tree block size being greater than 32, signaling the third flag in the SPS, a value of the third flag being set to one when the maximum transform size of the luma samples is equal to 64, and the value of the third flag being set to zero when the maximum transform size of the luma samples is equal to 32; wherein if the value of the coding tree block size is equal to or less than 32, the third flag is not signaled in the bitstream, and a value of the third flag is inferred to be equal to zero.

2. The method of claim 1, wherein the third flag comprises: sps_max_luma_transform_size_64_flag.

3. The method of claim 1, further comprising: the first flag being sps_transform_skip_enabled_flag and the second flag being slice_ts_residual_coding_disabled_flag.

4. The method of claim 1, wherein the value of the coding tree block size is determined based on a parameter of a specific luma coding tree block size of each coding tree unit.

5. A method of decoding video data, comprising: receiving a bitstream associated with a coding tree block, receiving, in a sequence parameter set (SPS) in the bitstream, a first flag, the first flag indicating whether a transform skip mode is enabled; in response to the first flag indicating that the transform skip mode is not enabled, disabling decoding of a second flag, the second flag indicating whether transform skip residual coding is used for resolving residual samples of a transform skip block of a current slice; determining a value of a coding tree block size associated with the SPS based on a parameter received in the sequence parameter set (SPS) in the bitstream; and determining whether to decode a third flag associated with a maximum transform size of a plurality of luma samples based on the value of the coding tree block size; in response to the value of the coding tree block size being greater than 32, decoding the third flag in the SPS, a value of the third flag being one when the maximum transform size of the luma samples is equal to 64, and the value of the third flag being zero when the maximum transform size of the luma samples is equal to 32; wherein, If the value of the coding tree block size is equal to or smaller than 32, the third flag is skipped from being decoded, and a value of the third flag is inferred to be equal to 0.

6. The method of claim 5, wherein the value of the coding tree block size is determined based on a parameter of a specific luma coding tree block size of each coding tree unit.

7. The method of claim 5, wherein the third flag comprises sps max luma transform size 64 flag.

8. The method of claim 5, wherein, the first flag is sps transform skip enabled flag and the second flag is slice ts residual coding disabled flag.

9. The method of claim 5, further comprising decoding the bitstream according to a Versatile Video Coding (VVC / H.266) standard.

10. A non-transitory computer readable medium storing a set of instructions and a video bitstream, the set of instructions executable by a processor to perform a method to generate the bitstream, the method comprising: receiving a video frame to be encoded; signaling, in a sequence parameter set (SPS), a first flag for a basic processing unit of the video frame to be encoded, the first flag indicating whether a transform skip mode is enabled; in response to the first flag indicating that the transform skip mode is not enabled, disabling signaling of a second flag, the second flag indicating whether transform skip residual coding is used for resolving residual samples of a transform skip block of a current slice; determining, based on a parameter signaled in the sequence parameter set (SPS), a value of a size of a coding tree block associated with the SPS, and based on the value of the coding tree block size, determining whether to signal a third flag, the third flag indicating a maximum transform size of a plurality of luma samples; wherein determining whether to signal the third flag comprises, in response to the value of the coding tree block size being greater than 32, signaling the third flag in the SPS, the value of the third flag being set to 1 when the maximum transform size of the luma samples is equal to 64, and the value of the third flag being set to 0 when the maximum transform size of the luma samples is equal to 32; wherein, if the value of the coding tree block size is equal to or smaller than 32, the third flag is not signaled in the bitstream, and a value of the third flag is inferred to be equal to 0.

11. The non-transitory computer readable medium of claim 10, wherein the third flag comprises sps max luma transform size 64 flag.

12. The non-transitory computer readable medium of claim 10, wherein the bitstream is encoded according to a Versatile Video Coding (VVC / H.266) standard.

13. The non-transitory computer readable medium of claim 10, wherein, The first flag is sps_transform_skip_enabled_flag, and the second flag is slice_ts_residual_coding_disabled_flag.