Method of signaling maximum transform size and residual coding

By signaling the maximum transform size and the residual coding method, the video coding process is optimized, the problem of insufficient coding efficiency in the prior art is solved, and more efficient storage and transmission performance is achieved.

CN120751150AActive Publication Date: 2025-10-03ALIBABA (CHINA) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511142512.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-01-29
Publication Date
2025-10-03
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Existing video coding standards have not yet fully utilized the signaling of maximum transform size and residual coding methods in high-efficiency video coding technology, resulting in insufficient coding efficiency.

Method used

The encoding process is optimized by determining the value of the coding tree block size by receiving a bitstream, and signaling a flag of the maximum transform size based on the value, and determining whether to signal a second flag of the residual coding method based on the first flag value for enabling the transform skip mode.

Benefits of technology

It improves the coding efficiency of video coding, reduces the requirements for storage space and transmission bandwidth, and enhances the flexibility and fault tolerance of the coding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751150A_ABST
    Figure CN120751150A_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for signaling a maximum transform size, as well as residual encoding methods. In accordance with certain disclosed embodiments, the method comprises: receiving a bitstream comprising a set of images; determining a value of a coding tree block size according to the received bit stream; and determining whether to signal a flag indicating a maximum transform size of a luma sample based on a value of the coding tree block size.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 980,117, filed on February 21, 2020, which is hereby incorporated by reference in its entirety. Technical Field

[0002] The present disclosure relates generally to video processing, and more particularly, to methods and apparatus for signaling a maximum transform size and a residual encoding method. Background Art

[0003] A video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission, and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are many video coding formats that use standardized video coding techniques, the most common of which are based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats. As more and more advanced video coding technologies are adopted in video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and apparatus for signaling a maximum transform size. In some exemplary embodiments, the method includes: receiving a bitstream comprising a set of images; determining a value for a coding tree block size based on the received bitstream; and determining whether to signal a flag indicating a maximum transform size for luma samples based on the value of the coding tree block size.

[0005] The apparatus may include a memory storing an instruction set; and one or more processors configured to execute the instruction set to cause the apparatus to perform: receiving a bitstream comprising a set of images, determining a value of a coding treeblock size based on the received bitstream; and determining whether to signal a flag of a maximum transform size for luma samples based on the value of the coding treeblock size.

[0006] An embodiment of the present disclosure also provides a non-transitory computer-readable medium storing an instruction set, which can be executed by at least one processor of a computer to cause the computer to perform a method for signaling a maximum transform size, the method comprising: receiving a bitstream including a set of images, determining a value of a coding tree block size based on the received bitstream; and determining whether to signal a flag indicating a maximum transform size for luma samples based on the value of the coding tree block size.

[0007] The present disclosure also provides a method and apparatus for signaling a residual coding method. In some exemplary embodiments, the method includes: receiving a bitstream including a set of images; determining, based on the received bitstream, a value of a first flag indicating whether a transform skip mode is enabled; and determining, based on the value of the first flag, whether to signal a second flag indicating a residual coding method.

[0008] The apparatus may include a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the apparatus to: receive a bitstream comprising a set of images; determine, based on the received bitstream, a value of a first flag indicating whether a transform skip mode is enabled; and determine, based on the value of the first flag, whether to signal a second flag indicating a residual coding method.

[0009] An embodiment of the present disclosure also provides a non-transitory computer-readable medium storing an instruction set, which can be executed by at least one processor of a computer to cause the computer to perform a method for signaling a residual encoding method, the method comprising: receiving a bit stream including a set of images; determining a value of a first flag based on the received bit stream, the value of the first flag indicating whether a transform skip mode is enabled; and determining whether to signal a second flag indicating a residual encoding method based on the value of the first flag. BRIEF DESCRIPTION OF THE DRAWINGS

[0010]

[0011] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings.The various features shown in the drawings are not drawn to scale.

[0011] Figure 1 is a structural diagram of an exemplary video sequence consistent with some embodiments of the present disclosure.

[0012] Figure 2A is a diagram illustrating an exemplary encoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0013] Figure 2B is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0014] Figure 3A is a diagram illustrating an exemplary decoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0015] Figure 3B is a diagram illustrating another exemplary decoding process of a hybrid video coding system consistent with some embodiments of the present disclosure.

[0016] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video consistent with some embodiments of the present disclosure.

[0017] Figure 5A is an exemplary method for signaling a maximum transform size consistent with some embodiments of the present disclosure.

[0018] Figure 5B is an exemplary method for signaling a maximum transform size consistent with some embodiments of the present disclosure.

[0019] Figure 6A is an exemplary method for signaling a residual encoding method consistent with some embodiments of the present disclosure.

[0020] Figure 6B is an exemplary method for signaling a residual encoding method consistent with some embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with aspects related to the present disclosure as described in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0022] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0023] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding technologies are incorporated into JEM, JEM achieves higher coding performance than HEVC.

[0024] The VVC standard was developed recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0025] A video is a set of static images (or "frames") arranged in time to store visual information. These images can be captured and stored in time using a video capture device (e.g., a camera), and displayed in time using a video playback device (e.g., a television, computer, smartphone, tablet, video player, or any end-user terminal with a display). Furthermore, in some applications, the video capture device can send the captured video to a video playback device (e.g., a computer with a monitor) in real time, for example, for monitoring, conferencing, or live broadcasting.

[0026] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The module used for compression is generally referred to as an "encoder," and the module used for decompression is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented as any of various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, and the like. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder."

[0027] The video encoding process can identify and retain useful information that can be used to reconstruct the image, while ignoring unimportant reconstruction information. If ignoring unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy." Otherwise, it can be called "lossless." Most encoding processes are lossy, as a trade-off to reduce the required storage space and transmission bandwidth.

[0028] Useful information about the image being coded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously coded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with position changes being of primary interest. A change in the position of a group of pixels representing an object can reflect the object's motion between the reference image and the current image.

[0029] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture coded using a previous picture as a reference is called a "P-picture," and a picture coded using both a previous picture and a future picture as reference is called a "B-picture" (the reference is "bidirectional").

[0030] Figure 1 The structure of an example video sequence 100 according to some embodiments of the present disclosure is shown. The video sequence 100 can be live video or video that has been captured and archived. The video 100 can be real-life video, computer-generated video (e.g., computer game video), or a combination of the two (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), an archive containing previously captured video (e.g., a video file stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.

[0031] like Figure 1 As shown, video sequence 100 may include a series of images arranged temporally along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, with more images between images 106 and 108. Figure 1 , picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after the picture. For example, the reference picture of picture 104 may be the picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and the present disclosure is not limited to such pictures. Figure 1 Example of a reference image shown.

[0032] Typically, video codecs do not encode or decode an entire image at once due to the computational complexity of the encoding and decoding tasks. Instead, they can divide the image into basic segments and encode or decode the image segments one by one. In this disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, Figure 1Structure 110 in shows an example structure of an image of video sequence 100 (e.g., any of images 102-108). In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing units may be referred to as “macroblocks” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as “coding tree units” (“CTUs”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units may have variable sizes in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of arbitrary shape and size. The size and shape of the basic processing units may be selected for the image based on a balance between coding efficiency and the level of detail to be maintained in the basic processing units. A CTU is the largest block unit and may include up to 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). A CTU may be further partitioned into coding units (CUs) using a quadtree, a binary tree, a ternary tree, or a combination thereof.

[0033] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color image may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0034] Video encoding has multiple stages of operation, examples of which are Figures 2A-2B and Figures 3A-3BAs shown. For each stage, the size of the basic processing unit may still be too large for processing, so it can be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-unit may be referred to as a "block" in some video coding standards (e.g., MPEG family, H.261, H.263 or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit may have the same size as the basic processing unit or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing units.

[0035] For example, in the mode decision phase (an example of which is Figure 2B As shown in FIG, the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the prediction type for each individual basic processing sub-unit.

[0036] For another example, in the prediction phase (the example is Figures 2A-2B ), the encoder can perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further divide the basic processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which prediction operations can be performed.

[0037] For another example, in the transformation phase (the example is Figures 2A-2B), the encoder can perform transform operations on the residual basic processing sub-unit (e.g., CU). However, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further divide the basic processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level transform operations can be performed. It should be noted that the division scheme of the same basic processing sub-unit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0038] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are shown by dotted lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.

[0039] In some embodiments, in order to provide parallel processing capabilities and error resilience for video encoding and decoding, an image can be divided into regions for processing so that the encoding or decoding process for a region of the image does not depend on information from any other region of the image. In other words, each region of the image can be processed separately. By doing so, the codec can process different regions of the image in parallel, thereby improving coding efficiency. In addition, when data in one region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing error resilience. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different images in the video sequence 100 can have different partitioning schemes for dividing the image into regions.

[0040] For example, in Figure 1 In FIG, the structure 110 is divided into three regions 114, 116 and 118, whose boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and structural areas in 110 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0041] Figure 2A Schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by an encoder. Figure 2AAs shown, the encoder may encode the video sequence 202 into a video bitstream 228 according to process 200A. Figure 1 The video sequence 100 in FIG. 2 may include a set of images (referred to as “original images”) arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , each original image of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, wherein the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for each region (e.g., regions 114-118) of the original image of the video sequence 202.

[0042] refer to Figure 2A , the encoder may feed the basic processing units of the original images of the video sequence 202 (referred to as "original BPUs") to the prediction stage 204 to generate prediction data 206 and prediction BPUs 208. The encoder may subtract the predicted BPUs 208 from the original BPUs to generate residual BPUs 210. The encoder may feed the residual BPUs 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to the binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path." During process 200A, after the quantization stage 214, the encoder may feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224, which is used in the prediction phase 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as a "reconstruction path." The reconstruction path can be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0043] The encoder may iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder may proceed to encode the next image in the video sequence 202.

[0044] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data in any manner.

[0045] In the prediction phase 204, at the current iteration, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate the prediction data 206 and the predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.

[0046] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for these differences, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the value of the pixel of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of this subtraction between the corresponding pixel of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without a noticeable loss in quality. Therefore, the original BPU is compressed.

[0047] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "base patterns". Each base pattern is associated with a "transform coefficient". The base patterns can have the same size (e.g., the size of the residual BPU 210), and each base pattern can represent a frequency-varying (e.g., frequency of brightness variation) component of the residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the base images are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are similar to the coefficients associated with the basis functions.

[0048] Different transform algorithms can use different basic patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform at the transform stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to restore the pixels of the residual BPU 210, the inverse transform may be multiplying the values ​​of corresponding pixels of the basic pattern by the corresponding associated coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basic pattern). Therefore, the encoder can only record the transform coefficients, from which the decoder can reconstruct the residual BPU 210 without receiving the basic pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. As a result, the residual BPU 210 is further compressed.

[0049] The encoder can further compress the transform coefficients in the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of change (e.g., the frequency of brightness changes). Because the human eye is generally better at detecting low-frequency changes, the encoder can ignore information about high-frequency changes without causing noticeable quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After this operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. This quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed as transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0050] Because the encoder ignores the remainder of this division during rounding operations, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer bits are required to quantize the transform coefficients 216. To achieve different levels of information loss, the encoder can use different values ​​for the quantization parameter or any other parameters of the quantization process.

[0051] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type at the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0052] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0053] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.

[0054] Figure 2B A schematic diagram of another example encoding process 200B according to an embodiment of the present disclosure is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B also includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B also includes a loop filter stage 232 and a buffer 234.

[0055] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-frame image prediction or "intra-frame prediction") can use pixels from one or more already encoded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more already encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.

[0056] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For an original BPU of the image being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder may generate a predicted BPU 208 by interpolating the neighboring BPUs. Interpolation techniques may include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder may perform interpolation at the pixel level, for example, by interpolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation may be located in various directions relative to the original BPU, such as vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or in any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 may include, for example, the location (eg, coordinates) of the used neighboring BPUs, the size of the used neighboring BPUs, interpolation parameters, the direction of the used neighboring BPUs relative to the original BPU, and the like.

[0057] As another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed BPU by BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs for the same image have been generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within a range of reference images (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image with the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When the encoder identifies (e.g., using a PEL recursive algorithm, a block matching algorithm, etc.) an area similar to the original BPU in the search window, the encoder may determine such an area as a matching region. The matching region may have a different size than the original BPU (e.g., smaller, equal, larger, or a different shape). Because the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 ), so the matching area can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When using multiple reference images (e.g., Figure 1 The encoder may search for matching regions and determine the motion vector associated with each reference image. In some embodiments, the encoder may assign weights to the pixel values ​​of the matching regions of the respective matching reference images.

[0058] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, weights associated with the reference images, etc.

[0059] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of ​​the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., Figure 1In some embodiments, if the encoder has assigned weights to the pixel values ​​of the matching regions of the reference image, the encoder may add the weighted sum of the pixel values ​​of the moved matching regions.

[0060] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 The picture 104 in is a unidirectional inter-frame predicted picture, where the reference picture (i.e., picture 102) precedes picture 04. Bidirectional inter-frame prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, Figure 1 Picture 106 in is a bidirectional inter-predicted picture, where the reference pictures (ie, pictures 104 and 08) are relative to picture 104 in both temporal directions.

[0061] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction and inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0062] In the reconstruction path of process 200B, if intra-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If inter-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232. At this stage, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced by inter-prediction. The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive compensation, adaptive loop filtering, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., as inter-frame prediction reference pictures for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder may encode the parameters of the loop filter (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information at the binary encoding stage 226.

[0063] Figure 3A A schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure is shown. Process 300A may correspond to Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figures 2A-2B quantization stage 214 in), typically, the video stream 304 is different from the video sequence 202. Figures 2A-2B 200A and 200B, the decoder may perform process 300A at a basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0064] like Figure 3AAs shown, the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit (referred to as a "coding BPU") for an encoded picture to a binary decoding stage 302, where the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to the prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.

[0065] The decoder may iteratively perform process 300A to decode each coded BPU of a coded picture and generate a prediction reference 224 for the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.

[0066] In the binary decoding stage 302, the decoder can perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, the transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packet form over the network, the decoder can depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.

[0067] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0068] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a decoded coded image (referred to as a "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder used intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPUs, parameters of interpolation, the direction of the neighboring BPU relative to the original BPU, etc. For another example, if the encoder used inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference images, one or more motion vectors associated with the matching regions respectively, and the like.

[0069] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra-frame prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter-frame prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in detail in Figure 2B After performing such spatial prediction or temporal prediction, the decoder may generate a predicted BPU 208, which may be added to the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown in FIG. Figure 3A As described in.

[0070] In process 300B, the decoder may feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current picture). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder may feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2BThe loop filter is applied to the prediction reference 224 in the manner shown. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-prediction reference picture for a future coded picture in the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-frame prediction was used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength).

[0071] Figure 4 FIG is a block diagram of an example apparatus 400 for encoding or decoding a video according to an embodiment of the present disclosure. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for video encoding or decoding. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), a field programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0072] The apparatus 400 may further include a memory 404 configured to store data (eg, instruction sets, computer code, intermediate data, etc.). Figure 4As shown, the stored data may include program instructions (e.g., for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).

[0073] The bus 410 may be a communication device that transmits data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0074] For ease of explanation and to avoid ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, separate module, or may be fully or partially integrated into any other component of device 400.

[0075] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0076] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.

[0077] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instances that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0078] In a first reference embodiment, a sequence parameter set (SPS) syntax element sps_max_luma_transform_size_64flag is signaled to specify the maximum transform size. sps_max_luma_transform_size_64 flag equal to 1 specifies that the maximum transform size in luma samples is equal to 64. sps_max_luma_transform_size_64 flag equal to 0 specifies that the maximum transform size in luma samples is equal to 32. In the first reference embodiment, sps_max_luma_transform_size_64 flag is always signaled regardless of the value of the coding tree block size (CtbSizeY). However, if CtbSizeY is less than 64, then sps_max_luma_transform_size_64_flag does not need to be signaled and can be inferred to be 0.

[0079] In addition, in the second reference embodiment, a slice level residual coding selection method can be used. The following is an exemplary semantics of slice_ts_residual_coding_disabled_flag:

[0080] In the second reference embodiment, slice_ts_residual_coding_disabled_flag is always signaled. However, since slice_ts_residual_coding_disabled_flag specifies the residual coding method for transform skip mode, if transform skip mode is disabled at the SPS level, slice_ts_residual_coding_disabled_flag does not need to be signaled in the slice header and can be inferred to be 0.

[0081] The present disclosure provides a method and apparatus for reducing the above-mentioned coding redundancy.

[0082] Figure 5A is an exemplary method for signaling a maximum transform size consistent with some embodiments of the present disclosure. In some embodiments, method 500A may be performed by an encoder, a decoder, and an apparatus (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 500A may be performed by a processor 402 of the computer. In some embodiments, the method 500A may be implemented by a computer program product contained in a computer-readable medium, the computer program product comprising a computer (e.g., Figure 4 The method may include the following steps:

[0083] In step 501, a bitstream comprising a set of images is received. As described, a basic processing unit of a color image may include a luma component (Y) representing achromatic luma information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, wherein the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0084] In step 503, the value of the coding tree block size is determined based on the received bitstream. The value of the coding tree block size is signaled in the received bitstream. For example, the value of CtbSizeY in the exemplary semantics is determined.

[0085] In step 505, a determination is made based on the value of the coding tree block size whether to signal a flag indicating the maximum transform size for luma samples. For example, the sequence parameter set (SPS) syntax element sps_max_luma_transform_size_64flag may be used to specify the maximum transform size. sps_max_luma_transform_size_64flag equal to 1 specifies that the maximum transform size in luma samples is equal to 64. sps_max_luma_transform_size_64flag equal to 0 specifies that the maximum transform size in luma samples is equal to 32.

[0086] In some embodiments, step 505 may include: Figure 5BSteps 505-1, 505-3 and 505-5 are shown.

[0087] In step 505 - 1 , it is determined whether the value of the coding tree block size satisfies the first condition or the second condition according to the received bitstream.

[0088] In some embodiments, the first condition may be a value greater than 32, and the second condition may be a value less than or equal to 32. Satisfying the first condition may result in signaling a flag indicating the maximum transform size. Otherwise, satisfying the second condition may result in determining that the flag is not signaled.

[0089] In step 505 - 3 , in response to the value of the coding treeblock size being greater than 32, the flag is signaled.

[0090] An exemplary SPS syntax is described below in Table 1. In some exemplary embodiments for signaling the maximum transform size, sps_max_luma_transform_size_64_flag is signaled only when CtbSizeY is greater than 32. Table 1 shows a portion of an exemplary SPS syntax table for signaling the maximum transform size according to some disclosed embodiments. In Table 1, the row with "if(CtbSizeY>32)" shows a change to the syntax used for the first reference embodiment. As described, in the first reference embodiment, sps_max_luma_transform_size_64_flag is always signaled regardless of the value of the coding tree block size (CtbSizeY). In contrast, according to some embodiments of the present disclosure, sps_max_luma_transform_size_64_flag is not always signaled. As shown in Table 1, the signaling of sps_max_luma_transform_size_64_flag is conditional on CtbSizeY. Table 1: Example SPS syntax table for the disclosed method for signaling maximum transform size

[0091] In step 505 - 5 , in response to the value of the coding treeblock size being less than or equal to 32, it is determined that the flag is not signaled.

[0092] In a signaling method consistent with Table 1, if CtbSizeY is not greater than 32, the value of sps_max_luma_transform_size_64_flag is inferred to be 0. The semantics of deriving CtbSizeY are as follows.

[0093] In some embodiments, the first condition and the second condition may be different from the above-described case. In the following example, the first condition may be a value not equal to 32, and the second condition may be a value equal to 32.

[0094] In an alternative step to step 505 - 3 , in response to the value of the coding treeblock size not being equal to 32, a flag is signaled.

[0095] In an alternative step to step 505 - 5 , in response to the value of the coding treeblock size being equal to 32, it is determined that the flag is not signaled.

[0096] An exemplary syntax is described below. The proposed syntax changes shown in Table 1 can also be implemented using the condition "CtbSizeY!=32" instead of "CtbSizeY>32". In this case: if CtbSizeY is not equal to 32, then sps_max_luma_transform_size_64 flag is signaled; if CtbSizeY is equal to 32, then sps_max_luma_transform_size_64 flag is inferred to be 0.

[0097] An exemplary syntax is described below. The proposed syntax changes shown in Table 1 can also be implemented using the condition "CtbSizeY>=64" instead of "CtbSizeY>32". In this case: if CtbSizeY is greater than or equal to 64, sps_max_luma_transform_size_64flag is signaled; if CtbSizeY is less than 64, sps_max_luma_transform_size_64flag is inferred to be 0. In other words, sps_max_luma_transform_size_64flag needs to be signaled only when CtbSizeY is greater than 32. If CtbSizeY is not greater than 32, sps_max_luma_transform_size_64flag is not signaled and is inferred to be equal to 0. Alternatively, sps_max_luma_transform_size_64flag is signaled only when CtbSizeY is greater than or equal to 64. If CtbSizeY is less than 64, sps_max_luma_transform_size_64_flag is not signaled and is inferred to be equal to 0.

[0098] In some exemplary embodiments for signaling the maximum transform size, notification of the sps_max_luma_transform_size_64_flag is conditional on sps_log2_ctu_size_minus5 instead of using CtbSizeY. As described above, sps_log2_ctu_size_minus5 plus 5 specifies the luma coding tree block size for each coding tree unit (CTU). A CTU can be 128×128 luma samples (plus corresponding chroma samples depending on the chroma format). Table 2 below shows a portion of an exemplary SPS syntax table for signaling the maximum transform size according to some disclosed embodiments. In Table 2, the row with "if(sps_log2_ctu_size_minus5>0)" shows a change to the syntax used for the first reference embodiment. As described, in the first reference embodiment, sps_max_luma_transform_size_64_flag is always signaled regardless of the value of the coding tree block size (e.g., CtbSizeY). In contrast, according to some embodiments of the present disclosure, sps_max_luma_transform_size_64_flag is not always signaled. As shown in Table 2, in some embodiments, the signaling of sps_max_luma_transform_size_64_flag is conditional on sps_log2_ctu_size_minus5. As shown in Table 2, if sps_log2_ctu_size_minus5 is greater than 0, then sps_max_luma_transform_size_64_flag is signaled; if sps_log2_ctu_size_minus5 is 0, then the value of sps_max_luma_transform_size_64_flag is inferred to be 0.

[0099] In some embodiments, the value of the coding tree block size in step 501 may further include a value associated with the luma coding tree block size of each coding tree unit. This value (eg, sps_log2_ctu_size_minus5) may be determined based on the received bitstream.

[0100] In some embodiments, in an alternative step to step 505 - 1 , it may be determined whether the value associated with the luma coding tree block size of each coding tree unit satisfies the first condition (eg, greater than 0) or the second condition (eg, equal to 0).

[0101] In some embodiments, in an alternative to step 505 - 3 , in response to the value associated with the luma coding tree block size of each coding tree unit satisfying a first condition (eg, greater than 0), the flag (eg, sps_max_luma_transform_size_64_flag) is signaled.

[0102] In some embodiments, in an alternative step to step 505-5, in response to the value associated with the luma coding tree block size of each coding tree unit satisfying the second condition (e.g., being equal to 0), determining that the flag (e.g., sps_max_luma_transform_size_64_flag) is not signaled. Table 2: Example SPS syntax table for the disclosed method for signaling maximum transform size

[0103] The above syntax changes shown in Table 2 can also be implemented using the condition “sps_log2_ctu_size_minus5 != 0”. In this case: if sps_log2_ctu_size_minus5 is not equal to 0, then sps_max_luma_transform_size_64_flag is signaled; if sps_log2_ctu_size_minus5 is equal to 0, then sps_max_luma_transform_size_64_flag is inferred to be 0.

[0104] In some embodiments, the first condition may also be a value not equal to 0 (e.g., sps_log2_ctu_size_minus5), as shown in the above example. In an alternative step to step 505-3, in response to the value associated with the luma coding tree block size of each coding tree unit satisfying the first condition (e.g., sps_log2_ctu_size_minus5 != 0), a flag (e.g., sps_max_luma_transform_size_64_flag) is signaled.

[0105] FIG6 is an exemplary method for signaling a residual encoding method consistent with some embodiments of the present disclosure. In some embodiments, method 600A may be performed by an encoder, a decoder, and a device (e.g., Figure 4 For example, a processor (e.g., Figure 4The method 600A may be performed by a processor 402 of the computer. In some embodiments, the method 600A may be implemented by a computer program product contained in a computer-readable medium, the computer program product including a computer (e.g., Figure 4 The method may include the following steps:

[0106] In step 601, a bitstream comprising a set of images is received. As described, a basic processing unit of a color image may include a luma component (Y) representing achromatic luma information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, wherein the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0107] In step 603, a value of a first flag indicating whether transform skip mode is enabled is determined based on the received bitstream. The value of the first flag may be signaled in the received bitstream. For example, the value of sps_transform_skip_enabled_flag may be determined.

[0108] In step 605, a determination is made as to whether a second flag indicating a residual coding method is signaled based on the value of the first flag. For example, the second flag may be a slice-level residual coding flag. The slice-level residual coding flag, slice_ts_residual_coding_disabled_flag, specifies the residual coding method for transform skip mode. If transform skip mode is disabled at the SPS level, it does not need to be signaled in the slice header and can be inferred to be 0.

[0109] In some embodiments, step 605 may include steps 605-1, 605-3, and 605-5. Figure 6B shown.

[0110] In step 605 - 1 , it is determined based on the received bit stream whether the value of the first flag satisfies the first condition or the second condition.

[0111] In some embodiments, the first condition may be a first flag having a value of 1, and the second condition may be a first flag having a value of 0. Satisfying the first condition may result in the second flag being signaled. Otherwise, satisfying the second condition may result in determining that the second flag is not signaled.

[0112] In step 605 - 3 , in response to the value of the first flag (eg, sps_transform_skip_enabled_flag) being 1, the second flag (eg, slice_ts_residual_coding_disabled_flag) is signaled.

[0113] In step 605 - 5 , in response to the value of the first flag (eg, sps_transform_skip_enabled_flag) being 0, it is determined that the second flag (eg, slice_ts_residual_coding_disabled_flag) is not signaled.

[0114] An exemplary SPS syntax is described below in Table 3. In some exemplary embodiments for signaling the residual coding method, if sps_transform_skip_enabled_flag is equal to 1, then the slice-level residual coding flag is signaled. If sps_transform_skip_enabled_flag is equal to 0, then the value of slice_ts_residual_coding_disabled_flag is inferred to be 0. The semantics of sps_transform_skip_enabled_flag are as follows.

[0115] Table 3 shows a portion of an exemplary slice header syntax table for signaling a residual coding method according to some disclosed embodiments. In Table 3, the row with "if(sps_transform_skip_enabled_flag)" shows a change to the syntax used for the second reference embodiment. As described above, in the second reference embodiment, slice_ts_residual_coding_disabled_flag is always signaled. In contrast, according to some embodiments of the present disclosure, slice_ts_residual_coding_disabled_flag is not always signaled. Slice_ts_residual_coding_disabled_flag is signaled only when sps_transform_skip_enabled_flag is equal to 1. If sps_transform_skip_enabled_flag is equal to 0, slice_ts_residual_coding_disabled_flag is not signaled and is inferred to be equal to 0. The proposed change can reduce signaling overhead at the slice header level. Table 3: Example slice header syntax table for the disclosed method for signaling residual coding method

[0116] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge storage, and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0117] The disclosed embodiments may be further described using the following terms: 1. A method for signaling video data, comprising: receiving a bitstream comprising a set of images; Determining a value of a coding tree block size based on a received bitstream; and Based on the value of the coding treeblock size, it is determined whether to signal a flag that indicates a maximum transform size for a plurality of luma samples. 2. The method of clause 1, wherein determining whether to signal the flag comprises: In response to the value of the coding tree block size being greater than 32, signaling the flag; or In response to the value of the coding tree block size being less than or equal to 32, determining not to signal the flag. 3. The method of clause 1, wherein determining whether to signal the flag comprises: In response to the value of the coding tree block size not being equal to 32, signaling the flag; or In response to the value of the coding treeblock size being equal to 32, it is determined not to signal the flag. 4. The method of clause 1, wherein the value of the coding tree block size further comprises a value associated with a luma coding tree block size of a coding tree unit. 5. The method of clause 4, wherein determining whether to signal the flag comprises: In response to a value associated with the luma coding treeblock size being greater than zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 6. The method of clause 4, wherein determining whether to signal the flag comprises: In response to the value associated with the luma coding treeblock size being not equal to zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 7. A method according to any of clauses 1-6, wherein the flag indicates that the maximum transform size is 64. 8. A method for signaling video data, comprising: receiving a bitstream comprising a set of images; determining, based on the received bitstream, a value of a first flag indicating whether a transform skip mode is enabled, and Based on a value of the first flag, it is determined whether to signal a second flag, the second flag indicating a residual encoding method. 9. A method according to clause 8, wherein determining whether to signal the second flag comprises: In response to the value of the first flag being 1, signaling the second flag; or In response to the value of the first flag being 0, it is determined that the second flag is not signaled. 10. A method according to any one of clauses 8 and 9, wherein: The first flag is signaled in a sequence parameter set, and The second flag is signaled in the slice header. 11. A device for signaling video data, comprising: a memory storing an instruction set; and one or more processors configured to execute the set of instructions to cause the apparatus to: receiving a bitstream comprising a set of images; Determining a value of a coding tree block size based on a received bitstream; and Based on the value of the coding treeblock size, it is determined whether to signal a flag that indicates a maximum transform size for luma samples. 12. The apparatus of clause 11, wherein upon determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further: In response to the value of the coding tree block size being greater than 32, signaling the flag; or In response to the value of the coding tree block size being less than or equal to 32, determining not to signal the flag. 13. The apparatus of clause 11, wherein upon determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further: In response to the value of the coding tree block size not being equal to 32, signaling the flag; or In response to the value of the coding tree block size being equal to 32, it is determined not to signal the flag. 14. The apparatus of clause 11, wherein the value of the coding tree block size further comprises a value associated with a luma coding tree block size of a coding tree unit. 15. The apparatus of clause 14, wherein upon determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further: In response to a value associated with the luma coding treeblock size being greater than zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 16. The apparatus of clause 14, wherein upon determining whether to signal the flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further: In response to the value associated with the luma coding treeblock size being not equal to zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 17. Apparatus according to any of clauses 11-16, wherein the flag indicates that the maximum transform size is 64. 18. A video data signal notification device, comprising: a memory storing an instruction set; and one or more processors configured to execute the set of instructions to cause the apparatus to: receiving a bitstream comprising a set of images; determining, based on the received bitstream, a value of a first flag indicating whether a transform skip mode is enabled, and Based on a value of the first flag, it is determined whether to signal a second flag, the second flag indicating a residual encoding method. 19. The apparatus of clause 18, wherein upon determining whether to signal the second flag, the one or more processors are configured to execute the set of instructions to cause the apparatus to further: In response to the value of the first flag being 1, signaling the second flag; or In response to the value of the first flag being 0, it is determined not to signal the second flag. 20. The apparatus according to any one of clauses 18 and 19, wherein: The first flag is signaled in a sequence parameter set, and The second flag is signaled in the slice header. 21. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a computer to cause the computer to perform a method for signaling video data, the method comprising: receiving a bitstream comprising a set of images; Determining a value of a coding tree block size based on a received bitstream; and Based on the value of the coding treeblock size, it is determined whether to signal a flag that indicates a maximum transform size for luma samples. 22. The non-transitory computer-readable medium of clause 21, wherein determining whether to signal the flag comprises: In response to the value of the coding tree block size being greater than 32, signaling the flag; or In response to the value of the coding tree block size being less than or equal to 32, determining not to signal the flag. 23. The non-transitory computer-readable medium of clause 21, wherein determining whether to signal the flag comprises: In response to the value of the coding tree block size not being equal to 32, signaling the flag; or In response to the value of the coding tree block size being equal to 32, it is determined not to signal the flag. 24. The non-transitory computer-readable medium of clause 21, wherein the value of the coding tree block size further comprises a value associated with a luma coding tree block size of a coding tree unit. 25. The non-transitory computer-readable medium of clause 24, wherein determining whether to signal the flag comprises: In response to a value associated with the luma coding treeblock size being greater than zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 26. The non-transitory computer-readable medium of clause 24, wherein determining whether to signal the flag comprises: In response to the value associated with the luma coding treeblock size being not equal to zero, signaling the flag; or In response to the value associated with the luma coding treeblock size being equal to 0, it is determined not to signal the flag. 27. The non-transitory computer-readable medium of any of clauses 21-26, wherein the flag indicates that the maximum transform size is 64. 28. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a computer to cause the computer to perform a method for signaling video data, the method comprising: receiving a bitstream comprising a set of images; determining, based on the received bitstream, a value of a first flag indicating whether a transform skip mode is enabled, and Based on a value of the first flag, it is determined whether to signal a second flag, the second flag indicating a residual encoding method. 29. The non-transitory computer-readable medium of clause 28, wherein determining whether to signal the second flag comprises: In response to the value of the first flag being 1, signaling the second flag; or In response to the value of the first flag being 0, it is determined not to signal the second flag. 30. The non-transitory computer-readable medium of any one of clauses 28 and 29, wherein: The first flag is signaled in a sequence parameter set, and The second flag is signaled in the slice header.

[0118] It should be noted that relational terms such as "first" and "second" herein are used only to distinguish an entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include," "have," "include," and "includes" and other similar forms are equivalent in meaning and are open-ended, in that one or more items following any of these words is not intended to be an exhaustive list of such one or more items, or to be limited to the listed one or more items.

[0119] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless otherwise feasible. For example, if it is stated that a database may include either A or B, then unless explicitly stated otherwise or otherwise feasible, the database may include either A, or B, or A and B. As a second example, if it is stated that a database may include either A, B, or C, then unless explicitly stated otherwise or otherwise feasible, the database may include either A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0120] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. The software can perform the disclosed method when executed by a processor. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.

[0121] In the foregoing description, embodiments have been described with reference to many specific details, which may vary from implementation to implementation. Certain modifications and variations may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure disclosed herein. This description and embodiments are to be considered exemplary only, with the true scope and spirit of the invention being indicated by the appended claims. The sequence of steps shown in the accompanying drawings is for illustrative purposes only and is not intended to be limiting to any particular sequence of steps. Therefore, it will be understood by those skilled in the art that these steps may be performed in a different order while implementing the same method.

[0122] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are used, they are used only in a generic and descriptive sense and not for purposes of limitation.

Claims

1. A video decoding method for decoding a bitstream of a set of images, wherein the basic processing unit of the images includes a luminance component representing achromatic luminance information, comprising: decoding a first flag indicating whether a transform skip mode is enabled; determining, based on the first flag, whether to decode a second flag indicating whether transform skip residual coding is used to parse one or more residual samples of a transform skip block of a target slice, and In response to the first flag having a first value, decoding of the second flag is selectively bypassed.

2. The method according to claim 1, wherein The first value is equal to 0.

3. The method according to claim 1, wherein The first flag is a sequence-level flag, and the second flag is a stripe-level flag.

4. The method according to claim 1, wherein In response to the first flag having a first value, it is determined that the transform skip mode is not enabled.

5. The method according to claim 1, further comprising: In response to the first flag having a first value, determining that the transform skip residual encoding is for parsing one or more residual samples of the transform skip block of the target slice.

6. The method according to claim 1, further comprising: In response to the first flag having a first value, the second flag is determined to be equal to 0.

7. The method according to claim 1, further comprising: In response to the first flag having a second value, the second flag is decoded. The method according to claim 7 , wherein the second value is equal to 1.

9. A video encoding method for encoding a set of images, wherein the basic processing unit of the images includes a luminance component representing achromatic luminance information, the method comprising: encoding a first flag into a bitstream indicating whether a transform skip mode is enabled; determining whether to encode a second flag into the bitstream based on the value of the first flag, the second flag indicating whether transform skip residual coding is used to parse one or more residual samples of a transform skip block of a target slice; as well as In response to the first flag having a first value, encoding of the second flag is selectively bypassed. The method of claim 9 , wherein the first value is equal to 0.

11. The method according to claim 9, wherein The first flag is a sequence level flag, and the second flag is a slice level flag.

12. The method according to claim 9, further comprising: In response to the transform skip mode being not enabled, the first flag is set to have the first value.

13. The method according to claim 9, further comprising: In response to the first flag having a first value, determining that the transform skip residual encoding is for parsing one or more residual samples of the transform skip block of the target slice.

14. The method according to claim 9, further comprising: In response to the first flag having a second value, the second flag is encoded into the bitstream. The method of claim 14 , wherein the second value is equal to 1.

16. A non-transitory computer-readable storage medium storing an instruction set and a video bitstream, wherein the instruction set is executable by one or more processors to generate a video encoding method for encoding a set of images, wherein a basic processing unit of the image includes a luminance component representing achromatic luminance information, the method comprising: encoding a first flag into a bitstream indicating whether a transform skip mode is enabled; determining whether to encode a second flag into the bitstream based on the value of the first flag, the second flag indicating whether transform skip residual coding is used to parse one or more residual samples of a transform skip block of a target slice; as well as In response to the first flag having a first value, encoding of the second flag is selectively bypassed. The non-transitory computer-readable storage medium of claim 16 , wherein the first value is equal to 0.

18. The non-transitory computer-readable storage medium of claim 16, wherein: The first flag is a sequence level flag, and the second flag is a slice level flag.

19. The non-transitory computer-readable storage medium of claim 16 , further comprising: In response to the first flag having a first value, it is determined that the transform skip mode is not enabled.

20. The non-transitory computer-readable storage medium of claim 16, further comprising: In response to the first flag having a first value, determining that the transform skip residual encoding is for parsing one or more residual samples of the transform skip block of the target slice.

Citation Information

Patent Citations

  • Method, apparatus and system for encoding and decoding the transform units of a coding unit

    CN104782125A

  • Image processing method, and device for same

    CN110402580A

  • Signaling coding schemes for residual values in transform skip for video coding

    CN114503590A

  • Devices and methods for modifications of syntax related to transform SKIP for high efficiency video coding (HEVC)

    US20140146894A1

  • Method, apparatus and system for encoding and decoding the transform units of a coding unit

    WO2014071439A1