Method and system for cross-component adaptive loop filters - Patents.com

JP2024523435A5Pending Publication Date: 2025-07-03ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023578193
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-24
Filing Date
2022-06-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Current video coding standards like VVC face challenges in achieving optimal compression performance due to the use of a single filter shape for cross-component adaptive loop filters (CCALF) that does not adapt to the characteristics of different video content, and restrictions on filter coefficients, leading to suboptimal encoding efficiency.

Method used

Adaptive selection of filter shapes and coefficients for CCALF based on content characteristics, allowing for shape-adaptive cross-component filtering, removing power-of-two constraints on coefficient values, and applying block-level filter adaptation to improve compression performance.

Benefits of technology

Enhances video coding efficiency by optimizing filter shapes and coefficients for different video content types, resulting in improved compression performance and reduced bit rate without compromising quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method of video processing using a cross-component adaptive loop filter (CCALF) is provided, the method including filtering decoded video content using the CCALF, where the CCALF is a 24-tap 9x9 filter.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims benefit of priority to U.S. Provisional Patent Application No. 63 / 215,521, filed June 28, 2021, U.S. Provisional Patent Application No. 63 / 235,111, filed August 19, 2021, and U.S. Provisional Patent Application No. 17 / 808,933, filed June 24, 2022, which are incorporated by reference in their entireties herein.

[0002] Technical Field The present disclosure relates generally to video processing, and more specifically, to methods and systems for cross-component loop adaptive filters for video coding. [Background technology]

[0003] background

[0003] A video is a collection of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, that specify specific video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of the new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]

[0004] Disclosure Summary

[0004] An embodiment of the present disclosure provides a method of video processing using a cross-component adaptive loop filter (CCALF), the method including filtering decoded video content using the CCALF, where the CCALF is a 24-tap 9x9 filter.

[0005]

[0005] An embodiment of the present disclosure provides a method of video processing using a cross-component adaptive loop filter (CCALF), the method including filtering video content using the CCALF, where the CCALF is a 24-tap 9x9 filter.

[0006]

[0006] An embodiment of the present disclosure provides a non-transitory computer-readable medium that stores a bitstream, the bitstream including a first index associated with encoded video data, the first index identifying a cross-component adaptive loop filter (CCALF), the CCALF being a 24-tap 9x9 filter, the first index causing a decoder to filter the decoded video content using the 24-tap 9x9 filter.

[0007] BRIEF DESCRIPTION OF THE DRAWINGS

[0001] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings, in which:

[0001] Various features shown are not drawn to scale. [Brief description of the drawings]

[0008] [Figure 1]

[0002] FIG. 1 is a schematic diagram illustrating an example structure of a video sequence according to some embodiments of the present disclosure. [Figure 2A] 1 is a schematic diagram illustrating an example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B] 4 is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 3A]1 is a schematic diagram illustrating an example decoding process for a hybrid video coding system consistent with an embodiment of the present disclosure. [Figure 3B]

[0006] FIG. 2 is a schematic diagram illustrating another exemplary decoding process for a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 4] 1 is a block diagram of an exemplary device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Diagram 5] 1 illustrates a cross-component adaptive loop filter (CCALF) process, according to some embodiments of the present disclosure. [Figure 6]

[0009] 1 illustrates an example CCALF filter shape used in versatile video coding (VVC), in accordance with some embodiments of the present disclosure. [Figure 7]

[0010] 1 shows a flowchart of an example method for CCALF, according to some embodiments of the present disclosure. [Figure 8]

[0011] 5 illustrates five example shapes for CCALF, according to some embodiments of the present disclosure. [Figure 9]

[0012] 1 illustrates a flowchart of an example method for signaling a best CCALF filter shape for each coding tree block (CTB) in a CCALF-enabled slice, in accordance with some embodiments of the present disclosure. [Figure 10]

[0013] 1 illustrates example semantics of syntax elements related to the updated CCALF, in accordance with some embodiments of the present disclosure. [Figure 11]

[0014] 1 illustrates an example adaptive loop filter (ALF) signaled within an adaptation parameter set (APS) syntax in accordance with some embodiments of the present disclosure. [Figure 12]

[0015] 1 illustrates an example updated VVC specification for the CCALF process, according to some embodiments of the present disclosure. [Figure 13]

[0016] 1 illustrates an example modification of semantics related to CCALF, according to some embodiments of the present disclosure. [Figure 14]

[0017] 1 illustrates an example APS syntax table according to some embodiments of the present disclosure. [Figure 15]

[0018] 1 illustrates a flowchart of an example method for signaling parameters related to CCALF, according to some embodiments of the present disclosure. [Figure 16]

[0019] 1 illustrates an example method of block classification according to some embodiments of the present disclosure. [Figure 17]

[0020] 1 illustrates four calculations of gradients for block classification according to some embodiments of the present disclosure. [Figure 18]

[0021] 1 illustrates an example of APS syntax for signaling filter coefficients for one or more CCALF classes, according to some embodiments of the present disclosure. [Figure 19]

[0022] 1 illustrates an example semantic modification according to some embodiments of the present disclosure. [Figure 20]

[0023] 1 illustrates an example updated VVC specification according to some embodiments of the present disclosure. [Figure 21]

[0024] 1 illustrates a flowchart of a method for signaling filter coefficients using class merging in accordance with some embodiments of the present disclosure. [Figure 22]

[0025] 1 illustrates APS syntax for a method of signaling filter coefficients using class merging, according to some embodiments of the present disclosure. [Diagram 23]

[0026] 1 illustrates an example semantic modification according to some embodiments of the present disclosure. [Figure 24]

[0027] 3 illustrates an example 24-tap 9×9 filter according to some embodiments of the present disclosure. [Diagram 25]

[0028] 25 illustrates an example updated VVC specification for the CCALF process using the filter shown in FIG. 24, according to some embodiments of the present disclosure. [Figure 26]

[0029] 4 illustrates an example 28-tap filter according to some embodiments of the present disclosure. [Figure 27]

[0030] 1 illustrates an example 36-tap filter according to some embodiments of the present disclosure. [Figure 28]

[0031] 1 illustrates an example APS syntax table according to some embodiments of the present disclosure. [Figure 29]

[0032] 1 illustrates a flowchart of an example method for signaling the best filter for each CTB, according to some embodiments of the present disclosure. [Diagram 30]

[0033] 1 illustrates a flowchart of an exemplary method of video processing according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Detailed Description

[0034] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different figures represent the same or similar elements unless otherwise specified. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present invention. Rather, the implementations are merely examples of devices and methods consistent with aspects related to the present invention, as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions set forth herein shall control.

[0010]

[0035] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the Joint Video Experts Team (JVET) of the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.

[0011]

[0036] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing techniques that go beyond HEVC using the Joint Search Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0012]

[0037] The VVC standard is a more recent development and continues to include more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0013]

[0038] A video is a collection of static pictures (or "frames") arranged in a chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display function) can be used to display such pictures in chronological order. Furthermore, in some applications, a video capture device can transmit captured videos in real time to a video playback device (e.g., a computer with a monitor) for surveillance, conferencing, live broadcast, etc.

[0014]

[0039] To reduce the storage space and transmission bandwidth required by such applications, video may be compressed before storage and transmission, and decompressed before display. This compression and decompression may be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The module for compression is generally referred to as an "encoder," and the module for decompression is generally referred to as a "decoder." The encoder and decoder may collectively be referred to as a "codec." The encoder and decoder may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder may include circuits such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder may include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be referred to as a "transcoder."

[0015]

[0040] A video coding process can identify and keep useful information that can be used to reconstruct a picture and ignore information that is not important for the reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, then such a coding process can be called "lossy". Otherwise, such a coding process can be called "lossless". Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.

[0016]

[0041] Useful information of the picture being coded (called the "current picture") includes changes with respect to a reference picture (e.g. a previously coded and reconstructed picture). Such changes may include pixel position changes, luminance changes or color changes, of which position changes are the most important. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.

[0017]

[0042] A picture that is coded without reference to another picture (i.e., such a picture is its own reference picture) is called an "I-picture." If some or all of the blocks in a picture (e.g., blocks that generally refer to a portion of a video picture) are predicted using intra- or inter-prediction with one reference picture (e.g., unidirectional prediction), the picture is called a "P-picture." If at least one block in a picture is predicted with two reference pictures (e.g., bidirectional prediction), the picture is called a "B-picture."

[0018]

[0043] 1 illustrates an example structure of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0019]

[0044] As shown in FIG. 1, a video sequence 100 may include a series of pictures arranged in time along a timeline including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture and its reference picture is picture 102 itself. Picture 104 is a P-picture and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture and its reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the embodiments of reference pictures to the example shown in FIG.

[0020]

[0045] Typically, a video codec does not encode or decode an entire picture at once because such a task is computationally complex. Rather, a video codec may divide a picture into elementary segments and encode or decode a picture segment by segment. In this disclosure, such elementary segments are referred to as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Basic processing units can have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between efficiency of coding and the level of detail one wishes to preserve within the basic processing unit.

[0021]

[0046] A basic processing unit may be a logical unit that may include various types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) that represents achromatic luminance information, one or more chroma components (e.g., Cb and Cr) that represent color information, and related syntax elements, and the luma and chroma components may have the same size basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated on each of its luma and chroma components.

[0022]

[0047] There are multiple operational stages in video coding, examples of which are shown in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process, and therefore may be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing sub-units are also logical units that may include various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated on each of its luma and chroma components. Note that such division can be performed to further levels, depending on the processing needs. Note also that different stages can divide the basic processing unit using different schemes.

[0023]

[0048] For example, in a mode decision stage (one example of which is shown in FIG. 2B), an encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0024]

[0049] In another example, in the prediction stage (one example of which is shown in FIG. 2A-FIG. 2B), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.

[0025]

[0050] In another example, in the transform stage (one example of which is shown in FIG. 2A-2B), the encoder can perform a transform operation on the residual elementary processing sub-unit (e.g., CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder can further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0026]

[0051] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.

[0027]

[0052] In some implementations, to provide video encoding and decoding with parallel processing and error resilience capabilities, a picture can be divided into regions for processing, so that the encoding or decoding process does not have to depend on information for one region of the picture from any other region of the picture. In other words, each region of the picture can be processed independently. In this way, the codec can process different regions of the picture in parallel, thus increasing the efficiency of the coding. Furthermore, if data for a region is corrupted in processing or lost in network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience capabilities. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that various pictures in the video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0028]

[0053] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure is not limited to such embodiments.

[0029]

[0054] FIG. 2A illustrates a schematic diagram of an example of an encoding process 200A consistent with embodiments of the present disclosure. For example, the encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 may be divided by the encoder into elementary processing units, elementary processing sub-units, or regions for processing. In some embodiments, the encoder may perform the process 200A at the level of the elementary processing units for each original picture of the video sequence 202. For example, the encoder may perform the process 200A in an iterative manner, and the encoder may encode a elementary processing unit in one iteration of the process 200A. In some embodiments, the encoder may perform process 200A in parallel for each original picture region of video sequence 202 (eg, regions 114-118).

[0030]

[0055] In FIG. 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, the encoder may feed quantized transform coefficients 216, after quantization stage 214, to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a predicted BPU 208 to generate a prediction reference 224 that is used in the prediction stage 204 of the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0031]

[0056] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate predicted references 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0032]

[0057] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.

[0033]

[0058] In the prediction step 204, the encoder in the current iteration may receive the original BPU and a prediction reference 224 and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting the prediction data 206, which may be used to reconstruct the original BPU from the prediction data 206 and the prediction reference 224 as a predicted BPU 208.

[0034]

[0059] Ideally, the predicted BPU 208 may be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 generally differs slightly from the original BPU. To record such differences, the encoder may generate the predicted BPU 208 and then subtract it from the original BPU to generate the residual BPU 210. For example, the encoder may subtract the values ​​(e.g., grayscale values ​​or RGB values) of pixels of the predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and the predicted BPU 208, each pixel of the residual BPU 210 may have a residual value. Compared to the original BPU, the prediction data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant loss of quality. Thus, the original BPU is compressed.

[0035]

[0060] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," where each basis pattern is associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0036]

[0061] Different transform algorithms may use different basis patterns. For example, different transform algorithms may be used in transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform in transform stage 212 is invertible. That is, the encoder may reconstruct residual BPU 210 by the inverse operation of the transform (referred to as "inverse transform"). For example, to reconstruct pixels of residual BPU 210, the inverse transform may be to multiply the values ​​of corresponding pixels of the basis pattern by the associated respective coefficients and add the products to result in a weighted sum. In a video coding standard, both the encoder and the decoder may use the same transform algorithm (and thus the same basis pattern). Thus, the encoder may record only the transform coefficients, and the decoder may reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they may be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.

[0037]

[0062] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scale factor") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and the transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, which further compresses the transform coefficients. The quantization process is also invertible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0038]

[0063] The quantization stage 214 may be lossy because the encoder ignores the remainder of such a division in a rounding operation. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0039]

[0064] In the binary coding stage 226, the encoder may use a binary coding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm, to encode the prediction data 206 and the quantized transform coefficients 216. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as, for example, a prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transformation in the transformation stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. The encoder may use the output data of the binary coding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0040]

[0065] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0041]

[0066] It should be noted that other variations of the process 200A can be used to encode the video sequence 202. In some embodiments, the encoder can perform the stages of the process 200A in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be separated into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages in FIG. 2A.

[0042]

[0067] 2B shows a schematic diagram of another example encoding process 200B consistent with an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0043]

[0068] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels of one or more neighboring BPUs already coded in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions of one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0044]

[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that are encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the encoder may perform extrapolation at a pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. In intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0045]

[0070] In another example, in the temporal prediction stage 2044, the encoder may perform inter prediction. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform an operation of "motion estimation" to look for a matching region within the range of the reference picture (called a "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine the region as a match region. The match region may have different (e.g., smaller, equal, larger, or differently shaped) dimensions than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the match region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., as in picture 106 in FIG. 1), the encoder can look for a match region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the match region of each matching reference picture.

[0046]

[0071] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the match region, a motion vector associated with the match region, a number of reference pictures, weights associated with the reference pictures, etc.

[0047]

[0072] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation." Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move the matching regions of the reference picture according to the motion vectors, in which the encoder may predict the original BPU of the current picture. If multiple reference pictures are used (e.g., as in picture 106 of FIG. 1), the encoder may move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values ​​of the matching regions. In some embodiments, if the encoder assigns weights to the pixel values ​​of the matching regions of the respective matching reference pictures, the encoder may add a weighted sum of the pixel values ​​of the moved matching regions.

[0048]

[0073] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter predicted picture in which reference pictures (e.g., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0049]

[0074] Continuing to refer to the forward path of the process 200B, after the spatial prediction step 2042 and the temporal prediction step 2044, in a mode decision step 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of the process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the prediction mode selected, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0050]

[0075] In the reconstruction path of the process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU being encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). The encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused during the coding of the prediction reference 224. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset (SAO), adaptive loop filter (ALF), etc. The loop filtered reference picture may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., for use as an inter-prediction reference picture for future pictures of the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in a binary coding stage 226.

[0051]

[0076] FIG. 3A shows a schematic diagram of an example of a decoding process 300A consistent with an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression process (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0052]

[0077] In FIG. 3A, the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a "coded BPU") to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may feed a prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.

[0053]

[0078] The decoder may iteratively perform the process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0054]

[0079] In the binary decoding stage 302, the decoder may perform an inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as, for example, the prediction mode, parameters of the prediction operation, the type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.

[0055]

[0080] 3B shows a schematic diagram of another example 300B of a decoding process consistent with an embodiment of the present disclosure. The process 300B may be modified from the process 300A. For example, the process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., H.26x series). Compared to the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0056]

[0081] In process 300B, for a coded elementary processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights respectively associated with the reference pictures, positions (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, etc.

[0057]

[0082] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.

[0058]

[0083] In the process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of the process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture to which all BPUs are decoded), the decoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., for use as an inter-prediction reference picture for a future encoded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU, the prediction data includes the loop filter parameters.

[0059]

[0084] FIG. 4 is a block diagram of an example of a device 400 for encoding or decoding video consistent with an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes instructions described herein, the device 400 may become a dedicated machine for encoding or decoding video. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), application specific integrated circuits (ASICs), and the like. In some embodiments, processor 402 may be a set of processors grouped together as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors including processor 402a, processor 402b, and processor 402n.

[0060]

[0085] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in a process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, and the like. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical component.

[0061]

[0086] Bus 410 , such as an internal bus (eg, a CPU memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), etc., may be a communication device that transfers data between components within device 400 .

[0062]

[0087] For ease of explanation and without inviting ambiguity, the present disclosure collectively refers to the processor 402 and other data processing circuitry as the "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module, or may be fully or partially combined within any other component of the device 400.

[0063]

[0088] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0064]

[0089] In some embodiments, device 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like.

[0065]

[0090] It should be noted that a video codec (e.g., a codec that executes processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0066]

[0091] The present disclosure provides a method for signaling and deriving parameters related to a cross-component adaptive loop filter (CCALF).

[0067]

[0092] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luma component, one filter selected from 25 filters based on local gradient direction and activity is used for each 4×4 block. In addition to the ALF, a cross-component adaptive loop filter (CCALF) is adopted in VVC. The CCALF is designed to operate in parallel with the luma ALF. In the CCALF, a linear filter is used to filter the luma sample values ​​and generate residual corrections for the chroma samples. FIG. 5 illustrates a cross-component adaptive loop filter (CCALF) process according to some embodiments of the present disclosure. As shown in FIG. 5, the reconstructed luma samples 511 output from the luma sample adaptive offset (SAO) 510 are received as inputs of the luma (Y) ALF 520, the chroma (Cb) CC-ALF 530, and the chroma (Cr) CC-ALF 540, respectively. The reconstructed chroma samples 551, 561 output from the chroma SAOs 550, 560 are received as inputs of the chroma ALF 570. Finally, the luma sample Y is obtained after the luma ALF 520. The chroma sample Cb is obtained by adding the residual correction 531 output from the Cb CC-ALF 530 and the output 571 of the chroma ALF 570. The chroma sample Cr is obtained by adding the residual correction 541 output from the Cr CC-ALF 540 and the output 572 of the chroma ALF 570.

[0068]

[0093] 6 illustrates an example CCALF filter shape in VVC, according to some embodiments of the present disclosure. As shown in FIG 6, in VVC, an 8-tap hexagonal-shaped filter 600 is used in the CCALF process.

[0069]

[0094] However, the current VVC design of CCALF cannot achieve optimal compression performance due to the following reasons: A single filter shape may not be optimal for all types of content. In VVC, the filter coefficients are restricted to have only values ​​in the form of powers of two from the following set: {-64,-32,-16,-8,-4,-2,-1,0,1,2,4,8,16,32,64). Furthermore, in VVC, the same filter coefficients are applied to all types of blocks. However, the correlation between adjacent pixels may depend on the block characteristics (edge ​​direction, activity, etc.).

[0070]

[0095] To improve the compression performance of the CCALF filter, this disclosure provides a method for shape-adaptive cross-component filtering.

[0071]

[0096] In VVC, a single filter shape (8-tap hexagonal shape filter shown in FIG. 6) is used for filtering. However, the correlation between adjacent pixels may depend on the characteristics of the video content. A single filter shape is not optimal for all types of content. This disclosure proposes to adaptively select the filter shape based on the characteristics of the content. FIG. 7 shows a flowchart of an example method 700 for CCALF according to some embodiments of the present disclosure. The method 700 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 700. In some embodiments, the method 700 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 7, the method 700 may include steps S702 to S706.

[0072]

[0097] In step S702, statistics of the video content are collected and analyzed, for example, the gradient direction and activity of the luma samples.

[0073]

[0098] In step S704, an optimal filter shape and associated filter coefficients are selected from a set of predefined filter shapes based on statistics. For example, for each filter shape in the set of predefined filter shapes, the sum of absolute differences (i.e., SAD) is calculated by the filtered chroma samples and the original samples. The filter shape with the smallest SAD is selected as the optimal filter shape.

[0074]

[0099] In some embodiments, the filtered chroma samples may be obtained by the following steps: classifying the chroma samples into multiple classes using statistics, calculating filter coefficients associated with each class by minimizing the mean squared error of the reconstructed chroma samples relative to the original chroma samples, and applying the associated filter coefficients to the reconstructed chroma samples to obtain filtered chroma samples for each class.

[0075]

[0100] In some embodiments, the selection of the filter shape is content-adaptive, i.e., the shape of the filter is selected based on the content. The granularity of the shape selection can be sequence level, and / or frame level, and / or slice level, and / or block level. In some embodiments, multiple shape-adaptive cross-component filters are provided. FIG. 8 shows five exemplary shape-adaptive cross-component filters for CCALF according to some embodiments of the present disclosure. The proposed shape-adaptive cross-component filters can be applied with any shape / pattern and are not limited to the shapes listed in FIG. 8. With reference to FIG. 8, the number of taps of different filter shapes can be the same or different. Filter shape 4 is the same as filter 600 shown in FIG. 6. In this example, it is assumed that the filter shapes are predefined and known by both the encoder and the decoder before the encoding / decoding process is performed. In some embodiments, each filter shape can be represented by a shape index (e.g., filters_shape_idx) in the syntax. The number of filter taps associated with each shape corresponding to the shape index is indicated.

[0076] [Table 1]

[0077]

[0101] In step S706, one or more parameters indicating the selected filter shape and filter coefficients are signaled, for example, to the decoder. In the case of sequence level granularity, the encoder can signal one of the predefined filter shapes to the decoder. Thus, the same filter shape is used for filtering for the entire sequence. In the case of frame level granularity, the parameters are signaled for each frame. In the case of block level, the frame is divided into multiple non-overlapping blocks and the parameters are signaled for each block. The parameters signaled at the frame level or block level are further described below.

[0078]

[0102] In some embodiments, the encoder may signal the total number of predefined filter shapes to the decoder, or alternatively, the total number of filter shapes may be predefined and known to both the encoder and the decoder before starting the encoding / decoding process.

[0079]

[0103] For each frame / slice, N filters and associated shapes and coefficients are signaled to the decoder. For each CTB, the best filter shape (from the N filters signaled in the frame / slice) is selected and signaled to the decoder. FIG. 9 shows a flowchart of an example method for signaling the best filter shape of each CTB in a slice in CCALF, according to some embodiments of the present disclosure. The method 900 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 900. In some embodiments, the method 900 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIG. 9, the method may include steps S902 to S906.

[0080]

[0104] In step S902, the number of filters is signaled. For example, the number of filters is N. Thus, N filters are expected by the encoder and the decoder. The number of filter shapes is also known by both the encoder and the decoder.

[0081]

[0105] In step S904, a parameter describing each filter (eg, filter_shape_idx) is signaled, and the filter coefficients associated with each filter are signaled, so that the N filters with their associated filter coefficients are known by the decoder.

[0082]

[0106] In step S906, an index (e.g., i) is signaled indicating the filter to be used for each CTU in the frame / slice. The maximum value of i is N-1. The filter to be used is selected from the N filters. The index allows the filter shape and filter coefficients associated with the selected filter to be determined.

[0083]

[0107] In some embodiments, the filter shape is selected implicitly by both the encoder / decoder. For example, in the case of block-level implicit filter shape derivation, both the encoder / decoder can analyze the reconstructed block (before filtering), and the filter shape is implicitly derived based on the characteristics of the reconstructed block. That is, the filter shape is determined by both the encoder and the decoder based on the video content. Thus, after step S904, the filter to be used can be determined by the decoder.

[0084]

[0108] The details of the proposed shape adaptation method are further described. In the following example, it is assumed that five filter shapes (e.g., the five filter shapes shown in FIG. 8) are predefined and known at both the encoder and the decoder before the start of the decoding process. The maximum value of N is 20. The number of coefficients of the kth filter shape (e.g., noCoeff[k]) is also predefined as follows: noCoeff[]=[16,15,11,13,7]

[0085]

[0109] In VVC, filter coefficients are signaled in an adaptation parameter set (APS). In this example, the filter coefficients are also signaled in the APS. Note that the filter coefficients are not limited to being signaled only in the APS. Alternatively or in addition, the filter coefficients may be signaled in the picture header and / or slice header, etc. For each filter, a shape index is signaled in the bitstream by the APS to indicate the shape of the signaled filter. FIG. 10 illustrates example semantics of updated syntax elements according to some embodiments of the present disclosure. FIG. 11 illustrates example ALF APS syntax according to some embodiments of the present disclosure. As shown in FIG. 10, values ​​of syntax elements alf_cc_cb_filters_shape_idx[k] 1010 and alf_cc_cr_filters_shape_idx[k] 1020 are modified. 11, syntax elements alf_cc_cb_filters_shape_idx[k] and alf_cc_cr_filters_shape_idx[k] are signaled to indicate the shape of the kth filter (1110, 1120). The filter coefficients associated with each filter are further signaled (1130, 1140).

[0086]

[0110] FIG. 12 illustrates an exemplary updated VVC specification for the CCALF process, according to some embodiments of the present disclosure. Consistent with FIG. 9-FIG. 11, the italicized parts as shown in FIG. 12 indicate modifications compared to existing VVC methods. Before starting the CCALF process for each CTB, a variable CcAlfShape is derived (1210) that represents the filter shape of the current CTB. The filter coefficients and adjacent tap positions are adaptively selected (1220) based on the value of the variable CcAlfShape.

[0087]

[0111] This disclosure also provides a method for removing the power-of-two constraint on filter coefficient values.

[0088]

[0112] In VVC, the CCALF filter coefficients are restricted to values ​​in the form of powers of two from the following set: {-64,-32,-16,-8,-4,-2,-1,0,1,2,4,8,16,32,64). This disclosure proposes removing this constraint to improve compression performance. In some embodiments, the values ​​of the filter coefficients can be any integer value between -64 and +64. In some embodiments, the filter coefficients are coded using variable length codes (i.e., ue(v) coding), whereas in VVC, the mapped coefficients are coded using fixed length codes. FIG. 13 illustrates the semantic changes compared to VVC, according to some embodiments of the present disclosure. FIG. 14 illustrates an example APS syntax table, according to some embodiments of the present disclosure. With reference to FIG. 13, the constraint on the values ​​of the filter coefficients is removed (1310). With reference to FIG. 14, the filter coefficients are coded using variable length codes (i.e., ue(v) coding) (1410).

[0089]

[0113] This disclosure also provides a method for block-level filter coefficient adaptation.

[0090]

[0114] In VVC, filter coefficients are signaled at the slice level by the APS. The same set of filter coefficients may not be optimal for all parts of a frame / slice. In some embodiments, a slice is divided into multiple non-overlapping M×N blocks. Each M×N block is then classified into one of the predefined classes based on the characteristics of the reconstructed block (before filtering). The blocks are classified using the reconstructed samples (before filtering), which are already known to the decoder before the start of the CCALF process. Thus, no signaling is required to indicate the class of the block. FIG. 15 shows a flowchart of a method 1500 for signaling parameters related to CCALF according to some embodiments of the present disclosure. The method 1500 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B) or may be performed by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform method 1500. In some embodiments, method 1500 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., device 400 of FIG. 4). The method may include steps S1502-S1506.

[0091]

[0115] In step S1502, a number of classes are predefined. The classification method is further described below.

[0092]

[0116] In step S1504, statistics of all classes are collected for each frame / slice.

[0093]

[0117] In step S1506, an optimal filter is generated for each class. For example, the optimal filter may be generated based on the characteristics of the class.

[0094]

[0118] In step S1508, the filter coefficients associated with the optimal filter for each class are signaled in the bitstream at the slice / frame level.

[0095]

[0119] Additionally, a method for block classification is provided.

[0096]

[0120] FIG. 16 shows a diagram of a method for block classification according to some embodiments of the present disclosure. With reference to FIG. 16, an image is divided into 16 blocks. In this example, before the CCALF process begins, each M×N block of the luma component can be classified as one of the predefined classes. The maximum number of classes is predefined and known to both the encoder and the decoder before the process begins.

[0097]

[0121] An example of block classification is described below. Here, assume that the maximum number of classes is 15 and the classification block size is 4×4, as shown in FIG. 16. In this example, the classification method is similar to the classification of the adaptive loop filter of VVC. However, the classification method of the present disclosure is not limited.

[0098]

[0122] In this example, for the luma component, each 4x4 block is classified into one of 15 classes. The classification index C is determined by its directionality D and the quantized value of the activity

number

number

[0099]

[0123] D and

number

number

[0100]

[0124] To reduce the complexity of block classification, a subsampled 1-D Laplacian calculation is applied. Figure 17 shows four calculations of gradients for block classification according to some embodiments of the present disclosure. As shown in Figure 17, the same subsample position is used for all directional gradient calculations, e.g., vertical gradient 1710, horizontal gradient 1720, and two diagonal gradients 1730, 1740.

[0101]

[0125] Next, the maximum and minimum values ​​of the horizontal and vertical gradients D are calculated.

number

[0102]

[0126] The maximum and minimum values ​​of the gradients in the two diagonal directions are

number

[0103]

[0127] These values ​​are compared with each other and with two thresholds t1 and t2 to derive the value of the directionality D. Step 1.

number

number

number

number

[0104]

[0128] Activity value A

number

number

[0105]

[0129] The above classification method is merely an exemplary method, and the specific block size and classification method are not limited in this disclosure. For example, any block classification method known in the art can be used. Alternatively, the classification method of the Luma ALF method can be directly reused without performing the classification of CCALF itself.

[0106]

[0130] The present disclosure also provides a method for signaling filter coefficients.

[0107]

[0131] In some embodiments, for each frame / slice, the filter coefficients of each class are signaled in the bitstream. For example, in the above example, for each slice, the encoder can generate a total of 15 filters (one for each class) and signal their parameters to the decoder. Figure 18 shows an example of APS syntax for signaling filter coefficients of all classes according to some embodiments of the present disclosure. As shown in Figure 18, 15 filters of 15 classes are signaled with associated filter coefficients of CC-Cb filter 1810 and CC-Cr filter 1820, respectively. It can be understood that the number of classes is not limited to 15. Figure 19 shows semantic changes compared to VVC according to some embodiments of the present disclosure. In some embodiments, each luma pixel is classified as one of the predefined classes. Figure 20 shows details of specification changes compared to VVC according to some embodiments of the present disclosure.

[0108]

[0132] As shown in FIG. 18, the filter coefficients of each filter corresponding to each class are signaled, and such signaling may cause significant overhead signaling bits. In order to reduce the overhead signaling bits, an alternative approach to signaling filter coefficients is proposed. In some embodiments, the encoder uses a class merging method in which the number of classes is adaptively selected based on the content. FIG. 21 shows a flowchart of a method 2100 of signaling filter coefficients using class merging according to some embodiments of the present disclosure. The method 2100 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B) or may be performed by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 2100. In some embodiments, the method 2100 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., device 400 of FIG. 4). As shown in FIG. 21, the method 2100 may include steps S2102 and S2104.

[0109]

[0133] In step S2102, the filters are merged into one or more merge classes. Each merge class includes one or more filters. Each merge class has its own index. For example, the number of classes is N (i.e., the number of filters is N), and the merge class can include M classes (i.e., filters), where N is an integer equal to or greater than M. For example, in the example discussed above, 15 filters are generated for 15 classes. The maximum number of merge classes can be 15 with indexes from 0 to 14. In some embodiments, for example, 15 filters can be merged into 6 merge classes. Each merge class can include one or more filters, and the filters can be merged based on the similarity of the filter coefficients of the filters. The filters of the merge class can be used for processing.

[0110]

[0134] In step S2104, the filter coefficients of each merge class are signaled. In some embodiments, the filter 600 shown in Figure 6 is used, and the number of filter coefficients is 7.

[0111]

[0135] Since the number of merged classes is less than or equal to the number of classes, this method can reduce overhead signaling bits for both the filters and the filter coefficients.

[0112]

[0136] For example, in syntax, a filter set may be represented by the syntax element filter_set, which may be defined as follows:

[0113] [Table 2]

[0114]

[0137] Figure 22 illustrates the APS syntax of a method for signaling filter coefficients using class merging, according to some embodiments of the present disclosure. Referring to Figure 22, the merge class and the corresponding filter are signaled (2210, 2220). Figure 23 illustrates the semantic changes compared to VVC, according to some embodiments of the present disclosure.

[0115]

[0138] In some embodiments, a single shape CCALF filter is used for all CTUs of a slice. Figure 24 illustrates an example 24-tap 9x9 filter 2400 according to some embodiments of the present disclosure. As shown in Figure 24, the filter 2400 has a roughly cross shape. Specifically, the 24-tap 9x9 filter 2400 is defined as follows:

number

[0116]

[0139] 25 illustrates modifications of the CCALF process using filter 2400 compared to the VVC specification, according to some embodiments of the present disclosure. The italicized parts indicate the modifications compared to existing VVC methods. With reference to FIG. 25, the number of filter coefficients 2510 is modified to 24, and the variation sum 2520 is modified accordingly.

[0117]

[0140] In some embodiments, more filters are provided. Figures 26 and 27 show an example 28-tap filter 2600 and an example 36-tap filter 2700, respectively, according to some embodiments of the present disclosure. It can be understood that if the 28-tap filter 2600 or the 36-tap filter 2700 is used, the number of filter coefficients and the variation sum need to be modified accordingly.

[0118]

[0141] This disclosure also provides a method for signaling filter coefficients. In VVC, CCALF filter coefficients are restricted to have only values ​​in the form of powers of two from the following set: {-64, -32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32, 64). In the VVC specification, fixed length codes are used to signal the index of the filter coefficients.

[0119]

[0142] In some embodiments, the values ​​of the filter coefficients are restricted to powers of two. However, instead of using fixed-length codes, variable-length codes (more specifically, unary coding) are used to signal the filter coefficients. Figure 28 shows an APS syntax table of the proposed method according to some embodiments of the present disclosure. With reference to Figure 28, the signaled filter coefficients are coded with a variable length (e.g., ue(v)) 2810. Thus, the filter coefficients can be more flexible, and the constraint of values ​​in the form of powers of two is removed.

[0120]

[0143] The present disclosure further provides a method for signaling CTB level-mapped filter index. In VVC, for each frame / slice, the number and coefficients of N filters are signaled. For each CTB, the encoder can select the best filter, which is signaled to the decoder. Figure 29 shows a flowchart of an example method for signaling the best filter for each CTB according to some embodiments of the present disclosure.

[0121]

[0144] In step S2902, the number of filters is signaled for each frame / slice. For example, the number of filters is N. Thus, N is known by both the encoder and the decoder. The number of filter shapes is also known by both the encoder and the decoder.

[0122]

[0145] In step S2904, the filter coefficients of each filter are signaled.

[0123]

[0146] In step S2906, a parameter indicating the best filter for each CTB in the slice is signaled. For example, the parameter may be a filter index (e.g., filter_idc), and the maximum value of the index is N-1. In some embodiments, if the maximum value of the index is N and the filter index (e.g., filter_idc) is equal to 0, no filtering is applied.

[0124]

[0147] In the method provided by the present disclosure, for each CTB, instead of directly signaling the filter index (e.g., filter_idc), the mapped filter index is signaled towards the decoder. FIG. 30 shows a flowchart of an example method 3000 of video processing according to some embodiments of the present disclosure. The method 3000 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B) or may be performed by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform the method 3000. In some embodiments, the method 3000 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). As shown in FIG. 30, the method 3000 may include steps S3002-S3008.

[0125]

[0148] In step S3002, a predefined parameter (eg, filter_idc_pos0) is signaled. The value of the predefined parameter (eg, filter_idc_pos0) may be any value from 0 to N, where N is the number of filters to be signaled.

[0126]

[0149] In step S3004, a first filter index (eg, filter_idc) is signaled in response to the default parameter being equal to 0. That is, method 2900 is performed.

[0127]

[0150] In step S3006, in response to the predefined parameter not being equal to 0, a mapped filter index (e.g., mapped_filter_idc) is derived from the first filter index (e.g., filter_idc). In some embodiments, the predefined parameter is assumed to be equal to 1 and is known by both the encoder and the decoder. The mapped filter index may be derived by the following steps: in response to the first filter index being equal to 0, a value of the mapped filter index is set as a value of the predefined parameter (e.g., 1), in response to the first filter index being less than or equal to the predefined parameter, a value of the mapped filter index is set as a value of the first filter index minus 1, and in other cases, a value of the mapped filter index is set as a value of the first filter index. For example, pseudocode may be written as follows:

[0128] [Table 3]

[0129]

[0151] In step S3008, the mapped_filter_idc is signaled in the bitstream.

[0130]

[0152] Since filter_idc is not equal binarized, if filter_idc is not equal to 0, the number of bits is reduced.

[0131]

[0153] At the decoder side, a first filter index is derived from the mapped filter index. In some embodiments, the default parameter is assumed to be equal to 1 and is known by both the encoder and the decoder. The first filter index is derived by the following steps: in response to the value of the mapped filter index being equal to the value of the predetermined parameter, set the first filter index as 0, i.e., no filtering is applied; in response to the value of the mapped filter index being less than the value of the predetermined parameter, set the value of the first filter index as the value of the mapped filter index plus 1; and for other values, set the value of the first filter index as the value of the mapped filter index. For example, the pseudo code can be written as follows:

[0132] [Table 4]

[0133]

[0154] In the above pseudocode, it is assumed that filter_idc_pos0 is fixed (equal to 1) for the coded video sequence. In some embodiments, the value of filter_idc_pos0 may also be adaptively derived in each CTB. One way to derive the value of filter_idc_pos0 is based on previously decoded adjacent CTBs. For example, if the filter_idc of the upper or left CTB is a non-zero value, the filter_idc_pos0 of the current CTB may be set to 1, otherwise filter_idc_pos0 is equal to 0. For example, the pseudocode may be written as follows:

[0134] [Table 5]

[0135]

[0155] This disclosure also provides a method to improve the context coding of filter indexes. In VVC, two syntax elements alf_ctb_cc_cb_idc and alf_ctb_cc_cr_idc are signaled in the bitstream to indicate the index of the filter used for the target CTB. For both alf_ctb_cc_cb_idc and alf_ctb_cc_cr_idc, only the first bin is context coded and the remaining bins are bypass coded.

[0136]

[0156] In some embodiments, both the first bin and the second bin are context coded, and the remaining bins are bypass coded.

[0137]

[0157] For those skilled in the art of video coding, the above embodiments may be combined into one embodiment.

[0138]

[0158] The embodiments may be further described using the following clauses: 1. A method of video processing using a cross-component adaptive loop filter (CCALF), comprising filtering decoded video content using the CCALF, the CCALF being a 24-tap 9x9 filter. 2.24 tap 9x9 filter is

number

number

number

number

number

[0139]

[0159] In some embodiments, a non-transitory computer readable storage medium is also provided. In some embodiments, the medium may store all or part of a video bitstream having indexes, flags and / or syntax elements indicating parameters related to CCALF. In some embodiments, the medium may store instructions that may be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions of the same. The apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces and / or memory.

[0140]

[0160] It should be noted that relative terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Moreover, the terms "comprise," "have," "contain," and "include," as well as other similar forms, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive listing of such items or to be limited to only the items listed.

[0141]

[0161] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if it is stated that a database may include A or B, the database may include A, B, A and B, unless otherwise specified or impracticable. As a second example, if it is stated that a database may include A, B or C, the database may include A, B, C, A and B, A and C, B and C, A and B and C, unless otherwise specified or impracticable.

[0142]

[0162] It will be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple ones of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0143]

[0163] In the above specification, the embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as examples only, with a true scope and spirit of the invention being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. As such, one skilled in the art can appreciate that the steps may be performed in different orders while implementing the same method.

[0144]

[0164] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications thereto may be made, and therefore, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

A method for decoding a bitstream, comprising: receiving the bitstream, and decoding the bitstream to output a video sequence , wherein the decoding comprises: filtering the decoded video content using a cross-component adaptive loop filter (CCALF), the CCALF being a 24-tap 9×9 filter. Claim 2 The 24-tap 9×9 filter is 【Number 1】 having a cross shape defined as, C 0 ~C 24 wherein C~C are filter coefficients of the 24-tap 9×9 filter, the method according to claim 1. Claim 3 The decoded video content includes video slices, and the 24-tap 9×9 filter is applied to all coding tree units (CTUs) of the video slices, the method according to claim 1. Claim 4 The values of the filter coefficients are integer values from -64 to +64, the method according to claim 1. Claim 5 The filter coefficients are encoded using variable-length codes, the method according to claim 1. A method for encoding a video sequence, comprising: receiving the video sequence, and encoding the video sequence by filtering video content using a cross-component adaptive loop filter (CCALF) , the CCALF being a 24-tap 9×9 filter. Claim 6 The 24-tap 9×9 filter is 【Number 2】 having a cross shape defined as, C 0 ~C 24 where C 0 to C 24 are filter coefficients of the 24-tap 9×9 filter, the method according to claim 6. Claim 7 The video content includes video slices, and the 24-tap 9×9 filter is applied to all coding tree units (CTUs) of the video slices, the method according to claim 6. Claim 8 The values of the filter coefficients are integer values from -64 to +64, the method according to claim 6. Claim 9 The filter coefficients are encoded using variable-length codes, the method according to claim 6. A method for storing a bitstream of a video sequence, comprising: receiving the video sequence, encoding one or more pictures of the video sequence, generating a bitstream, and storing the bitstream in a non-transitory computer-readable storage medium , wherein the encoding comprises: filtering video content using a cross-component adaptive loop filter (CCALF), the CCALF being a 24-tap 9×9 filter. Claim 11 The 24-tap 9×9 filter is [Number 3] having a cross shape defined as, C 0 ~C 24 is the filter coefficient of the 24-tap 9×9 filter, the method according to claim 11. The method according to claim 11, wherein the video content includes video slices, and the 24-tap 9×9 filter is applied to all coded tree units (CTUs) of the video slices. The method according to claim 11, wherein the value of the filter coefficient is an integer value from -64 to +64. The method according to claim 11, wherein the filter coefficient is coded using a variable length code.