Extended taps based on multiple input sources for adaptive loop filters in video coding

By using an adaptive loop filter (ALF) with an extended tap in video encoding and decoding, and using the intermediate filtering results of other filters as input, the problem of inefficient encoding and decoding in the prior art is solved, and the compression of bitstreams and video quality is improved.

CN120051987APending Publication Date: 2025-05-27DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072845.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively utilize the intermediate filter results of other filters as input to the adaptive loop filter (ALF) in video encoding and decoding, resulting in low encoding and decoding efficiency.

Method used

It is proposed to apply an adaptive loop filter (ALF) with an extended tap to video data processing, use the intermediate filtering result of the second filter as input to the extended tap, and perform conversion between visual media data and bit stream based on the ALF.

Benefits of technology

Through this method, the efficiency of video encoding and decoding is improved, the size of the bitstream is reduced, and the quality of video data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051987A_ABST
    Figure CN120051987A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. The mechanism includes determining to apply an adaptive loop filter (ALF) with an extended tap to a picture in a video. The intermediate filtering result of the second filter is used as the input of the extension tap. A conversion between the visual media data and the bitstream is performed based on the ALF.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority and benefit of International Patent Application No. PCT / CN2022 / 124841, filed on October 12, 2022. The content of the aforementioned patent application is incorporated herein by reference in its entirety. Technical field

[0003] The present disclosure relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background art

[0004] Digital video occupies the largest bandwidth used on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the invention

[0005] The first aspect relates to a method for processing video data, which includes: determining to apply an adaptive loop filter (ALF) with extended taps to a picture in a video, where the intermediate filtering result of a second filter is used as the input to the extended taps; and performing a conversion between visual media data and a bitstream based on the ALF.

[0006] The second aspect relates to a device for processing video data, which includes: a processor; and a non - transitory memory having instructions thereon, where the instructions, when executed by the processor, cause the processor to perform any of the foregoing aspects.

[0007] The third aspect relates to a non - transitory computer - readable medium, which includes a computer program product for use in a video codec device. The computer program product includes computer - executable instructions stored on the non - transitory computer - readable medium, such that when executed by a processor, the video codec device performs the method of any of the foregoing aspects.

[0008] The fourth aspect relates to a non - transitory computer - readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device. The method includes: determining to apply an adaptive loop filter (ALF) with extended taps to a picture in a video, where the intermediate filtering result of a second filter is used as the input to the extended taps; and generating the bitstream based on the determination.

[0009] The fifth aspect relates to a method for storing a bitstream of a video, which includes: determining to apply an adaptive loop filter (ALF) with extended taps to a picture in a video, where the intermediate filtering result of a second filter is used as the input to the extended taps; generating the bitstream based on the determination; and storing the bitstream in a non - transitory computer - readable recording medium.

[0010] The sixth aspect relates to a method, apparatus or system described in this patent document.

[0011] For clarity, any of the foregoing embodiments may be combined with one or more of the other foregoing embodiments to produce new embodiments within the scope of the present disclosure.

[0012] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more fully understand the present disclosure, reference is now made to the following brief description in conjunction with the drawings and the specific embodiments, in which like reference numerals represent like components.

[0014] Figure 1 An example of the nominal vertical and horizontal positions of 4:2:2 luminance and chrominance samples in a picture is shown.

[0015] Figure 2 An example encoder block diagram is shown.

[0016] Figure 3 An example picture divided into raster scan strips is shown.

[0017] Figure 4 An example picture divided into rectangular scan strips is shown.

[0018] Figure 5 An example picture partitioned into bricks is shown.

[0019] FIG. 6 shows an example of a coding tree block (CTB) across a picture boundary.

[0020] Figure 7 An example of an intra prediction mode is shown.

[0021] Figure 8 An example of a block boundary in a picture is shown.

[0022] Figure 9 An example of pixels involved in filter use is shown.

[0023] Figure 10 An example of the filter shape of ALF is shown.

[0024] Figure 11 An example of transform coefficients supported by a 5×5 rhombus filter is shown.

[0025] Figure 12 An example of relative coordinates supported by a 5×5 rhombus filter is shown.

[0026] Figure 13 and Figure 14 shows an example shape of an airspace tap.

[0027] Figures 15 to 18 shows an example ALF filter with extended taps.

[0028] Figure 19 is a block diagram showing an example video processing system.

[0029] Figure 20 is a block diagram of an example video processing device.

[0030] Figure 21 is a flowchart of an example method for video processing.

[0031] Figure 22 is a block diagram showing an example video codec system.

[0032] Figure 23 is a block diagram showing an example encoder.

[0033] Figure 24 is a block diagram showing an example decoder.

[0034] Figure 25 is a schematic diagram of an example encoder. Detailed Description

[0035] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, any number of techniques may be used to implement the disclosed systems and / or methods, whether currently known or to be developed. The present disclosure should not in any way be limited to the exemplary implementations, figures, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the full scope of the appended claims and their equivalents.

[0036] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. Additionally, the techniques described herein are applicable to other video codec protocols and designs.

[0037] 1. Initial Discussion

[0038] This document relates to video codec technology. Specifically, it relates to loop filters and other codec tools in image / video coding. These ideas can be applied, either alone or in various combinations, to video codecs such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video coding technologies.

[0039] 2. Abbreviations

[0040] The present disclosure includes the following abbreviations. Advanced Video Coding (Rec. ITU-T H.264|ISO / IEC 14496-10) (AVC), Coding Picture Buffer (CPB), Clean Random Access (CRA), Coding Tree Unit (CTU), Coding Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraint Information (GCI), International Organization for Standardization (ISO), International Electrotechnical Commission (IEC), High Efficiency Video Coding (also known as Rec. ITU-T H.265|ISO / IEC 23008-2, (HEVC)), Joint Exploration Model (JEM), Motion Constraint Tile Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Tier and Level (PTL), Picture Unit (PU), Reference Picture Resampling (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Coding Layer (VCL), Video Parameter Set (VPS), Versatile Video Coding (also known as Rec. ITU-T H.266|ISO / IEC 23090-3) (VVC), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Coding Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Coding Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).

[0041] 3. Video Coding Standards

[0042] Video coding standards have evolved mainly through the development of International Telecommunication Union - Telecommunication Standardization Sector (ITU-T) and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed Moving Picture Experts Group (MPEG-1) and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially launched, the Joint Video Exploration Team (JVET) was renamed the Joint Video Experts Team (JVET). VVC is a coding standard aimed at reducing the bitrate by 50% compared to HEVC. The VVC working draft and the VVC Test Model (VTM) are continuously updated.

[0043] An example version of the VVC draft (i.e., Versatile Video Coding (Draft 10)) can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. An example version of the reference software for VVC (called VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.

[0044] The ITU-T VCEG and ISO / IEC MPEG Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 are studying the potential need to standardize future video coding technologies that will significantly exceed the current VVC standard in terms of compression capabilities. Such future standardization actions could take the form of an extension of VVC or a completely new standard. These groups are jointly conducting this development activity in a joint collaborative effort called JVET to evaluate the compression technology designs proposed by their experts in this field. The first Exploration Experiment (EE) was established by JVET, and a reference software called the Enhanced Compression Model (ECM) is in use. The test model ECM is continuously updated.

[0045] 3.1 Color Space and Chroma Subsampling

[0046] A color space (also known as a color model (or color system)) is a mathematical model that describes a range of colors as a tuple of numbers, for example as 3 or 4 values or color components (e.g., RGB). Generally speaking, a color space is a concrete form of a coordinate system and a subspace. For video compression, the most commonly used color spaces are luminance, blue-difference chrominance, and red-difference chrominance (YCbCr) and red, green, and blue (RGB).

[0047] YCbCr, Y′CbCr, or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a collective term for a series of color spaces that are used in video systems and digital photography as part of the color image processing pipeline. Y' is the luminance component, and CB and CR are the blue-difference chrominance component and the red-difference chrominance component. Y′ (with an apostrophe) is different from Y, which means that the light intensity is non-linearly encoded based on the gamma-corrected RGB primary colors.

[0048] Chroma subsampling is the practice of encoding an image by achieving a lower resolution for chrominance information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 3.1.1. 4:4:4

[0050] In 4:4:4, each of the three Y′CbCr components has the same sampling rate. Therefore, there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2. 4:2:2

[0052] In 4:2:2, the two chrominance components are sampled at half the sampling rate of luminance. The horizontal chrominance resolution is halved while the vertical chrominance resolution remains the same. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference. Figure 1 Examples of the nominal vertical and horizontal positions in the 4:2:2 color format in a picture are shown. 3.1.3 4:2:0

[0054] In 4:2:0, compared to 4:1:1, the horizontal sampling is doubled. However, since the Cb and Cr channels are sampled only on alternate rows in this scheme, the vertical resolution is halved, and the data rate thus remains unchanged. Cb and Cr are each downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, each with different horizontal and vertical sampling positions. In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are located between pixels vertically (i.e., interleaved). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are arranged in an interleaved manner, at the middle positions between alternate luma samples. In 4:2:0 DV, Cb and Cr are horizontally co-located. Vertically, they are co-located on alternate rows.

[0055]

[0056]

[0057] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and Separate_colour_plane_flag

[0058] 3.2 Example encoding / decoding processes of video codecs

[0059] Figure 2 An example of the encoder block diagram of VVC is shown, which includes three loop filter modules: the deblocking filter (DF), sample adaptive offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets respectively and by applying a finite impulse response (FIR) filter. The encoding / decoding side information is used to signal this offset and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to try to capture and fix the artifacts generated by the previous stages.

[0060] 3.3 Definition of video / encoding / decoding units

[0061] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs that cover a rectangular area of the picture. A slice can be divided into one or more tiles, each tile including a certain number of CTU rows within the slice. A slice that is not divided into multiple tiles can also be called a tile. However, a tile that is a proper subset of a slice cannot be called a slice. A strip contains several slices of a picture or several tiles of a slice.

[0062] Supports two stripe modes, namely, the raster scan stripe mode and the rectangular stripe mode. In the raster scan stripe mode, a stripe contains a series of slices in the slice raster scan of a picture. In the rectangular stripe mode, a stripe contains multiple tiles of a picture, and these tiles together form a rectangular area of the picture. The tiles within a rectangular stripe are arranged in the order of the tile raster scan of the stripe. Figure 3 Shows an example of raster scan stripe segmentation of a picture (with 18 by 12 luminance CTUs), where the picture is divided into 12 slices and 3 raster scan stripes.

[0063] Figure 4 Shows an example of rectangular stripe segmentation of a picture (with 18 by 12 luminance CTUs), where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular stripes.

[0064] Figure 5 Shows an example of a picture segmented into slices, tiles, and rectangular stripes, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular stripes.

[0065] 3.3.1 CTU / CTB Sizes

[0066] In VVC, the CTU size (signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2) can be as small as 4×4.

[0067] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0068]

[0069]

[0070]

[0071] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:

[0072] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)

[0073] CtbSizeY = 1 << CtbLog2SizeY (7-10)

[0074] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2(7-11)

[0075] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)

[0076] MinTbLog2SizeY = 2 (7-13)

[0077] MaxTbLog2SizeY = 6 (7-14)

[0078] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)

[0079] MaxTbSizeY = 1 << <MaxTbLog2SizeY (7-16)

[0080] PicWidthInCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY) (7-17)

[0081] It should be noted that there seems to be a small error in the original text at "MaxTbSizeY = 1<<MaxTbLog2SizeY (7-16)", which should probably be "MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)". The translation is done based on the provided text as accurately as possible.PicHeightInCtbsY = Ceil(pic_height_in_luma_samples ÷ CtbSizeY) (7-18)

[0082] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)

[0083] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)

[0084] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)

[0085] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)

[0086] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)

[0087] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)

[0088] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)

[0089] 3.3.2 CTUs in a Picture

[0090] Assume that the CTB / LCU size is represented by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For Figures 6A to 6C those CTBs shown in Figure 6A , the CTB size is still equal to M×N. However, as Figure 6B shown, the lower boundary of the CTB is outside the picture, or as Figure 6C shown, the right boundary of the CTB is outside the picture, or as

[0091] 3.4 Intra Prediction

[0092] To capture any edge directions presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. Figure 7 The extended directional modes are shown in, and the planar and DC modes remain unchanged. These denser directional intra prediction modes apply to all block sizes as well as for luma and chroma intra prediction.

[0093] The angular intra prediction directions can be defined as from 45 degrees to -135 degrees in the clockwise direction, as Figure 7 shown. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled and remapped to the indices of the wide-angle modes after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode coding and decoding remain the same.

[0094] In HEVC, each intra-coded block has a square shape and the length of each side of the block is a power of 2. Thus, no division operation is required to generate the intra predictor using the DC mode. In VVC, the blocks can have rectangular shapes and in general, a division operation must be used for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.

[0095] 3.5 Inter Prediction

[0096] For each inter-predicted CU, the motion parameters include motion vectors, reference picture indices, reference picture list use indices, and extension information of the new coding features of VVC for sample generation for inter prediction. The motion parameters can be signaled in an explicit or implicit way. When a CU is coded using the skip mode, the CU is associated with a PU and there are no significant residual coefficients, no coded motion vector differences, and / or reference picture indices. The Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates and the extended schedule introduced in VVC. The Merge mode can be applied to any inter-predicted CU, not just limited to the skip mode. An alternative to the Merge mode is to explicitly send the motion parameters, where for each CU, the motion vectors, the corresponding reference picture indices for each reference picture list, the reference picture list use flags, and other useful information are explicitly signaled.

[0097] 3.6 Deblocking Filter

[0098] Deblocking filtering is an example loop filter in video codecs. In VVC, the deblocking filtering process is applied to CU boundaries, transform sub-block boundaries, and prediction sub-block boundaries. Prediction sub-block boundaries include prediction unit boundaries introduced by sub-block based temporal motion vector prediction (SbTMVP) and affine modes. Transform sub-block boundaries include transform unit boundaries introduced by sub-block transform (SBT) and intra-subdivision (ISP) modes, as well as transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first horizontally filtering the vertical edges of the entire picture, and then vertically filtering the horizontal edges. This specific order enables the application of multiple horizontal filtering or vertical filtering processes in parallel threads. The filtering process can also be implemented on a per-CTB basis with very little processing delay.

[0099] First, the vertical edges in the picture are filtered. Then, using the samples modified by the vertical edge filtering process as input, the horizontal edges in the picture are filtered. The vertical and horizontal edges in the CTB of each CTU are processed separately according to the coding unit. Filtering of the vertical edges of the coding blocks in the coding unit is performed in the geometric order of the coding blocks, starting from the edge on the left side of the coding block and proceeding towards the edge on the right side of the coding block. Filtering of the horizontal edges of the coding blocks in the coding unit is performed in the geometric order of the coding blocks, starting from the edge on the top of the coding block and proceeding towards the edge on the bottom of the coding block.

[0100] Figure 8 is an illustration 800 of sample 802 within the 8×8 block of sample 804. As shown, illustration 800 includes horizontal block boundaries and vertical block boundaries on 8×8 grids 806, 808 respectively. Additionally, illustration 800 shows non-overlapping blocks of 8×8 samples 810 that can be deblocked in parallel.

[0101] 3.6.1 Boundary Decision

[0102] Filtering is applied to 8×8 block boundaries. Additionally, such a boundary must be a transform block boundary or a coding sub-block boundary, for example, a transform block boundary or a coding sub-block boundary due to the use of affine motion prediction (ATMVP). For other boundaries, the deblocking filter is disabled.

[0103] 3.6.2 Boundary Strength Calculation

[0104] For a transform block boundary / coding sub-block boundary, if the boundary is within the 8×8 grid, the boundary can be filtered, and the setting of bS[xDi][yDj] for that edge (where [xDi][yDj] represents coordinates) is defined in Tables 2 and 3 respectively.

[0105]

[0106] Table 2. Boundary strength (when SPS IBC is disabled)

[0107]

[0108] Table 3. Boundary strength (when SPS IBC is enabled)

[0109] 3.6.3 Deblocking decision for the luminance component

[0110] Figure 9 These are examples of the pixels involved in the filter on / off decision and the strong / weak filter selection. A wider and stronger luminance filter is used only when Conditions 1, 2, and 3 are all true. Condition 1 is the "large block condition". This condition detects whether the samples on the P side and the Q side belong to large blocks, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.

[0111] bSidePisLargeBlk = ((the edge type is vertical and p 0 belongs to a CU with width >= 32) || (the edge type is horizontal and p 0 belongs to a CU with height >= 32))? true : false

[0112] bSideQisLargeBlk = ((the edge type is vertical and q 0 belongs to a CU with width >= 32) || (the edge type is horizontal and q 0 belongs to a CU with height >= 32))? true : false

[0113] Based on bSidePisLargeBlk and bSideQisLargeBlk, Condition 1 is defined as follows:

[0114] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? true : false

[0115] Next, if Condition 1 is true, then Condition 2 is further checked. First, the following variables are derived:

[0116] First, dp0, dp3, dq0, and dq3 are derived in the HEVC manner

[0117] if (the p side is greater than or equal to 32)

[0118] dp0 = (dp0 + Abs(p50 - 2 * p40 + p30) + 1) >> 1

[0119] dp3 = (dp3 + Abs(p53 - 2 * p43 + p33) + 1) >> 1

[0120] if (q side is greater than or equal to 32)

[0121] dq0 = (dq0 + Abs(q50 - 2 * q40 + q30) + 1) >> 1

[0122] dq3 = (dq3 + Abs(q53 - 2 * q43 + q33) + 1) >> 1

[0123] Condition 2 = (d < β)? true : false

[0124] where d = dp0 + dq0 + dp3 + dq3.

[0125] If both Condition 1 and Condition 2 are satisfied, further check whether any block uses sub - blocks:

[0126]

[0127]

[0128] Finally, if both Condition 1 and Condition 2 are satisfied, the de - blocking method will check Condition 3 (the strong filtering condition for large blocks), which is defined as follows. In Condition 3 StrongFilterCondition, the following variables are derived:

[0129] Derive dpq in the manner of HEVC.

[0130] Derive sp3 = Abs(p3 - p0) in the manner of HEVC

[0131]

[0132] Derive sq3 = Abs(q0 - q3) in the manner of HEVC

[0133]

[0134] According to HEVC, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (3 * β >> 5), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1)? true : false.

[0135] 3.6.4 Stronger Deblocking Filter for Luminance

[0136] When the sample points on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge >= 32 and when the height of the horizontal edge >= 32, the sample points are defined as belonging to a large block. The bilinear filter is listed as follows. Block boundary sample points pi (i = 0 to Sp - 1) and qi (i = 0 to Sq - 1), in the above HEVC deblocking, pi and qi are the i-th sample points in the row for filtering the vertical edge, or the i-th sample points in the column for filtering the horizontal edge, and then are replaced by linear interpolation as follows:

[0137] p i ′ = (f i * Middle s,t + (64 - f i ) * P s + 32) >> 6), clipped to p i ± tcPD i

[0138] q j ′ = (g j * Middle s,t + (64 - g j ) * Q s + 32) >> 6), clipped to q j ± tcPD j

[0139] where the tcPD i and tcPD j terms are the position - related clipping described above, and g j 、f i 、Middle s,t 、P s and Q s are given as follows.

[0140] 3.6.5 Chrominance Deblocking Decision

[0141] A chrominance strong filter is used on both sides of the block boundary. Here, when both sides of the chrominance edge are greater than or equal to 8 (chrominance position), the chrominance filter is selected, and the decision satisfies the following three conditions: The first decision is the boundary strength and the decision for large blocks. When the block width or height orthogonal to the block edge is equal to or greater than 8 in the chrominance sample domain, this filter can be applied. The second decision and the third decision are basically the same as the HEVC luma deblocking decision, which are the on - off decision and the strong filter decision respectively.

[0142] In the first decision, for chrominance filtering, the boundary strength (bS) is modified and the conditions are checked sequentially. If the condition is met, the remaining conditions with lower priority are skipped. Chrominance deblocking is performed when bS is equal to 2, or bS is equal to 1 when a large block boundary is detected. The second and third conditions are basically the same as the HEVC luma strong filter decision, as follows.

[0143] Under the second condition, d is derived according to HEVC luma deblocking. The second condition will be true when d is less than β. Under the third condition, StrongFilterCondition is derived as follows:

[0144] dpq is derived in the HEVC manner.

[0145] sp is derived in the HEVC manner 3 = Abs(p3 - p 0 )

[0146] sq is derived in the HEVC manner 3 = Abs(q 0 - q 3 )

[0147] According to the HEVC design, StrongFilterCondition = (dpq is less than (β >> 2), sp3 + sq3 is less than (β >> 3), and Abs(p0 - q0) is less than (5 * tC + 1) >> 1).

[0148] 3.6.6 Strong Deblocking Filter for Chrominance

[0149] The strong deblocking filter for chrominance is defined as follows:

[0150] p2′ = (3 * p3 + 2 * p2 + p1 + p0 + q0 + 4) >> 3

[0151] p1′ = (2 * p3 + p2 + 2 * p1 + p0 + q0 + q1 + 4) >> 3

[0152] p0′ = (p3 + p2 + p1 + 2 * p0 + q0 + q1 + q2 + 4) >> 3

[0153] An example chrominance filter performs deblocking on a 4×4 chrominance sample grid.

[0154] 3.6.7 Position-Dependent Clipping

[0155] The position-dependent clipping tcPD is applied to the output samples of the luminance filtering process, involving strong and long filters that modify 7, 5, and 3 samples at the boundaries. Assuming the quantization error distribution, the clipping values of the samples can be increased, and these samples are expected to have higher quantization noise, so the deviation between the reconstructed sample value and the true sample value is expected to be larger.

[0156] For each P or Q boundary filtered with an asymmetric filter, according to the result of the decision-making process, a position-dependent threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below), and these two tables are provided to the decoder as side information:

[0157] Tc7 = {6, 5, 4, 3, 2, 1, 1}; Tc3 = {6, 4, 2};

[0158] tcPD = (Sp == 3)? Tc3 : Tc7;

[0159] tcQD = (Sq == 3)? Tc3 : Tc7;

[0160] For P or Q boundaries filtered with a short symmetric filter, a lower-amplitude position-dependent threshold is applied:

[0161] Tc3 = {3, 2, 1};

[0162] After defining the thresholds, the filtered p’i and q’i sample values are clipped according to the tcP and tcQ clipping values:

[0163] p”i = Clip3(p’i + tcPi, p’i – tcPi, p’i);

[0164] q”j = Clip3(q’j + tcQj, q’j – tcQj, q’j);

[0165] Where p’i and q’j are the filtered sample values, p”i and q”j are the clipped output sample values, and tcPi is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. The function Clip3 is the clipping function specified in VVC.

[0166] 3.6.8. Sub-block Deblocking Adjustment

[0167] To achieve parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on the side where sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)) is used, as shown in the luminance control of the long filter. Extending this, the sub-block deblocking is adjusted such that the sub-block boundaries on the 8×8 grid near the CU or implicit TU boundary are restricted to modifying at most 2 samples on each side.

[0168] The following applies to sub - block boundaries that are not aligned with the CU boundary.

[0169] If (pattern block Q == SUBBLOCKMODE && edge!= 0) {

[0170] if (!(implicit TU && (edge == (64 / 4))))

[0171] if (edge == 2 || edge == (orthogonalLength - 2) || edge == (56 / 4) || edge == (72 / 4))

[0172] Sp = Sq = 2;

[0173] else

[0174] Sp = Sq = 3;

[0175] else

[0176] Sp = Sq = bSideQisLargeBlk? 5 : 3

[0177] }

[0178] Where edge equal to 0 corresponds to the CU boundary, edge equal to 2 or equal to orthogonalLength - 2 corresponds to the sub - block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of the TU is used, then implicit TU is true.

[0179] 3.7. Sample Adaptive Offset

[0180] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after deblocking filtering by using the offsets specified by the encoder for each Coding Tree Block (CTB). The video encoder first decides whether to apply SAO processing to the current slice. If SAO is applied to the slice, each CTB is classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding offsets to pixels in each category. SAO operations include: Edge Offset (EO), which uses edge attributes to classify pixels in SAO types 1 to 4; and Band Offset (BO), which uses pixel intensity to classify pixels in SAO type 5. Each applicable CTB has SAO parameters, including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag is equal to 1, the current CTB re-uses the SAO type and offsets of the left CTB. If sao_merge_up_flag is equal to 1, the current CTB re-uses the SAO type and offsets of the upper CTB.

[0181] SAO type Sampling point adaptive offset type to be used Number of categories 0 None 0 1 1-D 0-degree mode edge offset 4 2 1-D 90-degree mode edge offset 4 3 1-D 135-degree mode edge offset 4 4 1-D 45-degree mode edge offset 4 5 Band offset 4

[0182] Table 4. Specification of SAO Types

[0183] 3.8. Adaptive Loop Filter

[0184] Adaptive loop filtering for video coding and decoding minimizes the mean squared error between the original samples and the decoded samples by using a Wiener-based adaptive filter. ALF is at the last processing stage of each picture and can be regarded as a tool to capture and fix artifacts from previous stages. Appropriate filter coefficients are determined by the encoder and signaled explicitly to the decoder. To achieve better coding and decoding efficiency, especially for high-resolution video, local adaptivity is used for the luminance signal by applying different filters to different regions or blocks in the picture. In addition to filter adaptivity, filter on / off control at the Coding Tree Unit (CTU) level also helps to improve coding and decoding efficiency. Syntactically, the filter coefficients are sent in picture-level header information called the Adaptive Parameter Set, and the filter on / off flags for CTUs are interleaved at the CTU level in the slice data. This syntax design not only supports picture-level optimization but also achieves lower encoding latency.

[0185] 3.8.1. Signaling of Parameters

[0186] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF Adaptive Parameter Set (APS). The ALF APS can include up to 8 chroma filter and one luma filter set, for a total of up to 25 filters. Each of the 25 luma classes also includes an index. Classes with the same index share the same filter. By combining different classes, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficients is represented using a 0th order exponential Golomb code, followed by the sign bit of the non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APSs simultaneously.

[0187] The filter control syntax elements for ALF in VTM include two types of information. First, the ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. The chroma ALF can be enabled at the picture and slice levels only if the luma ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice, and CTB levels, the filter usage information is signaled at that level. If all slices within a picture use the same APS, the referenced ALF APS ID is encoded / decoded at the slice level or picture level. The luma component can reference up to 7 ALF APSs, while the chroma component can reference up to 1 ALF APS. For luma CTBs, an index is signaled to indicate which ALF APS or pre-trained luma filter set is used. For chroma CTBs, the index indicates which filter in the referenced APS is used.

[0188] The ALF data syntax elements associated with the LUMA component in VTM are listed below:

[0189]

[0190] When alf_luma_filter_signal_flag equals 1, it specifies that the set of luminance filters is signaled. When alf_luma_filter_signal_flag equals 0, it specifies that the set of luminance filters is not signaled. When alf_luma_clip_flag equals 0, it specifies that linear adaptive loop filtering is applied to the luminance component. When alf_luma_clip_flag equals 1, it specifies that non-linear adaptive loop filtering can be applied to the luminance component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of classes of adaptive loop filters for luminance coefficients that can be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters - 1 (inclusive of the endpoints). alf_luma_coeff_delta_idx[filtIdx] indicates the index of the signaled adaptive loop filter luminance coefficient increment for the filter class indicated by filtIdx, where filtIdx ranges from 0 to NumAlfFilters - 1. When alf_luma_coeff_delta_idx[filtIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1 + 1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1 (inclusive of the endpoints).

[0191] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luminance filter indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range of 0 to 128 (inclusive of the endpoints). alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luminance coefficient of the filter indicated by sfIdx, as follows:

[0192] If alf_luma_coeff_sign[sfIdx][j] equals 0, the corresponding luminance filter coefficient has a positive value.

[0193] Otherwise (alf_luma_coeff_sign[sfIdx][j] equals 1), the corresponding luma filter coefficient has a negative value.

[0194] When alf_luma_coeff_sign[sfIdx][j] does not exist, it is inferred to be equal to 0.

[0195] alf_luma_clip_idx[sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luma filter transmitted through the signal indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in VTM are listed below:

[0196]

[0197] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] being equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb).

[0198] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to 0. alf_use_aps_flag being equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag being equal to 1 specifies that the filter set from APS is applied to the luma CTB. When alf_use_aps_flag does not exist, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx should be in the range of 0 to sh_num_alf_aps_ids_luma - 1 (including the endpoints). When alf_luma_prev_filter_idx does not exist, it is inferred to be equal to 0.

[0199] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY], which specifies the filter set index of the luma CTB at the position (xCtb, yCtb), is derived as follows:

[0200] If alf_use_aps_flag is equal to 0, then AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to alf_luma_fixed_filter_idx.

[0201] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set to be equal to 16 + alf_luma_prev_filter_idx.

[0202] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15 (including the endpoints).

[0203] Based on the ALF design of VTM, the ALF design of ECM further introduces the concept of alternative filter sets into the luma filter. Based on the updated luma CTU ALF on / off decision for each alternative / round, multiple alternatives / rounds of training are performed on the luma filter. In this way, multiple filter sets are associated with each training alternative, and the class combination results of each filter set may be different. Each CTU can select the best filter set determined by RDO and signal the relevant alternative information. The data syntax elements of the ALF associated with the luma component in ECM are listed as follows:

[0204]

[0205]

[0206] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3 (inclusive of the endpoints). alf_luma_clip_flag[altIdx] being equal to 0 specifies that the linear adaptive loop filter is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] being equal to 1 specifies that the non-linear adaptive loop filter may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for which the luma coefficients can be signalled to the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters - 1 (inclusive of the endpoints).

[0207] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the adaptive loop filter luma coefficient increment for filter class indicated by filtIdx for the alternative luma filter set with index altIdx. The range of filtIdx is from 0 to NumAlfFilters–1. When alf_luma_coeff_delta_idx[filtIdx][altIdx] does not exist, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range (including endpoints) from 0 to alf_luma_num_filters_signalled_minus1[altIdx]. alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the luma filter signalled by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range (including endpoints) from 0 to 128.

[0208] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx of the alternative luma filter set with index altIdx, as follows:

[0209] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.

[0210] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.

[0211] When alf_luma_coeff_sign[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0.

[0212] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luma filter signaled by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] does not exist, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in the ECM are listed below:

[0213]

[0214] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luma filter of the codec tree block for the luma component of the codec tree unit at the luma position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] does not exist, it is inferred to be equal to zero.

[0215] 3.8.2. Filter Shape

[0216] In JEM, up to three diamond filter shapes can be selected for the luma component (as Figure 10 shown). The filter shape used for the luma component is signaled by an index at the picture level. Each square represents a sample, and Ci (i is 0 to 6 (left), 0 to 12 (middle), 0 to 20 (right)) represents the coefficient to be applied to that sample. For the chroma component in the picture, a 5×5 diamond shape is always used. In VVC, a 7×7 diamond shape is always used for luma, and a 5×5 diamond shape is always used for chroma.

[0217] 3.8.3 Classification of ALF

[0218] Each 2×2 (or 4×4) block is classified as one of 25 classes. The classification index C is derived based on its directionality D and the quantized value of activity as follows:

[0219]

[0220] To calculate D and First, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a 1D Laplacian operator:

[0221]

[0222] The indices i and j refer to the coordinates of the upper left sample in a 2×2 block, and R(i,j) indicates the reconstructed sample at the coordinates (i,j). The maximum and minimum values of the gradients in the horizontal and vertical directions are set to:

[0223]

[0224] And the maximum and minimum values of the gradients in the two diagonal directions are set to:

[0225]

[0226] To derive the value of the directionality D, these values are compared with each other and with two thresholds t 1 and t 2 as follows:

[0227] Step 1. If and are both true, then set D to 0.

[0228] Step 2. If then proceed to Step 3; otherwise proceed to Step 4;

[0229] Step 3. If then set D to 2; otherwise, set D to 1.

[0230] Step 4. If then set D to 4; otherwise, set D to 3.

[0231] The activity value A is calculated as:

[0232]

[0233] A is further quantized to the range from 0 to 4 (including the endpoints), and the quantized value is denoted as For the two chrominance components in the picture, no classification method is employed (i.e., a set of ALF coefficients is applied to each chrominance component).

[0234] 3.8.4. Geometric Transformations of Filter Coefficients

[0235] Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k,l) associated with the coordinates (k,l) according to the gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make the different blocks more similar by aligning the directionality of the ALF-applied blocks.

[0236] Three geometric transformations are introduced, including diagonal transformation, vertical flipping, and rotation:

[0237] Diagonal: f D (k, l) = f(l, k),

[0238] Vertical flip: f V (k, l) = f(k, K - l - 1),

[0239] Rotation: f R (k, l) = f(K - l - 1, k).

[0240] Where K is the size of the filter, and 0 ≤ k, l ≤ K - 1 are the coefficient coordinates, such that the position (0, 0) is at the upper left corner and the position (K - 1, K - 1) is at the lower right corner. The transformation is applied to the filter coefficient f(k, l) according to the gradient value calculated for the block. Table 5 summarizes the relationship between the transformation and the four gradients in the four directions. Figure 11 Shows the transformation coefficients for each position based on a 5×5 diamond.

[0241] Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 And g v <g h > Diagonal <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation

[0242] Table 5. Mapping of gradients and transformations calculated for a block.

[0243] 3.8.5. Filtering process

[0244] On the decoder side, when ALF is enabled for a block, each sample R(i, j) within the block is filtered to obtain the sample value R′(i, j) as shown below, where L represents the filter length, f m,n represents the filter coefficient, and f(k, l) represents the decoded filter coefficient.

[0245]

[0246] Figure 12 Shows an example of the relative coordinates for 5×5 diamond filter support, assuming the coordinates (i, j) of the current sample are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient.

[0247] 3.8.6. Non - linear filtering reformulation

[0248] Linear filtering can be reformulated into the following expression without affecting the encoding and decoding efficiency:

[0249]

[0250] Where w(i, j) are the same filter coefficients.

[0251] VVC introduces non - linearity to make the ALF more efficient by reducing the influence of these neighboring sample values when the neighboring sample values (I(x + i, y + j)) differ too much from the filtered current sample value (I(x, y)) through the use of a simple clipping function. More specifically, the ALF filter is modified as follows:

[0252]

[0253] where K(d, b)=min(b, max( - b, d)) is the clipping function, and k(i, j) is the clipping parameter, which depends on the (i, j) filter coefficients. The encoder performs optimization to find the best k(i, j).

[0254] A clipping parameter k(i, j) is specified for each ALF filter, and each filter coefficient signals a clipping value. This means that at most 12 clipping values can be signaled in the bitstream for each luminance filter, and at most 6 clipping values can be signaled in the bitstream for each chrominance filter. To limit the signaling cost and encoder complexity, only 4 fixed values are used, and these fixed values are the same for inter - frame and intra - frame stripes.

[0255] Since the variance of local differences in luminance is generally higher than that in chrominance, two different sets are applied for the luminance filter and the chrominance filter. The maximum sample value in each set (here 1024 for a 10 - bit bit - depth) is also introduced so that clipping can be disabled when not needed. These 4 values are selected by roughly equally dividing the entire range of sample values of luminance (encoded and decoded in 10 bits) in the logarithmic domain and the range of 4 to 1024 for chrominance. More precisely, the luminance table of clipping values is obtained by the following formula:

[0256] where M = 2 10 and N = 4

[0257] Similarly, the chrominance table of clipping values can be obtained according to the following formula:

[0258] where M = 2 10 , N = 4 and A = 4

[0259] 3.9. Bilateral Loop Filter

[0260] 3.9.1. Bilateral Image Filter

[0261] The bilateral image filter is a non - linear filter that smooths noise while preserving the edge structure. Bilateral filtering is a technique that makes the filter weights decrease not only with the distance between samples but also with the increase in intensity difference. In this way, over - smoothing of edges can be improved. The weights are defined as

[0262]

[0263] where Δx and Δy are the distances in the vertical and horizontal directions respectively, and ΔI is the intensity difference between the samples.

[0264] The edge-preserving denoising bilateral filter uses a low-pass Gaussian filter for both the domain filter and the range filter. The domain low-pass Gaussian filter assigns higher weights to pixels that are spatially closer to the central pixel. The range low-pass Gaussian filter assigns higher weights to pixels that are similar to the central pixel. Combining the range filter and the domain filter, the bilateral filter at the edge pixels becomes an elongated Gaussian filter that is oriented along the edge and significantly reduced in the gradient direction. This is why the bilateral filter can smooth the noise while preserving the edge structure.

[0265] 3.9.2. Bilateral Filter in Video Coding and Decoding

[0266] The bilateral filter in video coding and decoding is a coding and decoding tool for VVC. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and SAO act on the same input samples, and each filter produces an offset, which is then added to the input samples to produce the output samples, which enter the next stage after clipping. The spatial filtering strength σ d is determined by the block size, where the smaller the block, the greater the filtering strength, and the intensity filtering strength σ r is determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four closest samples are used, so the filtered sample intensity I F can be calculated as

[0267]

[0268] where I C represents the intensity of the central sample, ΔI A = I A - I C represents the intensity difference between the central sample and the upper sample. ΔI B , ΔI L and ΔI R represent the intensity differences between the central sample and the lower, left, and right samples respectively.

[0269] 4. Technical Problems Solved by the Disclosed Technical Solution

[0270] Example designs of the adaptive loop filter (ALF) in video coding and decoding have the following problems:

[0271] In some ALF designs, only the spatially reconstructed samples after other filtering (such as deblocking filtering (DBF), sample adaptive offset (SAO), and bilateral filtering (BF)) are used for filter training and filtering. However, other valuable information, such as samples filtered / generated by one or more predefined filters, can potentially be utilized.

[0272] In some ALF designs, only the spatially reconstructed samples after other filtering (such as deblocking filtering, SAO, and BF) are used for filter training and filtering. However, other valuable information, such as samples before DBF, SAO, or other stages, can potentially be utilized.

[0273] 5. List of solutions and embodiments

[0274] To solve the above problems, the methods summarized below are disclosed. The embodiments should be regarded as examples for explaining general concepts and should not be interpreted narrowly. In addition, these embodiments can be applied individually or combined in any way. It should be noted that the disclosed methods can be used as loop filters or post-processing. In the present disclosure, a video unit may refer to a sequence, picture, sub-picture, slice, CTU, block, and / or region. A video unit may include one color component or multiple color components. In the present disclosure, an ALF processing unit may refer to a sequence, picture, sub-picture, slice, CTU, block, region, or sample. An ALF processing unit may include one color component or multiple color components.

[0275] In the following disclosure, the filtered sample values are represented as the output of BF. For example, in the process described below:

[0276]

[0277] where I F represents the output of BF.

[0278] In the following disclosure, BF represents an example of bilateral filtering in video coding and decoding, which typically uses predefined filtering parameters to produce the output of BF. For example, there is no online training or signal transmission for the filtering parameters used in BF.

[0279] In the following disclosure, adaptive BF represents an improved version of BF on top of bilateral filtering in video coding and decoding. Adaptive BF involves online parameter training and signal transmission of parameters. For example, multiple filtered samples generated based on BF are further combined with parameters that are signaled and trained online.

[0280] Example 1

[0281] It is proposed to use the intermediate filtering result of a filter as the input of one or more extended taps. In one example, the intermediate filtering result of an offline-trained ALF filter can be used as the input of one or more extended taps. In one example, the intermediate filtering result can be generated by the reconstruction before ALF and the offline-trained filter of ALF. In one example, the intermediate filtering result can be generated by the reconstruction before DBF and the offline-trained filter of ALF. In one example, the intermediate filtering result of an online-trained ALF filter can be used as the input of one or more extended taps. In one example, the intermediate filtering result of other predefined filters can be used as the input of one or more extended taps.

[0282] In one example, a Gaussian filter can be applied. In one example, a bilateral filter can be applied. In one example, a guided filter can be applied. In one example, a median filter can be applied. In one example, a local median filter can be applied. In one example, a non-local median filter can be applied. In one example, a filter with low-pass property can be applied. In one example, a filter with high-pass property can be applied.

[0283] In one example, the intermediate filtering result of other online-trained filters can be used as the input of one or more extended taps. In one example, the input of intermediate filtering can be the reconstructed samples at different coding / decoding stages. In one example, the reconstruction before / after ALF of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after SAO / CCSAO (Cross-Component Sample Adaptive Offset) of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the bilateral filter (BIF) of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after DBF of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after any other stage of the current / reference frame can be used to generate the intermediate filtering result.

[0284] Example 2

[0285] It is proposed to use the reconstructed samples before / after different coding / decoding stages of the current frame as the input of one or more extended taps. In one example, the reconstruction before / after DBF of the current frame can be used as the input of one or more extended taps. In one example, the reconstruction before / after SAO / CCSAO of the current frame can be used as the input of one or more extended taps. In one example, the reconstruction before / after BIF of the current frame can be used as the input of one or more extended taps. In one example, the reconstruction before / after other stages of the current frame can be used as the input of one or more extended taps.

[0286] Example 3

[0287] It is proposed to use one or more samples within the coded / decoded pictures as the input source for one or more extended taps. In one example, the previously coded frame may be a reference frame in the reference picture list (RPL) or reference picture set (RPS) associated with the block / current strip / frame. In one example, the previously coded frame may be a short-term reference picture of the block / current strip / frame. In one example, the previously coded frame may be a long-term reference picture of the block / current strip / frame.

[0288] Example 4

[0289] In one example, the previously coded frame may not be a reference frame, but it is still stored in the decoded picture buffer (DPB).

[0290] Example 5

[0291] In one example, at least one indicator is signaled to indicate which previously coded frame(s) to use. In one example, an indicator is signaled to indicate which reference picture list to use. In one example, at least one indicator may be signaled to indicate a reference index. In one example, the indicator may be signaled conditionally, e.g., depending on how many reference pictures are included in the RPL / RPS. In one example, the indicator may be signaled conditionally, e.g., depending on how many previously decoded pictures are included in the DPB.

[0292] Example 6

[0293] In one example, it is determined on-the-fly which frames to utilize. In one example, the extended taps may obtain information from one or more previously coded frames in the DPB. In one example, the extended taps may obtain information from one / more reference frames in list 0. In one example, the extended taps may obtain information from one / more reference frames in list 1. In one example, the extended taps may obtain information from reference frames in both list 0 and list 1. In one example, the extended taps may obtain information from the reference frame closest to the current frame (e.g., having the smallest picture order count (POC) distance from the current strip / frame).

[0294] In one example, the extended taps may obtain information from a reference frame in the reference list whose reference index is equal to K (e.g., K = 0). In one example, K may be predefined. In one example, K may be derived on-the-fly based on the reference picture information. In one example, K may be signaled.

[0295] In one example, the extended tap can obtain information from the co-located frame. In one example, the frame to be utilized can be determined by decoding the information. In one example, the frame to be utilized can be defined as the top N (e.g., N = 1) most frequently used reference pictures of the samples within the current strip / frame. In one example, the frame to be utilized can be defined as the top N (e.g., N = 1) most frequently used reference pictures for each list of reference pictures (if available) of the samples within the current strip / frame. In one example, the frame to be utilized can be defined as the picture with the top N (e.g., N = 1) smallest POC distance / absolute POC distance relative to the current picture.

[0296] Example 7

[0297] In one example, whether to obtain information from a previously decoded frame can depend on the decoded information (e.g., coding mode / statistics / characteristics) of at least one region of the block to be filtered. In one example, the encoder can signal to the decoder whether to obtain information from a previously decoded frame of at least one region of the block to be filtered. In one example, whether to obtain information from a previously decoded frame can depend on the strip / picture type. In one example, it may only apply to inter-coded strips / pictures (e.g., P or B strips / pictures).

[0298] In one example, whether to obtain information from a previously decoded frame can depend on the availability of the reference pictures. In one example, whether to obtain information from a previously decoded frame can depend on the reference picture information or the picture information in the DPB. In one example, if the minimum POC distance (e.g., the minimum POC distance between the reference picture / picture in the DPB and the current picture) is greater than a threshold, it is disabled. In one example, whether to obtain information from a previously decoded frame can depend on the temporal layer index and / or QP and / or the dimension of the picture. In one example, it may apply to blocks with a given temporal layer index (e.g., the highest temporal layer). In one example, if the block to be filtered contains a portion of samples decoded in a non-inter mode, the extended tap cannot use the information from a previously decoded frame to filter the block.

[0299] In one example, the non-inter mode can be defined as the intra mode. In one example, the non-inter mode can be defined as a set of coding modes including but not limited to the intra / IBC / palette modes. In one example, the distortion between the current block and the matching block is calculated and used to decide whether to obtain information from a previously decoded frame to filter the current block.

[0300] In one example, the distortion between the parity block in a previously encoded / decoded frame and the current block can be used to determine whether to obtain information from the previously encoded / decoded frame to filter the current block. In one example, motion estimation can be first used to find a matching block from at least one previously encoded / decoded frame. In one example, when the distortion is greater than a predefined threshold, the information from the previously encoded / decoded frame cannot be used.

[0301] Example 8

[0302] In one example, the information includes two reference blocks and / or parity blocks of the current block, one from the first reference frame in list-0 and the other from the first reference frame in list-1.

[0303] Example 9

[0304] In one example, the disclosed method can be used for post-processing and / or pre-processing.

[0305] Example 10

[0306] In one example, the methods mentioned above can be used jointly.

[0307] Example 11

[0308] In one example, the methods mentioned above can be used separately.

[0309] Example 12

[0310] In one example, the proposed / described extended taps for the ALF method can be applied to any loop filtering tool, pre-processing or post-processing filtering method in video coding / decoding (including but not limited to ALF / cross-component adaptive loop filter (CCALF) or any other filtering method). In one example, the proposed extended tap method can be applied to the loop filtering method. In one example, the proposed extended tap method can be applied to ALF. In one example, the proposed extended tap method can be applied to CCALF. In one example, the proposed extended tap method can be applied to other loop filtering methods. In one example, the proposed extended tap method can be applied to the pre-processing filtering method. In one example, the proposed extended tap method can be applied to the post-processing filtering method.

[0311] Example 13

[0312] In the above examples, a video unit may refer to a sequence / picture / sub-picture / strip / slice / coding tree unit (CTU) / CTU row / group of CTUs / coding unit (CU) / prediction unit (PU) / transformation unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transformation block (TB) / any other region containing more than one luma or chroma sample / pixel.

[0313] Example 14

[0314] Whether and / or how to apply the methods disclosed above may be signaled in the bitstream. In one example, they may be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, e.g., in the sequence header, picture header, SPS, VPS, DPS, decoder capability information (DCI), PPS, APS, strip header, and slice group header. In one example, they may be signaled at a PB / TB / CB / PU / TU / CU / virtual pipeline data unit (VPDU) / CTU / CTU row / strip / slice / sub-picture / other type of region containing more than one sample or pixel.

[0315] Example 15

[0316] Whether and / or how to apply the methods disclosed above may depend on coding information, such as block size, color format, single / double tree splitting, color component, strip / picture type.

[0317] Other examples

[0318] In one example, at least one extended tap of the ALF filter further enhances the efficiency of the ALF. In one example, at least one extended tap may be different from the spatial taps in the ALF filter, which only utilize the information of the spatial neighboring samples of the target component. In one example, the spatial neighboring samples may come from the reconstruction after DBF / SAO / BF. In one example, the spatial tap may only use the spatial neighboring luma samples to filter the central luma sample inside an ALF filter. In one example, the spatial tap may only use the spatial neighboring chroma samples to filter the central chroma sample inside an ALF filter.

[0319] In one example, at least one extended tap and at least one spatial tap coexist in an ALF filter. In one example, the ALF filter may include both spatial taps and extended taps. In one example, the ALF filter may include M (e.g., M>0) spatial taps and N (e.g., N>0) extended taps.

[0320] In one example, one or more extended taps of the ALF filter may use one or more input sources. In one example, one or more extended taps of the ALF filter may use one input source. (For example, the reconstruction before DBF or the intermediate filtering result of a predefined filter or the reconstruction before SAO / BF). In one example, one or more extended taps of the ALF filter may use multiple input sources. (For example, the reconstruction before DBF and the intermediate filtering result of a predefined filter). The input sources of the extended taps may also be used by other taps in the ALF. The input sources of the extended taps cannot be used by other taps in the ALF. The input source of the extended tap may be a predicted sample. The input source of the extended tap may be a filter predicted sample. The input source of the extended tap may be a weighted sum or filtering of multiple sources. The input source of the extended tap may be the output of a function of multiple sources. The input source of the extended tap may be derived based on samples in different color components.

[0321] In one example, for different color formats and / or different color components, whether and / or how to apply a filter with at least one extended tap may be different. In one example, the ALF filter with at least one extended tap may be applied only to process the luminance component. In one example, the ALF filter with at least one extended tap may be applied only to process one of the chrominance components (e.g., the Cb or Cr component). In one example, the ALF filter with at least one extended tap may be applied to process all chrominance components (e.g., the Cb and Cr components). In one example, the ALF filter with at least one extended tap may be applied to filter both the luminance and chrominance components (e.g., the Y, Cb, and Cr components).

[0322] In one example, the coefficients of the extended taps inside the ALF filter may correspond to one or more input samples. In one example, the coefficients of the extended taps inside the ALF filter may correspond to only one input sample. In one example, the coefficients of the extended taps inside the ALF filter may correspond to N input samples (e.g., N = 2). In one example, the N input samples may be designed in a symmetric manner. In one example, the N input samples may be designed in an asymmetric manner. In one example, the coefficients of the extended taps inside the ALF filter may be shared by multiple inputs based on geometric information. In one example, multiple input samples may be located at symmetric positions.

[0323] In one example, an ALF filter with at least one extended tap can be of different shapes or sizes. In one example, the ALF filter can include spatial taps and extended taps of different shapes. In one example, inside the ALF filter, the shape / size for one or more spatial taps can be different from the shape / size for one or more extended taps. In one example, inside the ALF filter, the shape / size for one or more spatial taps can be the same as the shape / size for one or more extended taps. In one example, inside the ALF filter with at least one extended tap, the filter shape for one or more spatial taps is as described below. In one example, the filter shape for the spatial tap can be a diamond shape. In one example, the filter shape for the spatial tap can be a square shape. In one example, the filter shape for the spatial tap can be a cross shape. In one example, the filter shape for the spatial tap can be a symmetric shape. In one example, the filter shape for the spatial tap can be an asymmetric shape. In one example, the filter shape for the spatial tap can be other designed shapes. In one example, the filter shape for the spatial tap can be determined / derived in real time / through signal transmission.

[0324] In one example, inside the ALF filter with at least one extended tap, the filter shape for one or more extended taps is as described below. In one example, the filter shape for the extended tap can be a diamond shape. In one example, the filter shape for the extended tap can be a square shape. In one example, the filter shape for the extended tap can be a cross shape. In one example, the filter shape for the extended tap can be a symmetric shape. In one example, the filter shape for the extended tap can be an asymmetric shape. In one example, the filter shape for the extended tap can be other designed shapes. In one example, the filter shape for the extended tap can be determined / derived in real time / through signal transmission.

[0325] In one example, an ALF filter including at least one spatial tap and at least one extended tap can be designed as follows. In one example, one or more spatial taps can be used for filtering (the spatial tap can be regarded as using the reconstruction before the ALF as the input). Figures 13 to 14 Examples of the shapes of the spatial taps are shown. In one example, it can be Figure 13 designed in the shape of the spatial tap. In one example, it can be Figure 14 designed in the shape of the spatial tap.

[0326] In one example, one or more extended taps can be applied to filtering. In one example, an ALF filter with one or more spatial taps and one or more extended taps can be designed as follows. Figures 15 to 18 An example ALF filter with extended taps is shown. In one example, it can be designed Figure 15 to have an ALF filter with spatial taps and extended taps. In one example, it can be designed Figure 16 to have an ALF filter with spatial taps and extended taps. In one example, it can be designed Figure 17 to have an ALF filter with spatial taps and extended taps. In one example, it can be designed Figure 18 to have an ALF filter with spatial taps and extended taps.

[0327] In one example, a transformation based on geometric information can be applied. In one example, the transformation based on geometric information can be applied independently to one or more spatial taps. In one example, the transformation based on geometric information can be applied independently to one or more extended taps. In one example, the transformation based on geometric information can be applied jointly to one or more spatial taps and one or more extended taps.

[0328] In one example, one or more input sources can be used for one or more extended taps. In one example, Figures 15 to 18 input sources A, B, C, D, E as shown in the example of Figures 15 to 18 can be different. In another example,

[0329] input sources A, B, C, D, E as shown in the example of

[0330] can be the same. In one example, a possible input source for one or more extended taps can be the reconstruction before the ALF. In one example, a possible input source for one or more extended taps can be the reconstruction before the DBF. In one example, a possible input source for one or more extended taps can be the intermediate filtering result generated by the reconstruction before the ALF and the offline-trained filter of the ALF.

[0329] In one example, a specific offline-trained filter bank can be applied. In one example, the offline-trained filter bank indicator can be signaled / pre-defined / derived in real time. In one example, a specific offline-trained filter classifier can be applied. In one example, the offline-trained filter classifier indicator can be signaled / pre-defined / derived in real time. In one example, a possible input source for one or more extended taps can be the intermediate filtering result generated by the reconstruction before the DBF and the offline-trained filter of the ALF.

[0330] In one example, a specific offline-trained filter bank can be applied. In one example, an offline-trained filter bank indicator can be transmitted / instantly predefined / deduced via a signal. In one example, a specific offline-trained filter classifier can be applied. In one example, an offline-trained filter classifier indicator can be transmitted / instantly predefined / deduced via a signal. In one example, an indicator for the input source can be predefined / deduced in the APS / instantly.

[0331] In one example, Figures 15 to 18 The input sources of the ALF filter represented in can be arranged in order as follows. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-0 and the offline-trained filter bank corresponding to the filter bank indicator transmitted via a signal can be applied as input source A. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-1 and the offline-trained filter bank corresponding to the filter bank indicator transmitted via a signal can be applied as input source B. In one example, the reconstruction before the DBF of the current picture can be applied as input source C. In one example, the intermediate filtering result generated by feeding the reconstruction before DBF to classifier-0 and offline-trained filter bank-0 can be applied as input source D. In one example, the intermediate filtering result generated by feeding the reconstruction before DBF to classifier-0 and offline-trained filter bank-1 can be applied as input source E. In one example, the coded picture named reference image 0 inside the forward reference picture list (reference list 0) can be applied as input source D. In one example, the coded picture named reference image 0 inside the backward reference picture list (reference list 1) can be applied as input source E. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-0 and offline-trained filter bank-0 can be applied as input source D. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-0 and offline-trained filter bank-1 can be applied as input source E. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-0 and the offline-trained filter bank corresponding to the opposite indicator of the filter bank indicator transmitted via a signal can be applied as input source D. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to classifier-1 and the offline-trained filter bank corresponding to the opposite indicator of the filter bank indicator transmitted via a signal can be applied as input source E. In one example, the total number of extended taps inside the ALF filter can be jointly deduced based on shape, filter length, and symmetry constraints.

[0332] In one example, whether and / or how to apply at least one extended tap in the ALF may depend on the classification in the ALF. In one example, the classification may be based on the gradient information of the input source. In one example, the input source may be the reconstruction before the ALF. In one example, the input source may be the reconstruction before the DBF. In one example, the input source may be the intermediate filtering result generated by an offline-trained filter and the reconstruction before the ALF. In one example, the input source may be the intermediate filtering result generated by an offline-trained filter and the reconstruction before the DBF. In one example, the classification may be based on the band information of the input source. In one example, the input source may be the reconstruction before the ALF. In one example, the input source may be the reconstruction before the DBF. In one example, the input source may be the intermediate filtering result generated by an offline-trained filter and the reconstruction before the ALF. In one example, the input source may be the intermediate filtering result generated by an offline-trained filter and the reconstruction before the DBF.

[0333] In one example, the input samples of the extended tap can be filled between picture / strip / slice / CTU boundaries. In one example, the filling can be applied to any boundary during the encoding / decoding process. In one example, the filling can be applied to the picture / sub-picture boundary. In one example, the filling can be applied to the strip / slice boundary. In one example, the filling can be applied to the CTU / CTB boundary. In one example, the filling can be applied to the CU / TU / PU boundary. In one example, the filling can be applied to the block boundary. In one example, the filling can be applied to the cell boundary. In one example, the filling can be applied to the virtual boundary. In one example, the filling can be applied to any other type of boundary. In one example, different filling methods can be applied to the boundary. In one example, extended filling can be applied. In one example, mirror filling can be applied. In one example, repeat filling can be applied. In one example, other filling methods can be applied.

[0334] In one example, a first syntax element may be signaled to indicate whether a filter with at least one extended tap is enabled. In one example, the first syntax element may be coded by arithmetic coding. In one example, the first syntax element may be coded with at least one context. The context may depend on coding information of a current block or an adjacent block. The context may depend on a filtering shape of at least one adjacent block. In one example, the first syntax element may be coded by bypass coding. In one example, the first syntax element may be binarized by unary code, or truncated unary code, or fixed-length code, or exponential Golomb code, truncated exponential Golomb code, etc. In one example, the first syntax element may be signaled conditionally. For example, the first syntax element may be signaled only when the extended tap is available. The first syntax element may be coded in a predictive manner. The first syntax element may be predicted by an on / off decision of extended taps of at least one adjacent block. The first syntax element may be signaled independently for different color components. In an example, the first syntax element may be signaled and shared for different color components. In an example, the first syntax element may be signaled for a first color component, but not signaled for a second color component. The syntax element may be signaled in SPS / PPS / picture header / strip header / APS / CTU / CU / etc.

[0335] In one example, a first syntax element may be signaled to indicate which / what input sources are used for the extended taps inside the ALF filter. In one example, the first syntax element may be coded by arithmetic coding. In one example, the first syntax element may be coded with at least one context. The context may depend on the coding information of the current block or neighboring blocks. The context may depend on the filtering shape of at least one neighboring block. In one example, the first syntax element may be coded by bypass coding. In one example, the first syntax element may be binarized by unary code, or truncated unary code, or fixed-length code, or exponential Golomb code, truncated exponential Golomb code, etc. In one example, the first syntax element may be signaled conditionally. For example, the first syntax element may be signaled only when the extended taps are available. The first syntax element may be coded in a predictive manner. The first syntax element may be predicted by the on / off decision of the extended taps of at least one neighboring block. The first syntax element may be signaled independently for different color components. In an example, the first syntax element may be signaled and shared for different color components. In an example, the first syntax element may be signaled for the first color component, but not signaled for the second color component. The syntax element may be signaled in the SPS / PPS / picture header / slice header / APS / CTU / CU / etc. In one example, the syntax element may be signaled in the APS of the ALF filter for signaling.

[0336] The coefficients of at least one extended tap inside the ALF filter may be signaled in a syntax element structure such as APS. In one example, the coefficients of the extended taps may be included in the APS. In one example, the clipping parameters of the extended taps may be included in the APS. In one example, the class merging results of the extended taps may be included in the APS. In one example, the coefficients of the extended taps may be coded in a predictive manner. In one example, the coefficients of the extended taps may be coded by arithmetic coding using at least one context. In one example, the coefficients of the extended taps may be coded by bypass coding. In one example, the coefficients of the extended taps may be jointly coded with the coefficients of the spatial taps. In one example, other parameters of the extended taps may be included in the APS. In one example, the coefficients of the extended taps may be predefined fixed values.

[0337] Figure 19FIG. 0 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed format or an encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0338] System 4000 may include a codec component 4004 that may implement various codec or encoding methods described in this document. The codec component 4004 may reduce the average bit rate of the video from the input 4002 to the codec component 4004 to the output to produce a codec representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via a connected communication as shown by component 4006. The stored or transmitted bitstream representation (or codec representation) of the video received at input 4002 may be used by component 4008 to generate pixel values or a displayable video, which is sent to a display interface 4010. The process of generating a user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are called "encoding" operations or tools, it should be understood that encoding tools or operations are used in an encoder, and a decoder will perform the corresponding decoding tools or operations for the reverse process of encoding.

[0339] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0340] Figure 20FIG. 0 is a block diagram of an example video processing device 4100. The device 4100 can be used to implement one or more of the methods described herein. The device 4100 can be implemented as a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor 4102 can be configured to implement one or more of the methods described in this document. The memory(ies) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be at least partially included in the processor 4102 (e.g., a graphics co-processor).

[0341] Figure 21 FIG. 4 is a flow chart of an example method 4200 for video processing. The method 4200 includes determining at step 4202 to apply an ALF with extended taps to a picture in a video. An intermediate filtering result of a second filter is used as an input to the extended taps. Performing a conversion between visual media data and a bit stream based on the ALF at step 4204. According to an example, the conversion of step 4204 can include encoding at an encoder or decoding at a decoder.

[0342] It should be noted that the method 4200 can be implemented in a device for processing video data, the device including a processor and a non-transitory memory having instructions thereon, the device such as, a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to execute the method 4200. Additionally, the method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium such that when executed by a processor, cause the video codec device to execute the method 4200.

[0343] Figure 22 FIG. 11 is a block diagram of an example video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and the source device can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and the destination device can be referred to as a video decoding device.

[0344] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a series of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to the destination device 4320 via the I / O interface 4316 over the network 4330. The encoded video data may also be stored on a storage medium / server 4340 for access by the destination device 4320.

[0345] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain the encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, which may be configured to interface with an external display device.

[0346] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the HEVC standard, the VVC standard, and other current and / or future standards.

[0347] Figure 23 is a block diagram showing an example of a video encoder 4400, which may be the Figure 22 video encoder 4314 in the system 4300 shown. The video encoder 4400 may be configured to perform any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0348] The functional components of video encoder 4400 may include a splitting unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a buffer 4413, and an entropy coding unit 4414.

[0349] In other examples, video encoder 4400 may include more, fewer, or different functional components. In one example, prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0350] Furthermore, some components, such as motion estimation unit 4404 and motion compensation unit 4405, may be highly integrated, but are shown separately for explanatory purposes in the example of video encoder 4400.

[0351] Splitting unit 4401 may split a picture into one or more video blocks. Video encoder 4400 and video decoder 4500 may support various video block sizes.

[0352] Mode selection unit 4403 may, for example, select one of the coding / decoding modes based on error results, and provide the resulting intra- or inter-coded block to residual generation unit 4407 to generate residual block data and to reconstruction unit 4412 to reconstruct the coded / decoded block for use as a reference picture. In some examples, mode selection unit 4403 may select a Combined Intra and Inter Prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. Mode selection unit 4403 may also select the resolution of the motion vector for a block in the case of inter prediction (e.g., sub-pixel or integer pixel accuracy).

[0353] To perform inter prediction on a current video block, motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from buffer 4413 (rather than the picture associated with the current video block).

[0354] Motion estimation unit 4404 and motion compensation unit 4405 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice.

[0355] In some examples, the motion estimation unit 4404 may perform uni - directional prediction on a current video block, and the motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 or list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0356] In other examples, the motion estimation unit 4404 may perform bi - directional prediction on a current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 and may also search for another reference video block for the current video block in the reference pictures of list 1. Then, the motion estimation unit 4404 may generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0357] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. More precisely, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0358] In one example, the motion estimation unit 4404 may indicate a certain value in the syntax structure associated with the current video block, and this value indicates to the video decoder 4500 that the current video block has the same motion information as another video block.

[0359] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0360] As discussed above, the video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0361] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0362] The residual generation unit 4407 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0363] In other examples, such as in the skip mode, there may be no residual data for the current video block of the current video block, and the residual generation unit 4407 may not perform a subtraction operation.

[0364] The transform processing unit 4408 may generate a transform coefficient video block for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0365] After the transform processing unit 4408 generates the transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0366] The dequantization unit 4410 and the inverse transform unit 4411 may respectively apply dequantization and inverse transform to the transform coefficient video block to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the cache 4413.

[0367] After reconstructing a video block in the reconstruction unit 4412, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0368] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.

[0369] Figure 24 is a block diagram showing an example of a video decoder 4500, which may be Figure 22 the video decoder 4324 in the system 4300 shown. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0370] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a cache 4507. In some examples, the video decoder 4500 may perform a decoding process that is substantially inverse to the encoding process described with respect to the video encoder 4400.

[0371] The entropy decoding unit 4501 may extract the encoded bitstream. The encoded bitstream may include entropy coded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 may decode the entropy coded video data, and the motion compensation unit 4502 may determine motion information based on the entropy decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.

[0372] The motion compensation unit 4502 may generate a motion compensated block, and thus may perform interpolation based on an interpolation filter. The identity of the interpolation filter used with sub-pixel precision may be included in the syntax element.

[0373] The motion compensation unit 4502 may use the interpolation filter used by the video encoder 4400 during the encoding of the video block to calculate the interpolation of the sub-integer pixels of the reference block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and use the interpolation filter to generate a prediction block.

[0374] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks for encoding the frames and / or slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0375] The intra prediction unit 4503 can form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes the video block coefficients that are provided in the bitstream and decoded by the entropy decoding unit 4501, i.e., dequantizes them. The inverse transform unit 4505 applies an inverse transform.

[0376] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to eliminate blocking artifacts. Then, the decoded video block is stored in the buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also generates the decoded video for presentation on a display device.

[0377] Figure 25 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely, the deblocking filter (DF) 4602, the sample adaptive offset (SAO) 4604, and the adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses predefined filters, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by respectively adding offsets and applying a finite impulse response (FIR) filter and signaling the offsets and filter coefficients using coded side information. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and fix the artifacts created by the previous stages.

[0378] The encoder 4600 also includes an intra prediction component 4608 configured to receive an input video and a motion estimation / compensation (ME / MC) component 4610. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using reference pictures obtained from the reference picture buffer 4612. Residual blocks from the inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting an image to the DF 4602, SAO 4604, and ALF 4606 for filtering before these pictures are stored in the reference picture buffer 4612.

[0379] Next, a list of preferred solutions by some examples is provided.

[0380] The following solutions illustrate examples of the techniques discussed herein.

[0381] 1. A method for processing video data, comprising: determining to apply an adaptive loop filter (ALF) with extended taps to a picture in a video, wherein an intermediate filtering result of a second filter is used as an input to the extended taps; and performing a conversion between visual media data and a bitstream based on the ALF.

[0382] 2. The method according to solution 1, wherein an intermediate filtering result of an offline-trained ALF, an online-trained ALF, a predefined filter, or other online-trained filters is used as an input to the extended taps.

[0383] 3. The method according to any one of solutions 1 to 2, wherein the intermediate filtering result is generated by reconstruction before the ALF and the offline-trained ALF.

[0384] 4. The method according to any one of solutions 1 to 3, wherein the intermediate filtering result is generated by reconstruction before a deblocking filter DBF and the offline-trained ALF.

[0385] 5. The method according to any one of solutions 1 to 4, wherein a Gaussian filter, a bilateral filter, a guided filter, a median filter, a filter with low-pass attributes, or a filter with high-pass attributes is applied.

[0386] 6. The method according to any one of Solutions 1 to 5, wherein the intermediate filtering result of the ALF filter trained online is used as the input of the extended tap.

[0387] 7. The method according to any one of Solutions 1 to 6, wherein the input for the intermediate filtering result includes reconstructed samples at different coding and decoding stages.

[0388] 8. The method according to any one of Solutions 1 to 7, wherein the intermediate filtering result is generated by: reconstruction before ALF of the reference frame, reconstruction before sample adaptive offset (SAO) or cross-component SAO (CCSAO) of the reference frame, reconstruction before bilateral filter (BIF) of the reference frame, or reconstruction before deblocking filter (DBF) of the reference frame.

[0389] 9. The method according to any one of Solutions 1 to 8, wherein the intermediate filtering result is generated by: reconstruction after ALF of the reference frame, reconstruction after SAO or CCSAO of the reference frame, reconstruction after BIF of the reference frame, or reconstruction after DBF of the reference frame.

[0390] 10. The method according to any one of Solutions 1 to 9, wherein the extended tap receives inputs from: reconstruction before ALF of the current frame, reconstruction before sample adaptive offset (SAO) or cross-component SAO (CCSAO) of the current frame, or reconstruction before bilateral filter (BIF) of the current frame.

[0391] 11. The method according to any one of Solutions 1 to 10, wherein the extended tap receives inputs from: reconstruction after ALF of the current frame, reconstruction after SAO or CCSAO of the current frame, or reconstruction after BIF of the current frame.

[0392] 12. The method according to any one of Solutions 1 to 11, wherein the samples inside the coded and decoded picture are used as the input source of the extended tap.

[0393] 13. The method according to any one of Solutions 1 to 12, wherein the samples inside the reference frames in the reference picture list (RPL) or reference picture set (RPS) associated with the current block, current strip or current frame are used as the input source of the extended tap.

[0394] 14. The method according to any one of Solutions 1 to 13, wherein the reference frame is a long-term reference picture or a short-term reference picture of the current picture, current strip or current block.

[0395] 15. The method according to any one of Solutions 1 to 14, wherein the samples inside the frames in the decoded picture buffer are used as the input source of the extended tap.

[0396] 16. The method according to any one of Solutions 1 to 15, wherein a coded picture including samples serving as an input source for the extended tap is indicated by a signal transmission indicator.

[0397] 17. The method according to any one of Solutions 1 to 16, wherein the indicator indicates a reference picture list or a reference index, and wherein the indicator is conditionally signaled according to the number of reference pictures in the reference picture list or the number of decoded pictures in the decoded picture buffer.

[0398] 18. The method according to any one of Solutions 1 to 17, wherein the coded picture includes samples serving as an input source for the extended tap, and wherein the coded picture is dynamically determined.

[0399] 19. The method according to any one of Solutions 1 to 18, wherein the extended tap obtains inputs from frames in the decoded picture buffer, frames in reference picture list 0, frames in reference picture list 1, the reference frame closest to the current frame, the reference frame with a specific index, the co-located frame, or a combination thereof.

[0400] 20. The method according to any one of Solutions 1 to 19, wherein the frame serving as an input for the extended tap is determined based on decoding information.

[0401] 21. The method according to any one of Solutions 1 to 20, wherein whether to obtain information from a previously coded frame to serve as an input source for the extended tap depends on the decoding information of at least one region of the block to be filtered.

[0402] 22. The method according to any one of Solutions 1 to 21, wherein a syntax element indicates whether to obtain information from a previously coded frame to serve as an input source for the extended tap.

[0403] 23. The method according to any one of Solutions 1 to 22, wherein whether to obtain information from a previously coded frame to serve as an input source for the extended tap depends on the slice type or the picture type.

[0404] 24. The method according to any one of Solutions 1 to 23, wherein whether to obtain information from a previously coded frame to serve as an input source for the extended tap depends on the reference picture information or the picture information in the decoded picture buffer.

[0405] 25. The method according to any one of Solutions 1 to 24, wherein whether to obtain information from a previously coded frame to serve as an input source for the extended tap depends on the temporal layer index, the quantization parameter, or the size of the picture.

[0406] 26. The method according to any one of Solutions 1 to 25, wherein when the block contains samples encoded or decoded in a non-inter prediction mode, the extended taps do not use information from a previously encoded or decoded frame to filter the block.

[0407] 27. The method according to any one of Solutions 1 to 26, wherein the distortion between the current block and the matching block is used to determine whether to use information from a previously encoded or decoded frame to filter the current block.

[0408] 28. The method according to any one of Solutions 1 to 27, wherein the information includes two reference blocks or co-located blocks of the current block, one block from the first reference frame in List-0 and the other block from the first reference frame in List-1.

[0409] 29. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Solutions 1 to 28.

[0410] 30. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, cause the video codec device to execute the method according to any one of Solutions 1 to 28.

[0411] 31. A non-transitory computer-readable recording medium that stores a bitstream of a video generated by a method executed by a video processing device, wherein the method includes: determining to apply extended taps in an Adaptive Loop Filter (ALF); and generating the bitstream based on the determination.

[0412] 32. A method for storing a bitstream of a video, comprising: determining to apply extended taps in an Adaptive Loop Filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0413] 33. A method, apparatus, or system described in this patent document.

[0414] In the solutions described herein, an encoder can conform to format rules by generating an encoded or decoded representation according to the format rules. In the solutions described herein, a decoder can use the format rules to parse syntax elements in the encoded or decoded representation and, based on the presence and absence of syntax elements known from the format rules, to produce a decoded video.

[0415] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation (and vice versa). For example, the bitstream representation of a current video block may correspond to bits that are co-located or at different positions distributed in the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on the transformed and encoded error residual values and also using bits in the header and other fields in the bitstream. Further, during the conversion, the decoder may parse the bitstream based on a determination and knowing that certain fields may or may not be present, as described in the above solutions. Similarly, the encoder may determine whether to include certain syntax fields and generate the encoded representation accordingly by including or excluding the syntax fields in the encoded representation.

[0416] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry or in computer software, firmware, hardware, or a combination of one or more of them, the computer software, firmware, or hardware including the structures disclosed in this document and their equivalent structures. The disclosed and other embodiments may be implemented as one or more computer program products encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus, i.e., one or more modules of computer program instructions. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs being discussed, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that has been generated to encode information for transmission to a suitable receiver device.

[0417] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form (including as a stand-alone program or as a module, component, subroutine, or any other unit suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers, which may be located at one site or distributed across multiple sites and interconnected by a communication network.

[0418] The processes or logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), and the apparatus can also be implemented as special-purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0419] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read only memory (CD ROM) and digital versatile disc read only memory (DVD-ROM) disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0420] Although this disclosure contains many details, these details should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features specific to particular embodiments that may be specific to a particular technology. Certain features that are described in the context of separate embodiments in this disclosure may also be implemented combinatorially in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.

[0421] Similarly, although operations are shown in the figures in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order, or that all of the shown operations be performed to achieve a desired result. Additionally, the separation of various system components described in this disclosure should not be construed as requiring such separation in all embodiments.

[0422] Only a few embodiments and examples have been described, and other embodiments, enhancements, and changes may be made based on what is described and shown in this disclosure.

[0423] A first component is directly coupled to a second component when there is no intermediate component other than a line, trace, or another medium between the first component and the second component. A first component is indirectly coupled to a second component when there is an intermediate component other than a line, trace, or another medium between the first component and the second component. The term "coupled" and its variants include both direct coupling and indirect coupling. Unless otherwise stated, the use of the term "about" means including a range of ±10% of the subsequent value.

[0424] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. This example is considered illustrative rather than restrictive and is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0425] Additionally, various technologies, systems, subsystems, and methods described and shown as discrete or separate in the various embodiments may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations will be apparent to those skilled in the art, and such changes, substitutions, and alterations may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for processing video data, comprising: determining to apply an adaptive loop filter (ALF) with extended taps to a picture in a video, wherein an intermediate filtering result of a second filter is used as an input to the extended taps; and performing a conversion between visual media data and a bitstream based on the ALF.

2. The method according to claim 1, wherein an intermediate filtering result of an offline-trained filter of the ALF, an online-trained filter of the ALF, a predefined filter, or another online-trained filter is used as an input to the extended taps.

3. The method according to any one of claims 1 to 2, wherein the intermediate filtering result is generated by reconstruction before the ALF and an offline-trained filter of the ALF.

4. The method according to any one of claims 1 to 3, wherein the intermediate filtering result is generated by reconstruction before a deblocking filter (DBF) and the offline-trained filter of the ALF.

5. The method according to any one of claims 1 to 4, wherein the predefined filter includes a Gaussian filter.

6. The method according to any one of claims 1 to 5, wherein the predefined filter includes a bilateral filter, a guided filter, a median filter, a filter with a low-pass property, or a filter with a high-pass property.

7. The method according to any one of claims 1 to 6, wherein an intermediate filtering result of an online-trained filter of the ALF is used as an input to the extended taps.

8. The method according to any one of claims 1 to 7, wherein inputs for generating the intermediate filtering result include reconstructed samples at different codec stages.

9. The method according to any one of claims 1 to 8, wherein the intermediate filtering result is generated using reconstructed samples before or after the ALF of a reference frame.

10. The method according to any one of claims 1 to 9, wherein the intermediate filtering result is generated using reconstructed samples before or after a deblocking filter (DBF) of a reference frame.

11. The method according to any one of claims 1 to 10, wherein the intermediate filtering result is generated using reconstructed samples before or after sample adaptive offset (SAO) or cross-component SAO (CCSAO) of a reference frame, or reconstructed samples before or after a bilateral filter (BIF) of a reference frame.

12. The method according to any one of claims 1 to 11, wherein reconstructed samples at different codec stages of a current frame are used as an input to the extended taps.

13. The method according to any one of claims 1 to 12, wherein reconstructed samples before or after the DBF of the current frame are used as an input to the extended taps.

14. The method according to any one of claims 1 to 13, wherein reconstructed samples before the ALF of the current frame are used as an input to the extended taps, or wherein reconstructed samples before or after SAO or CCSAO of the current frame are used as an input to the extended taps, or wherein reconstructed samples before or after BIF of the current frame are used as an input to the extended taps.

15. The method according to any one of claims 1 to 14, wherein the samples within the coded or decoded picture are used as the input source for the extended taps.

16. The method according to any one of claims 1 to 15, wherein the samples within a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with the current block, current strip, or current frame are used as the input source for the extended taps.

17. The method according to any one of claims 1 to 16, wherein the reference frame is a long-term reference picture or a short-term reference picture of the current picture, current strip, or current block.

18. The method according to any one of claims 1 to 17, wherein the samples within a frame in the decoded picture buffer are used as the input source for the extended taps.

19. The method according to any one of claims 1 to 18, wherein a coded or decoded picture containing the samples used as the input source for the extended taps is indicated by a signaling indicator.

20. The method according to any one of claims 1 to 19, wherein the indicator indicates a reference picture list or a reference index, and wherein the indicator is signaled conditionally based on the number of reference pictures in the reference picture list or the number of decoded pictures in the decoded picture buffer.

21. The method according to any one of claims 1 to 20, wherein the coded or decoded picture contains the samples used as the input source for the extended taps, and wherein the coded or decoded picture is determined dynamically.

22. The method according to any one of claims 1 to 21, wherein the extended taps receive input from a frame in the decoded picture buffer, a frame in reference picture list 0, a frame in reference picture list 1, the reference frame closest to the current frame, a reference frame with a specific index, a co-located frame, or a combination of the above frames.

23. The method according to any one of claims 1 to 22, wherein the frame used as the input for the extended taps is determined based on the decoded information.

24. The method according to any one of claims 1 to 23, wherein whether information is obtained from a previously coded or decoded frame to be used as the input source for the extended taps depends on the decoded information of at least one region of the block to be filtered.

25. The method according to any one of claims 1 to 24, wherein a syntax element indicates whether information is obtained from a previously coded or decoded frame to be used as the input source for the extended taps.

26. The method according to any one of claims 1 to 25, wherein whether information is obtained from a previously coded or decoded frame to be used as the input source for the extended taps depends on the strip type or picture type.

27. The method according to any one of claims 1 to 26, wherein obtaining information from a previously coded or decoded frame to be used as the input source for the extended taps is applicable only to inter-coded strips or pictures.

28. The method according to any one of claims 1 to 27, wherein whether information is obtained from a previously coded or decoded frame to be used as the input source for the extended taps depends on the reference picture information or the picture information in the decoded picture buffer.

29. The method according to any one of claims 1 to 28, wherein whether to obtain information from a previously decoded frame to be used as an input source for the extended tap depends on a temporal layer index, a quantization parameter, or the size of a picture.

30. The method according to any one of claims 1 to 29, wherein when a block contains samples decoded in a non-inter prediction mode, the extended tap does not use information from a previously decoded frame to filter the block.

31. The method according to any one of claims 1 to 30, wherein the distortion between a current block and a matching block is used to determine whether to use information from a previously decoded frame to filter the current block.

32. The method according to any one of claims 1 to 31, wherein the information includes two reference blocks or co-located blocks of a current block, one block from a first reference frame in list-0 and the other block from a first reference frame in list-1.

33. The method according to any one of claims 1 to 32, further comprising: determining to apply the extended tap in a loop filter, and performing a conversion between the visual media data and the bitstream based on the loop filter.

34. The method according to any one of claims 1 to 33, wherein the loop filter includes a cross-component ALF (CCALF) filter.

35. The method according to any one of claims 1 to 34, wherein the conversion includes decoding the visual media data from the bitstream.

36. The method according to any one of claims 1 to 34, wherein the conversion includes encoding the visual media data into the bitstream.

37. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 36.

38. A non-transitory computer-readable medium including a computer program product for use in a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to perform the method according to any one of claims 1 to 36.

39. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method executed by a video processing apparatus, wherein the method comprises: determining to apply an extended tap in an adaptive loop filter (ALF) to a picture in the video, wherein an intermediate filtering result of a second filter is used as an input to the extended tap; and generating the bitstream based on the determination.

40. A method for storing a bitstream of a video, comprising: determining to apply an extended tap in an adaptive loop filter (ALF) to a picture in the video, wherein an intermediate filtering result of a second filter is used as an input to the extended tap; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.