Side information preparation for adaptive loop filter in video coding

By using adaptive loop filter (ALF) in video encoding and decoding, and using residual sample points for local adaptive filtering, the problems of artifacts and low encoding and decoding efficiency in high-resolution videos are solved, and more efficient video compression and quality improvement are achieved.

CN120530645APending Publication Date: 2025-08-22DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007772.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-12
Filing Date
2024-01-11
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When using high-resolution video, it is difficult to effectively reduce artifacts and improve encoding and decoding efficiency. Especially in the application of adaptive loop filters (ALFs), the prior art has not fully utilized edge information for optimization.

Method used

The adaptive loop filter (ALF) is used to receive the residual sample points of the current picture as input, and the mean square error between the original sample points and the decoded sample points is minimized through the Wiener-based adaptive filter, and the local adaptive filter application is performed during the encoding and decoding process. Combined with the filter on/off control at the level of the encoding and decoding tree unit (CTU) to optimize the encoding and decoding efficiency.

Benefits of technology

Improves the encoding and decoding efficiency of high-resolution videos, reduces artifacts, improves video quality and reduces encoding delays, and achieves more efficient video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120530645A_ABST
    Figure CN120530645A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. The mechanism includes determining to employ an adaptive loop filter (ALF) that receives a residual sample of a current picture as side information for use as input. A conversion is performed between the visual media data and the bitstream based on the ALF.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2023 / 071910, filed on January 12, 2023, the teachings and disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to the generation, storage and use of digital audio and video media information in file format. Background Art

[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is likely to continue to grow. Summary of the Invention

[0005] A first aspect relates to a method for processing video data, comprising: determining to employ an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; and performing conversion between visual media data and a bitstream based on the ALF.

[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the above aspects.

[0007] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of the above aspects.

[0008] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; and generating a bitstream based on the determination.

[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; generating a bitstream based on the determination; and storing the bitstream in a non-temporary computer-readable recording medium.

[0010] For purposes of clarity, any of the above-described embodiments may be combined with any one or more of the other preceding embodiments to create new embodiments within the scope of the present disclosure.

[0011] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0013] Figure 1 An example of nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture is shown.

[0014] Figure 2 An example encoder block diagram is shown.

[0015] Figure 3 An example picture segmented into raster scan strips is shown.

[0016] Figure 4 An example picture segmented into rectangular scan strips is shown.

[0017] Figure 5 An example picture segmented into bricks is shown.

[0018] Figures 6A-6C An example of a codec tree block (CTB) crossing a picture boundary is shown.

[0019] Figure 7 An example of an intra prediction mode is shown.

[0020] Figure 8 Examples of block boundaries in a picture are shown.

[0021] Figure 9 An example of pixels involved in filter usage is shown.

[0022] Figure 10 An example of a filter shape of an adaptive loop filter (ALF) is shown.

[0023] Figure 11 An example of transform coefficients supported by a 5×5 diamond filter is shown.

[0024] Figure 12 An example of relative coordinates supported by a 5×5 diamond filter is shown.

[0025] Figure 13 is a block diagram illustrating an example video processing system.

[0026] Figure 14 is a block diagram of an example video processing device.

[0027] Figure 15 is a flow chart of an example method of video processing.

[0028] Figure 16 is a block diagram illustrating an example video encoding and decoding system.

[0029] Figure 17 is a block diagram illustrating an example encoder.

[0030] Figure 18 is a block diagram illustrating an example decoder.

[0031] Figure 19 is a schematic diagram of an example encoder.

[0032] Figure 20 is a flow chart of an example method of video processing. DETAILED DESCRIPTION

[0033] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0034] The section headings used in this document are for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section only. In addition, the techniques described herein are applicable to other video codec protocols and designs.

[0035] 1. Preliminary Discussion

[0036] This document relates to video codecs. Specifically, it relates to loop filters and other codec tools in image / video codecs. These ideas can be applied alone or in various combinations to video codecs such as High Efficiency Video Codec (HEVC), Versatile Video Codec (VVC), or other video codecs.

[0037] 2. Abbreviation

[0038] This disclosure includes the following abbreviations: Advanced Video Codec (ITU-T Rec. H.264 | ISO / IEC 14496-10) (AVC), Coded Picture Buffer (CPB), Pure Random Access (CRA), Codec Tree Unit (CTU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraint Information (GCI), High Efficiency Video Codec (HEVC, also known as ITU-T Rec. H.265 | ISO / IEC 23008-2), Joint Exploration Model (JEM), Motion Constrained Slice Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Tier and Level (PTL), Picture Unit (PU), Reference Picture Resample (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Codec Layer (VCL), Video Parameter Set (VPS), Versatile Video Codec (VVC, also known as ITU-T Recommendation H.266 | ISO / IEC 23090-3), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Codec Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Codec Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).

[0039] 3. Video codec standards

[0040] Video codec standards have evolved primarily through the development of standards by the ITU-T and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC [1] standard. Starting with H.262, video codec standards are based on a hybrid video codec structure that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM) [2]. When the Versatile Video Codec (VVC) project was officially launched, JVET was renamed the Joint Video Experts Team (JVET). VVC is a codec standard that aims to reduce bitrate by 50% compared to HEVC. The VVC working draft and VVC test model (VTM) are continuously updated.

[0041] A sample version of the VVC draft, Versatile Video Codec (Draft 10), can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip A sample version of the VVC reference software, called VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2 .

[0042] The Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) are investigating the potential need for standardization of future video codec technologies with compression capabilities significantly exceeding the current VVC standard. This future standardization effort could take the form of extended VVC extensions or entirely new standards. These groups are jointly conducting this exploratory activity through a joint collaborative effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. JVET has established the first Exploration Experiment (EE) and uses reference software called the Enhanced Compression Model (ECM). The ECM test model is continuously updated.

[0043] 3.1 Color Space and Chroma Downsampling

[0044] A color space, also known as a color model (or color system), is a mathematical model that describes the range of colors as a tuple of numbers, such as 3 or 4 values ​​or color components (e.g., RGB). Generally speaking, a color space is a refinement of a coordinate system and subspaces. For video compression, the most commonly used color spaces are luminance, blue-difference chrominance, and red-difference chrominance (YCbCr) and red, green, blue (RGB).

[0045] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, and CB and CR are the blue-difference and red-difference chroma components. Y' (with a prime) is distinguished from Y, which is luma, meaning that light intensity is encoded nonlinearly based on the gamma-corrected RGB primaries.

[0046] Chroma downsampling is the practice of encoding an image at a lower resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to differences in color than to luminance. 3.1.1 4:4:4

[0048] In 4:4:4, each of the three components Y'CbCr has the same sampling rate. Therefore, there is no chroma downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2 4:2:2

[0050] In 4:2:2, the two chroma components are sampled at half the sampling rate of luma. The horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with almost no visual difference. Figure 1 An example of nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture is shown. 3.1.3 4:2:0

[0052] In 4:2:0, horizontal sampling is doubled compared to 4:1:1, but vertical resolution is halved because the Cb and Cr channels are sampled only on alternate lines. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 both horizontally and vertically. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positions.

[0053] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (in interstitial positions). In Joint Photographic Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are located in interstitial positions, in the middle of alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternate lines.

[0054] chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1

[0055] Table 1 SubWidthC and SubHeightC values ​​derived from chroma_format_idc and separate_colour_plane_flag

[0056] 3.2 Example Codec Streams for Video Codecs

[0057] Figure 2 An example encoder block diagram, such as in VVC, is shown. The encoder contains three loop filtering blocks: a deblocking filter (DF), sample adaptive offset (SAO), and an ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original and reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, and using the encoded side information to signal the offset and filter coefficients. The ALF is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts caused by previous stages.

[0058] 3.3 Definition of Video / Codec Unit

[0059] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of ​​the picture. A slice can be divided into one or more bricks, each brick consisting of multiple CTU rows within the slice. Slices that are not partitioned into multiple bricks can also be called bricks. However, bricks that are proper subsets of a slice cannot be called slices. A slice contains multiple slices of a picture or multiple bricks of a slice.

[0060] Two striping modes are supported: raster scan striping and rectangular striping. In raster scan striping, a strip contains a sequence of slices from a raster scan of the image. In rectangular striping, a strip contains multiple tiles that together form a rectangular region of the image. The tiles within a rectangular strip are in the order in which the tiles in the strip were raster scanned. Figure 3An example picture partitioned into raster scan strips is shown. In the example, the picture is partitioned according to raster scan stripe partitioning, where the picture comprises 18x12 luma CTUs and is divided into 12 slices and 3 raster scan stripes.

[0061] Figure 4 An example image is shown that is segmented into rectangular scan strips. For example, Figure 4 An example of rectangular slice partitioning of a picture with 18x12 luma CTUs is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.

[0062] Figure 5 An example image segmented into bricks is shown. For example, Figure 5 An example of a picture partitioned into slices, bricks, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the upper left slice contains 1 brick, the upper right slice contains 5 bricks, the lower left slice contains 2 bricks, and the lower right slice contains 3 bricks), and 4 rectangular strips.

[0063] 3.3.1 CTU / CTB Size

[0064] In VVC, the CTU size signaled in the sequence parameter set (SPS) through the syntax element log2_ctu_size_minus2 can be as small as 4×4.

[0065] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0066]

[0067]

[0068]

[0069] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size for each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:

[0070] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)

[0071] CtbSizeY = 1 << CtbLog2SizeY (7-10)

[0072] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)

[0073] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)

[0074] MinTbLog2SizeY=2 (7-13)

[0075] MaxTbLog2SizeY = 6 (7-14)

[0076] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)

[0077] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)

[0078] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)

[0079] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(7-18)

[0080] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)

[0081] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)

[0082] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)

[0083] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)

[0084] PicSizeInSamplesY=pic_width_in_luma_samples*pic_height_in_luma_samples(7-23)

[0085] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)

[0086] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)

[0087] 3.3.2 CTUs in a Picture

[0088] Figures 6A-6C Examples of CTB spanning picture boundaries are shown. Figure 6A CTBs spanning the bottom picture boundary are shown. Figure 6B CTBs spanning the right picture boundary are shown. Figure 6C CTBs spanning the bottom-right picture boundary are shown. Assume that the CTB / largest coding unit (LCU) size is indicated by M×N (usually M equals N), and for a CTB located at a picture boundary (or slice or strip or other type of boundary, with the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For example Figures 6A-6C For those CTBs depicted in , the CTB size is still equal to M×N, however, the bottom / right borders of the CTB are outside the picture.

[0089] 3.4 Intra-frame Prediction

[0090] Figure 7 An example of intra prediction mode is shown in Figure 2. To capture arbitrary edge directions present in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The extended directional modes are shown in Figure 2. Figure 7 As shown, planar and DC modes remain unchanged. These more densely packed directional intra prediction modes are applicable to all block sizes and for luma and chroma intra prediction.

[0091] like Figure 7 As shown, the angular intra prediction direction can be defined as from 45 degrees to -135 degrees clockwise. In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled and remapped to wide-angle mode indices after parsing. The total number of intra prediction modes remains unchanged, for example, 67, and the intra mode encoding and decoding remains unchanged.

[0092] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is required to generate intra prediction values ​​using DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid the use of division operations for DC prediction, only the longer side is used to calculate the average value of non-square blocks.

[0093] 3.5 Inter-frame prediction

[0094] For each inter-frame predicted CU, the motion parameters include motion vector, reference picture index, reference picture list usage index, and extended information for the new coding features of VVC that will be used to generate inter-frame prediction samples. The motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, encoded motion vector increments and / or reference picture indexes. Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs including spatial and temporal candidates and the extended scheduling introduced in VVC. Merge mode can be applied to any inter-frame predicted CU, not just for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, where for each CU, the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list usage flag and other useful information are explicitly transmitted through signals.

[0095] 3.6 Deblocking Filter

[0096] Deblocking filtering is an example loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform sub-block boundaries, and predicted sub-block boundaries. Prediction sub-block boundaries include prediction unit boundaries introduced by sub-block based temporal motion vector prediction (SbTMVP) and affine mode. Transform sub-block boundaries include transform unit boundaries introduced by sub-block transform (SBT) and intra sub-partitioning (ISP) mode, as well as transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first filtering the vertical edges of the entire picture horizontally, and then filtering the horizontal edges vertically. This specific order enables multiple horizontal filtering or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with small processing delay.

[0097] Vertical edges in the picture are filtered first. Horizontal edges in the picture are then filtered using the samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTBs of each CTU are processed separately on a codec unit basis. Vertical edges of the codec blocks in a codec unit are filtered, starting with the left-hand edge of the codec block and filtering in their geometric order to the right-hand edge of the codec block. Horizontal edges of the codec blocks in a codec unit are filtered, starting with the top edge of the codec block and filtering in their geometric order to the bottom edge of the codec block.

[0098] Figure 8 An example of block boundaries in a picture is shown. For example, Figure 8 Picture samples and horizontal and vertical block boundaries on an 8x8 grid are shown, as well as non-overlapping blocks of 8x8 samples, which can be deblocked in parallel.

[0099] 3.6.1 Boundary Decision

[0100] Filtering is applied on 8x8 block boundaries. Additionally, such boundaries must be transform block boundaries or codec subblock boundaries, for example due to the use of affine motion prediction (ATMVP). For other boundaries, deblocking filtering is disabled.

[0101] 3.6.2 Boundary strength calculation

[0102] For transform block boundaries / codec sub-block boundaries, if the boundary lies in the 8x8 grid, the boundary can be filtered and the settings of bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) of the edge are defined as in Table 2 and Table 3 respectively.

[0103]

[0104] Table 2 Boundary Strength (When SPS Intra Block Copy (IBC) is Disabled)

[0105]

[0106]

[0107] Table 3 Boundary Strength (When SPS IBC is Enabled)

[0108] 3.6.3 Deblocking Decision for Luminance Component

[0109] Figure 9 An example of pixels involved in filter usage is shown. For example, Figure 9 The pixels involved in the filter on / off decision and strong / weak filter selection are shown. The wider and stronger luma filter is used only when conditions 1, 2, and 3 are all true. Condition 1 is the "large block condition." This condition detects whether the samples on the P and Q sides belong to large blocks, which are represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.

[0110] bSidePisLargeBlk = ((edge ​​type is vertical and p0 belongs to a CU with width >= 32) || (edge ​​type is horizontal and p0 belongs to a CU with height >= 32))? True: False

[0111] bSideQisLargeBlk = ((edge ​​type is vertical and q0 belongs to a CU with width >= 32) || (edge ​​type is horizontal and q0 belongs to a CU with height >= 32))? True: False

[0112] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:

[0113] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? True: False

[0114] Next, if condition 1 is true, condition 2 will be checked. First, the following variables are derived:

[0115]

[0116]

[0117] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (large block strong filter condition), which is defined as follows. In condition 3 StrongFilterCondition, the following variables are derived:

[0118]

[0119] According to HEVC, StrongFilterCondition = (dpq less than (β>>2), sp3+sq3 less than (3*β>>5), and Abs(p0-q0) less than (5*tC+1)>>1)? True: False.

[0120] 3.6.4 Stronger Deblocking Filter for Luma

[0121] When the samples on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge is >= 32, and when the height of the horizontal edge is >= 32, the sample is defined as belonging to a large block. The bilinear filter is listed as follows. Then, the block boundary samples pi (i=0 to Sp-1) and qi (j=0 to Sq-1) in the HEVC deblocking described above (pi and qi are the i-th sample in the row for filtering the vertical edge, or the i-th sample in the column for filtering the horizontal edge) are replaced by the following linear interpolation:

[0122] p i ′=(f i *Middle s,t +(64-f i )*P s +32)>>6),clipped to p i ±tcPD i

[0123] q j ′=(g j *Middle s,t +(64-g j )*Q s +32)>>6), clipped to q j ±tcPD j

[0124] tcPD i and tcPD j The term is position-dependent clipping as described above, and g j 、f i 、Middle s,t 、P s and Q s is given as follows:

[0125] 3.6.5 Deblocking Decisions for Chroma

[0126] A chroma strong filter is applied on both sides of the block boundary. Here, the chroma filter is selected when the distance between the two sides of the chroma edge is greater than or equal to 8 (chroma position) and the following decision with three conditions is met: The first decision is for boundary strength and large blocks. The filter can be applied when the block width or height orthogonally spanning the block edge in the chroma sample domain is equal to or greater than 8. The second and third decisions are essentially the same as those used for HEVC luma deblocking, which are the on / off decision and the strong filter decision, respectively.

[0127] In the first decision, the boundary strength (bS) is modified for chroma filtering, and the conditions are checked sequentially. If the condition is met, the remaining conditions with lower priority are skipped. When bS is equal to 2, or when bS is equal to 1 when a large block boundary is detected, chroma deblocking is performed. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.

[0128] In the second condition, d is then derived in the same way as HEVC luma deblocking. When d is less than β, the second condition will be true. In the third condition, StrongFilterCondition is derived as follows:

[0129] Derived dpq in the same way as HEVC.

[0130] Derived according to HEVC method sp3 = Abs (p3-p0)

[0131] According to the HEVC method, sq3 = Abs(q0-q3)

[0132] According to HEVC design, StrongFilterCondition = (dpq less than (β>>2), sp3+sq3 less than (β>>3), and Abs(p0-q0) less than (5*tC+1)>>1)

[0133] 3.6.6 Strong Deblocking Filter for Chroma

[0134] The following strong deblocking filter for chroma is defined:

[0135] p2′=(3*p3+2*p2+p1+p0+q0+4)>>3

[0136] p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3

[0137] p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3

[0138] The example chroma filter performs deblocking on a 4x4 grid of chroma samples.

[0139] 3.6.7 Position-dependent limiting

[0140] Position-dependent clipping (tcPD) is applied to the output samples of the luminance filtering process involving strong and long filters that modify 7, 5, and 3 samples at the boundaries. Assuming the quantization error distribution, the clipping value can be increased for samples that are expected to have higher quantization noise, and thus higher deviations of the reconstructed sample values ​​from the true sample values.

[0141] For each P or Q boundary filtered with the asymmetric filter, a position-dependent threshold table is selected from two tables (e.g., Tc7 and Tc3 in the following table) provided to the decoder as side information, depending on the result of the decision process:

[0142] Tc7={6,5,4,3,2,1,1}; Tc3={6,4,2};

[0143] tcPD=(Sp==3)? Tc3:Tc7;

[0144] tcQD=(Sq==3)? Tc3:Tc7;

[0145] For P or Q boundaries filtered with a short symmetric filter, a lower magnitude position-dependent threshold is applied:

[0146] Tc3={3,2,1};

[0147] After defining the thresholds, the filtered p'i and q'i sample values ​​are clipped according to the tcP and tcQ clipping values:

[0148] p”i=Clip3(p’i+tcPi,p’i–tcPi,p’i);

[0149] q”j=Clip3(q’j+tcQj,q’j–tcQj,q’j);

[0150] where p'i and q'i are filtered sample values, p"i and q"j are clipped output sample values, and tcPitcPi is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. Function Clip3 is the clipping function as it is specified in VVC.

[0151] 3.6.8 Sub-block Deblocking Adjustment

[0152] To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on one side using sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)), as shown in the luma control for the long filter. By extension, sub-block deblocking is adjusted so that sub-block boundaries on the 8x8 grid near CU or implicit TU boundaries are restricted to modifying at most two samples on each side.

[0153] The following applies to sub-block boundaries that are not aligned with CU boundaries.

[0154]

[0155]

[0156] Where edge equals 0 corresponds to a CU boundary, edge equals 2 or equal to orthogonalLength-2 corresponds to a sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TUs is used, implicitTU is true.

[0157] 3.7 Sample Adaptive Compensation

[0158] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after the deblocking filter using an offset specified by the encoder for each CTB. The video encoder first determines whether to apply the SAO process to the current slice. If SAO is applied to the slice, each CTB is classified into one of the five SAO types shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding an offset to the pixels in each category. SAO operations include edge offset (EO) and band offset (BO). EO uses edge attributes for pixel classification in SAO types 1 to 4, and BO uses pixel intensity for pixel classification in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag is 1, the current CTB will reuse the offset and SAO type of the CTB to the left. If sao_merge_up_flag is 1, the current CTB will reuse the offset and SAO type of the CTB above it.

[0159] SAO Type The type of sample adaptive backoff to be used Number of categories 0 none 0 1 1-D 0-degree pattern edge compensation 4 2 1-D 90-degree pattern edge compensation 4 3 1-D 135 degree pattern edge compensation 4 4 1-D 45-degree pattern edge compensation 4 5 With compensation 4

[0160] Table 4 Specifications of SAO types

[0161] 3.8 Adaptive Loop Filter

[0162] Adaptive loop filtering for video coding and decoding minimizes the mean square error between the original samples and the decoded samples by using an adaptive filter based on Wiener. ALF is located in the last processing stage of each picture and can be regarded as a tool to capture and repair artifacts from the previous stage. The appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder through a signal. In order to achieve better coding efficiency, especially for high-resolution video, local adaptation is used for luminance signals by applying different filters to different areas or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also helps to improve coding efficiency. In terms of syntax, the filter coefficients are sent in a picture level header called an adaptive parameter set, and the filter on / off flag of the CTU is interleaved at the CTU level in the slice data. This syntax design not only supports picture level optimization, but also achieves low coding delay.

[0163] 3.8.1 Parameter Signaling

[0164] According to the ALF design in VTM, filter coefficients and clipping indexes are carried in the ALF adaptive parameter set (APS). The ALF APS can include up to 8 chroma filters and a luminance filter set with up to 25 filters. Each of the 25 luminance categories also includes an index. Categories with the same index share the same filter. By merging different categories, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficient is represented by a 0th order Exp-Golomb code followed by a sign bit for the non-zero coefficient. When clipping is enabled, a two-bit fixed-length code is also used to transmit the clipping index by signaling for each filter coefficient. Up to 8 ALF APSs can be used simultaneously by the decoder.

[0165] The filter control syntax elements of ALF in VTM include two types of information. First, the ALF on / off flag is transmitted through signals at the sequence, picture, slice and CTB levels. Chroma ALF can be enabled at the picture and slice levels only when luma ALF is enabled at the corresponding level. Second, if ALF is enabled at this level, the filter usage information is transmitted through signals at the picture, slice and CTB levels. If all slices within a picture use the same APS, the referenced ALF APSID is encoded and decoded at the slice level or picture level. The luma component can reference up to 7 ALF APSs, and the chroma component can reference 1 ALF APS. For the luma CTB, an index is transmitted through signals, which indicates which ALF APS or offline trained luma filter set is used. For the chroma CTB, the index indicates which filter in the referenced APS is used.

[0166] The data syntax elements of the ALF associated with the luma (LUMA) component in the VTM are listed as follows:

[0167]

[0168]

[0169] alf_luma_filter_signal_flag equal to 1 specifies that the luma filter set is signaled. alf_luma_filter_signal_flag equal to 0 specifies that the luma filter set is not signaled. alf_luma_clip_flag equal to 0 specifies that linear adaptive loop filtering is applied to the luma component. alf_luma_clip_flag equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the luma component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of adaptive loop filter categories for which luma coefficients may be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters–1, inclusive. alf_luma_coeff_delta_idx[filtIdx] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter category indicated by filtIdx, which shall be in the range of 0 to NumAlfFilters–1. When alf_luma_coeff_delta_idx[filtIdx] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1+1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1, inclusive.

[0170] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] is not present, it is presumed to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range of 0 to 128, inclusive. alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx, as follows:

[0171] If alf_luma_coeff_sign[sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.

[0172] Otherwise (alf_luma_coeff_sign[sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.

[0173] When alf_luma_coeff_sign[sfIdx][j] is not present, it is inferred to be equal to 0.

[0174] alf_luma_clip_idx[sfIdx][j] specifies the clip index of the clip value to be used before multiplying the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree unit syntax elements for the ALF associated with the luma (LUMA) component in the VTM are listed as follows:

[0175]

[0176]

[0177] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb).

[0178] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to 0. alf_use_aps_flag equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag equal to 1 specifies that the filter set from the APS is applied to the luma CTB. When alf_use_aps_flag is not present, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx shall be in the range of 0 to sh_num_alf_aps_ids_luma–1, inclusive. When alf_luma_prev_filter_idx is not present, it is inferred to be equal to 0.

[0179] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] that specifies the filter set index of the luma CTB at position (xCtb, yCtb) is derived as follows:

[0180] If alf_use_aps_flag is equal to 0, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to alf_luma_fixed_filter_idx.

[0181] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to 16+alf_luma_prev_filter_idx.

[0182] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15, inclusive.

[0183] Based on the ALF design of VTM, the ALF design of ECM further introduces the concept of alternative filter sets into the luma filter. The luma filter is trained for multiple alternatives / rounds based on the updated luma CTU ALF on / off decision for each alternative / round. In this way, there will be multiple filter sets associated with each training alternative, and the class merging results for each filter set can be different. Each CTU can select the best filter set through RDO, and the relevant alternative information will be transmitted through the signal. The data syntax elements of the ALF associated with the luma (LUMA) component in ECM are listed as follows:

[0184]

[0185] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3, inclusive. alf_luma_clip_flag[altIdx] equal to 0 specifies that linear adaptive loop filtering is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for which luma coefficients may be signaled, where the luma coefficients are the luma coefficients of the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters–1, inclusive.

[0186] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter category indicated by filtIdx in the range 0 to NumAlfFilters–1, for the alternative luma filter set with index altIdx. When alf_luma_coeff_delta_idx[filtIdx][altIdx] is not present, it is presumed to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[altIdx], inclusive. alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx, of the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] is not present, it is presumed to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range of 0 to 128, inclusive.

[0187] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx of the alternative luma filter set with index altIdx, as follows:

[0188] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.

[0189] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.

[0190] When alf_luma_coeff_sign[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0.

[0191] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clip index of the clip value to be used before multiplying the j-th coefficient of the signaled luma filter indicated by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree unit syntax elements for the ALF associated with the luma (LUMA) component in the ECM are listed as follows:

[0192]

[0193]

[0194] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luma filter applied to the codec tree block of the luma component of the codec tree unit at luma position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to 0.

[0195] 3.8.2 Filter Shape

[0196] Figure 10 An example of filter shape for ALF is shown. In JEM, up to three diamond filter shapes (such as Figure 10 As shown) can be selected for the luma component. The index is signaled at the picture level to indicate the filter shape used for the luma component. Each square represents a sample, and Ci (i is 0 to 6 (left), 0 to 12 (center), 0 to 20 (right)) represents the coefficient to be applied to the sample. For the chroma components in a picture, a 5×5 diamond shape is always used. In VVC, a 7×7 diamond shape is always used for luma, and a 5×5 diamond shape is always used for chroma.

[0197] 3.8.3 Classification of ALF

[0198] Each 2×2 (or 4×4) block is classified into one of 25 categories. The category index C is based on the quantized values ​​of its directionality D and activity is derived as follows:

[0199]

[0200] To calculate D and First, use the 1-D Laplacian to calculate the gradient in the horizontal, vertical, and two diagonal directions:

[0201]

[0202]

[0203] The indices i and j refer to the coordinates of the top left sample in the 2×2 block, and R(i,j) indicates the reconstructed sample at coordinate (i,j). The maximum and minimum values ​​of the gradients D in the horizontal and vertical directions are then set to:

[0204]

[0205] And the maximum and minimum values ​​of the gradients in the two diagonal directions are set to:

[0206]

[0207] To derive the value of the directivity D, these values ​​are compared with each other and with two thresholds t1 and t2:

[0208] Step 1. If and are both true, then D is set to 0.

[0209] Step 2. If Then continue with step 3; otherwise, continue with step 4.

[0210] Step 3. If Then D is set to 2; otherwise D is set to 1.

[0211] Step 4. If Then D is set to 4; otherwise D is set to 3.

[0212] The activity value A is calculated as:

[0213]

[0214] A is further quantized to the range of 0 to 4 (inclusive), and the quantized value is expressed as For the two chroma components in a picture, no classification method is applied, ie, a single set of ALF coefficients is applied for each chroma component.

[0215] 3.8.4 Geometric Transformation of Filter Coefficients

[0216] Before filtering each 2×2 block, geometric transformations (such as rotations or diagonal and vertical flips) are applied to the filter coefficients f(k, l) associated with the coordinates (k, l), depending on the gradient values ​​calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks more similar by aligning the directionality of the blocks to which the ALF is applied.

[0217] Three geometric transformations are introduced, including diagonal, vertical flip and rotation:

[0218] Diagonal: f D (k,l)=f(l,k),

[0219] Flip vertically: f V (k,l)=f(k,Kl-1),

[0220] Rotation: f R (k,l)=f(Kl-1,k).

[0221] Where K is the size of the filter, and 0≤k,l≤K-1 are the coefficient coordinates, such that position (0,0) is in the upper left corner and position (K-1,K-1) is in the lower right corner. The transformation is applied to the filter coefficients f(k,l) depending on the gradient value calculated for the block. The relationship between the transformation and the four gradients in the four directions is summarized in Table 5.

[0222] Figure 11 An example of transform coefficients for a relative harmonizer supported by a 5×5 diamond filter is shown. For example, Figure 11 The transform coefficients for each position based on a 5x5 diamond are shown.

[0223]

[0224]

[0225] Table 5 Mapping of gradients and transformations calculated for a block

[0226] 3.8.5 Filtering process

[0227] At the decoder side, when ALF is enabled for a block, each sample R(i,j) within the block is filtered, resulting in the following sample value r′(i,j), where L represents the filter length and f m,n denotes the filter coefficient, and f(k, l) denotes the decoded filter coefficient.

[0228]

[0229] Figure 12An example of relative coordinates supported by a 5×5 diamond filter is shown, assuming that the coordinates (i, j) of the current sample point are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient.

[0230] 3.8.6 Nonlinear Filter Reconstruction (Reformulation)

[0231] Linear filtering can be reconstructed as follows without affecting the encoding and decoding efficiency:

[0232]

[0233] where w(i,j) are the same filter coefficients.

[0234] VVC introduces nonlinearity by using a simple clipping function to reduce the influence of neighboring sample values ​​(I(x+i,y+j)) when they differ too much from the current sample value being filtered (I(x,y), making the ALF more efficient. More specifically, the ALF filter is modified as follows:

[0235]

[0236] where K(d,b)=min(b,max(-b,d)) is the clipping function and k(i,j) is the clipping parameter, which depends on the (i,j) filter coefficients. The encoder performs an optimization to find the best k(i,j).

[0237] A clipping parameter k(i,j) is specified for each ALF filter, with one clipping value signaled for each filter coefficient. This means that a maximum of 12 clipping values ​​per luma filter and a maximum of 6 clipping values ​​per chroma filter can be signaled in the bitstream. To limit signaling cost and encoder complexity, only four fixed values ​​are used that are the same for inter and intra slices.

[0238] Because the variance of local differences in luma is generally higher than in chroma, two different sets of filters for luma and chroma are applied. A maximum sample value in each set (here 1024 for a bit depth of 10 bits) is also introduced so that clipping can be disabled if not necessary. The four values ​​are chosen by dividing the full range of sample values ​​for luma (in 10-bit codec) and the range from 4 to 1024 for chroma roughly equally in the logarithmic domain. More precisely, the luma clipping value table has been obtained by the following formula:

[0239] Where M = 2 10 And N=4

[0240] Similarly, the chroma clipping value table is obtained according to the following formula:

[0241] Where M = 2 10 , N=4 and A=4

[0242] 3.9 Bilateral Loop Filter

[0243] 3.9.1 Bilateral Image Filter

[0244] Bilateral image filter is a nonlinear filter that smooths noise while preserving edge structure. Bilateral filtering is a technique that reduces the filter weight not only with the distance between samples but also with the increase of intensity difference. In this way, the over-smoothing of edges can be improved. The weight is defined as

[0245]

[0246] where Δx and Δy are the distances in the vertical and horizontal directions, and ΔI is the intensity difference between the samples.

[0247] The edge-preserving denoising bilateral filter uses low-pass Gaussian filters for both the domain filter and the range filter. The domain low-pass Gaussian filter assigns higher weight to pixels spatially close to the center pixel. The range low-pass Gaussian filter assigns higher weight to pixels similar to the center pixel. By combining the range and domain filters, the bilateral filter at edge pixels becomes a thin Gaussian filter oriented along the edge, with the gradient direction significantly reduced. This is why the bilateral filter can smooth noise while preserving edge structure.

[0248] 3.9.2 Bilateral Filters in Video Codecs

[0249] The bilateral filter in video codec is a codec tool for VVC [2]. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and the SAO act on the same input samples, each filter produces an offset, and these offsets are then added to the input samples to produce output samples, which are passed to the next stage after clipping. Spatial filter strength (strength) σ d Determined by the block size, where smaller blocks are filtered more strongly, and the intensity filter strength σ r Determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four most recent samples are used, so the filtered sample intensity (intensity) I F can be calculated as

[0250]

[0251] Among them I C Indicates the intensity of the central sample point, ΔI A =I A -I C Indicates the intensity difference between the center sample point and the sample point above it. ΔI B , ΔI L andΔI R Represents the intensity difference between the center sample point and the samples below, to the left, and to the right, respectively.

[0252] 4. Technical problems solved by the disclosed technical solution

[0253] The following problems exist in the exemplary design of an adaptive loop filter (ALF) in video codecs.

[0254] First, in the example ALF design, only spatially reconstructed samples after other filtering (such as deblocking filter / SAO / BF) are used for filter training and filtering. However, there is other valuable information that can be potentially exploited, such as prediction samples or residual samples.

[0255] Second, in the example ALF design, the side information is used directly without any modification. However, further adjustments or preparations for the side information can reduce the storage cost.

[0256] 5. List of solutions and implementation examples

[0257] In order to solve the above problems, the following methods are disclosed. The embodiments should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.

[0258] It should be noted that the disclosed method can be used as a loop filter or as post-processing.

[0259] In this disclosure, a video unit may refer to a sequence, a picture, a sub-picture, a slice, a CTU, a block and / or a region.A video unit may include one color component or multiple color components.

[0260] In this disclosure, an ALF processing unit may refer to a sequence, a picture, a sub-picture, a slice, a CTU, a block, a region, or a sample. An ALF processing unit may include one color component or multiple color components. In addition, side information may include codec information other than the samples to be filtered by the current filter, and such side information may be used to determine the use of the current filter for the samples to be filtered.

[0261] 1) It is proposed to use modified and / or prepared side information for ALF.

[0262] a. In one example, the following side information can be used for ALF.

[0263] a) In one example, residual samples can be used for ALF.

[0264] 1. In one example, the residual samples of the current picture can be used.

[0265] 2. In one example, residual samples of other coded pictures can be used.

[0266] b) In one example, the prediction samples may be used for ALF.

[0267] 1. In one example, the predicted samples of the current picture can be used.

[0268] 2. In one example, prediction samples from other coded pictures may be used.

[0269] c) In one example, reconstruction at different stages can be used for ALF.

[0270] 1. In one example, DBF previous reconstruction can be used.

[0271] 2. In one example, reconstruction prior to SAO, cross-component SAO (CCSAO), BF, and / or Hadamard transform domain filter (HTDF) may be used.

[0272] 3. In one example, reconstruction before ALF or cross-component ALF (CCALF) may be used.

[0273] 4. In one example, any other stages / tools prior to reconstruction may be used.

[0274] d) In one example, coded reference pictures may be used for ALF.

[0275] 1. In one example, forward reference pictures can be used for ALF.

[0276] 2. In one example, backward reference pictures can be used for ALF.

[0277] 3. In one example, long-term reference pictures can be used for ALF.

[0278] e) In one example, an intervening picture or frame may be used for ALF.

[0279] 1. In one example, an intervening picture / frame generated from data inside the current group of pictures (GOP) may be used.

[0280] 2. In one example, an intervening picture or frame generated from data outside the current GOP may be used.

[0281] f) In one example, mode and / or motion information may be used for ALF.

[0282] 1. Intra prediction mode can be used for ALF.

[0283] 2. Codec modes (such as Inter, Intra, Merge, Skip, Affine, DMVR, etc.) can be used for ALF.

[0284] 3. At least one reference index or list may be used for ALF.

[0285] 4. At least one MV can be used for ALF.

[0286] 5. At least one transform type can be used for ALF.

[0287] g) In one example, the output of any predefined filter can be used for ALF.

[0288] 1. In one example, the output of the offline trained filter of the ALF may be used.

[0289] 2. In one example, the output of the Gaussian filter can be used for ALF.

[0290] 3. In one example, the output of a Sobel filter, a Prewitt filter, a Roberts filter, and / or a Canny filter can be used for the ALF.

[0291] 4. In one example, the output of the Hadamard transform domain filter can be used for ALF.

[0292] 5. In one example, the output of the bilateral filter can be used for ALF.

[0293] 6. In one example, the output of the low-pass filter can be used for ALF.

[0294] 7. In one example, the output of the high-pass filter can be used for ALF.

[0295] 8. In one example, the output of any other filter can be used for ALF.

[0296] 9. In one example, the above side information can be used as input to the mentioned predefined filters.

[0297] h) Alternatively, transform domain coefficients can be used for ALF.

[0298] 1. In one example, the transform may be a discrete cosine transform (DCT).

[0299] 2. In one example, the transform may be a discrete wavelet transform (DWT).

[0300] 3. In one example, the transform may be a low frequency non-separable transform (LFNST).

[0301] 4. In one example, the transform may be a non-separable primary transform (NSPT).

[0302] 5. In one example, the transform may be a Hadamard transform.

[0303] 6. In one example, the transformation may be any other transformation.

[0304] b. In one example, the following modification / adjustment method may be applied to the side information to generate modified / prepared side information.

[0305] a) In one example, the side information can be used directly without modification.

[0306] b) In one example, the side information may be clipped to a range / bit depth.

[0307] 1. In one example, side information may be clipped into a predefined, signaled, and / or derived clipping range.

[0308] a. In one example, the residual samples may be clipped into a predefined, signaled and / or derived clipping range.

[0309] 2. In one example, side information may be clipped to a predefined, signaled and / or derived N-bit depth (eg, N=10).

[0310] a. In one example, the residual samples may be clipped to a predefined, signaled and / or derived N-bit depth (eg, N=10).

[0311] c) In one example, the side information can be scaled to range or bit depth.

[0312] 1. In one example, side information can be scaled into a predefined, signaled, and / or derived range.

[0313] 2. In one example, the side information can be scaled to a predefined, signaled and / or derived N-bit depth (eg, N=10).

[0314] d) In one example, the side information may be transformed into range, domain and / or bit depth.

[0315] 1. In one example, side information can be transformed into a predefined, signaled and / or derived scope.

[0316] 2. In one example, side information can be transformed into a predefined, signaled, and / or derived domain.

[0317] 3. In one example, the side information may be transformed into a predefined, signaled and / or derived N-bit depth (eg, N=10).

[0318] e) In one example, side information can be filtered into range, domain and / or bit depth.

[0319] 1. In one example, side information can be filtered into a predefined, signaled and / or derived range.

[0320] 2. In one example, side information can be filtered into a predefined, signaled and / or derived domain.

[0321] 3. In one example, side information may be filtered into a predefined, signaled and / or derived N-bit depth (eg, N=10).

[0322] f) In one example, side information can be used in a fusion manner.

[0323] 1. In one example, multiple side information can be fused through weighted sum.

[0324] 2. In one example, multiple side information can be fused through online trained coefficients.

[0325] 3. In one example, multiple side information can be fused via offline trained and / or predefined coefficients.

[0326] 2) It is proposed to use the above side information in different stages of ALF.

[0327] a. In one example, side information can be used for ALF classification.

[0328] a) In one example, the original side information can be used for classification.

[0329] b) In one example, the modified or prepared side information can be used for classification.

[0330] 1. In one example, the clipped residual samples can be used for classification.

[0331] 2. In one example, the scaled residual samples can be used for classification.

[0332] b. In one example, side information can be used for ALF filtering.

[0333] a) In one example, the original side information can be used for filtering.

[0334] b) In one example, the modified or prepared side information can be used for filtering.

[0335] 1. In one example, the clipped residual samples can be used for filtering.

[0336] 2. In one example, the scaled residual samples can be used for filtering.

[0337] c. In one example, side information can be used for CCALF filtering.

[0338] a) In one example, the original side information can be used for filtering.

[0339] b) In one example, the modified or prepared side information can be used for filtering.

[0340] 1. In one example, the clipped residual samples can be used for filtering.

[0341] 2. In one example, the scaled residual samples can be used for filtering.

[0342] d. In one example, side information can be used for SAO.

[0343] e. In one example, side information can be used for CCSAO.

[0344] f. In one example, side information can be used for BF.

[0345] g. In one example, side information can be used in HTDF.

[0346] h. In one example, side information can be used for DBF.

[0347] i. In one example, side information can be used in any other stages / tools.

[0348] 3) In one example, whether and / or how to modify or prepare the side information may be signaled in the bitstream.

[0349] a. In one example, a first syntax element may be signaled to indicate whether modified or prepared side information is enabled / used.

[0350] a) In one example, the first syntax element may be encoded by arithmetic coding.

[0351] 1. In one example, a first syntax element may be encoded or decoded using at least one context.

[0352] a. The context can depend on the codec information of the current block or neighboring blocks.

[0353] b. The context may depend on the filter shape of at least one neighboring block.

[0354] 2. In one example, the first syntax element may be encoded using bypass codec.

[0355] b) In one example, the first syntax element may be binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, a truncated exponential Golomb code, or the like.

[0356] c) In one example, the first syntax element may be conditionally signaled.

[0357] 1. For example, the first syntax element may be signaled only if side information is available.

[0358] d) The first syntax element can be coded in a predictive manner.

[0359] 1. The first syntax element can be predicted by an on / off decision prepared by side information of at least one neighboring block.

[0360] e) The first syntax element may be signaled independently for different color components.

[0361] 1. Alternatively, the first syntax element may be signaled and shared for different color components.

[0362] 2. Alternatively, the first syntax element may be signaled for the first color component but not for the second color component.

[0363] f) Syntax elements can be signaled in SPS, PPS, picture header, slice header, APS, CTU, CU, etc.

[0364] 4) In one example, coded reference pictures can be accessed during different coding stages.

[0365] a. In one example, coded reference pictures may be accessed during the prediction loop stage.

[0366] b. In one example, the coded reference pictures can be accessed during the loop filter stage.

[0367] a) In one example, the coded reference pictures may be accessed before DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filter process.

[0368] b) In one example, the coded reference pictures may be accessed during DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filter process.

[0369] c) In one example, the coded reference pictures may be accessed after DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filter process.

[0370] c. In one example, the coded reference pictures can be accessed after the loop filter stage.

[0371] a) In one example, the coded reference pictures can be accessed after the loop filter stage through a motion compensation based padding method.

[0372] 5) In one example, the disclosed method can be used for post-processing and / or pre-processing.

[0373] 6) The above side information can be in the same color component as the sample to be filtered, or in a different color component. The side information can be as follows:

[0374] a. In one example, residual samples can be used.

[0375] b. In one example, prediction samples may be used.

[0376] c. In one example, reconstruction at different stages may be used.

[0377] d. In one example, coded reference pictures may be used.

[0378] e. In one example, an insert picture / frame may be used.

[0379] f. In one example, the output of any predefined filter can be used.

[0380] g. Alternatively, transform domain coefficients may be used.

[0381] 7) The above side information can be at the same location as the sample point to be filtered, or within the range around the sample point to be filtered. The side information can be as follows:

[0382] a. In one example, residual samples can be used.

[0383] b. In one example, prediction samples may be used.

[0384] c. In one example, reconstruction at different stages may be used.

[0385] d. In one example, coded reference pictures may be used.

[0386] e. In one example, an insert picture / frame may be used.

[0387] f. In one example, the output of any predefined filter can be used.

[0388] g. Alternatively, transform domain coefficients may be used.

[0389] 8) In one example, the above methods can be used in combination.

[0390] 9) Alternatively, the above methods can be used alone.

[0391] 10) In one example, the proposed preparation of side information for ALF methods can be applied to any loop filtering tools, pre-processing or post-processing filtering methods in video codecs (including but not limited to ALF, CCALF or any other filtering methods).

[0392] a. In one example, the preparation of the proposed side information method can be applied to the loop filtering method.

[0393] a) In one example, the preparation of the proposed side information method can be applied to ALF.

[0394] b) In one example, the preparation of the proposed side information method can be applied to CCALF.

[0395] c) Alternatively, the preparation of the proposed side information method can be applied to other loop filtering methods.

[0396] b. In one example, the preparation of the proposed side information method can be applied to the pre-processing filtering method.

[0397] c. In one example, the preparation of the proposed side information method can be applied to a post-processing filtering method.

[0398] 11) In the above examples, a video unit may refer to a sequence, a picture, a sub-picture, a slice, a codec tree unit (CTU), a CTU row or CTU group, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.

[0399] 12) Whether and / or how the above disclosed methods are applied can be signaled in the bitstream.

[0400] a. In one example, they can be signaled at the sequence level, group of picture level, picture level, slice level, or slice group level, such as in a sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, slice header, or slice group header.

[0401] b. In one example, they can be signaled at PB, TB, CB, PU, ​​TU, CU, Virtual Pipeline Data Unit (VPDU), CTU, CTU row, slice, slice, sub-picture, or other kinds of regions containing more than one sample or pixel.

[0402] 13) Whether and / or how to apply the above disclosed method may depend on coded information such as block size, color format, single-tree partitioning or dual-tree partitioning, color component, slice type or picture type.

[0403] 14) The syntax elements disclosed above can be binarized as flags, fixed length codes, EG(x) codes, unary codes, truncated unary codes, truncated binary codes, etc. It can be signed or unsigned.

[0404] 15) The syntax elements disclosed above can be encoded or decoded using at least one context model, or they can be bypassed.

[0405] 16) The syntax elements (SE) disclosed above may be signaled in a conditional manner.

[0406] a. SE is signaled only if the corresponding function is applicable.

[0407] b. SE is signaled only if the dimensions (width and / or height) of the block meet the conditions.

[0408] 17) The syntax elements disclosed above may be signaled at a block level, a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level, such as in the codec structure of a CTU, a CU, a TU, a PU, a CTB, a CB, a TB, a PB, a sequence header, a picture header, an SPS, a VPS, a DPS, a DCI, a PPS, an APS, a slice header, or a slice group header.

[0409] 18) The proposed method(s) may be combined with another codec tool such as affine, multiple transform selection (MTS), LFNST, Merge with motion vector difference (MMVD), matrix-based intra prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), DMVR, history-based motion vector prediction (HMVP), template matching, IBC, palette, etc.

[0410] 19) The proposed method(s) may be compatible with another codec tool, such as affine, MTS, LFNST, MMVD, MIP, ISP, CCLM, CCCM, SMVD, BDOF, DMVR, HMVP, template matching, IBC, palette, etc.

[0411] a. In one example, if the proposed method(s) are used, the excluded codec tools are implicitly disabled without signaling.

[0412] b. In one example, if the excluded codec is used, the proposed method(s) are implicitly disabled without signaling.

[0413] 6. References

[0414] [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson, and R. Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in Proceedings of the Institute of Electrical and Electronics Engineers (IEEE) Picture Coding Symposium (PCS), November 2019.

[0415] Figure 13is a block diagram illustrating an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical networks (PONs), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0416] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this document. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to generate a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or transmitted via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video stored or communicated at input 4002 can be used by component 4008 to generate pixel values ​​or displayable video that is sent to display interface 4010. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it is understood that the codec tools or operations are used by the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0417] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0418] Figure 144 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this document. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, such as a graphics coprocessor.

[0419] Figure 15 4 is a flow chart of an example method 4200 for video processing. Method 4200 includes determining, at step 4202, to use side information as input to an adaptive loop filter (ALF). At step 4204, conversion is performed between visual media data and a bitstream based on the ALF. According to an example, the conversion at step 4204 may include encoding at an encoder or decoding at a decoder.

[0420] It should be noted that method 4200 can be implemented in an apparatus for processing video data that includes a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium that includes a computer program product for use with a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs method 4200.

[0421] Figure 16 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where destination device 4320 can be referred to as a video decoding device.

[0422] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of the picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to target device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0423] The target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320, or may be external to the target device 4320, wherein the target device 4320 may be configured to interface with an external display device.

[0424] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other existing and / or future standards.

[0425] Figure 17 is a block diagram illustrating an example of a video encoder 4400, wherein the video encoder 4400 may be Figure 16 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0426] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414, wherein the prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406.

[0427] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0428] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but are represented separately in the example of the video encoder 4400 for purposes of explanation.

[0429] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.

[0430] The mode selection unit 4403 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0431] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 4413 (other than the picture associated with the current video block).

[0432] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0433] In some examples, motion estimation unit 4404 may perform unidirectional prediction on the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0434] In other examples, motion estimation unit 4404 may perform bidirectional prediction on the current video block. Motion estimation unit 4404 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 4404 may then generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0435] In some examples, motion estimation unit 4404 can output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 can reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 4404 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0436] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.

[0437] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0438] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0439] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0440] The residual generation unit 4407 can generate residual data for the current video block by subtracting the (one or more) prediction video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0441] In other examples, such as in skip mode, there may be no residual data for the current video block and the residual generation unit 4407 may not perform a subtraction operation.

[0442] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0443] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0444] Inverse quantization unit 4410 and inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in buffer 4413.

[0445] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0446] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0447] Figure 18 is a block diagram illustrating an example of a video decoder 4500, wherein the video decoder 4500 may be Figure 16 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0448] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the inverse of the encoding process described with respect to video encoder 4400.

[0449] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by implementing AMVP and Merge modes.

[0450] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0451] The motion compensation unit 4502 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information, and the motion compensation unit 4502 may use the interpolation filters to generate a prediction block.

[0452] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0453] The intra prediction unit 4503 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0454] Reconstruction unit 4506 can sum the residual block with the corresponding prediction block generated by motion compensation unit 4502 or intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0455] Figure 19 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples, by adding an offset and by applying a finite impulse response (FIR) filter, respectively, taking advantage of the encoded side information of the offset and filter coefficients transmitted through the signal. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts caused by previous stages.

[0456] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding component 4618. The entropy coding component 4618 entropy encodes the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these images are stored in the reference picture cache 4612 .

[0457] Figure 20 4700 is a flow chart of an example method 4700 for video processing. The method 4700 includes determining, at step 4702, to employ an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as input side information. At step 4704, conversion is performed between visual media data and a bitstream based on the ALF. According to an example, the conversion at step 4704 may include encoding at an encoder or decoding at a decoder.

[0458] It should be noted that method 4700 can be implemented in an apparatus for processing video data that includes a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4700. Furthermore, method 4700 can be performed by a non-transitory computer-readable medium that includes a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs method 4700.

[0459] A list of some example preferred solutions is provided below.

[0460] The following solutions illustrate examples of the techniques discussed herein.

[0461] 1. A method for processing video data, comprising: determining to use side information as input of an adaptive loop filter (ALF); and performing conversion between visual media data and a bitstream based on the ALF.

[0462] 2. The method of solution 1, wherein the side information comprises residual samples, prediction samples, reconstruction, coded reference pictures, inserted pictures, intra prediction modes, coding modes, reference indices, reference lists, motion vectors, transform types, outputs from another filter, transform domain coefficients, or a combination thereof.

[0463] 3. The method according to any of solutions 1-2, wherein the side information is not modified when being input into the ALF.

[0464] 4. A method according to any of solutions 1-3, wherein the side information is clipped, scaled, transformed, filtered or fused into a range or bit depth.

[0465] 5. A method according to any of solutions 1-4, wherein the range or bit depth is defined, signaled or derived.

[0466] 6. The method according to any one of solutions 1-5, wherein the side information is used for classification of the ALF or filtering of the ALF.

[0467] 7. The method according to any of solutions 1-6, wherein the ALF is a cross-component ALF (CCALF).

[0468] 8. A method according to any one of solutions 1-7, wherein the side information is used as input to sample adaptive offset (SAO), cross-component SAO (CCSAO), bilateral filter (BF), Hadamard transform domain filter (HTDF), deblocking filter (DBF), pre-processing filter, post-processing filter or a combination thereof.

[0469] 9. The method of any of solutions 1-8, wherein the use of the side information is signaled in the bitstream.

[0470] 10. The method according to any one of solutions 1-9, wherein a syntax element indicates whether the side information is used.

[0471] 11. The method according to any one of solutions 1-10, wherein the syntax element is context-coded, and wherein the context depends on the coding information of the block or the filter shape of the block.

[0472] 12. A method according to any one of solutions 1-11, wherein the syntax elements are coded using bypass coding, binarization coding, conditional signaling, predictive coding, independent coding for different color components, or a combination thereof.

[0473] 13. A method according to any of solutions 1-12, wherein the ALF receives input from a coded reference picture, and wherein the coded reference picture is accessed during the prediction loop stage, during the loop filter stage, after the loop filter, or a combination thereof.

[0474] 14. The method according to any of solutions 1-13, wherein the side information is received from a different color component than the sample to be filtered.

[0475] 15. The method according to any one of solutions 1 to 14, wherein the side information is received from the same position as the sample point to be filtered or a range of positions around the sample point to be filtered.

[0476] 16. The method of any of solutions 1-15, wherein the use of the side information is signaled in the bitstream.

[0477] 17. The method according to any of the solutions 1-16, wherein the use of the side information depends on the encoded and decoded information.

[0478] 18. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-17.

[0479] 19. A non-transitory computer-readable medium comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs a method according to any one of solutions 1-17.

[0480] 20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining to use side information as an input of an adaptive loop filter (ALF); and generating the bitstream based on the determination.

[0481] 21. A method for storing a bitstream of a video, comprising: determining to employ side information as an input of an adaptive loop filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0482] 22. A method, apparatus or system as described in this document.

[0483] The following solutions illustrate further examples of the techniques discussed herein.

[0484] 1. A method for processing video data, comprising: determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; and performing conversion between visual media data and a bitstream based on the ALF.

[0485] 2. The method according to solution 1, wherein the ALF receives the predicted samples of the current picture as side information used as input.

[0486] 3. The method according to any of solutions 1-2, wherein the ALF receives as input the reconstructed samples before applying a deblocking filter (DBF).

[0487] 4. The method according to any of the solutions 1-3, wherein the ALF receives as input side information reconstructed samples before applying the ALF or before applying a cross-component ALF (CCALF).

[0488] 5. The method according to any of solutions 1-4, wherein the ALF receives a forward reference picture or a backward reference picture as side information used as input.

[0489] 6. The method according to any one of solutions 1-5, wherein the ALF receives side information as input, the side information comprising: the output of an offline trained filter, the output of a Gaussian filter, the output of a low-pass filter, or a combination thereof.

[0490] 7. The method according to any one of solutions 1-6, wherein the side information is used directly without being modified.

[0491] 8. A method according to any of solutions 1-7, wherein the side information is filtered into range, domain, bit depth or a combination thereof.

[0492] 9. The method according to any one of solutions 1-8, wherein the side information is used for classification or filtering in the ALF or CCALF.

[0493] 10. The method according to any one of solutions 1-9, wherein the side information includes clipped residual samples used for filtering in the ALF or CCALF.

[0494] 11. The method according to any of solutions 1-10, wherein a first syntax element is included in the bitstream to indicate whether modified side information or prepared side information is enabled or used.

[0495] 12. A method according to any one of solutions 1-11, wherein the first syntax element is binarized by a unary code, a truncated unary code, a fixed-length code, an exponential Golomb code, a truncated exponential Golomb code, or a combination thereof.

[0496] 13. The method according to any of solutions 1-12, wherein the first syntax element is signaled independently for different color components.

[0497] 14. A method according to any one of solutions 1-13, wherein the ALF receives input from a coded reference picture, and wherein the coded reference picture is accessed during application of: a deblocking filter (DBF), sample adaptive offset (SAO), cross-component SAO (CCSAO), a bilateral filter (BF), a chroma BF (ChromaBF), the ALF, the CCALF or a combination thereof.

[0498] 15. A method according to any one of solutions 1-14, wherein the ALF receives side information as input, the side information including residual samples, reconstructed samples at different stages, coded reference pictures, output from the filter, and wherein the side information is received from the same color component or a different color component as the samples to be filtered.

[0499] 16. A method according to any one of solutions 1 to 15, wherein the ALF receives side information as input, the side information including residual samples, reconstructed samples at different stages, coded reference pictures, and outputs from filters, and wherein the side information is obtained from the same position as the sample to be filtered or from a position within a range around the sample to be filtered.

[0500] 17. The method according to any one of solutions 1-16, wherein the side information is applied in the ALF or CCALF.

[0501] 18. A method according to any one of solutions 1-17, wherein the side information includes residual samples of other coded pictures, predicted samples of other coded pictures, reconstructed samples before applying SAO, CCSAO, BF or Hadamard transform domain filter (HTDF), reconstructed samples before applying any filter, long-term reference pictures, inserted pictures generated from data inside the current GOP, inserted pictures generated from data outside the current GOP, intra-frame prediction mode, coding mode, reference index, reference list, motion vector, transform type, output from Sobel filter, Prewitt filter, Roberts filter, Canny filter, HTDF, BF, high-pass filter or any other filter, transform domain coefficients for a transform including discrete cosine transform (DCT), discrete wavelet transform (DWT), low frequency non-separable transform (LFNST), non-separable primary transform (NSPT), HTDF or any other transform, or a combination thereof.

[0502] 19. A method according to any one of solutions 1-18, wherein the side information or residual samples are clipped to a predefined, signaled or derived clipping range and to a predefined, signaled or derived N-bit depth.

[0503] 20. A method according to any one of solutions 1-19, wherein the side information is: scaled to a predefined, signaled or derived range, scaled to a predefined, signaled or derived N-bit depth, transformed to a predefined, signaled or derived range, transformed to a predefined, signaled or derived domain, transformed to a predefined, signaled or derived N-bit depth, filtered to a predefined, signaled or derived range, filtered to a predefined, signaled or derived domain, filtered to a predefined, signaled or derived N-bit depth, fused by weighted sum, fused by online trained coefficients, fused by offline trained or predefined coefficients, or a combination thereof.

[0504] 21. The method according to any one of solutions 1-20, wherein modified or prepared side information, including clipped residual samples or scaled residual samples, is used for classification in the ALF, filtering in the ALF or filtering in CCALF.

[0505] 22. The method according to any of the solutions 1-21, wherein the side information is used as input to SAO, CCSAO, BF, HTDF, DBF, pre-processing filter, post-processing filter or a combination thereof.

[0506] 23. The method of any of solutions 1-22, wherein the use of the side information is signaled via a syntax element in the bitstream.

[0507] 24. The method according to any of solutions 1-23, wherein the syntax element is context coding or bypass coding, and wherein the context depends on the coding information of the block or the filter shape of the block.

[0508] 25. A method according to any one of solutions 1-24, wherein the syntax element: is turned on by signal transmission when side information is available, is predicted by an on / off decision prepared by side information of at least one neighboring block, is signaled and shared by different color components, is signaled for a first color component and not signaled for a second color component, is signaled in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptation parameter set (APS), a codec tree unit (CTU), a codec unit (CU) or a combination thereof.

[0509] 26. A method according to any of solutions 1-25, wherein the ALF receives input from a coded reference picture, and wherein the coded reference picture is accessed during the prediction loop stage, during the loop filter stage, after the loop filter, or a combination thereof.

[0510] 27. A method according to any one of solutions 1-26, wherein the encoded reference picture is accessed before or after applying DBF, SAO, CCSAO, BF, ChromaBF, the ALF, CCALF, or wherein the encoded reference picture is accessed through motion compensation based padding after the loop filter stage.

[0511] 28. A method according to any of solutions 1-27, wherein the side information is applied to a pre-processing filter or a post-processing filter of the video.

[0512] 29. A method according to any one of solutions 1-28, wherein the ALF receives side information as input, the side information including predicted samples, inserted pictures or frames, or transform coefficients, wherein the side information is received from the same color component or a different color component as the sample to be filtered, or wherein the side information is obtained from the same position as the sample to be filtered or from a position within a range around the sample to be filtered.

[0513] 30. The method according to any one of solutions 1-29, wherein the method is used in combination or alone.

[0514] 31. The method according to any of the solutions 1-30, wherein the preparation of the side information for the ALF is applied as part of a loop filter tool, a pre-processing method or a post-processing method.

[0515] 32. A method according to any one of solutions 1-31, wherein the filter is applied to a video unit, and wherein the video unit is a sequence, a picture, a sub-picture, a slice, a CTU, a CTU row, a CTU group, a CU, a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.

[0516] 33. The method of any of solutions 1-32, wherein use of the method is signaled in the bitstream.

[0517] 34. A method according to any one of solutions 1-33, wherein the use of the method is transmitted by signal at the sequence level, picture group level, picture level, slice level or slice group level, and is included in a sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header or slice group header, or wherein the use of the method is transmitted by signal at a prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, slice, slice, sub-picture or other area containing more than one sample or pixel.

[0518] 35. A method according to any one of solutions 1-34, wherein the application of the method depends on the encoded and decoded information, and the encoded and decoded information includes block size, color format, single tree partitioning, dual tree partitioning, color component, slice type or picture type.

[0519] 36. A method according to any one of solutions 1-35, wherein the syntax elements: are binarized as flags, fixed-length codes, exponential-Golomb codes, unary codes, truncated unary codes or truncated binary codes, are signed or unsigned, are encoded or decoded using a context model, are bypassed, are conditionally transmitted through a signal, are transmitted through a signal only when the corresponding function is applied, are transmitted through a signal when the height or width of the block meets a condition, or a combination thereof.

[0520] 37. A method according to any of solutions 1-36, wherein the syntax elements are transmitted through signals at the sequence level, picture group level, picture level, slice level or slice group level, in a sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, slice header or slice group header.

[0521] 38. A method according to any one of solutions 1-37, wherein the method is used in combination with or excluded from use with affine, multiple transform selection (MTS), LFNST, Merge with Motion Vector Difference (MMVD), matrix-based intra prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra block copy (IBC) or palette.

[0522] 39. A method according to any of solutions 1-38, wherein the excluded tools are implicitly disabled without signaling, or wherein the method is implicitly disabled without signaling when the excluded tools are used.

[0523] 40. The method of any of solutions 1-39, wherein the converting comprises encoding the visual media data into the bitstream.

[0524] 41. The method of any of solutions 1-40, wherein the converting comprises decoding the visual media data from the bitstream.

[0525] 42. An apparatus for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Solutions 1-41.

[0526] 43. A non-transitory computer-readable medium comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs a method according to any one of Solutions 1-41.

[0527] 44. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; and generating the bitstream based on the determination.

[0528] 45. A method for storing a bitstream of a video, comprising: determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of a current picture as side information used as input; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0529] In the described solution, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the described solution, a decoder can parse syntax elements in the codec representation according to the format rules using known information about the presence and absence of syntax elements to produce decoded video.

[0530] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits spread across the same position in the bitstream or at different positions as defined by the syntax. For example, a macroblock may be encoded based on transformed and encoded error residual values, and bits in headers and other fields in the bitstream may also be used. Furthermore, during conversion, a decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solution. Similarly, an encoder may determine whether to include or not include particular syntax fields, and generate a codec representation accordingly by including the syntax fields or excluding the syntax fields from the codec representation.

[0531] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0532] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that preserves other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to a related program, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers that are located at a site or are distributed across multiple sites and interconnected by a communication network.

[0533] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0534] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks. However, a computer need not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0535] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features unique to particular embodiments of particular technologies. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. In addition, although features may function in certain combinations as described above and may even be initially claimed in this manner, in some cases, one or more features in a claimed combination may be omitted from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.

[0536] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed sequentially in the particular order or sequence shown, or that all illustrated operations be performed to achieve desired results. Furthermore, the partitioning of various system components in the embodiments described in this patent document should not be understood as requiring such partitioning in all embodiments.

[0537] Only a few implementations and examples are described, and other implementations, improvements, and variations can be made based on what is described and illustrated in this patent document.

[0538] A first component is directly coupled to a second component when there are no intervening components (other than lines, traces, or other intermediaries between the first and second components). A first component is indirectly coupled to a second component when there are intervening components other than lines, traces, or other intermediaries between the first and second components. The term "coupled" and its variations include both direct and indirect couplings. The use of the term "about" is intended to include a range of ±10% of the subsequent figure unless otherwise indicated.

[0539] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples should be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0540] In addition, the techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled can be directly connected, or can be indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or other). Other examples of variations, substitutions, and alterations can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for processing video data, comprising: Determining to use an adaptive loop filter (ALF), wherein the ALF receives residual samples of the current picture as side information used as input; as well as Conversion between visual media data and a bitstream is performed based on the ALF.

2. The method according to claim 1, wherein The ALF receives the predicted samples of the current picture as side information used as input.

3. The method according to any one of claims 1 to 2, wherein The ALF receives as input the reconstructed samples before applying a deblocking filter (DBF).

4. The method according to any one of claims 1 to 3, wherein The ALF receives as input side information reconstructed samples before applying the ALF or before applying a cross-component ALF (CCALF).

5. The method according to any one of claims 1 to 4, wherein The ALF receives a forward reference picture or a backward reference picture as side information as input.

6. The method according to any one of claims 1 to 5, wherein The ALF receives as input side information including: an output of an offline trained filter, an output of a Gaussian filter, an output of a low-pass filter, or a combination thereof.

7. The method according to any one of claims 1 to 6, wherein The side information is used directly without modification.

8. The method according to any one of claims 1 to 7, wherein The side information is filtered into range, domain, bit depth, or a combination thereof.

9. The method according to any one of claims 1 to 8, wherein The side information is used for classification or filtering in the ALF or CCALF.

10. The method according to any one of claims 1 to 9, wherein The side information includes clipped residual samples used for filtering in the ALF or CCALF.

11. The method according to any one of claims 1 to 10, wherein A first syntax element is included in the bitstream to indicate whether modified side information or prepared side information is enabled or used.

12. The method according to any one of claims 1 to 11, wherein The first syntax element is binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, a truncated exponential Golomb code, or a combination thereof.

13. The method according to any one of claims 1 to 12, wherein The first syntax element is signaled independently for different color components.

14. The method according to any one of claims 1 to 13, wherein: The ALF receives input from a coded reference picture, and wherein the coded reference picture is accessed during application of: a deblocking filter (DBF), sample adaptive offset (SAO), cross-component SAO (CCSAO), a bilateral filter (BF), a chroma BF (ChromaBF), the ALF, the CCALF, or a combination thereof.

15. The method according to any one of claims 1 to 14, wherein The ALF receives as input side information, the side information comprising residual samples, reconstructed samples at different stages, coded reference pictures, and outputs from filters, wherein the side information is received from the same color component as the samples to be filtered or a different color component.

16. The method according to any one of claims 1 to 15, wherein The ALF receives side information as input, the side information including residual samples, reconstructed samples at different stages, coded reference pictures, and output from a filter, wherein the side information is obtained from the same position as the sample to be filtered or from a position within a range around the sample to be filtered.

17. The method according to any one of claims 1 to 16, wherein: The side information is applied in the ALF or CCALF.

18. The method according to any one of claims 1 to 17, wherein The side information includes residual samples of other coded pictures, predicted samples of other coded pictures, reconstructed samples before applying SAO, CCSAO, BF or Hadamard transform domain filter (HTDF), reconstructed samples before applying any filter, long-term reference pictures, inserted pictures generated from data inside the current GOP, inserted pictures generated from data outside the current GOP, intra-frame prediction mode, coding mode, reference index, reference list, motion vector, transform type, output from Sobel filter, Prewitt filter, Roberts filter, Canny filter, HTDF, BF, high-pass filter or any other filter, transform domain coefficients for a transform including discrete cosine transform (DCT), discrete wavelet transform (DWT), low frequency non-separable transform (LFNST), non-separable primary transform (NSPT), HTDF or any other transform, or a combination thereof.

19. The method according to any one of claims 1 to 18, wherein The side information or residual samples are clipped to a predefined, signaled or derived clipping range and to a predefined, signaled or derived N-bit depth.

20. The method according to any one of claims 1 to 19, wherein The side information is: scaled to a predefined, signaled or derived range, scaled to a predefined, signaled or derived N-bit depth, transformed to a predefined, signaled or derived range, transformed to a predefined, signaled or derived domain, transformed to a predefined, signaled or derived N-bit depth, filtered to a predefined, signaled or derived range, filtered to a predefined, signaled or derived domain, filtered to a predefined, signaled or derived N-bit depth, fused by weighted sum, fused by online trained coefficients, fused by offline trained or predefined coefficients, or a combination thereof.

21. The method according to any one of claims 1 to 20, wherein The modified or prepared side information, including the clipped residual samples or the scaled residual samples, is used for classification in the ALF, filtering in the ALF, or filtering in the CCALF.

22. The method according to any one of claims 1 to 21, wherein The side information is used as input to SAO, CCSAO, BF, HTDF, DBF, pre-processing filter, post-processing filter, or a combination thereof.

23. The method according to any one of claims 1 to 22, wherein The use of the side information is signaled via syntax elements in the bitstream.

24. The method according to any one of claims 1 to 23, wherein The syntax element is context coding or bypass coding, and wherein the context depends on coding information of a block or a filter shape of a block.

25. The method according to any one of claims 1 to 24, wherein The syntax element is: turned on by signal transmission when side information is available, predicted by an on / off decision prepared by side information of at least one neighboring block, signaled and shared by different color components, signaled for a first color component and not signaled for a second color component, and signaled in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptation parameter set (APS), a codec tree unit (CTU) or a codec unit (CU) or a combination thereof.

26. The method according to any one of claims 1 to 25, wherein The ALF receives input from coded reference pictures, and wherein the coded reference pictures are accessed during a prediction loop stage, during a loop filter stage, after a loop filter, or a combination thereof.

27. The method according to any one of claims 1 to 26, wherein The coded reference picture is accessed before or after applying DBF, SAO, CCSAO, BF, ChromaBF, the ALF, CCALF, or wherein the coded reference picture is accessed through motion compensation based padding after the loop filter stage.

28. The method according to any one of claims 1 to 27, wherein The side information is applied to a pre-processing filter or a post-processing filter of the video.

29. The method according to any one of claims 1 to 28, wherein The ALF receives as input side information, the side information comprising prediction samples, interpolated pictures or frames, or transform coefficients, wherein the side information is received from the same color component as or a different color component than the sample to be filtered, or wherein the side information is obtained from the same position as or from a position within a range around the sample to be filtered.

30. The method according to any one of claims 1 to 29, wherein The methods are used in combination or individually.

31. The method according to any one of claims 1 to 30, wherein The preparation of side information for the ALF is applied as part of a loop filter tool, a pre-processing method or a post-processing method.

32. The method according to any one of claims 1 to 31, wherein The filter is applied to a video unit, and wherein the video unit is a sequence, a picture, a sub-picture, a slice, a CTU, a CTU row, a CTU group, a CU, a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.

33. The method according to any one of claims 1 to 32, wherein The use of the method is signaled in the bitstream.

34. The method according to any one of claims 1 to 33, wherein Use of the method is signaled at a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level, and is included in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header, or wherein, The use of the method is signaled at a prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline decoding unit (VPDU), codec tree unit (CTU), CTU row, slice, slice, sub-picture, or other region containing more than one sample or pixel.

35. The method according to any one of claims 1 to 34, wherein The application of the method depends on the coded information including block size, color format, single-tree partitioning, dual-tree partitioning, color component, slice type or picture type.

36. The method according to any one of claims 1 to 35, wherein Syntax element: binarized as a flag, fixed-length code, Exponential Golomb code, unary code, truncated unary code or truncated binary code, signed or unsigned, encoded or decoded using a context model, bypassed, conditionally signaled, signaled only when the corresponding function is applied, signaled when the block height or width meets a condition, or a combination thereof.

37. The method according to any one of claims 1 to 36, wherein Syntax elements are signaled at the sequence level, group of picture level, picture level, slice level, or slice group level in a sequence header, picture header, SPS, VPS, DPS, decoding capability information (DCI), PPS, APS, slice header, or slice group header.

38. The method according to any one of claims 1 to 37, wherein The method is used in combination with or excluded from use with affine, multiple transform selection (MTS), LFNST, Merge with Motion Vector Difference (MMVD), matrix-based intra prediction (MIP), ISP, cross-component linear model (CCLM), convolutional cross-component model (CCCM), symmetric motion vector difference (SMVD), bidirectional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based motion vector prediction (HMVP), template matching, intra block copy (IBC), or palette.

39. The method according to any one of claims 1 to 38, wherein The excluded tools are implicitly disabled without signaling, or wherein the method is implicitly disabled without signaling when the excluded tools are used.

40. The method according to any one of claims 1 to 39, wherein The converting includes encoding the visual media data into the bitstream.

41. The method according to any one of claims 1 to 40, wherein The converting includes decoding the visual media data from the bitstream.

42. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-41.

43. A non-transitory computer-readable medium comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method according to any one of claims 1-41.

44. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein: The method comprises: Determining to use an adaptive loop filter (ALF), the ALF receiving residual samples of the current picture as side information used as input; and The bitstream is generated based on the determination.

45. A method for storing a bitstream of a video, comprising: Determining to use an adaptive loop filter (ALF), wherein the ALF receives residual samples of the current picture as side information used as input; generating the bitstream based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.