Multiple side information for adaptive loop filter in video coding
By introducing an adaptive loop filter (ALF) of extended taps and airspace taps in video encoding and decoding, the problem of inefficient encoding and decoding in the prior art is solved, and more efficient artifact repair and coding efficiency are achieved.
Patent Information
- Application Number
- CN202380090052.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-29
- Filing Date
- 2023-12-20
- Publication Date
- 2025-08-08
AI Technical Summary
When using high-resolution video, it is difficult to effectively reduce artifacts and improve encoding and decoding efficiency. Especially in the application of adaptive loop filters (ALFs), the filter design in the prior art is not flexible enough, resulting in low encoding efficiency.
Adaptive loop filter (ALF) with extended taps and airspace taps is adopted to dynamically adjust the filter parameters to adapt to the video characteristics of different regions by converting between video data and bit streams, thereby improving the encoding and decoding efficiency.
More efficient artifact repair and coding efficiency improvement are achieved, especially in high-resolution video processing, which significantly reduces encoding delay and improves video quality.
Smart Images

Figure CN120457697A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2022 / 143473 filed on December 29, 2022, the teachings and disclosures of which are incorporated herein by reference. Technical Field
[0003] This application document relates to the generation, storage and use of digital audio and video media information in file format. Background Art
[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demands for digital video usage are likely to continue to grow. Summary of the Invention
[0005] A first aspect relates to a method for processing video data, comprising: determining to employ an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; and performing conversion between visual media data and a bitstream based on the ALF.
[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the above aspects.
[0007] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of the above aspects.
[0008] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to adopt an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; and generating the bitstream based on the determination.
[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] For purposes of clarity, any of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.
[0011] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0013] Figure 1 An example of nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture is shown.
[0014] Figure 2 An example encoder block diagram is shown.
[0015] Figure 3 An example picture divided into raster scan strips is shown.
[0016] Figure 4 An example picture segmented into rectangular scan strips is shown.
[0017] Figure 5 An example picture segmented into bricks is shown.
[0018] Figures 6A-6C An example of a codec tree block (CTB) crossing a picture boundary is shown.
[0019] Figure 7 An example of an intra prediction mode is shown.
[0020] Figure 8 Examples of block boundaries in a picture are shown.
[0021] Figure 9 An example of pixels involved in filter usage is shown.
[0022] Figure 10 An example of a filter shape of an adaptive loop filter (ALF) is shown.
[0023] Figure 11 An example of transform coefficients supported by a 5×5 diamond filter is shown.
[0024] Figure 12 An example of relative coordinates supported by a 5×5 diamond filter is shown.
[0025] Figure 13 An example shape of at least one spatial tap is shown.
[0026] Figure 14 An example shape of at least one spatial tap is shown.
[0027] Figure 15 is a block diagram illustrating an example video processing system.
[0028] Figure 16 is a block diagram of an example video processing device.
[0029] Figure 17 is a flow chart of an example method of video processing.
[0030] Figure 18 is a block diagram illustrating an example video encoding and decoding system.
[0031] Figure 19 is a block diagram illustrating an example encoder.
[0032] Figure 20 is a block diagram illustrating an example decoder.
[0033] Figure 21 is a schematic diagram of an example encoder.
[0034] Figure 22 is a flow chart of an example method of video processing. DETAILED DESCRIPTION
[0035] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified without departing from the full scope of the appended claims and their equivalents.
[0036] The section headings used in this document are for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section only. In addition, the techniques described herein are applicable to other video codec protocols and designs.
[0037] 1. Preliminary Discussion
[0038] This document relates to video codecs. Specifically, it relates to loop filters and other codec tools used in image / video codecs. The concepts presented can be applied alone or in various combinations to video codecs such as High Efficiency Video Codec (HEVC), Versatile Video Codec (VVC), or other video codecs.
[0039] 2. Abbreviation
[0040] This disclosure includes the following abbreviations. Advanced Video Codec (Recommendation ITU-T H.264 | ISO / IEC 14496-10) (AVC), Coded Picture Buffer (CPB), Pure Random Access (CRA), Codec Tree Unit (CTU), Coded Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), General Constraint Information (GCI), High Efficiency Video Codec, also known as Recommendation ITU-T H.265 | ISO / IEC 23008-2, (HEVC), Joint Exploration Model (JEM), Motion Constrained Slice Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Tier, and Level (PTL), Picture Unit (PU), Reference Picture Resampled (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Codec Layer (VCL), Video Parameter Set (VPS), Versatile Video Codec (also known as Recommendation ITU-T H.266 | ISO / IEC 23090-3) (VVC), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Codec Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Codec Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).
[0041] 3. Video codec standards
[0042] Video codec standards have evolved primarily through the development of ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the MPEG-1 and MPEG-4 visual standards, and the two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC [1] standards. Starting with H.262, video codec standards are based on a hybrid video codec structure that utilizes temporal prediction plus transform coding. In order to explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG. JVET adopted many methods and put them into a reference software called the Joint Exploration Model (JEM) [2]. When the Versatile Video Codec (VVC) project was officially launched, JVET was renamed the Joint Video Experts Team (JVET). VVC is a codec standard that aims to reduce the bit rate by 50% compared to HEVC. The working draft of VVC and the VVC Test Model (VTM) are constantly updated.
[0043] A sample version of the VVC draft, Versatile Video Codec (Draft 10), can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip A sample version of VVC's reference software, called VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2 .
[0044] The Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) are studying the potential need to standardize future video codec technologies with compression capabilities significantly exceeding the current VVC standard. Such future standardization efforts could take the form of extensions to VVC or entirely new standards. These groups are jointly conducting this exploratory activity in a joint collaborative effort known as the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. The first Exploratory Experiment (EE) was established by JVET, and reference software, called the Enhanced Compression Model (ECM), is currently in use. The ECM test model is continuously updated.
[0045] 3.1 Color Space and Chroma Downsampling
[0046] A color space, also known as a color model (or color system), is a mathematical model that describes the range of colors as a tuple of numbers, such as 3 or 4 values or color components (e.g., RGB). Generally speaking, a color space is a refinement of a coordinate system and subspaces. For video compression, the most commonly used color spaces are luminance, blue-difference chrominance, and red-difference chrominance (YCbCr) and red, green, blue (RGB).
[0047] YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color imaging pipeline in video and digital photography systems. Y' is the luma component, and CB and CR are the blue-difference and red-difference chroma components. Y' (with a prime) is distinguished from luma Y, meaning that light intensity is encoded nonlinearly based on the gamma-corrected RGB primaries.
[0048] Chroma subsampling is the practice of encoding an image at a lower resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to differences in color than in luminance. 3.1.1 4:4:4
[0050] In 4:4:4, each of the three components, Y', Cb, and Cr, has the same sampling rate. Therefore, there is no chroma downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2 4:2:2
[0052] In 4:3:2, the two chroma components are sampled at half the sampling rate of luma. The horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with little visual difference.
[0053] Figure 1 An example of nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture is shown. 3.1.3 4:2:0
[0055] In 4:2:0, horizontal sampling is doubled compared to 4:1:1, but vertical resolution is halved because the Cb and Cr channels are sampled only on alternate lines in this scheme. Therefore, the data rate remains the same. Cb and Cr are downsampled by a factor of 2 both horizontally and vertically. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positions.
[0056] In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are located between pixels vertically (in interstitial positions). In Joint Photographic Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are located in interstitial positions, in the middle of alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally. Vertically, they are co-located on alternate lines.
[0057] chroma_format_idc separate_colour_plane_flag Chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0058] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and Separate_colour_plane_flag
[0059] 3.2 Example Encoding and Decoding Process of Video Codec
[0060] Figure 2An example encoder block diagram for VVC is shown, which contains three loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offset and applying a finite impulse response (FIR) filter, respectively. Coded side information is used to signal the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be seen as a tool that attempts to capture and repair artifacts produced by previous stages.
[0061] 3.3 Definition of Video / Codec Unit
[0062] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the picture. A slice can be divided into one or more bricks, each of which contains multiple CTU rows within the slice. Slices that are not divided into multiple bricks may also be called bricks. However, bricks that are proper subsets of a slice may not be called slices. A slice contains multiple slices of a picture or multiple bricks of a slice.
[0063] Two striping modes are supported: raster scan striping and rectangular striping. In raster scan striping, a strip consists of a sequence of slices from a raster scan of the image. In rectangular striping, a strip consists of multiple tiles that together form a rectangular region of the image. The tiles within a rectangular strip are in the order in which the tiles were raster scanned.
[0064] Figure 3 An example picture is shown that is divided into raster scan strips. For example, Figure 3 An example of raster scan stripe partitioning of a picture with 18x12 luma CTUs is shown, where the picture is divided into 12 slices and 3 raster scan stripes.
[0065] Figure 4 An example image is shown that is divided into rectangular scan strips. For example, Figure 4 An example of rectangular slice partitioning of a picture with 18x12 luma CTUs is shown, where the picture is partitioned and / or divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0066] Figure 5 An example picture is shown that is divided into bricks. For example, Figure 5 An example of a picture partitioned into slices, tiles, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 tiles (the upper left slice contains 1 tile, the upper right slice contains 5 tiles, the lower left slice contains 2 tiles, and the lower right slice contains 3 tiles), and 4 rectangular strips.
[0067] 3.3.1 CTU / CTB Size
[0068] In VVC, the CTU size signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2 can be as small as 4×4.
[0069] 7.3.2.3 Sequence Parameter Set RBSP Syntax
[0070]
[0071]
[0072]
[0073] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size for each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0074] CtbLog2SizeY=log2_ctu_size_minus2+2 (7-9)
[0075] CtbSizeY=1< <CtbLog2SizeY (7-10)
[0076] MinCbLog2SizeY=log2_min_luma_coding_block_size_minus2+2 (7-11)
[0077] MinCbSizeY=1< <MinCbLog2SizeY (7-12)
[0078] MinTbLog2SizeY=2 (7-13)
[0079] MaxTbLog2SizeY = 6 (7-14)
[0080] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0081] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0082] PicWidthInCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY) (7-17)
[0083] PicHeightInCtbsY = Ceil(pic_height_in_luma_samples ÷ CtbSizeY) (7-18)
[0084] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0085] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0086] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0087] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0088] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)
[0089] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0090] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0091] 3.3.2 CTUs in a Picture
[0092] Figures 6A-6C Examples of CTBs crossing the picture boundary are shown. Figure 6A CTBs crossing the bottom picture boundary are shown. Figure 6B CTBs crossing the right picture boundary are shown. Figure 6C CTBs crossing the bottom - right picture boundary are shown. Assume a CTB / largest coding unit (LCU) size indicated by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For Figures 6A-6C the CTBs depicted in , the CTB size still equals M×N. However, the bottom boundary / right boundary of the CTB is outside the picture.
[0093] 3.4 Intra - prediction
[0094] Figure 7 Examples of intra - prediction modes are shown. To capture any edge direction presented in natural videos, the number of directional intra - prediction modes is extended from 33 used in HEVC to 65. The extended directional modes are as Figure 7 shown, and the planar and DC modes remain unchanged. These denser directional intra - prediction modes apply to all block sizes as well as luminance and chrominance intra - prediction.
[0095] As Figure 7 shown, the angular intra - prediction direction can be defined as the clockwise direction from 45 degrees to - 135 degrees. In VTM, for non - square blocks, several angular intra - prediction modes are adaptively replaced by wide - angle intra - prediction modes. The replaced modes are signaled and remapped to the indices of the wide - angle modes after parsing. The total number of intra - prediction modes remains the same, e.g., 67, and the intra - mode coding and decoding remain unchanged.
[0096] In HEVC, each intra - coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is required to generate intra - prediction values using the DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non - square blocks.
[0097] 3.5 Inter - prediction
[0098] For each inter-frame predicted CU, the motion parameters include motion vector, reference picture index, reference picture list usage index and extended information for the new coding features of VVC that will be used for inter-frame prediction sample generation. The motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no encoded motion vector increments and / or reference picture indexes. Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs including spatial and temporal candidates and the extended scheduling introduced in VVC. Merge mode can be applied to any inter-frame predicted CU, not just for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, where for each CU, the motion vector, the corresponding reference picture index of each reference picture list, the reference picture list usage flag and other useful information are explicitly transmitted through signals.
[0099] 3.6 Deblocking Filter
[0100] Deblocking filtering is an example loop filter in video codecs. In VVC, the deblocking filtering process is applied to CU boundaries, transform sub-block boundaries, and predicted sub-block boundaries. Prediction sub-block boundaries include prediction unit boundaries introduced by sub-block based temporal motion vector prediction (SbTMVP) and affine mode. Transform sub-block boundaries include transform unit boundaries introduced by sub-block transform (SBT) and intra sub-partitioning (ISP) mode and transform, which is due to the implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first filtering the vertical edges of the entire picture horizontally, and then filtering the horizontal edges vertically. This specific order enables multiple horizontal filtering processes or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with only a small processing delay.
[0101] First, vertical edges in the image are filtered. Then, horizontal edges in the image are filtered, using the samples modified by the vertical edge filtering process as input. Vertical and horizontal edges in the CTB of each CTU are processed separately on a codec unit basis. Vertical edges of codec blocks in a codec unit are filtered starting from the left edge of the codec block and proceeding to the right edge of the codec block, following the geometric order of the codec blocks. Horizontal edges of codec blocks in a codec unit are filtered starting from the top edge of the codec block and proceeding to the bottom edge of the codec block, following the geometric order of the codec blocks.
[0102] Figure 8 An example of block boundaries in a picture is shown. For example, Figure 8Picture samples and horizontal and vertical block boundaries on an 8x8 grid are shown, as well as non-overlapping blocks of 8x8 samples that can be deblocked in parallel.
[0103] 3.6.1 Boundary Decision
[0104] Filtering is applied to 8x8 block boundaries. In addition, such boundaries must be transform block boundaries or codec subblock boundaries, for example, due to the use of affine motion prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0105] 3.6.2 Boundary strength calculation
[0106] For transform block boundaries / codec sub-block boundaries, if the boundary lies in an 8×8 grid, the boundary can be filtered and the settings of bS[xDi][yDj] (where [xDi][yDj] represents the coordinates) for the edge are defined as Table 2 and Table 3, respectively.
[0107]
[0108]
[0109] Table 2. Boundary Strength (When SPS IBC is Disabled)
[0110]
[0111] Table 3. Boundary Strength (When SPS IBC is Enabled)
[0112] 3.6.3 Deblocking Decision for Luminance Component
[0113] Figure 9 An example of pixels involved in filter usage is shown. For example, Figure 9 The pixels involved in the filter on / off decision and strong / weak filter selection are shown. The wider and stronger luma filter is used only when conditions 1, 2, and 3 are all true. Condition 1 is the "large block condition." This condition detects whether the samples on the P and Q sides belong to large blocks, which are represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0114] bSidePisLargeBlk = ((edge type is vertical and p0 belongs to a CU with width >= 32) || (edge type is horizontal and p0 belongs to a CU with height >= 32))? True: False
[0115] bSideQisLargeBlk = ((edge type is vertical and q0 belongs to a CU with width >= 32) || (edge type is horizontal and q0 belongs to a CU with height >= 32))? True: False
[0116] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0117] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? True: False
[0118] Next, if condition 1 is true, condition 2 will be checked. First, the following variables are derived:
[0119] First, derive dp0, dp3, dq0, dq3 according to the HEVC method
[0120] if (p side is greater than or equal to 32)
[0121] dp0=(dp0+Abs(p50-2*p40+p30)+1)>>1
[0122] dp3=(dp3+Abs(p53-2*p43+p33)+1)>>1
[0123] if (q side is greater than or equal to 32)
[0124] dq0=(dq0+Abs(q50-2*q40+q30)+1)>>1
[0125] dq3=(dq3+Abs(q53-2*q43+q33)+1)>>1
[0126] Condition 2 = (d < β)? True: False
[0127] Where d = dp0 + dq0 + dp3 + dq3.
[0128] If conditions 1 and 2 are valid, it further checks whether any block in the block uses a subblock:
[0129]
[0130] Finally, if both conditions 1 and 2 are valid, the deblocking method will check condition 3 (large block strong filter condition), which is defined as follows. In condition 3 StrongFilterCondition, the following variables are derived:
[0131] Derived dpq in the same way as HEVC.
[0132] Derived according to HEVC method sp3 = Abs (p3-p0)
[0133] if (p side is greater than or equal to 32)
[0134] if(Sp==5)
[0135] sp3=(sp3+Abs(p5-p3)+1)>>1
[0136] else
[0137] sp3=(sp3+Abs(p7-p3)+1)>>1
[0138] According to the HEVC method, sq3 = Abs(q0-q3)
[0139] if (q side is greater than or equal to 32)
[0140] If(Sq==5)
[0141] sq3=(sq3+Abs(q5-q3)+1)>>1
[0142] else
[0143] sq3=(sq3+Abs(q7-q3)+1)>>1
[0144] According to HEVC, StrongFilterCondition = (dpq less than (β>>2), sp3+sq3 less than (3*β>>5), and Abs(p0-q0) less than (5*tC+1)>>1)? True: False.
[0145] 3.6.4 Stronger Deblocking Filter for Luma
[0146] When the samples on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge is >= 32, and when the height of the horizontal edge is >= 32, the sample is defined as belonging to a large block. The bilinear filter is listed as follows. The block boundary samples pi (i=0 to Sp-1) and qi (j=0 to Sq-1) in the above HEVC deblocking (pi and qi are the i-th sample in the row for filtering vertical edges, or the i-th sample in the column for filtering horizontal edges) are then replaced by linear interpolation as follows:
[0147] p i ′=(fi i *Middle s,t +(64-f i )*P s +32)>>6), limit to p i ±tcPDi
[0148] q j ′=(g j *Middle s,t +(64-g j )*Q s +32)>>6), limit to q j ±tcPD j
[0149] tcPD i and tcPD j The term is the position-dependent clipping mentioned above, and g j 、f i 、Middle s,t 、P s and Q s Given below.
[0150] 3.6.5 Chroma Deblocking Decision
[0151] A chroma strong filter is applied on both sides of a block boundary. Here, the chroma filter is selected when the distance between both sides of the chroma edge is greater than or equal to 8 (chroma position) and the following decision with three conditions is met. The first is a decision for boundary strength and large blocks. The filter can be applied when the block width or height orthogonally crossing the block edge is equal to or greater than 8 in the chroma sample domain. The second and third decisions are essentially the same as those used for HEVC luma deblocking, which are the on / off decision and the strong filter decision, respectively.
[0152] In the first decision, the boundary strength (bS) is modified for chroma filtering, and the conditions are checked sequentially. If the condition is met, the remaining conditions with lower priority are skipped. When bS is equal to 2, or when bS is equal to 1 when a large block boundary is detected, chroma deblocking is performed. The second and third conditions are essentially the same as the HEVC luma strong filter decision below.
[0153] In the second condition, d is derived in the same way as HEVC luma deblocking. When d is less than β, the second condition will be true. In the third condition, StrongFilterCondition is derived as follows:
[0154] Derived dpq in the same way as HEVC.
[0155] Derived according to HEVC method sp3 = Abs (p3-p0)
[0156] According to the HEVC method, sq3 = Abs(q0-q3)
[0157] According to HEVC design, StrongFilterCondition = (dpq is less than (β>>2), sp3+sq3 is less than (β>>3), and Abs(p0-q0) is less than (5*tC+1)>>1).
[0158] 3.6.6 Strong Deblocking Filter for Chroma
[0159] The following strong deblocking filter for chroma is defined:
[0160] p2′=(3*p3+2*p2+p1+p0+q0+4)>>3
[0161] p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3
[0162] p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3
[0163] The example chroma filter performs deblocking on a 4x4 grid of chroma samples.
[0164] 3.6.7 Position-dependent limiting
[0165] Position-dependent clipping, tcPD, is applied to the output samples of the luminance filtering process involving strong and long filters that modify 7, 5, and 3 samples at the boundaries. Assuming the quantization error distribution, the clipping value can be increased for samples that are expected to have higher quantization noise, so the reconstructed sample values are expected to deviate more from the true sample values.
[0166] For each P or Q boundary filtered with an asymmetric filter, a position-dependent threshold table is selected from two tables provided to the decoder as side information (e.g., Tc7 and Tc3 in the following table) based on the result of the decision process:
[0167] Tc7={6,5,4,3,2,1,1}; Tc3={6,4,2};
[0168] tcPD=(Sp==3)? Tc3:Tc7;
[0169] tcQD=(Sq==3)? Tc3:Tc7;
[0170] For P or Q boundaries filtered with a short symmetric filter, a lower magnitude position-dependent threshold is applied:
[0171] Tc3={3,2,1};
[0172] After defining the thresholds, the filtered p'i and q'i sample values are clipped according to the tcP and tcQ clipping values:
[0173] p”i=Clip3(p’i+tcPi,p’i–tcPi,p’i);
[0174] q”j=Clip3(q’j+tcQj,q’j–tcQj,q’j);
[0175] where p'i and q'j are filtered sample values, p"i and q"j are output sample values after clipping, and tcPi and tcPi are clipping thresholds derived from the VVC tc parameters and tcPD and tcQD. Function Clip3 is the clipping function as specified by VVC.
[0176] 3.6.8 Sub-block Deblocking Adjustment
[0177] To enable parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on one side using sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)), as shown in the luma control of the long filter. By extension, sub-block deblocking is adjusted so that sub-block boundaries close to CU or implicit TU boundaries on the 8×8 grid are restricted to modifying at most two samples on each side.
[0178] The following applies to sub-block boundaries that are not aligned with CU boundaries.
[0179]
[0180] Where edge equals 0 corresponds to a CU boundary, edge equals 2 or equal to orthogonalLength-2 corresponds to a sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TUs is used, implicitTU is true.
[0181] 3.7 Sample Adaptive Compensation
[0182] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after the deblocking filter using an offset specified by the encoder for each CTB. The video encoder first determines whether the SAO process will be applied to the current slice. If SAO is applied to the slice, each CTB is classified into one of five SAO types, as shown in Table 4. The concept of SAO is to classify pixels into categories and reduce distortion by adding an offset to the pixels in each category. SAO operations include edge offset (EO) and band offset (BO), where EO uses edge attributes for pixel classification in SAO types 1 to 4, and BO uses pixel intensity for pixel classification in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag is equal to 1, the current CTB will reuse the SAO type and offset of the CTB to the left. If sao_merge_up_flag is equal to 1, the current CTB will reuse the SAO type and offset of the CTB above it.
[0183]
[0184]
[0185] Table 4. SAO type specifications
[0186] 3.8 Adaptive Loop Filter
[0187] Adaptive loop filtering for video codecs minimizes the mean square error between the original samples and the decoded samples by using a Wiener-based adaptive filter. ALF is located in the last processing stage of each picture and can be regarded as a tool to capture and repair artifacts from previous stages. The appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder through a signal. In order to achieve better codec efficiency, especially for high-resolution video, local adaptation is used for the luminance signal by applying different filters to different areas or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also helps to improve codec efficiency. In terms of syntax, the filter coefficients are sent in a picture-level header called an adaptation parameter set, and the filter on / off flag of the CTU is interleaved at the CTU level in the slice data. This syntax design not only supports picture-level optimization, but also achieves low coding delay.
[0188] 3.8.1 Parameter Signaling
[0189] According to the ALF design in VTM, the filter coefficients and clipping index are carried in the ALF adaptive parameter set (APS). The ALF APS can include up to 8 chroma filters and a luminance filter set with up to 25 filters. Each of the 25 luminance categories also includes an index. Categories with the same index share the same filter. By merging different categories, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficient is represented using a 0th order Exp-Golomb code, followed by a sign bit for non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APSs simultaneously.
[0190] The filter control syntax element of ALF in VTM includes two kinds of information. First, the ALF on / off flag is transmitted by signal at the sequence, picture, slice and CTB level. Chroma ALF can be enabled at the picture and slice level only when luma ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice and CTB level, the filter usage information is transmitted by signal at that level. If all slices within a picture use the same APS, the referenced ALFAPS ID is encoded and decoded at the slice level or picture level. The luma component can reference up to 7 ALF APSs, and the chroma component can reference 1 ALFAPS. For the luma CTB, an index is transmitted by signal to indicate which ALF APS or offline trained luma filter set to use. For the chroma CTB, the index indicates which filter in the referenced APS to use.
[0191] The data syntax elements of the ALF associated with the luma component in the VTM are listed as follows:
[0192]
[0193]
[0194] alf_luma_filter_signal_flag equal to 1 specifies that the luma filter set is signaled. alf_luma_filter_signal_flag equal to 0 specifies that the luma filter set is not signaled. alf_luma_clip_flag equal to 0 specifies that linear adaptive loop filtering is applied to the luma component. alf_luma_clip_flag equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the luma component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of adaptive loop filter categories for which luma coefficients may be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters-1, inclusive. alf_luma_coeff_delta_idx[filtIdx] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter category indicated by filtIdx, in the range of 0 to NumAlfFilters-1. When alf_luma_coeff_delta_idx[filtIdx] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1+1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1, inclusive.
[0195] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] is not present, it is presumed to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range of 0 to 128, inclusive. alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx, as follows:
[0196] If alf_luma_coeff_sign[sfIdx][j] is equal to 0, the corresponding luma filter coefficient is positive.
[0197] Otherwise (alf_luma_coeff_sign[sfIdx][j] is equal to 1), the corresponding luma filter coefficient is negative.
[0198] When alf_luma_coeff_sign[sfIdx][j] is not present, it is inferred to be equal to 0.
[0199] alf_luma_clip_idx[sfIdx][j] specifies the clip index of the clip value to be used before multiplying the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree unit syntax elements for the ALF associated with the luma component in the VTM are listed as follows:
[0200]
[0201] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 1 specifies that an adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 0 specifies that an adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at the luma position (xCtb, yCtb).
[0202] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to 0. alf_use_aps_flag equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag equal to 1 specifies that the filter set from the APS is applied to the luma CTB. When alf_use_aps_flag is not present, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx shall be in the range of 0 to sh_num_alf_aps_ids_luma-1, inclusive. When alf_luma_prev_filter_idx is not present, it is inferred to be equal to 0.
[0203] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] that specifies the filter set index of the luma CTB at position (xCtb, yCtb) is derived as follows:
[0204] If alf_use_aps_flag is equal to 0, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to alf_luma_fixed_filter_idx.
[0205] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to 16+alf_luma_prev_filter_idx.
[0206] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15 (inclusive).
[0207] Based on the ALF design of VTM, the ALF design of ECM further introduces the concept of alternative filter sets into the luma filter. The luma filter trains multiple alternatives / rounds based on the updated luma CTU ALF on / off decision for each alternative / round. In this way, there will be multiple filter sets associated with each trained alternative, and the class merging results for each filter set can be different. Each CTU can select the best filter set through RDO, and the relevant alternative information will be transmitted through the signal. The data syntax elements of the ALF associated with the luma component in ECM are listed as follows:
[0208]
[0209]
[0210] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3, inclusive. alf_luma_clip_flag[altIdx] equal to 0 specifies that linear adaptive loop filtering is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for the alternative luma filter set with index altIdx for which luma coefficients may be signaled. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters-1, inclusive.
[0211] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter class indicated by filtIdx in the range 0 to NumAlfFilters-1, of the alternative luma filter set with index altIdx. When alf_luma_coeff_delta_idx[filtIdx][altIdx] is not present, it is presumed to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[altIdx], inclusive. alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx, for the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] is not present, it is presumed to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range of 0 to 128, inclusive.
[0212] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx of the alternative luma filter set with index altIdx, as follows:
[0213] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luma filter coefficient is positive.
[0214] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luma filter coefficient is negative.
[0215] When alf_luma_coeff_sign[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0.
[0216] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clip index of the clip value to be used before multiplying the j-th coefficient of the signaled luma filter indicated by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree unit syntax elements for the ALF associated with the luma component in the ECM are listed as follows:
[0217]
[0218] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luma filter applied to the codec tree block of the luma component of the codec tree unit at luma position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to 0.
[0219] 3.8.2 Filter Shape
[0220] Figure 10 An example of filter shape for ALF is shown in FIG. In JEM, up to three diamond filter shapes (such as Figure 10 As shown in Figure 2. An index is signaled at the picture level to indicate the filter shape for the luma component. Each square represents a sample, and Ci (i is 0 to 6 (left), 0 to 12 (center), 0 to 20 (right)) represents the coefficient to be applied to the sample. For the chroma components in a picture, a 5×5 diamond is always used. In VVC, a 7×7 diamond is always used for luma, and a 5×5 diamond is always used for chroma.
[0221] 3.8.3 Classification of ALF
[0222] Each 2×2 (or 4×4) block is classified into one of 25 categories. The category index C is derived based on the quantized values of its directionality D and activity A^ as follows:
[0223]
[0224] To calculate D and First, use the 1-D Laplacian operator to calculate the gradient in the horizontal, vertical, and two diagonal directions:
[0225]
[0226] The indices i and j refer to the coordinates of the top left sample in the 2×2 block, and R(i,j) indicates the reconstructed sample at coordinate (i,j). The maximum and minimum values of the gradient D in the horizontal and vertical directions are set to:
[0227]
[0228] And the maximum and minimum values of the gradients in the two diagonal directions are set to:
[0229]
[0230] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0231] Step 1. If and are both true, then D is set to 0.
[0232] Step 2. If Then continue with step 3; otherwise, continue with step 4.
[0233] Step 3. If Then D is set to 2; otherwise D is set to 1.
[0234] Step 4. If Then D is set to 4; otherwise D is set to 3.
[0235] The activity value A is calculated as:
[0236]
[0237] A is further quantized to a range of 0 to 4 (inclusive), and the quantized value is expressed as For the two chroma components in a picture, no classification method is applied, ie, a single set of ALF coefficients is applied for each chroma component.
[0238] 3.8.4 Geometric Transformation of Filter Coefficients
[0239] Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k, l) associated with the coordinates (k, l) according to the gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make the directionality of different blocks to which the ALF is applied more similar by aligning them.
[0240] Three geometric transformations are introduced, including diagonal, vertical flip and rotation:
[0241] Diagonal:fD (k,l)=f(l,k),
[0242] Flip vertically:f V (k,l)=f(k,Kl-1),
[0243] Rotation:f R (k,l)=f(Kl-1,k).
[0244] Where K is the size of the filter, and 0≤k,l≤K-1 are the coefficient coordinates, such that position (0,0) is in the upper left corner and position (K-1,K-1) is in the lower right corner. The transform is applied to the filter coefficients f(k,l) according to the gradient value calculated for the block. The relationship between the transform and the four gradients in the four directions is summarized in Table 5.
[0245] Figure 11 An example of transform coefficients for a relative harmonizer supported by a 5×5 diamond filter is shown. For example, Figure 11 The transform coefficients for each position based on a 5x5 diamond are shown.
[0246] Gradient value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> diagonal <![CDATA[g d1 <g d2 And g h <g v ]]> Flip vertically <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation
[0247] Table 5. Mapping of gradients computed for a block to transformations
[0248] 3.8.5 Filtering process
[0249] At the decoder side, when ALF is enabled for a block, each sample R(i,j) within the block is filtered, resulting in the following sample value R′(i,j), where L represents the filter length and f m,n denotes the filter coefficient, and f(k, l) denotes the decoded filter coefficient.
[0250]
[0251] Figure 12 An example of relative coordinates for a 5×5 diamond filter support is shown assuming that the coordinates (i, j) of the current sample point are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient.
[0252] 3.8.6 Nonlinear Filter Reconstruction (Reformulation)
[0253] Linear filtering can be reconstructed as follows without affecting codec efficiency:
[0254]
[0255] where w(i,j) are the same filter coefficients.
[0256] VVC introduces nonlinearity by using a simple clipping function to reduce the influence of neighboring sample values when the neighboring sample values I(x+i,y+j)) differ too much from the current sample value I(x,y) being filtered, thereby making ALF more efficient. More specifically, the ALF filter is modified as follows:
[0257]
[0258] where K(d,b)=min(b,max(-b,d)) is the clipping function and k(i,j) is the clipping parameter, which depends on the (i,j) filter coefficients. The encoder performs an optimization to find the best k(i,j).
[0259] A clipping parameter k(i,j) is specified for each ALF filter, with one clipping value signaled for each filter coefficient. This means that for the luma filter, a maximum of 12 clipping values can be signaled in the bitstream, and for the chroma filters, a maximum of 6 clipping values can be signaled in the bitstream. To limit signaling cost and encoder complexity, only four fixed values are used that are the same for both inter and intra slices.
[0260] Because the variance of local differences in luma is generally higher than that of local differences in chroma, two different sets of values for the luma and chroma filters are applied. A maximum sample value in each set (here 1024 for a 10-bit bit depth) is also introduced so that clipping can be disabled when not necessary. Four values are chosen by dividing the entire range of sample values for luma (encoded on 10 bits) and the range from 4 to 1024 for chroma roughly equally in the logarithmic domain. More precisely, the luma clipping value table has been obtained by the following formula:
[0261] Where M = 2 10 And N=4
[0262] Similarly, the chroma limit value table is obtained according to the following formula:
[0263] Where M = 2 10 , N=4, A=4
[0264] 3.9 Bilateral Loop Filter
[0265] 3.9.1 Bilateral Image Filter
[0266] Bilateral image filtering is a nonlinear filter that smooths noise while preserving edge structure. Bilateral filtering is a technique that reduces the filter weight not only with the distance between samples but also with the increase of intensity difference. In this way, the over-smoothing of edges can be improved. The weight is defined as
[0267]
[0268] where Δx and Δy are the distances in the vertical and horizontal directions, and ΔI is the intensity difference between the samples.
[0269] The edge-preserving denoising bilateral filter uses low-pass Gaussian filters for both the domain filter and the range filter. The domain low-pass Gaussian filter assigns higher weights to pixels spatially close to the center pixel. The range low-pass Gaussian filter assigns higher weights to pixels similar to the center pixel. Combining the range and domain filters, the bilateral filter at edge pixels becomes a thin Gaussian filter that is oriented along the edge and has a strong gradient reduction. This is why the bilateral filter can smooth noise while preserving edge structure.
[0270] 3.9.2 Bilateral Filters in Video Codecs
[0271] The bilateral filter in video codec is a codec tool used for VVC [2]. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and the SAO act on the same input samples, each filter produces an offset, and these offsets are then added to the input samples to produce output samples, which are passed to the next stage after clipping. The spatial filter strength σ d Determined by the block size, where smaller blocks are filtered more strongly, and the intensity filter strength σ r Determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four most recent samples are used, so the filtered sample strength I F can be calculated as
[0272]
[0273] Among them I C Indicates the intensity of the central sample point, ΔI A =I A -I C Indicates the intensity difference between the center sample point and the sample point above. ΔI B , ΔI L andΔI R Represents the intensity difference between the center sample point and the samples below, to the left, and to the right, respectively.
[0274] 4. Technical problems solved by the disclosed technical solution
[0275] An example design of an adaptive loop filter (ALF) in video codec has the following problems.
[0276] In the example ALF design, only spatially reconstructed samples after other filtering (such as deblocking filter / SAO / BF) are used for filter training and filtering. However, there is other valuable information that can be potentially utilized, such as samples filtered / generated by one or more predefined filters.
[0277] In the example ALF design, only spatially reconstructed samples after other filtering (such as deblocking filter / SAO / BF) are used for filter training and filtering. However, there is other valuable information that can potentially be utilized, such as samples before deblocking filter (DBF), SAO or other stages.
[0278] In the example ALF design, only spatially reconstructed samples after other filtering (such as deblocking filter / SAO / BF) are used for filter training and filtering. However, there are other valuable information that can be potentially utilized, such as prediction / residual samples and coded reference pictures.
[0279] 5. List of solutions and implementation examples
[0280] In order to solve the above problems, the following methods are disclosed. The embodiments should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.
[0281] It should be noted that the disclosed method can be used as a loop filter or as post-processing.
[0282] In this disclosure, a video unit may refer to a sequence, a picture, a sub-picture, a slice, a CTU, a block and / or a region.A video unit may include one color component or multiple color components.
[0283] In this disclosure, an ALF processing unit may refer to a sequence, a picture, a sub-picture, a slice, a CTU, a block, a region, or a sample. An ALF processing unit may include one color component or multiple color components.
[0284] In the present disclosure, an extended tap may refer to a filter tap that is not used by the ALF in VVC. For example, the extended tap may be a filter input that includes information other than the sample values of spatial neighbors (e.g., spatially adjacent) to the current sample being filtered.
[0285] Example 1
[0286] In one example, at least one extended tap is used in the ALF filter to further improve the efficiency of the ALF.
[0287] Example 2
[0288] In one example, at least one extended tap may be different from a spatial tap in an ALF filter that utilizes only information of spatially neighboring samples of the target component (neighboring a center sample to be filtered).
[0289] In one example, the spatially neighboring samples may come from reconstruction after DBF / SAO / BF. In one example, the spatial taps may use only the spatially neighboring luma samples to filter the center luma sample within an ALF filter. In one example, the spatial taps may use only the spatially neighboring chroma samples to filter the center chroma sample within an ALF filter.
[0290] Example 3
[0291] In one example, at least one extension tap and at least one spatial domain tap may coexist within one ALF filter.
[0292] In one example, the ALF filter may include both spatial domain taps and extension taps. In one example, the ALF filter may include M (eg, M>0) spatial domain taps and N (eg, N>0) extension taps.
[0293] Example 4
[0294] In one example, one or more expansion taps of the ALF filter may use one or more input sources.
[0295] In one example, one or more expansion taps of the ALF filter may use one input source.
[0296] For example, the input source may be based on one of the following: reconstruction before DBF, intermediate filtering results of a predefined filter, reconstruction before SAO / BF, prediction samples, residual samples, co-located samples, any modification of the above sources, and / or any combination or fusion (e.g., weighted sum) of at least two sources.
[0297] In one example, one or more expansion taps of the ALF filter may use multiple input sources. For example, the multiple input sources may include more than one of the following: reconstruction before DBF, intermediate filtering results of a predefined filter, reconstruction before SAO / BF, prediction samples, residual samples, co-located samples, and / or any modifications of the above sources.
[0298] The input source of the extended tap may also be used by other taps in the ALF. The input source of the extended tap may not be used by other taps in the ALF. The input source of the extended tap may be filtered prediction samples. The input source of the extended tap may be filtered residual samples. The input source of the extended tap may be a weighted sum or filtering of multiple sources. The input source of the extended tap may be the output of a function of multiple sources. The input source of the extended tap may be derived based on samples in different color components.
[0299] Example 5
[0300] In one example, whether and / or how to apply a filter having at least one extended tap may be performed in different manners for different color formats and / or different color components.
[0301] In one example, an ALF filter having at least one extended tap may be applied only to process the luma component. In one example, an ALF filter having at least one extended tap may be applied only to process one of the chroma components (e.g., a Cb component or a Cr component). In one example, an ALF filter having at least one extended tap may be applied to process all chroma components. (e.g., a Cb component and a Cr component). In one example, an ALF filter having at least one extended tap may be applied to filter both luma and chroma components. (e.g., a Y component, a Cb component, and a Cr component).
[0302] Example 6
[0303] In one example, the coefficients of the expansion taps inside the ALF filter may correspond to one or more input samples.
[0304] In one example, the coefficients of the expansion taps inside the ALF filter may correspond to only one input sample.
[0305] In one example, the coefficients of the expansion taps within the ALF filter may correspond to N input samples (e.g., N=2). In one example, the N input samples may be designed in a symmetrical manner. In one example, the N input samples may be designed in an asymmetrical manner. In one example, the coefficients of the expansion taps within the ALF filter may be shared by multiple inputs based on geometric information.
[0306] In one example, the plurality of input samples may be located at symmetrical positions.
[0307] Example 7
[0308] In one example, an ALF filter with at least one extended tap may use a shape or size that is different from the shape or size used by an ALF without extended taps (such as the ALF in VVC).
[0309] In one example, the ALF filter may include different shapes for spatial domain taps and extension taps.
[0310] In one example, within the ALF filter, the shape / size used for one or more spatial taps may be different from the shape / size used for one or more expansion taps.
[0311] In one example, inside the ALF filter, the shape / size used for one or more spatial taps may be the same as the shape / size used for one or more expansion taps.
[0312] In one example, within an ALF filter having at least one extended tap, the filter shape used for one or more spatial domain taps may be as described below. In one example, the filter shape used for the spatial domain tap may be a rhombus. In one example, the filter shape used for the spatial domain tap may be a square. In one example, the filter shape used for the spatial domain tap may be a cross. In one example, the filter shape used for the spatial domain tap may be a symmetrical shape. In one example, the filter shape used for the spatial domain tap may be an asymmetrical shape. In one example, the filter shape used for the spatial domain tap may be a shape of other designs. In one example, the filter shape used for the spatial domain tap may be determined in real time / transmitted by signal / derived.
[0313] In one example, within an ALF filter having at least one extended tap, the filter shape used for one or more extended taps may be as described below. In one example, the filter shape used for the extended tap may be a rhombus. In one example, the filter shape used for the extended tap may be a square. In one example, the filter shape used for the extended tap may be a cross. In one example, the filter shape used for the extended tap may be a symmetrical shape. In one example, the filter shape used for the extended tap may be an asymmetrical shape. In one example, the filter shape used for the extended tap may be a shape of other designs. In one example, the filter shape used for the extended tap may be determined in real time / transmitted by a signal / derived.
[0314] In one example, an ALF filter including at least one spatial tap and at least one extension tap may be designed as follows.
[0315] In one example, one or more spatial taps can be used for filtering (the spatial taps can be considered as using the reconstruction before ALF as input). In one example, the shape for the spatial taps can be designed as Figure 13 . Figure 13An example shape for at least one spatial tap is shown. In one example, the shape for the spatial tap can be designed as Figure 14 . Figure 14 Example shapes for at least one spatial tap are shown.
[0316] Example 8
[0317] In one example, there may be one or more input sources to be used for one or more expansion taps.
[0318] In one example, the following input sources may be used jointly within one filter shape containing one or more expansion taps: pre-ALF reconstruction; pre-DBF reconstruction; filter output trained offline by feeding pre-ALF reconstruction; and coded reference pictures.
[0319] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed by the reconstruction before ALF; and output of a predefined filter (e.g., a Gaussian filter) fed by the reconstruction before DBF.
[0320] In one example, the following input sources can be used jointly within one filter shape containing one or more extended taps: pre-ALF reconstruction; pre-DBF reconstruction; filter output trained offline by feeding pre-ALF reconstruction; coded reference pictures; and prediction or residual samples.
[0321] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed by the reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by the reconstruction before DBF; and prediction or residual samples.
[0322] In one example, the following input sources can be used jointly within one filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed with reconstruction before ALF; coded reference pictures; and output of a predefined filter (e.g., a Gaussian filter) fed with prediction or residual samples.
[0323] In one example, the following input sources can be used jointly within one filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed by the reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by the reconstruction before DBF; and output of a predefined filter (e.g., a Gaussian filter) fed by prediction or residual samples.
[0324] In one example, the following input sources can be used jointly within one filter shape containing one or more extended taps: reconstruction before ALF; reconstruction before DBF; filter output trained offline by feeding reconstruction before ALF; coded reference pictures; and filter output trained offline by feeding prediction or residual samples.
[0325] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of a filter trained offline by feeding reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by reconstruction before DBF; and output of a filter trained offline by feeding prediction or residual samples.
[0326] In one example, the following input sources can be used jointly within one filter shape containing one or more extended taps: reconstruction before ALF; reconstruction before DBF; output of a filter trained offline by feeding reconstruction before ALF; coded reference pictures; output of a predefined filter (e.g., a Gaussian filter) fed with prediction or residual samples; and output of a filter trained offline by feeding prediction or residual samples.
[0327] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of a filter trained offline by feeding reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by reconstruction before DBF; output of a predefined filter (e.g., a Gaussian filter) fed by prediction or residual samples; output of a filter trained offline by feeding prediction or residual samples.
[0328] In one example, the following input sources can be used jointly within one filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed with reconstruction before ALF; coded reference pictures; output of a predefined filter (e.g., a Gaussian filter) fed with prediction or residual samples; and prediction or residual samples.
[0329] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of an offline trained filter fed by the reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by the reconstruction before DBF; output of a predefined filter (e.g., a Gaussian filter) fed by prediction or residual samples; and prediction or residual samples.
[0330] In one example, the following input sources can be used jointly within a filter shape containing one or more extended taps: reconstruction before ALF; reconstruction before DBF; filter output trained offline by feeding reconstruction before ALF; coded reference pictures; filter output trained offline by feeding prediction or residual samples; and prediction or residual samples.
[0331] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; filter output trained offline by feeding reconstruction before ALF; predefined filter (e.g., Gaussian filter) output fed by reconstruction before DBF; filter output trained offline by feeding prediction or residual samples; and prediction or residual samples.
[0332] In one example, the following input sources can be used jointly within one filter shape containing one or more extended taps: reconstruction before ALF; reconstruction before DBF; filter output trained offline by feeding reconstruction before ALF; coded reference pictures; filter output trained offline by feeding prediction or residual samples; predefined filter (e.g., Gaussian filter) output by feeding prediction or residual samples; and prediction or residual samples.
[0333] In one example, the following input sources can be used jointly within a filter shape containing one or more expansion taps: reconstruction before ALF; reconstruction before DBF; output of a filter trained offline by feeding reconstruction before ALF; output of a predefined filter (e.g., a Gaussian filter) fed by reconstruction before DBF; output of a filter trained offline by feeding prediction or residual samples; output of a predefined filter (e.g., a Gaussian filter) fed by prediction or residual samples; and prediction or residual samples.
[0334] In one example, the number of expansion taps belonging to different input sources may be different.
[0335] In one example, the shapes of expansion taps belonging to different input sources may be different.
[0336] In one example, a single tap shape may be applied.
[0337] In one example, a diamond shape may be applied.
[0338] In one example, a square shape may be applied.
[0339] In one example, a cross shape may be applied.
[0340] In one example, a symmetrical shape may be applied.
[0341] In one example, an asymmetrical shape may be applied.
[0342] In one example, the indicator for the input source may be signaled / predefined / derived on the fly in the APS.
[0343] In one example, the total number of extension taps inside the ALF filter can be jointly derived based on the shape, filter length, and symmetry constraints.
[0344] Example 9
[0345] In one example, a first syntax element may be signaled to indicate whether a filter having at least one extended tap is enabled.
[0346] In one example, the first syntax element may be encoded or decoded using arithmetic coding. In one example, the first syntax element may be encoded or decoded using at least one context. The context may depend on coding information of the current block or a neighboring block. The context may depend on a filter shape of at least one neighboring block. In one example, the first syntax element may be encoded or decoded using bypass coding. In one example, the first syntax element may be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential Golomb code, a truncated exponential Golomb code, or the like.
[0347] In one example, the first syntax element may be conditionally signaled. For example, the first syntax element may be signaled only if the extension tap is available.
[0348] The first syntax element may be coded or decoded in a predictive manner. The first syntax element may be predicted by an on / off decision of an extension tap of at least one neighboring block.
[0349] In one example, the first syntax element may be signaled independently for different color components. In one example, the first syntax element may be signaled and shared for different color components. In one example, the first syntax element may be signaled for a first color component but not for a second color component.
[0350] The syntax elements may be signaled in SPS / PPS / picture header / slice header / APS / CTU / CU / etc.
[0351] Example 10
[0352] In one example, a first syntax element may be signaled to indicate which / what input source is used for the expansion taps inside the ALF filter.
[0353] In one example, the first syntax element may be encoded using arithmetic coding. In one example, the first syntax element may be encoded using at least one context. The context may depend on codec information of the current block or a neighboring block. The context may depend on a filter shape of at least one neighboring block. In one example, the first syntax element may be encoded using bypass coding.
[0354] In one example, the first syntax element may be binarized using a unary code, a truncated unary code, a fixed-length code, an exponential Golomb code, a truncated exponential Golomb code, or the like.
[0355] In one example, the first syntax element may be conditionally signaled. For example, the first syntax element may be signaled only if the extension tap is available.
[0356] The first syntax element may be coded or decoded in a predictive manner. The first syntax element may be predicted by an on / off decision of an extension tap of at least one neighboring block.
[0357] The first syntax element may be signaled independently for different color components. In one example, the first syntax element may be signaled and shared for different color components. In one example, the first syntax element may be signaled for the first color component but not for the second color component.
[0358] The syntax elements may be signaled in an SPS / PPS / picture header / slice header / APS / CTU / CU / etc. In one example, the syntax elements may be signaled in an APS for a signaled ALF filter.
[0359] Example 11
[0360] The coefficients of at least one extended tap inside the ALF filter can be signaled in a syntax element structure (such as an APS). In one example, the coefficients of the extended taps can be included in the APS. In one example, the clipping parameters of the extended taps can be included in the APS. In one example, the category merging results of the extended taps can be included in the APS. In one example, the coefficients of the extended taps can be coded in a predictive manner. In one example, the coefficients of the extended taps can be coded using arithmetic coding with at least one context. In one example, the coefficients of the extended taps can be coded using bypass coding. In one example, the coefficients of the extended taps can be jointly coded with the coefficients of the spatial domain taps. In one example, other parameters of the extended taps can be included in the APS. In one example, the coefficients of the extended taps can be predefined fixed values.
[0361] Example 12
[0362] In one example, an intermediate filtering result of at least one fixed or adaptive filter is used as input to one or more expansion taps.
[0363] Example 13
[0364] In one example, an intermediate filter result of an offline trained filter may be used as an input to one or more expansion taps. In one example, the intermediate filter result may be generated by a pre-ALF reconstruction and an offline trained filter of the ALF. In one example, the intermediate filter result may be generated by a pre-DBF reconstruction and an offline trained filter of the ALF.
[0365] Example 14
[0366] In one example, intermediate filtering results of an online trained ALF filter may be used as input to one or more expansion taps.
[0367] Example 15
[0368] In one example, the intermediate filtering results of other predefined filters can be used as the input of one or more expansion taps. In one example, a Gaussian filter can be applied. In one example, a bilateral filter can be applied. In one example, a guided filter can be applied. In one example, a median filter can be applied. In one example, a local median filter can be applied. In one example, a non-local median filter can be applied. In one example, a filter with a low-pass property can be applied. In one example, a filter with a high-pass property can be applied.
[0369] Example 16
[0370] In one example, intermediate filtering results of other online trained filters may be used as input to one or more expansion taps.
[0371] Example 17
[0372] In one example, the input for intermediate filtering can be reconstructed samples at different encoding and decoding stages. In one example, the reconstruction before / after the adaptive loop filter (ALF) of the current frame / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the sample adaptive offset (SAO) / cross-component SAO (CCSAO) of the current frame / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the bilateral filter (BIF) of the current frame / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the deblocking filter (DBF) of the current frame / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after any other stage of the current frame / reference frame can be used to generate the intermediate filtering result.
[0373] Example 18
[0374] In one example, reconstructed samples before / after different encoding / decoding stages of the current frame are used as inputs to one or more expansion taps.
[0375] In one example, reconstructions before / after the DBF of the current frame may be used as input to one or more extension taps. In one example, reconstructions before / after the SAO / CCSAO of the current frame may be used as input to one or more extension taps. In one example, reconstructions before / after the BIF of the current frame may be used as input to one or more extension taps. In one example, reconstructions before / after other stages of the current frame may be used as input to one or more extension taps.
[0376] Example 19
[0377] In one example, the prediction / residual samples of the current frame are used as input to one or more extension taps. In one example, the prediction samples of the current frame can be used as input to one or more extension taps. In one example, the residual samples of the current frame can be used as input to one or more extension taps.
[0378] Example 20
[0379] In one example, samples within one or more coded pictures are used as input sources for one or more expansion taps.
[0380] Example 21
[0381] In one example, the previously coded frame may be a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with the block / current slice / frame. In one example, the previously coded frame may be a short-term reference picture for the block / current slice / frame. In one example, the previously coded frame may be a long-term reference picture for the block / current slice / frame.
[0382] Example 22
[0383] In one example, the previously decoded frame may not be a reference frame, but is stored in a decoded picture buffer (DPB).
[0384] Example 23
[0385] In one example, at least one indicator is signaled to indicate which previously coded frame(s) to use. In one example, one indicator is signaled to indicate which reference picture list to use. In one example, at least one indicator may be signaled to indicate a reference index. In one example, the indicator may be signaled conditionally, for example, depending on how many reference pictures are included in the RPL / RPS. In one example, the indicator may be signaled conditionally, for example, depending on how many previously decoded pictures are included in the DPB.
[0386] Example 24
[0387] In one example, which frame(s) to utilize is determined on the fly. In one example, the extension tap may obtain information from one / more previously coded frames in the DPB. In one example, the extension tap may obtain information from one / more reference frames in list 0. In one example, the extension tap may obtain information from one / more reference frames in list 1. In one example, the extension tap may obtain information from reference frames in both list 0 and list 1. In one example, the extension tap may obtain information from the reference frame that is closest (e.g., has the smallest picture order count (POC) distance to the current slice / frame) to the current frame.
[0388] In one example, the extended tap may obtain information from a reference frame in a reference list with a reference index equal to K (e.g., K=0). In one example, K may be predefined. In one example, K may be derived on the fly based on reference picture information. In one example, K may be transmitted via a signal.
[0389] In one example, the expansion taps may obtain information from the co-located frame.
[0390] In one example, which frame(s) are utilized may be determined by decoding information. In one example, which frame(s) are utilized may be defined as the most frequently used top N (e.g., N=1) reference pictures for samples within the current slice / frame. In one example, which frame(s) are utilized may be defined as the most frequently used top N (e.g., N=1) reference pictures for each reference picture list (if available) for samples within the current slice / frame. In one example, which frame(s) are utilized may be defined as the picture with the smallest top N (e.g., N=1) POC distance / absolute POC distance relative to the current picture.
[0391] Example 25
[0392] In one example, whether to obtain information from a previously coded frame may depend on decoded information (e.g., codec mode / statistics / characteristics) of at least one region of the block to be filtered. In one example, the encoder may signal to the decoder whether to obtain information from a previously coded frame of at least one region of the block to be filtered.
[0393] In one example, whether to obtain information from a previously coded frame may depend on the slice / picture type. In one example, this applies only to inter-coded slices / pictures (e.g., P or B slices / pictures). In one example, whether to obtain information from a previously coded frame may depend on the availability of reference pictures.
[0394] In one example, whether to obtain information from a previously coded frame may depend on reference picture information or picture information in the DPB. In one example, if the minimum POC distance (e.g., the minimum POC distance between a reference picture / picture in the DPB and the current picture) is greater than a threshold, it is disabled.
[0395] In one example, whether to obtain information from a previously coded frame may depend on the temporal layer index and / or QP and / or the dimension of the picture. In one example, it applies to blocks with a given temporal layer index (e.g., the highest temporal layer).
[0396] In one example, if the block to be filtered contains some samples that were coded in non-inter mode, the extended tap may not use information from a previously coded frame to filter the block. In one example, non-inter mode may be defined as intra mode. In one example, non-inter mode may be defined as a set of codec modes, including but not limited to intra / intra block copy (IBC) / palette mode.
[0397] In one example, the distortion between the current block and the matching block is calculated and used to determine whether to obtain information from a previously encoded frame to filter the current block. In one example, the distortion between the co-located block in the previously encoded frame and the current block can be used to determine whether to obtain information from the previously encoded frame to filter the current block. In one example, motion estimation can first be used to find a matching block from at least one previously encoded frame. In one example, when the distortion is greater than a predefined threshold, information from the previously encoded frame can be disregarded.
[0398] Example 26
[0399] In one example, the information includes two reference blocks and / or co-located blocks for the current block, one of which is from the first reference frame in list 0 and the other is from the first reference frame in list 1 .
[0400] Example 27
[0401] In one example, the coded reference pictures may be accessed during different coding stages. In one example, the coded reference pictures may be accessed during the prediction loop stage.
[0402] In one example, the coded reference pictures can be accessed during the loop filter stage. In one example, the coded reference pictures can be accessed before DBF / SAO / CCSAO / BF / Chroma BF / ALF / Cross-Component ALF (CCALF) or any other loop filtering process. In one example, the coded reference pictures can be accessed during DBF / SAO / CCSAO / BF / Chroma BF / ALF / CCALF or any other loop filtering process. In one example, the coded reference pictures can be accessed after DBF / SAO / CCSAO / BF / Chroma BF / ALF / CCALF or any other loop filtering process.
[0403] In one example, the coded reference pictures may be accessed after the loop filter stage. In one example, the coded reference pictures may be accessed after the loop filter stage by a motion compensation based padding method.
[0404] Example 28
[0405] In one example, the ALF filter may include more than one sub-filter. For example, one sub-filter may be based on samples in a reference picture. For example, one sub-filter may be based on samples in a current picture. For example, one sub-filter may be based on samples that are adjacent (e.g., adjacent and / or non-adjacent) to the current sample. For example, one sub-filter may be based on samples that are co-located with the current sample (e.g., in the current picture and / or in the reference picture).
[0406] For example, a sub-filter may be based on at least one intermediate filtered sample generated from "another filter." For example, the intermediate filtered samples may be generated by filtering a group of samples that are adjacent to and / or co-located with the current sample. For example, the "another filter" may be predefined. For example, the "another filter" may be trained offline. For example, the "another filter" may be transmitted online via a signal. For example, the intermediate filtered samples may have the same position as the current sample. For example, the intermediate filtered samples may be adjacent to (e.g., adjacent, non-adjacent, or co-located in time) the current sample.
[0407] For example, a sub-filter may be based on reconstructed samples. For example, a sub-filter may be based on predicted samples. For example, a sub-filter may be based on residual samples. For example, a sub-filter may be based on samples before deblocking. For example, a sub-filter may be based on samples after deblocking. For example, a sub-filter may be based on samples before DBF / SAO / CCSAO / BF. For example, a sub-filter may be based on samples after DBF / SAO / CCSAO / BF.
[0408] For example, the ALF filter may be constructed from one or more sub-filters from the following: a sub-filter calculated based on spatially neighboring reconstructed samples after SAO, a sub-filter calculated based on spatially neighboring reconstructed samples before deblocking, a sub-filter calculated based on spatially neighboring reconstructed samples after deblocking, a sub-filter calculated based on spatially neighboring prediction samples, a sub-filter calculated based on spatially neighboring residual samples, a sub-filter calculated based on temporal samples in a reference picture, a sub-filter calculated based on one or more median filtered samples, and / or a combination of sub-filters thereof.
[0409] For example, more than one sub-filter may be applied in cascade (eg, without selectivity determination). For example, all sub-filters may be applied in cascade (eg, without selectivity determination).
[0410] For example, at least one sub-filter may be used conditionally / adaptively. For example, it may be conditionally controlled by a syntax element. For example, it may be conditionally controlled by a predefined rule.
[0411] Example 29
[0412] In one example, the disclosed method can be used in post-processing and / or pre-processing.
[0413] Example 30
[0414] The above sources can be located in the same color component as the samples to be filtered, or in different color components. The source can be a reconstruction before ALF; a reconstruction before DBF; the output of a filter trained offline by feeding the reconstruction before ALF; a coded reference picture; the output of a filter trained offline by feeding prediction or residual samples; or prediction or residual samples.
[0415] Example 31
[0416] The above sources can be located at the same position as the samples to be filtered, or within the range around the samples to be filtered. The sources can be: reconstruction before ALF; reconstruction before DBF; filter output trained offline by feeding reconstruction before ALF; coded reference pictures; filter output trained offline by feeding prediction or residual samples; or prediction or residual samples.
[0417] Example 32
[0418] In one example, the above methods may be used in combination.
[0419] Example 33
[0420] In one example, the above method can be used alone.
[0421] Example 34
[0422] In one example, the described extended tap method for ALF can be applied to any loop filtering tool, pre-processing, or post-processing filtering method in a video codec (including but not limited to ALF / CCALF or any other filtering method).
[0423] Example 35
[0424] In one example, the extended tap method can be applied to the loop filtering method. In one example, the extended tap method can be applied to ALF. In one example, the extended tap method can be applied to CCALF. Alternatively, the extended tap method can be applied to other loop filtering methods.
[0425] Example 36
[0426] In one example, the expanded tap method may be applied to a pre-processing filtering method. In one example, the expanded tap method may be applied to a post-processing filtering method.
[0427] Example 37
[0428] In the above examples, a video unit may refer to a sequence / picture / sub-picture / slice / slice / codec tree unit (CTU) / CTU row / CTU group / codec unit (CU) / prediction unit (PU) / transform unit (TU) / codec tree block (CTB) / codec block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luma or chroma sample / pixel.
[0429] Example 38
[0430] Whether and / or how the above disclosed method is applied may be signaled in the bitstream.
[0431] In one example, whether and / or how to apply the method disclosed above can be signaled at the sequence level / picture group level / picture level / slice level / slice group level, such as in a sequence header, a picture header, an SPS, a VPS, a DPS, decoding capability information (DCI), a PPS, an APS, a slice header, and a slice group header.
[0432] In one example, whether and / or how to apply the methods disclosed above may be signaled at a PB, TB, CB, PU, TU, CU, virtual pipeline data unit (VPDU), CTU, CTU row, slice, slice, sub-picture, or other types of regions containing more than one sample or pixel.
[0433] Example 39
[0434] Whether and / or how to apply the above disclosed method may depend on coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type.
[0435] 6. References
[0436] [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson, and R. Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in Proceedings of the IEEE Picture Coding Symposium (PCS), November 2019.
[0437] Figure 154 is a block diagram illustrating an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.
[0438] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this document. Codec component 4004 can reduce the average bit rate of the video from the input 4002 of codec component 4004 to the output to generate a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or transmitted via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video received at input 4002 or the communication transmission can be used by component 4008 to generate pixel values or transmit to a displayable video of display interface 4010. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it is understood that the codec tools or operations are used by the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0439] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.
[0440] Figure 164 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this document. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, for example, a graphics coprocessor.
[0441] Figure 17 4 is a flow chart of an example method 4200 for video processing. Method 4200 includes, at step 4202, determining to employ an adaptive loop filter (ALF) with extended taps. The ALF with extended taps receives data from a plurality of input sources. At step 4204, conversion is performed between visual media data and a bitstream based on the ALF. According to an example, the conversion at step 4204 may include encoding at an encoder or decoding at a decoder.
[0442] It should be noted that method 4200 can be implemented in an apparatus for processing video data that includes a processor and non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs method 4200.
[0443] Figure 18 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where destination device 4320 can be referred to as a video decoding device.
[0444] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of the picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to target device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0445] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or may be external to target device 4320, which may be configured to interface with an external display device.
[0446] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or further standards.
[0447] Figure 19 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 18 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0448] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414.
[0449] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.
[0450] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but are represented separately in the example of the video encoder 4400 for purposes of explanation.
[0451] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0452] The mode selection unit 4403 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0453] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 4413 (other than the picture associated with the current video block).
[0454] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0455] In some examples, motion estimation unit 4404 may perform unidirectional prediction on the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0456] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block. The motion estimation unit 4404 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 4404 may then generate a reference index and a motion vector, where the reference index indicates the reference pictures in list 0 and list 1 that contain the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0457] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal motion information for the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0458] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.
[0459] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0460] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0461] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0462] The residual generation unit 4407 can generate residual data for the current video block by subtracting the (one or more) prediction video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0463] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform a subtraction operation.
[0464] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.
[0465] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0466] Inverse quantization unit 4410 and inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in buffer 4413.
[0467] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0468] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0469] Figure 20 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 18 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0470] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the inverse of the encoding process described for video encoder 4400.
[0471] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.
[0472] The motion compensation unit 4502 may generate a motion compensated block, and interpolation may be performed based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0473] The motion compensation unit 4502 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information, and the motion compensation unit 4502 may use the interpolation filters to generate a prediction block.
[0474] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0475] The intra prediction unit 4503 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0476] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. As desired, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0477] Figure 21 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples, by adding an offset and by applying a finite impulse response (FIR) filter, respectively, taking advantage of the encoded side information of the offset and filter coefficients transmitted through the signal. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts created by previous stages.
[0478] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed to an entropy coding component 4618. The entropy coding component 4618 performs entropy coding and decoding on the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these images are stored in the reference picture cache 4612 .
[0479] Figure 22 4700 is a flow chart of an example method 4700 for video processing. The method 4700 includes determining, at step 4702, to employ an ALF having at least one extension tap and at least one spatial tap. At step 4704, conversion is performed between visual media data and a bitstream based on the ALF. The conversion at step 4704 may include encoding at an encoder or decoding at a decoder, depending on the example.
[0480] It should be noted that method 4700 can be implemented in an apparatus for processing video data that includes a processor and non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4700. Furthermore, method 4700 can be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs method 4700.
[0481] A list of some example preferred solutions is provided below.
[0482] The following solutions illustrate examples of the techniques discussed herein.
[0483] 1. A method for processing video data, comprising: determining to use an adaptive loop filter (ALF) with extended taps, wherein the ALF with extended taps receives data from a plurality of input sources; and performing conversion between visual media data and a bitstream based on the ALF.
[0484] 2 . The method according to claim 1 , wherein the input sources are used jointly inside the shape of the ALF including the expansion taps.
[0485] 3. The method according to any one of claims 1-2, wherein the input source comprises a reconstruction before ALF, a reconstruction before a deblocking filter (DBF), a filter output trained offline by feeding the reconstruction before ALF, a coded reference picture, or a combination thereof.
[0486] 4. The method according to any one of claims 1-3, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding the pre-ALF reconstruction, a predefined filter output by feeding the pre-DBF reconstruction, or a combination thereof.
[0487] 5. The method according to any one of claims 1 to 4, wherein the input source comprises pre-ALF reconstruction, pre-DBF reconstruction, filter outputs trained offline by feeding pre-ALF reconstruction, coded reference pictures, prediction or residual samples, or a combination thereof.
[0488] 6. The method according to any one of claims 1 to 5, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output fed by the pre-ALF reconstruction, a predefined filter output fed by the pre-DBF reconstruction, prediction or residual samples, or a combination thereof.
[0489] 7. The method according to any one of claims 1 to 6, wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, an offline trained filter output by feeding a reconstruction before ALF, a coded reference picture, a predefined filter output by feeding prediction or residual samples, or a combination thereof.
[0490] 8. The method according to any one of claims 1 to 7, wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, an offline trained filter output by feeding the reconstruction before ALF, a predefined filter output by feeding the reconstruction before DBF, and a predefined filter output by feeding prediction or residual samples, or a combination thereof.
[0491] 9. The method according to any one of claims 1 to 8, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, an offline trained filter output by feeding prediction or residual samples, or a combination thereof.
[0492] 10. The method according to any one of claims 1 to 9, wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, an offline trained filter output by feeding the reconstruction before ALF, a predefined filter output by feeding the reconstruction before DBF, and an offline trained filter output by feeding prediction or residual samples, or a combination thereof.
[0493] 11. The method according to any one of claims 1 to 10, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, a predefined filter output by feeding a prediction or residual sample, an offline trained filter output by feeding a prediction or residual sample, or a combination thereof.
[0494] 12. The method according to any one of claims 1 to 11, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, a predefined filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0495] 13. The method according to any one of claims 1 to 12, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding the pre-ALF reconstruction, a predefined filter output by feeding the pre-DBF reconstruction, a predefined filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0496] 14. The method according to any one of claims 1 to 13, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, an offline trained filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0497] 15. The method according to any one of claims 1 to 14, wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, a filter output trained offline by feeding the reconstruction before ALF, a predefined filter output trained offline by feeding the reconstruction before DBF; a filter output trained offline by feeding prediction or residual samples, and prediction or residual samples, or a combination thereof.
[0498] 16. The method according to any one of claims 1 to 15, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, an offline trained filter output by feeding prediction or residual samples, a predefined filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0499] 17. The method according to any one of claims 1 to 16, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding the pre-ALF reconstruction, a predefined filter output by feeding the pre-DBF reconstruction, an offline trained filter output by feeding prediction or residual samples, a predefined filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0500] 18. The method according to any one of claims 1 to 17, wherein the number of expansion taps associated with different input sources is different.
[0501] 19. The method of any one of claims 1-18, wherein the shapes of the expansion taps associated with different input sources are different, and wherein the shapes include a single tap, a diamond, a square, a cross, a symmetrical shape, or an asymmetrical shape.
[0502] 20. The method of any one of claims 1-19, wherein the indicator for the input source is signaled, defined or derived.
[0503] 21. The method according to any one of claims 1-20, wherein the total number of extension taps inside an adaptive loop filter (ALF) filter is jointly derived based on shape, filter length and symmetry constraints.
[0504] 22. The method according to any one of claims 1 to 21, wherein the input source is in the same color component as the samples to be filtered or in a different color component than the samples to be filtered.
[0505] 23. The method according to any one of claims 1 to 22, wherein the input source comprises a pre-ALF reconstruction, a pre-DBF reconstruction, an offline trained filter output by feeding a pre-ALF reconstruction, a coded reference picture, an offline trained filter output by feeding prediction or residual samples, prediction or residual samples, or a combination thereof.
[0506] 24. The method according to any one of claims 1 to 23, wherein the input source is located at the same position as the sample point to be filtered or within a range of positions around the position of the sample point to be filtered.
[0507] 25. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-24.
[0508] 26. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method according to any one of claims 1-24.
[0509] 27. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to use an adaptive loop filter (ALF) with extended taps, wherein the ALF with extended taps receives data from multiple input sources; and generating the bitstream based on the determination.
[0510] 28. A method for storing a bitstream of a video, comprising: determining to employ an adaptive loop filter (ALF) with extended taps, wherein the ALF with extended taps receives data from a plurality of input sources; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0511] 29. A method, apparatus or system as described in this document.
[0512] The following solutions illustrate further examples of the techniques discussed herein.
[0513] 1. A method for processing video data, comprising: determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; and performing conversion between visual media data and a bitstream based on the ALF.
[0514] 2. The method of solution 1, wherein the spatial taps receive values of reconstructed spatial neighboring samples before or after applying a deblocking filter (DBF), a sample adaptive offset (SAO) filter, a bilateral filter (BF), a cross-component sample adaptive offset (CCSAO) filter, or a combination thereof.
[0515] 3. The method according to any one of solutions 1-2, wherein the spatial taps only receive neighboring luma samples for filtering a central luma sample inside the ALF.
[0516] 4. The method according to any one of solutions 1-3, wherein the spatial taps only receive neighboring chroma samples for filtering the center chroma sample inside the ALF.
[0517] 5. The method according to any one of solutions 1-4, wherein the ALF comprises M spatial taps and N extension taps, where N and M are integer values greater than zero.
[0518] 6. A method according to any one of solutions 1-5, wherein the expansion tap receives input from one or more input sources, wherein the one or more input sources include reconstructed samples before applying DBF, intermediate results of a predefined filter, residual samples, filtered residual samples, prediction samples, filtered prediction samples, or a combination thereof.
[0519] 7. The method according to any one of solutions 1-6, wherein the ALF filter with extended taps is applied only to process the luma component or only to process one chroma component.
[0520] 8. The method according to any of solutions 1-7, wherein the extended tap in the ALF corresponds to a single input sample point.
[0521] 9. The method according to any one of solutions 1-8, wherein the extended taps in the ALF correspond to N input samples, where N is equal to 2, and wherein positions of the N input samples are symmetrical.
[0522] 10. The method according to any one of solutions 1-9, wherein the ALF filter adopts a first shape for the spatial taps and a second shape for the extended taps, wherein the first shape and the second shape are different.
[0523] 11. The method according to any one of solutions 1-10, wherein the ALF filter uses a first size for the spatial taps and a second size for the extended taps, wherein the first size and the second size are different.
[0524] 12. The method according to any one of solutions 1 to 11, wherein the spatial tap is in a diamond shape or a cross shape, and wherein the extended tap is in a diamond shape or a cross shape.
[0525] 13. The method according to any one of solutions 1 to 12, wherein the spatial domain tap is in a diamond shape, and its height and width, including the samples to be filtered, are both nine samples.
[0526] 14. A method according to any one of solutions 1-13, wherein the spatial domain tap adopts a combination of a cross and a square, the height and width of the cross including the samples to be filtered are both thirteen samples, and the height and width of the square including the samples to be filtered are both five samples.
[0527] 15. The method according to any of the solutions 1-14, wherein input sources are used jointly inside a filter shape for the expansion taps.
[0528] 16. The method according to any of solutions 1-15, wherein the shapes of the extension taps associated with different input sources are different, and wherein the shapes include a single tap shape, a diamond shape, a cross shape, or a combination thereof.
[0529] 17. A method according to any one of solutions 1-16, wherein an indicator indicates an input source for the extended tap or the spatial tap, and wherein the indicator is signaled in an adaptive parameter set (APS), is predefined, or is derived on the fly.
[0530] 18. A method according to any one of solutions 1-17, wherein a first syntax element is transmitted via a signal to indicate whether the ALF with extended taps is enabled or which input source is used for the extended taps, and wherein the first syntax element is encoded using bypass encoding and decoding.
[0531] 19. A method according to any one of solutions 1-18, wherein the first syntax element is transmitted via a signal in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an APS, a codec tree unit (CTU) or a codec unit (CU).
[0532] 20. The method according to any of solutions 1-19, wherein the expansion taps receive intermediate results generated by reconstruction before applying ALF and offline trained filters of ALF or by reconstruction before applying DBF and offline trained filters of ALF.
[0533] 21. The method according to any of the solutions 1-20, wherein the expansion taps receive an intermediate result generated by applying a Gaussian filter.
[0534] 22. The method according to any one of solutions 1-21, wherein the expansion tap receives an intermediate result generated by a pre-DBF reconstruction or a post-DBF reconstruction of the current frame or the reference frame.
[0535] 23. The method according to any one of solutions 1-22, wherein the expansion tap receives residual samples of the current frame as input.
[0536] 24. The method according to any one of solutions 1-23, wherein the extended tap utilizes information other than spatially neighboring samples adjacent to the center sample to be filtered.
[0537] 25. The method according to any one of solutions 1-24, wherein at least one extension tap and at least one spatial domain tap coexist within an ALF filter, or wherein one or more extension taps of the ALF filter use one or more input sources.
[0538] 26. The method according to any of solutions 1-25, wherein one or more expansion taps of the ALF filter use an input source, and wherein the input source is based on a previous reconstruction of SAO or BF, a predicted sample, a co-located sample, or any modification thereof.
[0539] 27. A method according to any of solutions 1-26, wherein one or more expansion taps of the ALF filter use multiple input sources, and wherein the input sources include one or more of reconstruction, prediction samples, co-located samples or any modifications thereof before SAO or BF.
[0540] 28. A method according to any one of solutions 1-27, wherein the input source of the extended tap is used by other taps in the ALF, is not used by other taps in the ALF, is a filtered prediction sample, is a weighted sum of multiple sources or a filter of multiple sources, is the output of a function of multiple sources, is derived based on samples in different color components, or a combination thereof.
[0541] 29. A method according to any one of solutions 1-28, wherein whether or how a filter with at least one extended tap is applied is different for different color formats or different color components, or wherein an ALF filter with at least one extended tap is applied to process all chrominance components, or wherein an ALF filter with at least one extended tap is applied to filter both luminance and chrominance components.
[0542] 30. The method according to any of solutions 1-29, wherein the coefficients of the expansion taps inside the ALF filter correspond to N input samples, wherein the N input samples are designed in an asymmetric manner.
[0543] 31. A method according to any one of solutions 1-29, wherein the coefficients of the extended taps inside the ALF filter are shared by multiple inputs based on geometric information, or wherein multiple input samples are located at symmetrical positions, or wherein the ALF filter with at least one extended tap uses a shape or size different from the shape or size used by the ALF without an extended tap, or wherein within the ALF filter, the shape or size used for one or more spatial taps is the same as the shape or size used for one or more extended taps.
[0544] 32. A method according to any one of solutions 1-31, wherein a square, symmetrical, asymmetrical or other designed shape is used inside an ALF filter with at least one extended tap or at least one spatial tap, or wherein the spatial tap is determined on the fly, transmitted by a signal or derived.
[0545] 33. A method according to any of the solutions 1-32, wherein the ALF filter comprising at least one spatial tap and at least one extension tap uses a reconstruction before the ALF as input.
[0546] 34. A method according to any one of solutions 1-33, wherein one or more input sources are used for one or more extended taps, wherein the input sources are used jointly within a filter shape comprising one or more extended taps, and wherein the input sources include reconstruction before ALF, reconstruction before DBF, filter output of offline training by feeding reconstruction before ALF, filter output of offline training by feeding prediction or residual samples, coded reference pictures, predefined filter output of reconstruction before DBF, predefined filter output by feeding prediction or residual samples, prediction samples or residual samples.
[0547] 35. A method according to any one of solutions 1-34, wherein the number of extended taps belonging to different input sources is different, or wherein the shapes of the extended taps belonging to different input sources are different, or wherein a square, a symmetrical shape or an asymmetrical shape is applied, or wherein the total number of extended taps inside the ALF filter is jointly derived based on the shape, the filter length or the symmetry constraint.
[0548] 36. A method according to any one of solutions 1-35, wherein the first syntax element is encoded and decoded by arithmetic coding, encoded and decoded using at least one context, the context is encoded and decoded according to the coding and decoding information of the current block or a neighboring block, the context is encoded and decoded according to the filter shape of at least one neighboring block, encoded and decoded by binarization of a unary code, encoded and decoded by truncated unary code, encoded and decoded by a fixed length code, encoded and decoded by an exponential Golomb code, encoded and decoded by a truncated exponential Golomb code, conditionally transmitted through a signal, transmitted through a signal only when the extension tap is available, encoded and decoded in a predictive manner, predicted by the on or off decision of the extension tap of at least one neighboring block, independently transmitted through a signal for different color components, transmitted through a signal and shared for different color components, or transmitted through a signal for a first color component but not transmitted through a signal for a second color component.
[0549] 37. A method according to any one of solutions 1-36, wherein the coefficients of at least one extended tap inside the ALF filter are transmitted through a signal in a syntax element structure, the category merging result of the extended tap is included in the APS, the coefficients of the extended tap are encoded and decoded in a predictive manner, the coefficients of the extended tap are encoded and decoded using arithmetic coding with at least one context, the coefficients of the extended tap are encoded and decoded using bypass coding, the coefficients of the extended tap are jointly encoded and decoded with the coefficients of the spatial domain taps, other parameters of the extended tap are included in the APS, or the coefficients of the extended tap are predefined fixed values.
[0550] 38. A method according to any one of solutions 1-37, wherein the intermediate filtering result of at least one fixed or adaptive filter is used as the input of one or more expansion taps, and wherein the intermediate filtering result is the result of an offline trained filter, the result of an online trained ALF filter, the result of a bilateral filter, the result of a guided filter, the result of a median filter, the result of a local median filter, the result of a non-local median filter, the result of a filter with low-pass properties, the result of a filter with high-pass properties, the result of an online trained filter, reconstructed samples at different encoding and decoding stages, reconstruction before or after ALF of the current frame or reference frame, reconstruction before or after SAO or cross-component SAO (CCSAO) of the current frame or reference frame, reconstruction before or after bilateral filter (BIF) of the current frame or reference frame, or reconstruction before or after any other stage of the current frame or reference frame.
[0551] 39. A method according to any of solutions 1-38, wherein the input of one or more extended taps includes prediction or residual samples of the current frame, or samples within one or more coded pictures.
[0552] 40. A method according to any one of solutions 1-39, wherein the previously encoded frame is a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with a block, a current slice or a frame, and wherein the previously encoded frame is a short-term reference picture or a long-term reference picture.
[0553] 41. The method of any of solutions 1-40, wherein the previously decoded frame is not a reference frame and is stored in a decoded picture buffer (DPB).
[0554] 42. A method according to any one of solutions 1-41, wherein at least one indicator is signaled to indicate which previously encoded frame to use, or wherein one indicator is signaled to indicate which reference picture list to use, or wherein at least one indicator can be signaled to indicate a reference index, or wherein the indicator can be conditionally signaled depending on how many reference pictures are included in the RPL or RPS, or wherein the indicator is conditionally signaled depending on how many previously decoded pictures are included in the DPB, or wherein the frame to be utilized is determined on the fly.
[0555] 43. A method according to any one of solutions 1-42, wherein the extended taps obtain information from one or more previously encoded frames in the DPB, from one or more reference frames in list 0, list 1 or both, from a reference frame with a minimum picture order count (POC) distance to the current slice or frame, from a reference frame with a reference index equal to K in a reference list, or from a co-located frame, where K is transmitted by signal, predefined or derived on the fly based on reference picture information.
[0556] 44. A method according to any one of solutions 1-43, wherein which frame(s) to be used is determined based on decoding information, or wherein which frame(s) to be used is defined as the top N most frequently used reference pictures for the samples within the current slice or current frame, or wherein which frame(s) to be used can be defined as the top N most frequently used reference pictures of each reference picture list (when available) for the samples within the current slice or current frame, or wherein which frame(s) to be used is defined as the pictures with the top N smallest POC distances or absolute POC distances relative to the current picture.
[0557] 45. A method according to any of solutions 1-44, wherein whether to obtain information from a previously coded frame depends on decoded information of at least one region of the block to be filtered, or wherein whether to obtain information from a previously coded frame of at least one region of the block to be filtered is signaled, or wherein whether to obtain information from a previously coded frame depends on a slice or picture type, or wherein obtaining information from a previously coded frame is applicable to inter-coded slices / pictures, or wherein whether to obtain information from a previously coded frame depends on availability of reference pictures, or wherein whether to obtain information from a previously coded frame depends on reference picture information or picture information in a DPB, or wherein obtaining information from a previously coded frame is disabled when a minimum POC distance is greater than a threshold, or wherein whether to obtain information from a previously coded frame depends on a temporal layer index, a quantization parameter (QP), or a dimension of a picture, or wherein obtaining information from a previously coded frame is applicable to a block with a highest temporal layer index, or wherein when the block to be filtered contains a portion of samples coded in non-inter mode, the extended tap does not use information from a previously coded frame to filter the block, or wherein non-inter mode is defined as intra mode, or wherein non-inter mode is defined as a set of coding modes including intra, intra block copy (IBC), or palette mode, or wherein distortion between a current block and a matching block is calculated and used to decide whether to obtain information from a previously coded frame to filter the current block, or wherein distortion between a co-located block in a previously coded frame and the current block is used to decide whether to obtain information from a previously coded frame to filter the current block, or wherein motion estimation is first used to find a matching block from at least one previously coded frame, or wherein when the distortion is greater than a predefined threshold, information from a previously coded frame is not used, or wherein the information includes two reference blocks or co-located blocks for the current block, one of which is from a first reference frame in list 0 and the other is from a first reference frame in list 1.
[0558] 46. A method according to any one of solutions 1-45, wherein the coded reference pictures are accessed during different coding stages, or wherein the coded reference pictures are accessed during the prediction loop stage, or wherein the coded reference pictures are accessed during the loop filter stage, or wherein the coded reference pictures are accessed before DBF, SAO, CCSAO, BF, ChromaBF (ChromaBF), ALF, cross-component ALF (CCALF), or any other loop filtering process, or wherein the coded reference pictures are accessed during DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filtering process, or wherein the coded reference pictures are accessed after DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filtering process, or wherein the coded reference pictures are accessed after the loop filter stage, or wherein the coded reference pictures are accessed after the loop filter stage by a motion compensation based padding method.
[0559] 47. A method according to any one of solutions 1-46, wherein the ALF filter includes more than one sub-filter, or one of the sub-filters is based on samples in a reference picture, or one of the sub-filters is based on samples in a current picture, or one of the sub-filters is based on samples adjacent to the current sample, or one of the sub-filters is based on samples co-located with the current sample, or one of the sub-filters is based on at least one intermediate filtered sample generated from another filter, or the intermediate filtered samples are generated by filtering a group of samples adjacent to or co-located with the current sample, or the other filter is predefined, trained offline, or transmitted online through a signal, or the intermediate filtered samples have the same position as the current sample, or the intermediate filtered samples are adjacent to the current sample, or one of the sub-filters is based on reconstructed samples, predicted samples, residual samples, samples before deblocking, samples after deblocking, DBF, SAO, CCSA or samples before SAO or BF, or samples after DBF, SAO, CCSAO or BF, or wherein the ALF filter is constructed by one or more sub-filters, the one or more sub-filters including a sub-filter calculated based on spatially neighboring reconstructed samples after SAO, a sub-filter calculated based on spatially neighboring reconstructed samples before deblocking, a sub-filter calculated based on spatially neighboring reconstructed samples after deblocking, a sub-filter calculated based on spatially neighboring prediction samples, a sub-filter calculated based on spatially neighboring residual samples, a sub-filter calculated based on temporal samples in a reference picture, and a sub-filter calculated based on one or more median filtered samples, or wherein more than one sub-filter is applied via cascading without selective determination, or wherein all sub-filters are applied via cascading without selective determination, or wherein at least one sub-filter is used conditionally or adaptively, or wherein the sub-filter is conditionally controlled by a syntax element or a predefined rule.
[0560] 48. The method according to any of solutions 1-47, wherein the extended taps of the ALF are used for post-processing or pre-processing.
[0561] 49. A method according to any one of solutions 1-48, wherein the input source belongs to the same color component as the sample to be filtered or a different color component from the sample to be filtered, or wherein the input source is located at the same position as the sample to be filtered or within a range around the sample to be filtered, or wherein the input source includes a reconstruction before ALF, a reconstruction before DBF, an output of an offline trained filter obtained by feeding a reconstruction before ALF, a coded reference picture, an output of an offline trained filter obtained by feeding a prediction or residual sample, or a prediction or residual sample.
[0562] 50. The method according to any one of solutions 1-49, wherein the method is applied in combination or individually.
[0563] 51. A method according to any one of solutions 1-50, wherein the extended taps of the ALF are applied to any loop filtering tool, pre-processing or post-processing filtering in the video codec, including but not limited to ALF or CCALF or any other filtering, or wherein the extended taps of the ALF are applied to a loop filter, ALF, CCALF, other loop filters, pre-processing filters or post-processing filters.
[0564] 52. A method according to any of solutions 1-51, wherein the video unit is a sequence, a picture, a sub-picture, a slice, a CTU, a CTU row, a CTU group, a CU, a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.
[0565] 53. A method according to any of solutions 1-52, wherein the use of the method is signaled in the bitstream, or wherein the use of the method is signaled at the sequence level, picture level, picture level, slice level, slice group level, or is signaled in a sequence header, picture header, SPS, video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), PPS, APS, slice header or slice group header, or wherein the use of the method is signaled in a PB, TB, CB, PU, TU, CU, virtual pipeline data unit (VPDU), CTU, CTU row, slice, slice, sub-picture, or other type of area containing more than one sample or pixel.
[0566] 54. A method according to any of solutions 1-53, wherein the use of the method depends on coded information including block size, color format, single tree partitioning or dual tree partitioning, color components, or slice or picture type.
[0567] 55. The method of any of solutions 1-54, wherein the converting comprises encoding the visual media data into the bitstream.
[0568] 56. The method of any of solutions 1-54, wherein the converting comprises decoding the visual media data from the bitstream.
[0569] 57. A device for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Solutions 1-56.
[0570] 58. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs a method according to any one of Solutions 1-56.
[0571] 59. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; and generating the bitstream based on the determination.
[0572] 60. A method for storing a bitstream of a video, comprising: determining to employ an adaptive loop filter (ALF) having at least one extension tap and at least one spatial tap; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0573] In the described solution, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the described solution, a decoder can parse syntax elements in the codec representation according to the format rules using known information about the presence and absence of syntax elements to produce decoded video.
[0574] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are co-located or dispersed at different locations in the bitstream as defined by the syntax. For example, a macroblock may be encoded based on error residual values from a transform and a codec and may also be encoded using bits from headers and other fields in the bitstream. Furthermore, during conversion, a decoder may parse the bitstream based on this determination, using known information that some fields may or may not be present, as described in the above solution. Similarly, an encoder may determine whether to include or not include specific syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.
[0575] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0576] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file preserving other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers, which are located at a site or distributed across multiple sites and interconnected by a communication network.
[0577] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0578] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor that executes instructions and one or more memory devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0579] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features unique to particular embodiments of particular technologies. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may function in certain combinations as described above and may even be initially claimed in this manner, in some cases, one or more features in a claimed combination may be omitted from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.
[0580] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed sequentially in the particular order or sequence shown, or that all illustrated operations be performed to achieve desired results. Furthermore, the partitioning of various system components in the embodiments described in this patent document should not be understood as requiring such partitioning in all embodiments.
[0581] Only a few implementations and examples are described, and other implementations, improvements, and variations can be made based on what is described and illustrated in this patent document.
[0582] A first component is directly coupled to a second component when there are no intervening components between the first and second components other than a line, trace, or other medium. A first component is indirectly coupled to a second component when there are intervening components between the first and second components other than a line, trace, or other medium. The term "coupled" and its variations encompass both direct and indirect couplings. Unless otherwise specified, the use of the term "about" is intended to include a range of ±10% of the subsequent figure.
[0583] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0584] In addition, the techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled can be directly connected, or can be indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications can be determined by those skilled in the art and can be made without departing from the spirit and scope of the present disclosure.
Claims
1. A method for processing video data, comprising: Determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial domain tap; as well as Conversion between visual media data and a bitstream is performed based on the ALF.
2. The method of claim 1 , wherein the spatial taps receive values of reconstructed spatial neighboring samples before or after applying a deblocking filter (DBF), a sample adaptive offset (SAO) filter, a bilateral filter (BF), a cross-component sample adaptive offset (CCSAO) filter, or a combination thereof.
3. The method according to any one of claims 1-2, wherein the spatial tap receives only neighboring luma samples for filtering a central luma sample inside the ALF.
4. The method according to any one of claims 1 to 3, wherein the spatial domain taps only receive neighboring chroma samples for filtering a center chroma sample inside the ALF.
5. The method according to any one of claims 1 to 4, wherein the ALF comprises M spatial taps and N extension taps, wherein N and M are integer values greater than zero.
6. The method according to any one of claims 1 to 5, wherein the expansion tap receives input from one or more input sources, the one or more input sources comprising reconstructed samples before applying DBF, intermediate results of a predefined filter, residual samples, filtered residual samples, prediction samples, filtered prediction samples, or a combination thereof.
7. The method according to any one of claims 1 to 6, wherein the ALF filter with the extended taps is applied only to process a luma component or only to process one chroma component.
8. The method according to any one of claims 1 to 7, wherein the extended tap in the ALF corresponds to a single input sample.
9. The method according to any one of claims 1 to 8, wherein the extended taps in the ALF correspond to N input samples, where N is equal to 2, and wherein positions of the N input samples are symmetrical.
10. The method according to any one of claims 1 to 9, wherein the ALF filter adopts a first shape for the spatial domain taps and a second shape for the extended taps, wherein the first shape and the second shape are different.
11. The method according to any one of claims 1 to 10, wherein the ALF filter adopts a first size for the spatial taps and a second size for the extended taps, wherein the first size and the second size are different.
12. The method according to any one of claims 1 to 11, wherein the spatial tap is in a diamond shape or a cross shape, and wherein the extended tap is in a diamond shape or a cross shape.
13. The method according to any one of claims 1 to 12, wherein the spatial domain tap is in a diamond shape, and a height and a width thereof are both nine samples including the samples to be filtered.
14. The method according to any one of claims 1 to 13, wherein the spatial domain tap is a combination of a cross and a square, the height and width of the cross including the samples to be filtered are both thirteen samples, and the height and width of the square including the samples to be filtered are both five samples.
15. The method according to any of claims 1-14, wherein input sources are used jointly within a filter shape for the expansion taps.
16. The method of any one of claims 1-15, wherein the shapes of the extension taps associated with different input sources are different, and wherein the shapes comprise a single tap shape, a diamond shape, a cross shape, or a combination thereof.
17. The method according to any one of claims 1-16, wherein an indicator indicates an input source for the extension tap or the spatial tap, and wherein the indicator is signaled in an adaptive parameter set (APS), is predefined, or is derived on the fly.
18. The method according to any one of claims 1 to 17, wherein a first syntax element is signaled to indicate whether the ALF with the extension tap is enabled or which input source is used for the extension tap, and wherein the first syntax element is encoded using bypass codec.
19. The method according to any one of claims 1-18, wherein the first syntax element is signaled in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an APS, a codec tree unit (CTU), or a codec unit (CU).
20. The method according to any one of claims 1-19, wherein the expansion tap receives an intermediate result generated by reconstruction before applying ALF and an offline trained filter of ALF or by reconstruction before applying DBF and an offline trained filter of ALF.
21. The method of any one of claims 1-20, wherein the expansion taps receive an intermediate result generated by applying a Gaussian filter.
22. The method according to any one of claims 1 to 21, wherein the expansion tap receives an intermediate result generated by a pre-DBF reconstruction or a post-DBF reconstruction of a current frame or a reference frame.
23. The method according to any one of claims 1 to 22, wherein the expansion tap receives residual samples of the current frame as input.
24. The method according to any one of claims 1 to 23, wherein the expanded tap utilizes information other than spatially neighboring samples adjacent to a center sample to be filtered.
25. The method according to any one of claims 1 to 24, wherein At least one extended tap and at least one spatial domain tap coexist within one ALF filter, or wherein one or more extended taps of the ALF filter use one or more input sources.
26. The method according to any one of claims 1 to 25, wherein one or more expansion taps of the ALF filter use an input source, and wherein the input source is based on a previous reconstruction of SAO or BF, a predicted sample, a co-located sample, or any modification thereof.
27. The method according to any one of claims 1 to 26, wherein one or more expansion taps of the ALF filter use multiple input sources, and wherein the input sources include one or more of reconstruction before SAO or BF, predicted samples, co-located samples, or any modifications thereof.
28. The method according to any one of claims 1-27, wherein the input source of the extended tap is used by other taps in the ALF, is not used by other taps in the ALF, is a filtered prediction sample, is a weighted sum of multiple sources or a filter of multiple sources, is the output of a function of multiple sources, is derived based on samples in different color components, or a combination of the above.
29. A method according to any one of claims 1 to 28, wherein whether or how a filter with at least one extended tap is applied is different for different color formats or different color components, or wherein an ALF filter with at least one extended tap is applied to process all chrominance components, or wherein an ALF filter with at least one extended tap is applied to filter both luminance and chrominance components.
30. The method according to any one of claims 1-29, wherein the coefficients of the extended taps inside the ALF filter correspond to N input samples, wherein the N input samples are designed in an asymmetric manner.
31. A method according to any one of claims 1-29, wherein the coefficients of the extended taps inside the ALF filter are shared by multiple inputs based on geometric information, or wherein multiple input samples are located at symmetrical positions, or wherein the ALF filter with at least one extended tap uses a shape or size different from the shape or size used by the ALF without the extended tap, or wherein within the ALF filter, the shape or size used for one or more spatial taps is the same as the shape or size used for one or more extended taps.
32. The method of any one of claims 1-31, wherein a square, symmetrical, asymmetrical, or other designed shape is used inside an ALF filter having at least one extended tap or at least one spatial tap, or wherein the spatial taps are determined on the fly, transmitted by a signal, or derived.
33. The method according to any one of claims 1-32, wherein the ALF filter comprising at least one spatial tap and at least one extension tap uses a reconstruction prior to the ALF as input.
34. The method of any one of claims 1 to 33, wherein one or more input sources are used for one or more extension taps, wherein the input sources are used jointly within a filter shape comprising one or more extension taps, and wherein the input sources comprise pre-ALF reconstruction, pre-DBF reconstruction, offline trained filter outputs by feeding pre-ALF reconstruction, offline trained filter outputs by feeding prediction or residual samples, coded reference pictures, predefined filter outputs by feeding pre-DBF reconstruction, predefined filter outputs by feeding prediction or residual samples, prediction samples, or residual samples.
35. A method according to any one of claims 1-34, wherein the number of extension taps belonging to different input sources is different, or wherein the shapes of the extension taps belonging to different input sources are different, or wherein a square, a symmetrical shape or an asymmetrical shape is applied, or wherein the total number of extension taps inside the ALF filter is jointly derived based on the shape, the filter length or a symmetry constraint.
36. The method of claim 1 , wherein the first syntax element is encoded by arithmetic coding, encoded using at least one context, the context being encoded based on codec information of a current block or a neighboring block, the context being encoded based on a filter shape of at least one neighboring block, encoded by binarization with a unary code, encoded by a truncated unary code, encoded by a fixed length code, encoded by an exponential Golomb code, encoded by a truncated exponential Golomb code, conditionally signaled, signaled only if the extension tap is available, predicted, predicted by an on or off decision of an extension tap of at least one neighboring block, independently signaled for different color components, signaled and shared for different color components, or signaled for a first color component but not for a second color component.
37. A method according to any one of claims 1-36, wherein the coefficients of at least one extended tap inside the ALF filter are transmitted through a signal in a syntax element structure, the category merging result of the extended tap is included in the APS, the coefficients of the extended tap are encoded and decoded in a predictive manner, the coefficients of the extended tap are encoded and decoded using arithmetic coding with at least one context, the coefficients of the extended tap are encoded and decoded using bypass coding, the coefficients of the extended tap are jointly encoded and decoded with the coefficients of the spatial domain taps, other parameters of the extended tap are included in the APS, or the coefficients of the extended tap are predefined fixed values.
38. The method of any one of claims 1-37, wherein an intermediate filtering result of at least one fixed or adaptive filter is used as an input to one or more expansion taps, and wherein the intermediate filtering result is a result of an offline trained filter, a result of an online trained ALF filter, a result of a bilateral filter, a result of a guided filter, a result of a median filter, a result of a local median filter, a result of a non-local median filter, a result of a filter with a low-pass property, a result of a filter with a high-pass property, a result of an online trained filter, reconstructed samples at different codec stages, reconstruction before or after ALF of a current frame or a reference frame, reconstruction before or after SAO or cross-component SAO (CCSAO) of a current frame or a reference frame, reconstruction before or after bilateral filtering (BIF) of a current frame or a reference frame, or reconstruction before or after any other stage of a current frame or a reference frame.
39. The method according to any one of claims 1 to 38, wherein the input of the one or more extended taps comprises prediction or residual samples of the current frame, or samples within one or more coded pictures.
40. The method of any one of claims 1-39, wherein the previously coded frame is a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with a block, a current slice, or a frame, and wherein the previously coded frame is a short-term reference picture or a long-term reference picture.
41. The method of any one of claims 1-40, wherein a previously encoded frame is not a reference frame and is stored in the decoded picture buffer (DPB).
42. The method of any one of claims 1-41, wherein at least one indicator is signaled to indicate which previously coded frame to use, or wherein one indicator is signaled to indicate which reference picture list to use, or wherein at least one indicator may be signaled to indicate a reference index, or wherein the indicator may be conditionally signaled depending on how many reference pictures are included in the RPL or RPS, or wherein the indicator is conditionally signaled depending on how many previously decoded pictures are included in the DPB, or wherein the frame to be utilized is determined on the fly.
43. The method of any one of claims 1-42, wherein the extended taps obtain information from one or more previously coded frames in the DPB, from one or more reference frames in list 0, list 1, or both, from the reference frame having a minimum picture order count (POC) distance to the current slice or frame, from a reference frame with a reference index equal to K in a reference list, or from a co-located frame, where K is signaled, predefined, or derived on the fly based on reference picture information.
44. The method according to any one of claims 1 to 43, wherein the frame(s) to be used is determined based on decoding information, or wherein the frame(s) to be used is defined as the top N most frequently used reference pictures for the samples within the current slice or current frame, or wherein the frame(s) to be used can be defined as the top N most frequently used reference pictures in each available reference picture list for the samples within the current slice or current frame, or wherein the frame(s) to be used is defined as the pictures with the top N smallest POC distances or the smallest absolute POC distances with respect to the current picture.
45. The method according to any one of claims 1 to 44, wherein whether to obtain information from a previously coded frame depends on decoded information of at least one area of the block to be filtered, or wherein whether to obtain information from a previously coded frame of at least one area of the block to be filtered is transmitted by a signal, or wherein whether to obtain information from a previously coded frame depends on the slice type or the picture type, or wherein obtaining information from a previously coded frame is applicable to an inter-coded slice or picture, or wherein whether to obtain information from a previously coded frame depends on availability of a reference picture, or wherein whether to obtain information from a previously coded frame depends on the reference picture information or the picture information in the DPB, or wherein obtaining information from a previously coded frame is disabled when the minimum POC distance is greater than a threshold, or wherein whether to obtain information from a previously coded frame depends on the temporal layer index, a quantization parameter (QP), or the dimension of the picture, or wherein obtaining information from a previously coded frame is applicable to a block with a highest temporal layer index, or or wherein when the block to be filtered contains some samples coded in non-inter mode, the extended tap does not use information from a previously coded frame to filter the block, or wherein the non-inter mode is defined as an intra mode, or wherein the non-inter mode is defined as a set of coding modes including intra, intra block copy (IBC), or palette mode, or wherein a distortion between the current block and the matching block is calculated and used to decide whether to obtain information from a previously coded frame to filter the current block, or wherein the distortion between the co-located block in the previously coded frame and the current block is used to decide whether to obtain information from the previously coded frame to filter the current block, or wherein motion estimation is first used to find a matching block from at least one previously coded frame, or wherein when the distortion is greater than a predefined threshold, information from the previously coded frame is not used, or wherein the information includes two reference blocks or co-located blocks for the current block, one of which is from a first reference frame in list 0, and the other is from a first reference frame in list 1.
46. The method of any one of claims 1-45, wherein the coded reference pictures are accessed during different coding stages, or wherein the coded reference pictures are accessed during the prediction loop stage, or wherein the coded reference pictures are accessed during the loop filter stage, or wherein the coded reference pictures are accessed before DBF, SAO, CCSAO, BF, ChromaBF (ChromaBF), ALF, cross-component ALF (CCALF), or any other loop filtering process, or wherein the coded reference pictures are accessed during DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filtering process, or wherein the coded reference pictures are accessed after DBF, SAO, CCSAO, BF, ChromaBF, ALF, CCALF, or any other loop filtering process, or wherein the coded reference pictures are accessed after the loop filter stage, or wherein the coded reference pictures are accessed after the loop filter stage by a motion compensation based padding method.
47. The method according to any one of claims 1 to 46, wherein the ALF filter comprises more than one sub-filter, or one of the sub-filters is based on samples in a reference picture, or one of the sub-filters is based on samples in the current picture, or one of the sub-filters is based on samples adjacent to the current sample, or one of the sub-filters is based on samples co-located with the current sample, or one of the sub-filters is based on at least one intermediate filtered sample generated from another filter, or the intermediate filtered samples are generated by filtering a group of samples adjacent to or co-located with the current sample, or the another filter is predefined, trained offline, or signaled online, or the intermediate filtered samples have the same position as the current sample, or the intermediate filtered samples are adjacent to the current sample, or one of the sub-filters is based on reconstructed samples, predicted samples, residual samples, samples before deblocking, samples after deblocking, DBF, SA O, samples before CCSAO or BF, or samples after DBF, SAO, CCSAO or BF, or wherein the ALF filter is constructed by one or more sub-filters, the one or more sub-filters including a sub-filter calculated based on spatially neighboring reconstructed samples after SAO, a sub-filter calculated based on spatially neighboring reconstructed samples before deblocking, a sub-filter calculated based on spatially neighboring reconstructed samples after deblocking, a sub-filter calculated based on spatially neighboring prediction samples, a sub-filter calculated based on spatially neighboring residual samples, a sub-filter calculated based on temporal samples in a reference picture, and a sub-filter calculated based on one or more median filtered samples, or wherein more than one sub-filter is applied via cascading without selective determination, or wherein all sub-filters are applied via cascading without selective determination, or wherein at least one sub-filter is used conditionally or adaptively, or wherein the sub-filter is conditionally controlled by a syntax element or a predefined rule.
48. The method according to any one of claims 1-47, wherein the extended taps of the ALF are used for post-processing or pre-processing.
49. The method according to any one of claims 1 to 48, wherein the input source belongs to the same color component as the sample to be filtered or a different color component from the sample to be filtered, or wherein the input source is located at the same position as the sample to be filtered or is within a range around the sample to be filtered, or wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, an output of an offline trained filter obtained by feeding a reconstruction before ALF, a coded reference picture, an output of an offline trained filter obtained by feeding prediction or residual samples, or prediction or residual samples.
50. The method according to any one of claims 1-49, wherein the method is applied in combination or alone.
51. A method according to any one of claims 1-50, wherein the extended taps of the ALF are applied to any loop filtering tool, pre-processing or post-processing filtering in a video codec, including but not limited to ALF or CCALF or any other filtering, or wherein the extended taps of the ALF are applied to a loop filter, ALF, CCALF, other loop filter, pre-processing filter or post-processing filter.
52. The method of any one of claims 1-51, wherein a video unit is a sequence, a picture, a sub-picture, a slice, a CTU, a CTU row, a CTU group, a CU, a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), or any other region containing more than one luma or chroma sample or pixel.
53. The method of any one of claims 1-52, wherein use of the method is signaled in a bitstream, or wherein use of the method is signaled at a sequence level, a group of pictures level, a picture level, a slice level, a slice group level, or is signaled in a sequence header, a picture header, an SPS, a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a PPS, an APS, a slice header, or a slice group header, or wherein use of the method is signaled in a PB, a TB, a CB, a PU, a TU, a CU, a virtual pipeline data unit (VPDU), a CTU, a CTU row, a slice, a slice, a sub-picture, or other kind of region containing more than one sample or pixel.
54. The method of any one of claims 1-53, wherein use of the method depends on coded information including block size, color format, single or dual tree partitioning, color components, or slice or picture type.
55. The method of any one of claims 1-54, wherein the converting comprises encoding the visual media data into the bitstream.
56. The method of any one of claims 1-54, wherein the converting comprises decoding the visual media data from the bitstream.
57. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-56.
58. A non-transitory computer-readable medium comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method according to any one of claims 1-56.
59. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises: Determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial domain tap; as well as The bitstream is generated based on the determination.
60. A method for storing a bitstream of a video, comprising: Determining to use an adaptive loop filter (ALF) having at least one extension tap and at least one spatial domain tap; generating the bitstream based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.