Switchable input source based extended taps for adaptive loop filters in video coding
By introducing an extended tap into the ALF and combining the airspace tap, using the output information of multiple filters, the problem of limited ALF efficiency in the prior art is solved, and more efficient video encoding and decoding filtering performance is achieved.
Patent Information
- Application Number
- CN202380072843.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art fails to effectively utilize the output information of other filters in video encoding and decoding, resulting in limited efficiency of the adaptive loop filter (ALF).
Introduce an extended tap into the ALF and combines a space tap to enhance the filtering efficiency of the ALF using the output information of other filters such as de-blocking filter (DBF), sample point adaptive offset (SAO) and bilateral filter (BF).
By introducing extended taps and combining the output information of multiple filters, the efficiency of ALF is improved and the filtering performance during video encoding and decoding is enhanced.
Smart Images

Figure CN120035993A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2022 / 124822, filed on October 12, 2022. The contents of the foregoing patent application are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to the generation, storage and consumption of digital audio-visual media information in file format. Background Art
[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth requirements for digital video usage are likely to continue to grow. Summary of the invention
[0005] A first aspect relates to a method for processing video data, comprising: determining to apply an extended tap in an adaptive loop filter (ALF); and performing conversion between visual media data and a bitstream based on the ALF.
[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of the preceding aspects.
[0007] The third aspect relates to a non-transitory computer-readable medium, including a computer program product for use by a video codec device, the computer program product including computer executable instructions stored on the non-transitory computer-readable medium, so that when a processor executes the computer executable instructions, the video codec device performs the method of any of the aforementioned aspects.
[0008] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining to apply an extended tap in an adaptive loop filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] A sixth aspect relates to a method, device or system described in this patent document.
[0010] For clarity, any of the foregoing embodiments may be combined with one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0011] These and other features will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.
[0012] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 An example of nominal vertical and horizontal positions of luma and chroma samples in a 4:2:2 format in a picture is shown.
[0014] Figure 2 An example encoder block diagram is shown.
[0015] Figure 3 An example picture segmented into raster scan strips is shown.
[0016] Figure 4 An example picture segmented into rectangular scan strips is shown.
[0017] Figure 5 An example picture showing partitioning into bricks.
[0018] Figures 6A-6C An example of a codec tree block (CTB) crossing a picture boundary is shown.
[0019] Figure 7 An example of an intra prediction mode is shown.
[0020] Figure 8 Examples of block boundaries in a picture are shown.
[0021] Fig. 9 An example of pixels involved in filter usage is shown.
[0022] Fig.10 An example of a filter shape for an ALF is shown.
[0023] Fig.11 An example of transform coefficients supported by a 5×5 diamond filter is shown.
[0024] Fig.12 An example of relative coordinates supported by a 5×5 diamond filter is shown.
[0025] Fig.13 and Fig.14 Example shapes of spatial domain taps are shown.
[0026] Figures 15 to 18 An example ALF filter with extended taps is shown.
[0027] Fig.19 is a block diagram illustrating an example video processing system.
[0028] Fig. 20is a block diagram of an example video processing device.
[0029] Fig.21 is a flow chart of an example method of video processing.
[0030] Fig. 22 is a block diagram illustrating an example video codec system.
[0031] Fig.23 is a block diagram illustrating an example encoder.
[0032] Fig.24 is a block diagram illustrating an example decoder.
[0033] Fig.25 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0034] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should not be limited in any way to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the full scope of the appended claims and their equivalents.
[0035] The section titles used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the techniques described herein are applicable to other video codec protocols and designs.
[0036] 1. Initial Discussion
[0037] This document relates to video codec technology. In particular, this document relates to loop filters and other codec tools in image / video codecs. These ideas can be applied to video codecs such as High Efficiency Video Codec (HEVC), Versatile Video Codec (VVC), or other video codec technologies, either alone or in various combinations.
[0038] 2. Abbreviations
[0039] The present disclosure includes the following abbreviations. Advanced Video Codec (Rec. ITU-T H.264 | ISO / IEC 14496-10) (AVC), Codec Picture Buffer (CPB), Pure Random Access (CRA), Codec Tree Unit (CTU), Codec Video Sequence (CVS), Decoded Picture Buffer (DPB), Decoding Parameter Set (DPS), Generic Constraint Information (GCI), International Organization for Standardization (ISO), International Electrotechnical Commission (IEC), High Efficiency Video Codec (also known as Rec. ITU-T H.265 | ISO / IEC 23008-2, (HEVC)), Joint Exploration Model (JEM), Motion Constrained Slice Set (MCTS), Network Abstraction Layer (NAL), Output Layer Set (OLS), Picture Header (PH), Picture Parameter Set (PPS), Profile, Layer and Level (PTL), Picture Unit (PU), Reference Picture Resample (RPR), Raw Byte Sequence Payload (RBSP), Supplemental Enhancement Information (SEI), Slice Header (SH), Sequence Parameter Set (SPS), Video Codec Layer (VCL), Video Parameter Set (VPS), Versatile Video Codec (also known as Rec. ITU-T H.266 | ISO / IEC 23090-3) (VVC), VVC Test Model (VTM), Video Usability Information (VUI), Transform Unit (TU), Codec Unit (CU), Deblocking Filter (DF), Sample Adaptive Offset (SAO), Adaptive Loop Filter (ALF), Codec Block Flag (CBF), Quantization Parameter (QP), Rate-Distortion Optimization (RDO), and Bilateral Filter (BF).
[0040] 3. Video Codec Standards
[0041] Video codec standards have evolved primarily through the development of International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed Moving Picture Experts Group (MPEG-1) and MPEG-4 Vision, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards are based on a hybrid video codec structure that utilizes temporal prediction plus transform codec. In order to explore future video codec technologies beyond HEVC, the Video Codec Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many approaches and put them into a reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project was officially launched, the Joint Video Exploration Team (JVET) was renamed the Joint Video Experts Group (JVET). VVC is a coding standard that aims to reduce the bit rate by 50% compared to HEVC. The VVC Working Draft and the VVC Test Model (VTM) are continuously updated.
[0042] A sample version of the VVC draft (i.e., Universal Video Codec (Draft 10)) can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wg11 / JVET-S2001-v17.zip. A sample version of the reference software for VVC (named VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0043] ITU-T VCEG and ISO / IEC MPEG Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 are studying the potential need for standardization of future video codec technologies with compression capabilities that will significantly exceed the current VVC standard. This future standardization action may take the form of extensions of extensions of VVC or a completely new standard. These groups are working together on this development activity in a joint collaborative effort known as JVET to evaluate the compression technology designs proposed by their experts in the field. The first Exploratory Experiment (EE) was established by JVET and reference software named Enhanced Compression Model (ECM) is in use. The test model ECM is continuously updated.
[0044] 3.1 Color Space and Chroma Downsampling
[0045] A color space (also called a color model (or color system)) is a mathematical model that describes a range of colors as a tuple of numbers, for example as 3 or 4 values or color components (e.g., RGB). In general, a color space is a specialization of coordinate systems and subspaces. For video compression, the most commonly used color spaces are luminance, blue difference chrominance, and red difference chrominance (YCbCr) and red, green, blue (RGB).
[0046] YCbCr, Y′CbCr or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a general term for a family of color spaces that are used in video systems and digital photography as part of the color image processing pipeline. Y' is the luminance component, and CB and CR are the blue difference chrominance component and the red difference chrominance component. Y' (with a prime) is different from Y, which means that the light intensity is non-linearly encoded based on the gamma-corrected RGB primaries.
[0047] Chroma downsampling is the practice of encoding images by achieving a lower resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 3.1.1.4:4:4
[0049] In 4:4:4, each of the three components Y′CbCr has the same sampling rate. Therefore, there is no chroma downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 3.1.2.4:2:2
[0051] In 4:2:2, the two chroma components are sampled at half the sampling rate of luma. The horizontal chroma resolution is halved, while the vertical chroma resolution is unchanged. This reduces the bandwidth of the uncompressed video signal by one third, but there is little visual difference. Figure 1 Examples of nominal vertical and horizontal positions in a 4:2:2 color format are shown in FIG. 3.1.3 4:2:0
[0053] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but since the Cb and Cr channels are sampled only on alternate lines in this scheme, the vertical resolution is halved and the data rate remains the same. Cb and Cr are each downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, each with different horizontal and vertical sampling positions. In MPEG-2, Cb and Cr are co-aligned horizontally. Cb and Cr are located between pixels vertically (i.e., staggered). In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are arranged in an interlaced manner, located halfway between alternate luminance samples. In 4:2:0DV, Cb and Cr are co-aligned horizontally. Vertically, they are co-aligned on alternate lines.
[0054] chroma_format_idc separate_colour_plane_flag Chroma Format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0055] Table 1. SubWidthC and SubHeightC values derived from chroma_format_idc and Separate_colour_plane_flag
[0056] 3.2 Example Encoding and Decoding Process of Video Codec
[0057] Figure 2 An example of an encoder block diagram of VVC is shown, which contains three loop filtering modules: deblocking filter (DF), sample adaptive offset (SAO) and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively. The side information of the codec is used to transmit the offset and filter coefficients through the signal. ALF is located at the last processing stage of each picture and can be regarded as a tool that tries to capture and repair artifacts produced by previous stages.
[0058] 3.3 Definition of Video / Codec Unit
[0059] A picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of a picture. A slice can be divided into one or more bricks, each brick consisting of a certain number of CTU rows within the slice. A slice that is not divided into multiple bricks can also be called a brick. However, a brick that is a proper subset of a slice cannot be called a slice. A strip contains several slices of a picture or several bricks of a slice.
[0060] Two strip modes are supported, namely, raster scan strip mode and rectangular strip mode. In raster scan strip mode, a strip contains a series of slices in the slice raster scan of the picture. In rectangular strip mode, a strip contains a certain number of bricks of the picture, which together form a rectangular area of the picture. The bricks in a rectangular strip are arranged in the order of the brick raster scan of the strip. Figure 3 An example of raster scan stripe partitioning of a picture (with 18 by 12 luma CTUs) is shown, where the picture is divided into 12 slices and 3 raster scan stripes.
[0061] Figure 4 An example of rectangular slice partitioning of a picture (with 18 by 12 luma CTUs) is shown, where the picture is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular slices.
[0062] Figure 5 An example of a picture being partitioned into slices, bricks, and rectangular strips is shown, where the picture is divided into 4 slices (2 slice columns and 2 slice rows), 11 bricks (the upper left slice contains 1 brick, the upper right slice contains 5 bricks, the lower left slice contains 2 bricks, and the lower right slice contains 3 bricks), and 4 rectangular strips.
[0063] 3.3.1 CTU / CTB size
[0064] In VVC, the CTU size (signaled in the sequence parameter set (SPS) through the syntax element log2_ctu_size_minus2) can be as small as 4×4.
[0065] 7.3.2.3 Sequence parameter set RBSP syntax
[0066]
[0067]
[0068]
[0069] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size for each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC, and PicHeightInSamplesC are derived as follows:
[0070] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0071] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0072] MinCbLog2SizeY =log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0073] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0074] MinTbLog2SizeY = 2 (7-13)
[0075] MaxTbLog2SizeY = 6 (7-14)
[0076] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0077] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0078] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0079] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY )(7-18)
[0080] PicSizeInCtbsY = PicWidthInCtbsY * PicHeightInCtbsY (7-19)
[0081] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0082] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0083] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY (7-22)
[0084] PicSizeInSamplesY=pic_width_in_luma_samples * pic_height_in_luma_samples (7-23)
[0085] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0086] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0087] 3.3.2 CTUs in a Picture
[0088] Assume that the CTB / LCU size is represented by M×N (usually M equals N), and for a CTB located at the picture boundary (or slice or strip or other type of boundary, taking the picture boundary as an example), K×L samples are within the picture boundary, where K < M or L < N. For FIG. 6A to FIG. 6C those CTBs shown in Fig. 6A , the CTB size is still equal to M×N. However, as Figure 6B shown, the lower boundary of the CTB is outside the picture, or as Figure 6CAs shown, the lower / right border of the CTB is outside the picture.
[0089] 3.4 Intra-frame prediction
[0090] To capture arbitrary edge orientations present in natural videos, the number of directional intra modes is expanded to 65 from the 33 used in HEVC. Figure 7 The extended directional modes are shown in , and the planar and DC modes remain unchanged. These more dense directional intra prediction modes are applicable to all block sizes and luma and chroma intra prediction.
[0091] The angular intra prediction direction can be defined as from 45 degrees to -135 degrees in the clockwise direction, such as Figure 7 As shown in FIG. 1 . In VTM, for non-square blocks, several angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled and remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, e.g., 67, and the intra mode encoding and decoding remains unchanged.
[0092] In HEVC, each intra-coded block has a square shape, and the length of each side of the block is a power of 2. Therefore, no division operation is required to generate the intra predictor using the DC mode. In VVC, blocks can have a rectangular shape and in general a division operation must be used for each block. To avoid the division operation for DC prediction, only the longer sides are used to calculate the average for non-square blocks.
[0093] 3.5 Inter-frame prediction
[0094] For each inter-predicted CU, the motion parameters include motion vector, reference picture index, reference picture list usage index, and extended information of the new codec feature of VVC for sample generation for inter-prediction. Motion parameters can be transmitted by signal explicitly or implicitly. When the CU is encoded and decoded using skip mode, the CU is associated with one PU and has no significant residual coefficients, no encoded and decoded motion vector increments and / or reference picture indexes. Merge mode is defined as obtaining the motion parameters of the current CU from neighboring CUs, including spatial and temporal candidates and the extended schedule introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just skip mode. An alternative to Merge mode is to explicitly send motion parameters, where for each CU, motion vectors, corresponding reference picture indexes for each reference picture list, reference picture list usage flags, and other useful information are explicitly transmitted by signal.
[0095] 3.6 Deblocking Filter
[0096] Deblocking filtering is an example loop filter in a video codec. In VVC, the deblocking filtering process is applied to CU boundaries, transform subblock boundaries, and predicted subblock boundaries. Prediction subblock boundaries include prediction unit boundaries introduced by subblock-based temporal motion vector prediction (SbTMVP) and affine mode. Transform subblock boundaries include transform unit boundaries introduced by subblock transform (SBT) and intra sub-partitioning (ISP) mode, as well as transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as first filtering the vertical edges of the entire picture horizontally, and then filtering the horizontal edges vertically. This specific order enables the application of multiple horizontal filtering or vertical filtering processes in parallel threads. The filtering process can also be implemented CTB by CTB, and the processing delay is very small.
[0097] First, filter the vertical edges in the picture. Then, use the samples modified by the vertical edge filtering process as input to filter the horizontal edges in the picture. The vertical and horizontal edges in the CTB of each CTU are processed separately according to the codec unit. Filter the vertical edges of the codec blocks in the codec unit starting from the edge on the left side of the codec block and proceeding to the edge on the right side of the codec block in the geometric order of the codec blocks. Filter the horizontal edges of the codec blocks in the codec unit starting from the edge at the top of the codec block and proceeding to the edge at the bottom of the codec block in the geometric order of the codec blocks.
[0098] Figure 8 8 is a diagram 800 of samples 802 within an 8x8 block of samples 804. As shown, diagram 800 includes horizontal block boundaries and vertical block boundaries on an 8x8 grid 806, 808, respectively. In addition, diagram 800 shows non-overlapping blocks of 8x8 samples 810, which may be deblocked in parallel.
[0099] 3.6.1 Boundary Decision
[0100] The filtering is applied on 8×8 block boundaries. In addition, such boundaries must be transform block boundaries or codec subblock boundaries, for example, due to the use of affine motion prediction (ATMVP). For other boundaries, the deblocking filter is disabled.
[0101] 3.6.2 Boundary strength calculation
[0102] For transform block boundaries / codec sub-block boundaries, if the boundary is within the 8×8 grid, the boundary can be filtered and the settings of bS[xDi][yDj] for the edge (where [xDi][yDj] represents the coordinates) are defined as in Table 2 and Table 3, respectively.
[0103]
[0104] Table 2. Boundary Strength (When SPS IBC is disabled)
[0105]
[0106] Table 3. Boundary Strength (When SPS IBC is Enabled)
[0107] 3.6.3 Deblocking Decision for Luminance Component
[0108] Fig. 9 is an example of a pixel involved in the filter on / off decision and strong / weak filter selection. A wider and stronger luma filter is used only when conditions 1, 2, and 3 are all true (TRUE). Condition 1 is the "large block condition". This condition detects whether the samples on the P side and Q side belong to a large block, represented by the variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0109] bSidePisLargeBlk=((edge type is vertical and p 0 Belongs to CUs with width >= 32)||(edge type is horizontal and p 0 Belongs to CU with height >= 32))? True: False
[0110] bSideQisLargeBlk=((edge type is vertical and q 0 Belongs to CUs with width >= 32)||(edge type is horizontal and q 0 Belongs to CU with height >= 32))? True: False
[0111] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0112] Condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? True: False
[0113] Next, if condition 1 is true, we further check condition 2. First, we derive the following variables:
[0114] First, derive dp0, dp3, dq0, dq3 according to the HEVC method
[0115] if (side p is greater than or equal to 32)
[0116] dp0=(dp0+Abs(p50-2*p40+p30)+1)>>1
[0117] dp3=(dp3+Abs(p53-2*p43+p33)+1)>>1
[0118] if (q side is greater than or equal to 32)
[0119] dq0=(dq0+Abs(q50-2*q40+q30)+1)>>1
[0120] dq3=(dq3+Abs(q53-2*q43+q33)+1)>>1
[0121] Condition 2 = (d < β)? True: False
[0122] Where d=dp0+dq0+dp3+dq3.
[0123] If both conditions 1 and 2 are met, further check whether any block uses sub-blocks:
[0124]
[0125]
[0126] Finally, if both conditions 1 and 2 are met, the deblocking method will check condition 3 (large block strong filter condition), which is defined as follows. In condition 3 StrongFilterCondition, the following variables are derived:
[0127] Derive dpq in the same way as HEVC.
[0128] Derived according to HEVC method sp3 = Abs (p3-p0)
[0129]
[0130] According to the HEVC method, sq3 = Abs (q0-q3) is derived
[0131]
[0132] According to HEVC, StrongFilterCondition = (dpq is less than (β>>2), sp3+sq3 is less than (3*β>>5), and Abs(p0-q0) is less than (5*tC+1)>>1)? True: False.
[0133] 3.6.4 Stronger Deblocking Filter for Luma
[0134] When the samples on either side of the boundary belong to a large block, a bilinear filter is used. When the width of the vertical edge is >= 32 and when the height of the horizontal edge is >= 32, the sample is defined as belonging to a large block. The bilinear filter is listed as follows. The block boundary samples pi (i=0 to Sp-1) and qi (i=0 to Sq-1), in the above HEVC deblocking, pi and qi are the i-th sample in the row for filtering the vertical edge, or the i-th sample in the column for filtering the horizontal edge, and then replaced by linear interpolation as shown below:
[0135] p i ′=(f i *Middle s,t +(64-f i )*P s +32)>>6), limit to p i ±tcPD i
[0136] q j ′=(g j *Middle s,t +(64-g j )*Q s +32)>>6), limit to q j ±tcPD j
[0137] where tcPD i and tcPD j The term is the position-dependent clipping described above, and g j 、f i 、Middle s,t , P s and Q s Given below.
[0138] 3.6.5 Chroma Deblocking Decision
[0139] A chroma strong filter is used on both sides of the block boundary. Here, the chroma filter is selected when both sides of the chroma edge are greater than or equal to 8 (chroma position), and the decision that meets the following three conditions: The first decision is the decision of boundary strength and large blocks. This filter can be applied when the block width or height orthogonal to the block edge is equal to or greater than 8 in the chroma sample domain. The second and third decisions are basically the same as the HEVC luma deblocking decisions, which are on / off decisions and strong filter decisions, respectively.
[0140] In the first decision, for chroma filtering, the boundary strength (bS) is modified and the conditions are checked sequentially. If the condition is met, the remaining conditions with lower priority are skipped. Chroma deblocking is performed when bS is equal to 2, or bS is equal to 1 when a large block boundary is detected. The second and third conditions are basically the same as the HEVC luma strong filter decision, as shown below.
[0141] Under the second condition, d is derived according to HEVC luminance deblocking. When d is less than β, the second condition will be true. Under the third condition, StrongFilterCondition is derived as follows:
[0142] Derive dpq in the same way as HEVC.
[0143] Derivation of sp in the HEVC way 3 =Abs(p3-p 0 )
[0144] Derivation of sq in the HEVC way 3 =Abs(q 0 -q 3 )
[0145] According to HEVC design, StrongFilterCondition = (dpq is less than (β>>2), sp3+sq3 is less than (β>>3), and Abs(p0-q0) is less than (5*tC+1)>>1).
[0146] 3.6.6 Strong Deblocking Filter for Chroma
[0147] The strong deblocking filter for chroma is defined as follows:
[0148] p2′=(3*p3+2*p2+p1+p0+q0+4)>>3
[0149] p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3
[0150] p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3
[0151] The example chroma filter performs deblocking on a 4x4 grid of chroma samples.
[0152] 3.6.7 Position-dependent limiting
[0153] Position dependent clipping tcPD is applied to the output samples of the luminance filtering process, involving strong and long filters that modify 7, 5 and 3 samples at the boundaries. Assuming a quantization error distribution, the clipping value can be increased for samples that are expected to have higher quantization noise and therefore the reconstructed sample values are expected to deviate more from the true sample values.
[0154] For each P or Q boundary filtered with an asymmetric filter, a position-dependent threshold table is selected from two tables (e.g., Tc7 and Tc3 listed below) that are provided to the decoder as auxiliary information, based on the result of the decision-making process:
[0155] Tc7={6,5,4,3,2,1,1}; Tc3={6,4,2};
[0156] tcPD=(Sp==3)? Tc3:Tc7;
[0157] tcQD=(Sq==3)? Tc3:Tc7;
[0158] For P or Q boundaries filtered with a short symmetric filter, a lower magnitude position-dependent threshold is applied:
[0159] Tc3 = {3, 2, 1};
[0160] After defining the threshold, the filtered p′i and q′i sample values are clipped according to the tcP and tcQ clipping values:
[0161] p”i=Clip3(p’i+tcPi,p’i–tcPi,p’i);
[0162] q”j=Clip3(q’j+tcQj,q’j–tcQ j,q’j);
[0163] where p'i and q'j are the filtered sample values, p"i and q"j are the clipped output sample values, and tcPi is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. Function Clip3 is the clipping function as specified in VVC.
[0164] 3.6.8. Sub-block deblocking adjustment
[0165] To achieve parallel-friendly deblocking using both long filters and sub-block deblocking, the long filter is restricted to modifying at most 5 samples on one side using sub-block deblocking (AFFINE or ATMVP or decoder-side motion vector refinement (DMVR)), as shown in the luma control of the long filter. By extension, sub-block deblocking is adjusted so that sub-block boundaries on the 8×8 grid close to CU or implicit TU boundaries are restricted to modifying at most two samples on each side.
[0166] The following applies to sub-block boundaries that are not aligned with CU boundaries.
[0167]
[0168] Where edge equal to 0 corresponds to a CU boundary, edge equal to 2 or equal to orthogonalLength-2 corresponds to a sub-block boundary 8 samples away from the CU boundary, etc. If implicit partitioning of TUs is used, implicitTU is true.
[0169] 3.7. Sample Adaptive Offset
[0170] Sample adaptive offset (SAO) is applied to the reconstructed signal after the deblocking filter by using the offset specified by the encoder for each CTB. The video encoder first decides whether to apply SAO processing to the current slice. If SAO is applied to the slice, each CTB will be classified into one of five SAO types, as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding offsets to pixels of each category. SAO operations include: edge offset (EO), which uses edge attributes for pixel classification in SAO types 1 to 4; and band offset (BO), which uses pixel intensity for pixel classification in SAO type 5. Each applicable CTB has SAO parameters, including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag is equal to 1, the current CTB will reuse the SAO type and offset of the left CTB. If sao_merge_up_flag is equal to 1, the current CTB will reuse the SAO type and offset of the CTB above.
[0171]
[0172]
[0173] Table 4. SAO type regulations
[0174] 3.8. Adaptive Loop Filter
[0175] Adaptive loop filtering for video codecs minimizes the mean square error between the original samples and the decoded samples by using a Wiener-based adaptive filter. ALF is located at the last processing stage of each picture and can be regarded as a tool to capture and repair artifacts from previous stages. The appropriate filter coefficients are determined by the encoder and explicitly transmitted to the decoder through the signal. In order to achieve better codec efficiency, especially for high-resolution video, local adaptation is used for the luminance signal by applying different filters to different areas or blocks in the picture. In addition to filter adaptation, filter on / off control at the codec tree unit (CTU) level also helps to improve codec efficiency. Syntactically, the filter coefficients are sent in the header information at the picture level called the adaptive parameter set, and the filter on / off flag of the CTU is interleaved at the CTU level in the strip data. This syntax design not only supports picture-level optimization, but also achieves lower encoding latency.
[0176] 3.8.1. Signal transmission of parameters
[0177] According to the ALF design in VTM, filter coefficients and clipping indices are carried in the ALF Adaptive Parameter Set (APS). The ALF APS may include up to 8 chroma filters and one luma filter set, for a total of up to 25 filters. Each of the 25 luma classes also includes an index. Classes with the same index share the same filter. By merging different classes, the number of bits required to represent the filter coefficients is reduced. The absolute value of the filter coefficient is represented using a 0th order exponential Golomb code followed by a sign bit for non-zero coefficients. When clipping is enabled, a two-bit fixed-length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APS simultaneously.
[0178] The filter control syntax elements of ALF in VTM include two types of information. First, the ALF on / off flag is transmitted through signals at the sequence, picture, slice, and CTB levels. The chroma ALF can be enabled at the picture and slice levels only when the luma ALF is enabled at the corresponding level. Secondly, if ALF is enabled at the picture, slice, and CTB level, the filter usage information is transmitted through signals at that level. If all slices within a picture use the same APS, the referenced ALF APS ID is encoded and decoded at the slice level or picture level. The luma component can reference up to 7 ALF APSs, while the chroma component can reference up to 1 ALF APS. For the luma CTB, the signal transmission index indicates which ALF APS or offline trained luma filter set to use. For the chroma CTB, the index indicates which filter in the referenced APS to use.
[0179] The ALF data syntax elements associated with the luma (LUMA) component in the VTM are listed below:
[0180]
[0181]
[0182] alf_luma_filter_signal_flag equal to 1 specifies that the luma filter set is signaled. alf_luma_filter_signal_flag equal to 0 specifies that the luma filter set is not signaled. alf_luma_clip_flag equal to 0 specifies that linear adaptive loop filtering is applied to the luma component. alf_luma_clip_flag equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the luma component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of adaptive loop filter classes for which luma coefficients may be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters-1, inclusive. alf_luma_coeff_delta_idx[filtIdx] indicates the index of the signaled adaptive loop filter luma coefficient delta for the filter class indicated by filtIdx, where filtIdx ranges from 0 to NumAlfFilters-1. When alf_luma_coeff_delta_idx[filtIdx] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1+1)) bits. The value of alf_luma_coeff_delta_idx[filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1, inclusive.
[0183] alf_luma_coeff_abs[sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_coeff_abs[sfIdx][j] is not present, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[sfIdx][j] shall be in the range of 0 to 128, inclusive. alf_luma_coeff_sign[sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx, as follows:
[0184] If alf_luma_coeff_sign[sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.
[0185] Otherwise (alf_luma_coeff_sign[sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.
[0186] When alf_luma_coeff_sign[sfIdx][j] is not present, it is inferred to be equal to 0.
[0187] alf_luma_clip_idx[sfIdx][j] specifies the clip index of the clip value to be used before multiplying the jth coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_clip_idx[sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in the VTM are listed below:
[0188]
[0189] alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 1 specifies that the adaptive loop filter is applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at luma position (xCtb, yCtb). alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] equal to 0 specifies that the adaptive loop filter is not applied to the codec tree block of the color component indicated by cIdx of the codec tree unit at luma position (xCtb, yCtb).
[0190] When alf_ctb_flag[cIdx][xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to 0. alf_use_aps_flag equal to 0 specifies that one of the fixed filter sets is applied to the luma CTB. alf_use_aps_flag equal to 1 specifies that the filter set from the APS is applied to the luma CTB. When alf_use_aps_flag is not present, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx should be in the range of 0 to sh_num_alf_aps_ids_luma-1, inclusive. When alf_luma_prev_filter_idx is not present, it is inferred to be equal to 0.
[0191] The variable AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the filter set index for the luma CTB at position (xCtb, yCtb) derived as follows:
[0192] If alf_use_aps_flag is equal to 0, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to alf_luma_fixed_filter_idx.
[0193] Otherwise, AlfCtbFiltSetIdxY[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is set equal to 16+alf_luma_prev_filter_idx.
[0194] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15 (inclusive).
[0195] Based on the ALF design of VTM, the ALF design of ECM further introduces the concept of alternative filter sets into the luma filter. Based on the updated luma CTU ALF on / off decision of each alternative / round, the luma filter is trained for multiple alternatives / rounds. In this way, there will be multiple filter sets associated with each training alternative, and the class merging results of each filter set may be different. Each CTU can select the best filter set by RDO, and the relevant alternative information will be transmitted through the signal. The data syntax elements of the ALF associated with the luma component in the ECM are listed as follows:
[0196]
[0197]
[0198] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for the luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3, inclusive. alf_luma_clip_flag[altIdx] equal to 0 specifies that linear adaptive loop filtering is applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_clip_flag[altIdx] equal to 1 specifies that nonlinear adaptive loop filtering may be applied to the alternative luma filter set for the luma component with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter classes for which luma coefficients may be signaled to the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters-1, inclusive.
[0199] alf_luma_coeff_delta_idx[altIdx][filtIdx] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter class denoted by filtIdx for the alternative luma filter set with index altIdx, where filtIdx ranges from 0 to NumAlfFilters – 1. When alf_luma_coeff_delta_idx[filtIdx][altIdx] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[altIdx][filtIdx] is Ceil(Log2(alf_luma_num_filters_signalled_minus1[altIdx]+1)) bits. The value of alf_luma_coeff_delta_idx[altIdx][filtIdx] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[altIdx], inclusive. alf_luma_coeff_abs[altIdx][sfIdx][j] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx of the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[altIdx][sfIdx][j] shall be in the range of 0 to 128, inclusive.
[0200] alf_luma_coeff_sign[altIdx][sfIdx][j] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx of the alternative luma filter set with index altIdx as follows:
[0201] If alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 0, the corresponding luma filter coefficient has a positive value.
[0202] Otherwise (alf_luma_coeff_sign[altIdx][sfIdx][j] is equal to 1), the corresponding luma filter coefficient has a negative value.
[0203] When alf_luma_coeff_sign[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0.
[0204] alf_luma_clip_idx[altIdx][sfIdx][j] specifies the clip index of the clip value to be used before multiplying the j-th coefficient of the signaled luma filter represented by sfIdx of the alternative luma filter set indexed by altIdx. When alf_luma_clip_idx[altIdx][sfIdx][j] is not present, it is inferred to be equal to 0. The codec tree syntax elements associated with the luma component in the ECM are listed below:
[0205]
[0206] alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] specifies the index of the alternative luma filter of the codec tree block applied to the luma component of the codec tree unit at luma position (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not present, it is inferred to be equal to zero.
[0207] Filter shape
[0208] In JEM, up to three diamond filter shapes can be selected for the luminance component (e.g. Fig.10 ). The filter shape for the luma component is indicated at the picture level by a signal transmission index. Each square represents a sample, and Ci (i is 0 to 6 (left), 0 to 12 (center), 0 to 20 (right)) represents the coefficient to be applied to the sample. For the chroma components in a picture, a 5×5 diamond shape is always used. In VVC, a 7×7 diamond shape is always used for luma, and a 5×5 diamond shape is always used for chroma.
[0209] 3.8.3 Classification of ALF
[0210] Each 2×2 (or 4×4) block is classified into one of 25 classes. The classification index C is based on the quantized value of its directionality D and activity The derivation is as follows:
[0211]
[0212] In order to calculate D and First, use the 1-dimensional Laplacian operator to calculate the gradient in the horizontal, vertical and two diagonal directions:
[0213]
[0214] The indices i and j refer to the coordinates of the top left sample point in the 2×2 block, and R(i,j) indicates the reconstructed sample point at coordinate (i,j). The maximum and minimum values of the gradient D in the horizontal and vertical directions are set to:
[0215]
[0216] And the maximum and minimum values of the gradients in the two diagonal directions are set to:
[0217]
[0218] To derive the value of the directivity D, these values are compared with each other and with two thresholds t 1 and t 2 For comparison:
[0219] Step 1. If and If both are true, D is set to 0.
[0220] Step 2. If Then continue from step 3; otherwise continue from step 4;
[0221] Step 3. If If yes, set D to 2; otherwise, set D to 1.
[0222] Step 4. If Then set D to 4; otherwise, set D to 3.
[0223] The activity value A is calculated as:
[0224]
[0225] A is further quantized into a range of 0 to 4 (inclusive), and the quantized value is represented as For the two chrominance components in a picture, no classification method is used (ie, one set of ALF coefficients is applied to each chrominance component).
[0226] 3.8.4. Geometric transformation of filter coefficients
[0227] Before filtering each 2×2 block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f(k,l) associated with the coordinates (k,l) according to the gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks more similar by aligning their directionality, to which the ALF is applied.
[0228] Introduce three geometric transformations, including diagonal transformation, vertical flip and rotation:
[0229] Diagonal: f D (k,l)=f(l,k),
[0230] Flip vertically: f V (k,l)=f(k,Kl-1),
[0231] Rotation: f R (k,l)=f(Kl-1,k).
[0232] Where K is the size of the filter, and 0≤k,l≤K-1 are the coefficient coordinates, so that the position (0,0) is at the upper left corner, and the position (K-1,K-1) is at the lower right corner. The transform is applied to the filter coefficients f(k,l) according to the gradient value calculated for the block. Table 5 summarizes the relationship between the transform and the four gradients in the four directions. Fig.11 The transform coefficients for each position based on a 5×5 diamond are shown.
[0233] Gradient Value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transformation <![CDATA[g d2 <g d1 And g v <g h ]]> diagonal <![CDATA[g d1 <g d2 And g h <g v ]]> Flip vertically <![CDATA[g d1 <g d2 And g v <g h ]]> Rotation
[0234] Table 5. Mapping of gradients and transformations computed for a block.
[0235] 3.8.5. Filtering process
[0236] At the decoder side, when ALF is enabled for a block, each sample R(i,j) within the block is filtered to obtain the sample value R′(i,j) as shown below, where L represents the filter length and f m,n represents the filter coefficient, and f(k,l) represents the decoded filter coefficient.
[0237]
[0238] Fig.12 An example of relative coordinates for a 5×5 diamond filter support is shown, assuming that the coordinates (i, j) of the current sample point are (0, 0). Samples in different coordinates filled with the same color are multiplied by the same filter coefficient.
[0239] 3.8.6. Nonlinear Filter Reconstruction (reformulation)
[0240] Linear filtering can be reconstructed as the following expression without affecting the encoding and decoding efficiency:
[0241]
[0242] where w(i,j) are the same filter coefficients.
[0243] VVC introduces nonlinearity to make ALF more efficient by using a simple clipping function to reduce the impact of neighboring sample values (I(x+i,y+j)) when they differ too much from the current sample value being filtered (I(x,y)). More specifically, the ALF filter is modified as follows:
[0244]
[0245] where K(d,b)=min(b,max(-b,d)) is the clipping function and k(i,j) are the clipping parameters, which depend on the (i,j) filter coefficients. The encoder performs an optimization to find the best k(i,j).
[0246] A clipping parameter k(i,j) is specified for each ALF filter, with one clipping value signaled for each filter coefficient. This means that up to 12 clipping values can be signaled in the bitstream for each luma filter, and up to 6 clipping values can be signaled in the bitstream for each chroma filter. To limit signaling cost and encoder complexity, only 4 fixed values are used, which are the same for inter and intra slices.
[0247] Because the variance of local differences in luma is usually higher than that of local differences in chroma, two different sets of luma filters and chroma filters are applied. A maximum sample value in each set (here 1024 for a bit depth of 10 bits) is also introduced so that clipping can be disabled when not needed. These 4 values have been chosen by roughly evenly dividing the full range of sample values for luma (in 10-bit encoding) and the range of 4 to 1024 for chroma in the logarithmic domain. More precisely, the luma table of clipping values is obtained by the following formula:
[0248] Where M = 2 10 And N = 4
[0249] Similarly, the chromaticity table of the clipped value can be obtained according to the following formula:
[0250] Where M = 2 10 , N = 4 and A = 4
[0251] 3.9. Bilateral Loop Filter
[0252] 3.9.1. Bilateral Image Filter
[0253] Bilateral image filter is a nonlinear filter that smoothes noise while preserving edge structure. Bilateral filtering is a technique that makes the filter weight decrease not only with the distance between samples, but also with the increase of intensity difference. In this way, edge oversmoothing can be improved. The weight is defined as
[0254]
[0255] where Δx and Δy are the distances in the vertical and horizontal directions, respectively, and ΔI is the intensity difference between the sample points.
[0256] The edge-preserving denoising bilateral filter uses low-pass Gaussian filters for both the domain filter and the range filter. The domain low-pass Gaussian filter gives higher weights to pixels that are spatially close to the center pixel. The range low-pass Gaussian filter gives higher weights to pixels that are similar to the center pixel. Combining the range filter and the domain filter, the bilateral filter at the edge pixel becomes an elongated Gaussian filter oriented along the edge and greatly reduced in the gradient direction. This is why the bilateral filter can smooth out the noise while preserving the edge structure.
[0257] 3.9.2. Bilateral Filters in Video Codecs
[0258] The bilateral filter in video codec is a codec tool for VVC. The filter acts as a loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and SAO act on the same input samples, and each filter produces an offset, which is then added to the input samples to produce output samples, which enter the next stage after clipping. The spatial filter strength σ d Determined by the block size, the smaller the block, the greater the filter strength, and the intensity filter strength σ r Determined by the quantization parameter, where stronger filtering is used for higher QPs. Only the four closest samples are used, so the filtered sample strength I F can be calculated as
[0259]
[0260] Among them I C Indicates the intensity of the central sample point, ΔI A =I A -I C Indicates the intensity difference between the center sample point and the sample point above. ΔI B , ΔI L and ΔI R Represents the intensity difference between the center sample point and the samples below, to the left, and to the right, respectively.
[0261] 4. Technical problems solved by the disclosed technical solutions
[0262] The example design of adaptive loop filter (ALF) in video codec has the following problems:
[0263] In some ALF designs, only spatially reconstructed samples after other filtering, such as deblocking filtering (DBF), sample adaptive offset (SAO), and bilateral filtering (BF), are used for filter training and filtering. However, other valuable information, such as samples filtered / generated by one or more predefined filters, can potentially be utilized.
[0264] In some ALF designs, only spatially reconstructed samples after other filtering (such as deblocking filter, SAO and BF) are used for filter training and filtering. However, other valuable information, such as samples before DBF, SAO or other stages, can potentially be utilized.
[0265] 5. List of solutions and implementation examples
[0266] In order to solve the above problems, a method as summarized below is disclosed. The embodiments should be regarded as examples for explaining general concepts and should not be interpreted narrowly. In addition, these embodiments can be applied individually or combined in any way. It should be noted that the disclosed method can be used as a loop filter or post-processing. In the present disclosure, a video unit can refer to a sequence, a picture, a sub-picture, a strip, a CTU, a block and / or a region. A video unit may include one color component or multiple color components. In the present disclosure, an ALF processing unit may refer to a sequence, a picture, a sub-picture, a strip, a CTU, a block, a region or a sample. An ALF processing unit may include one color component or multiple color components.
[0267] In the following disclosure, the filtered sample value is represented as the output of BF. For example, in the process described below:
[0268]
[0269] Among them I F Represents the output of BF.
[0270] In the following disclosure, BF refers to an example of bilateral filtering in video codecs, which typically uses predefined filter parameters to generate the output of the BF. For example, there is no online training or signal transmission of the filter parameters used in the BF.
[0271] In the following disclosure, adaptive BF refers to an improved version of BF over bilateral filtering in video codecs. Adaptive BF involves online parameter training and signal transmission of parameters. For example, multiple filtered samples generated based on BF are further combined with parameters transmitted by signal and trained online.
[0272] Example 1
[0273] In one example, at least one extended tap of the ALF filter further enhances the efficiency of the ALF. In one example, at least one extended tap may be different from a spatial tap in the ALF filter, which utilizes only information of spatially neighboring samples of the target component. In one example, the spatially neighboring samples may come from reconstruction after DBF / SAO / BF. In one example, the spatial tap may filter a central luma sample inside an ALF filter using only spatially neighboring luma samples. In one example, the spatial tap may filter a central chroma sample inside an ALF filter using only spatially neighboring chroma samples.
[0274] Example 2
[0275] In one example, at least one extended tap and at least one spatial domain tap coexist in one ALF filter. In one example, the ALF filter may include both spatial domain taps and extended taps. In one example, the ALF filter may include M (e.g., M>0) spatial domain taps and N (e.g., N>0) extended taps.
[0276] Example 3
[0277] In one example, one or more extended taps of the ALF filter may use one or more input sources. In one example, one or more extended taps of the ALF filter may use one input source. (For example, reconstruction before DBF or intermediate filtering results of a predefined filter or reconstruction before SAO / BF). In one example, one or more extended taps of the ALF filter may use multiple input sources. (For example, reconstruction before DBF and intermediate filtering results of a predefined filter). The input source of the extended tap may also be used by other taps in the ALF. The input source of the extended tap may not be used by other taps in the ALF. The input source of the extended tap may be a predicted sample. The input source of the extended tap may be a filter predicted sample. The input source of the extended tap may be a weighted sum or filtering of multiple sources. The input source of the extended tap may be the output of a function of multiple sources. The input source of the extended tap may be derived based on samples in different color components.
[0278] Example 4
[0279] In one example, whether and / or how to apply a filter with at least one extended tap may be different for different color formats and / or different color components. In one example, an ALF filter with at least one extended tap may be applied only to process a luminance component. In one example, an ALF filter with at least one extended tap may be applied only to process one of the chrominance components (e.g., a Cb or Cr component). In one example, an ALF filter with at least one extended tap may be applied to process all chrominance components (e.g., Cb and Cr components). In one example, an ALF filter with at least one extended tap may be applied to filter luminance and chrominance components (e.g., Y, Cb, and Cr components).
[0280] Example 5
[0281] In one example, the coefficients of the extended taps inside the ALF filter may correspond to one or more input samples. In one example, the coefficients of the extended taps inside the ALF filter may correspond to only one input sample. In one example, the coefficients of the extended taps inside the ALF filter may correspond to N input samples (e.g., N=2). In one example, the N input samples may be designed in a symmetrical manner. In one example, the N input samples may be designed in an asymmetric manner. In one example, the coefficients of the extended taps inside the ALF filter may be shared by multiple inputs based on geometric information. In one example, multiple input samples may be located at symmetrical positions.
[0282] Example 6
[0283] In one example, the ALF filter with at least one extended tap may use different shapes or sizes. In one example, the ALF filter may include spatial taps and extended taps of different shapes. In one example, inside the ALF filter, the shape / size for one or more spatial taps may be different from the shape / size for one or more extended taps. In one example, inside the ALF filter, the shape / size for one or more spatial taps may be the same as the shape / size for one or more extended taps. In one example, inside the ALF filter with at least one extended tap, the filter shape for one or more spatial taps may be as described below. In one example, the filter shape for the spatial tap may be a diamond shape. In one example, the filter shape for the spatial tap may be a square shape. In one example, the filter shape for the spatial tap may be a cross shape. In one example, the filter shape for the spatial tap may be a symmetrical shape. In one example, the filter shape for the spatial tap may be an asymmetrical shape. In one example, the filter shape for the spatial tap may be other designed shapes. In one example, the filter shape for the spatial tap may be determined / transmitted / derived by signaling on the fly.
[0284] Example 7
[0285] In one example, inside an ALF filter having at least one extended tap, the filter shape for one or more extended taps may be as described below. In one example, the filter shape for the extended tap may be a diamond shape. In one example, the filter shape for the extended tap may be a square shape. In one example, the filter shape for the extended tap may be a cross shape. In one example, the filter shape for the extended tap may be a symmetrical shape. In one example, the filter shape for the extended tap may be an asymmetrical shape. In one example, the filter shape for the extended tap may be other designed shapes. In one example, the filter shape for the extended tap may be determined / transmitted / derived in real time via a signal.
[0286] Example 8
[0287] In one example, an ALF filter including at least one spatial tap and at least one extension tap may be designed as follows: In one example, one or more spatial taps may be used for filtering (the spatial taps may be viewed as using the reconstruction prior to ALF as input). Figure 13 to Figure 14 shows example shapes of spatial taps. In one example, Fig.13 Design the shape of the spatial tap. In one example, you can press Fig.14 Design the shape of the spatial taps.
[0288] In one example, one or more extension taps may be applied for filtering. In one example, an ALF filter having one or more spatial domain taps and one or more extension taps may be designed as follows. Figures 15 to 18 An example ALF filter with extended taps is shown. In one example, Fig.15 Design an ALF filter with spatial taps and extended taps. In one example, Fig.16 Design an ALF filter with spatial taps and extended taps. In one example, Fig.17 Design an ALF filter with spatial taps and extended taps. In one example, Fig.18 Design an ALF filter with spatial domain taps and extended taps.
[0289] Example 9
[0290] In one example, a transformation based on geometric information may be applied. In one example, a transformation based on geometric information may be applied independently to one or more spatial domain taps. In one example, a transformation based on geometric information may be applied independently to one or more extended taps. In one example, a transformation based on geometric information may be applied jointly to one or more spatial domain taps and one or more extended taps.
[0291] Example 10
[0292] In one example, one or more input sources can be used for one or more expansion taps. In one example, Figures 15 to 18 The input sources A, B, C, D, E shown in the example may be different. In another example, Figures 15 to 18 The input sources A, B, C, D, E shown in the example of may be the same. In one example, one possible input source for one or more extended taps may be a reconstruction before ALF. In one example, one possible input source for one or more extended taps may be a reconstruction before DBF. In one example, one possible input source for one or more extended taps may be an intermediate filtering result generated by a reconstruction before ALF and an offline trained filter of ALF.
[0293] In one example, a specific offline trained filter bank may be applied. In one example, an offline trained filter bank indicator may be signaled / predefined / derived on the fly. In one example, a specific offline trained filter classifier may be applied. In one example, an offline trained filter classifier indicator may be signaled / predefined / derived on the fly. In one example, one possible input source for one or more extension taps may be an intermediate filtering result generated by reconstruction prior to the DBF and an offline trained filter of the ALF.
[0294] In one example, a specific offline trained filter bank may be applied. In one example, an offline trained filter bank indicator may be signaled / predefined / derived on the fly. In one example, a specific offline trained filter classifier may be applied. In one example, an offline trained filter classifier indicator may be signaled / predefined / derived on the fly. In one example, an indicator for an input source may be signaled / predefined / derived on the fly in the APS.
[0295] In one example, Figures 15 to 18 The input sources of the ALF filter represented in can be arranged in order as follows. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group corresponding to the filter group indicator transmitted by the signal and trained offline can be applied as the input source A. In one example, the intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-1 and the filter group corresponding to the filter group indicator transmitted by the signal and trained offline can be applied as the input source B. In one example, the reconstruction before DBF of the current picture can be applied as the input source C. In one example, the intermediate filtering result generated by feeding the reconstruction before DBF to the classifier-0 and the filter group-0 trained offline can be applied as the input source D. In one example, the intermediate filtering result generated by feeding the reconstruction before DBF to the classifier-0 and the filter group-1 trained offline can be applied as the input source E. In one example, the coded image named reference image 0 inside the forward reference picture list (reference list 0) can be applied as the input source D. In one example, a coded image named reference image 0 inside the backward reference picture list (reference list 1) may be applied as an input source E. In one example, an intermediate filtering result generated by feeding the reconstruction before ALF into classifier-0 and filter bank-0 trained offline may be applied as an input source D. In one example, an intermediate filtering result generated by feeding the reconstruction before ALF into classifier-0 and filter bank-1 trained offline may be applied as an input source E. In one example, an intermediate filtering result generated by feeding the reconstruction before ALF into classifier-0 and a filter bank trained offline corresponding to the opposite indicator of the filter bank indicator transmitted by the signal may be applied as an input source D. In one example, an intermediate filtering result generated by feeding the reconstruction before ALF into classifier-1 and a filter bank trained offline corresponding to the opposite indicator of the filter bank indicator transmitted by the signal may be applied as an input source E. In one example, the total number of extension taps inside the ALF filter may be jointly derived based on the shape, filter length, and symmetry constraints.
[0296] Example 11
[0297] In one example, whether and / or how to apply at least one extended tap in the ALF may depend on the classification in the ALF. In one example, the classification may be based on the gradient information of the input source. In one example, the input source may be the reconstruction before the ALF. In one example, the input source may be the reconstruction before the DBF. In one example, the input source may be the intermediate filtering result generated by the filter trained offline and the reconstruction before the ALF. In one example, the input source may be the intermediate filtering result generated by the filter trained offline and the reconstruction before the DBF. In one example, the classification may be based on the frequency band information of the input source. In one example, the input source may be the reconstruction before the ALF. In one example, the input source may be the reconstruction before the DBF. In one example, the input source may be the intermediate filtering result generated by the filter trained offline and the reconstruction before the ALF. In one example, the input source may be the intermediate filtering result generated by the filter trained offline and the reconstruction before the DBF.
[0298] Example 12
[0299] In one example, the input samples of the extended tap may be padded between picture / strip / slice / CTU boundaries. In one example, padding may be applied to any boundary during the encoding / decoding process. In one example, padding may be applied to picture / sub-picture boundaries. In one example, padding may be applied to strip / slice boundaries. In one example, padding may be applied to CTU / CTB boundaries. In one example, padding may be applied to CU / TU / PU boundaries. In one example, padding may be applied to block boundaries. In one example, padding may be applied to unit boundaries. In one example, padding may be applied to virtual boundaries. In one example, padding may be applied to any other type of boundary. In one example, different padding methods may be applied to boundaries. In one example, extended padding may be applied. In one example, mirror padding may be applied. In one example, repeated padding may be applied. In one example, other padding methods may be applied.
[0300] Example 13
[0301] In one example, a first syntax element may be signaled to indicate whether a filter having at least one extended tap is enabled. In one example, the first syntax element may be encoded and decoded by arithmetic coding and decoding. In one example, the first syntax element may be encoded and decoded with at least one context. The context may depend on the coding and decoding information of the current block or the neighboring block. The context may depend on the filter shape of at least one neighboring block. In one example, the first syntax element may be encoded and decoded with bypass coding and decoding. In one example, the first syntax element may be binarized by a unary code, or a truncated unary code, or a fixed length code, or an exponential Golomb code, a truncated exponential Golomb code, etc. In one example, the first syntax element may be conditionally signaled. For example, the first syntax element may be signaled only when the extended tap is available. The first syntax element may be encoded and decoded in a predictive manner. The first syntax element may be predicted by an on / off decision of the extended tap of at least one neighboring block. The first syntax element may be independently signaled for different color components. In one example, the first syntax element may be signaled and shared for different color components. In one example, a first syntax element may be signaled for a first color component, but not for a second color component.The syntax element may be signaled in an SPS / PPS / picture header / slice header / APS / CTU / CU / etc.
[0302] Example 14
[0303] In one example, the first syntax element may be transmitted by signal to indicate which / what input sources are used for the extended taps inside the ALF filter. In one example, the first syntax element may be encoded and decoded by arithmetic coding and decoding. In one example, the first syntax element may be encoded and decoded with at least one context. The context may depend on the coding and decoding information of the current block or the adjacent block. The context may depend on the filtering shape of at least one adjacent block. In one example, the first syntax element may be encoded and decoded with bypass coding and decoding. In one example, the first syntax element may be binarized by a unary code, or a truncated unary code, or a fixed length code, or an exponential Golomb code, a truncated exponential Golomb code, etc. In one example, the first syntax element may be conditionally transmitted by signal. For example, the first syntax element may be transmitted by signal only when the extended tap is available. The first syntax element may be encoded and decoded in a predictive manner. The first syntax element may be predicted by the on / off decision of the extended tap of at least one adjacent block. The first syntax element may be independently transmitted by signal for different color components. In one example, the first syntax element may be transmitted by signal and shared for different color components. In one example, a first syntax element may be signaled for a first color component, but not for a second color component. The syntax element may be signaled in an SPS / PPS / picture header / slice header / APS / CTU / CU / etc. In one example, the syntax element may be signaled in an APS for a signaled ALF filter.
[0304] Example 15
[0305] The coefficients of at least one extended tap inside the ALF filter may be signaled in a syntax element structure such as an APS. In one example, the coefficients of the extended taps may be included in the APS. In one example, the clipping parameters of the extended taps may be included in the APS. In one example, the class merging results of the extended taps may be included in the APS. In one example, the coefficients of the extended taps may be encoded and decoded in a predictive manner. In one example, the coefficients of the extended taps may be encoded and decoded using arithmetic coding and decoding using at least one context. In one example, the coefficients of the extended taps may be encoded and decoded using bypass coding and decoding. In one example, the coefficients of the extended taps may be jointly encoded and decoded with the coefficients of the spatial domain taps. In one example, other parameters of the extended taps may be included in the APS. In one example, the coefficients of the extended taps may be predefined fixed values.
[0306] Example 16
[0307] It is proposed to use the intermediate filtering results of the filter as the input of one or more extension taps. In one example, the intermediate filtering results of the ALF filter trained offline can be used as the input of one or more extension taps. In one example, the intermediate filtering results can be generated by the reconstruction before ALF and the filter trained offline of ALF. In one example, the intermediate filtering results can be generated by the reconstruction before DBF and the filter trained offline of ALF. In one example, the intermediate filtering results of the ALF filter trained online can be used as the input of one or more extension taps. In one example, the intermediate filtering results of other predefined filters can be used as the input of one or more extension taps.
[0308] In one example, a Gaussian filter may be applied. In one example, a bilateral filter may be applied. In one example, a guided filter may be applied. In one example, a median filter may be applied. In one example, a local median filter may be applied. In one example, a non-local median filter may be applied. In one example, a filter with a low-pass property may be applied. In one example, a filter with a high-pass property may be applied.
[0309] In one example, the intermediate filtering results of other online trained filters can be used as inputs to one or more extended taps. In one example, the inputs to the intermediate filtering can be reconstructed samples at different encoding and decoding stages. In one example, the reconstruction before / after the ALF of the current / reference frame can be used to generate the intermediate filtering results. In one example, the reconstruction before / after the SAO / cross-component sample adaptive offset (CCSAO) of the current / reference frame can be used to generate the intermediate filtering results. In one example, the reconstruction before / after the bilateral filter (BIF) of the current / reference frame can be used to generate the intermediate filtering results. In one example, the reconstruction before / after the DBF of the current / reference frame can be used to generate the intermediate filtering results. In one example, the reconstruction before / after any other stage of the current / reference frame can be used to generate the intermediate filtering results.
[0310] Example 17
[0311] It is proposed to use the reconstructed samples before / after different encoding and decoding stages of the current frame as the input of one or more extension taps. In one example, the reconstruction before / after the DBF of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after the SAO / CCSAO of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after the BIF of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after other stages of the current frame can be used as the input of one or more extension taps.
[0312] Example 18
[0313] It is proposed to use samples inside one or more coded pictures as input sources for one or more extension taps. In one example, the previously coded frame may be a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with the block / current slice / frame. In one example, the previously coded frame may be a short-term reference picture of the block / current slice / frame. In one example, the previously coded frame may be a long-term reference picture of the block / current slice / frame.
[0314] Example 19
[0315] In one example, the previously decoded frame may not be a reference frame, but it is still stored in a decoded picture buffer (DPB).
[0316] Example 20
[0317] In one example, at least one indicator is signaled to indicate which previously coded frame(s) to use. In one example, an indicator is signaled to indicate which reference picture list to use. In one example, at least one indicator may be signaled to indicate a reference index. In one example, the indicator may be conditionally signaled, e.g., depending on how many reference pictures are included in the RPL / RPS. In one example, the indicator may be conditionally signaled, e.g., depending on how many previously decoded pictures are included in the DPB.
[0318] Example 21
[0319] In one example, it is determined on the fly which frames to utilize. In one example, the extension tap may obtain information from one or more previously coded frames in the DPB. In one example, the extension tap may obtain information from one / more reference frames in list 0. In one example, the extension tap may obtain information from one / more reference frames in list 1. In one example, the extension tap may obtain information from reference frames in both list 0 and list 1. In one example, the extension tap may obtain information from a reference frame that is closest to the current frame (e.g., having the smallest picture order count (POC) distance to the current slice / frame).
[0320] In one example, the extended tap may obtain information from a reference frame in a reference list with a reference index equal to K (e.g., K=0). In one example, K may be predefined. In one example, K may be derived on the fly based on reference picture information. In one example, K may be transmitted via a signal.
[0321] In one example, the extended tap may obtain information from the co-located frame. In one example, which frame to utilize may be determined by decoding information. In one example, the frame to utilize may be defined as the first N (e.g., N=1) most frequently used reference pictures for the samples within the current slice / frame. In one example, the frame to utilize may be defined as the first N (e.g., N=1) most frequently used reference pictures for each reference picture list (if available) for the samples within the current slice / frame. In one example, the frame to utilize may be defined as the picture with the first N (e.g., N=1) smallest POC distances / absolute POC distances relative to the current picture.
[0322] Example 22
[0323] In one example, whether to obtain information from a previously coded frame may depend on decoded information (e.g., codec mode / statistics / characteristics) of at least one region of the block to be filtered. In one example, the encoder may signal to the decoder whether to obtain information from a previously coded frame of at least one region of the block to be filtered. In one example, whether to obtain information from a previously coded frame may depend on the slice / picture type. In one example, it may only apply to inter-coded slices / pictures (e.g., P or B slices / pictures).
[0324] In one example, whether to obtain information from a previously coded frame may depend on the availability of a reference picture. In one example, whether to obtain information from a previously coded frame may depend on reference picture information or picture information in the DPB. In one example, if the minimum POC distance (e.g., the minimum POC distance between a reference picture / picture in the DPB and the current picture) is greater than a threshold, it is disabled. In one example, whether to obtain information from a previously coded frame may depend on the temporal layer index and / or QP and / or the size of the picture. In one example, it may be applicable to blocks with a given temporal layer index (e.g., the highest temporal layer). In one example, if the block to be filtered contains a portion of samples coded in a non-inter-frame mode, the extended tap cannot use information from a previously coded frame to filter the block.
[0325] In one example, non-inter mode may be defined as intra mode. In one example, non-inter mode may be defined as a set of codec modes including but not limited to intra / IBC / palette modes. In one example, the distortion between the current block and the matching block is calculated and used to decide whether to obtain information from a previously coded frame to filter the current block.
[0326] In one example, the distortion between the co-located block in the previously coded frame and the current block may be used to decide whether to obtain information from the previously coded frame to filter the current block. In one example, motion estimation may be first used to find a matching block from at least one previously coded frame. In one example, when the distortion is greater than a predefined threshold, information from the previously coded frame may not be used.
[0327] Example 23
[0328] In one example, the information includes two reference blocks and / or co-located blocks for the current block, one of which is from the first reference frame in list-0 and the other is from the first reference frame in list-1.
[0329] Example 24
[0330] In one example, the disclosed method can be used for post-processing and / or pre-processing.
[0331] Example 25
[0332] In one example, the above mentioned methods may be used in combination.
[0333] Example 26
[0334] In one example, the above-mentioned methods can be used alone.
[0335] Example 27
[0336] In one example, the proposed / described extended taps for the ALF method may be applied to any loop filtering tool, pre-processing or post-processing filtering method in the video codec (including but not limited to ALF / cross-component adaptive loop filter (CCALF) or any other filtering method). In one example, the proposed extended tap method may be applied to the loop filtering method. In one example, the proposed extended tap method may be applied to ALF. In one example, the proposed extended tap method may be applied to CCALF. In one example, the proposed extended tap method may be applied to other loop filtering methods. In one example, the proposed extended tap method may be applied to a pre-processing filtering method. In one example, the proposed extended tap method may be applied to a post-processing filtering method.
[0337] Example 28
[0338] In the above examples, a video unit may refer to a sequence / picture / sub-picture / slice / slice / codec tree unit (CTU) / CTU row / CTU group / codec unit (CU) / prediction unit (PU) / transform unit (TU) / codec tree block (CTB) / codec block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luma or chroma sample / pixel.
[0339] Example 29
[0340] Whether and / or how the methods disclosed above are applied may be signaled in the bitstream. In one example, they may be signaled at a sequence level / picture group level / picture level / slice level / slice group level, such as in a sequence header, picture header, SPS, VPS, DPS, decoder capability information (DCI), PPS, APS, slice header, and slice group header. In one example, they may be signaled in a PB / TB / CB / PU / TU / CU / virtual pipe data unit (VPDU) / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel.
[0341] Example 30
[0342] Whether and / or how to apply the above disclosed method may depend on codec information such as block size, color format, single / dual tree partitioning, color component, slice / picture type.
[0343] Other Examples
[0344] It is proposed to use the intermediate filtering results of the filter as the input of one or more extension taps. In one example, the intermediate filtering results of the ALF filter trained offline can be used as the input of one or more extension taps. In one example, the intermediate filtering results can be generated by the reconstruction before ALF and the filter trained offline of ALF. In one example, the intermediate filtering results can be generated by the reconstruction before DBF and the filter trained offline of ALF.
[0345] In one example, intermediate filtering results of an online trained ALF filter may be used as input to one or more expansion taps.
[0346] In one example, the intermediate filtering results of other predefined filters can be used as input to one or more extended taps. In one example, a Gaussian filter can be applied. In one example, a bilateral filter can be applied. In one example, a guided filter can be applied. In one example, a median filter can be applied. In one example, a local filter can be applied. In one example, a non-local median filter can be applied. In one example, a filter with a low-pass property can be applied. In one example, a filter with a high-pass property can be applied.
[0347] In one example, intermediate filtering results of other online trained filters may be used as input to one or more expansion taps.
[0348] In one example, the input of the intermediate filtering can be the reconstructed samples at different encoding and decoding stages. In one example, the reconstruction before / after the ALF of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the SAO / CCSAO of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the BIF of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after the DBF of the current / reference frame can be used to generate the intermediate filtering result. In one example, the reconstruction before / after any other stage of the current / reference frame can be used to generate the intermediate filtering result.
[0349] It is proposed to use the reconstructed samples before / after different encoding and decoding stages of the current frame as the input of one or more extension taps. In one example, the reconstruction before / after the DBF of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after the SAO / CCSAO of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after the BIF of the current frame can be used as the input of one or more extension taps. In one example, the reconstruction before / after other stages of the current frame can be used as the input of one or more extension taps.
[0350] It is proposed to use samples inside one or more codec pictures as input sources of one or more extension taps.
[0351] In one example, the previously coded frame may be a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with the block / current slice / frame. In one example, the previously coded frame may be a short-term reference picture of the block / current slice / frame. In one example, the previously coded frame may be a long-term reference picture of the block / current slice / frame.
[0352] In one example, the previously decoded frame may not be a reference frame, but it is still stored in a decoded picture buffer (DPB).
[0353] In one example, at least one indicator is signaled to indicate which previously coded frame(s) to use. In one example, an indicator is signaled to indicate which reference picture list to use. In one example, at least one indicator may be signaled to indicate a reference index. In one example, the indicator may be conditionally signaled, e.g., depending on how many reference pictures are included in the RPL / RPS. In one example, the indicator may be conditionally signaled, e.g., depending on how many previously decoded pictures are included in the DPB.
[0354] In one example, it is determined on the fly which frames to utilize. In one example, the extension tap may obtain information from one or more previously coded frames in the DPB. In one example, the extension tap may obtain information from one / more reference frames in list 0. In one example, the extension tap may obtain information from one / more reference frames in list 1. In one example, the extension tap may obtain information from reference frames in both list 0 and list 1. In one example, the extension tap may obtain information from a reference frame closest to the current frame (e.g., having the smallest POC distance to the current slice / frame). In one example, the extension tap may obtain information from a reference frame with a reference index equal to K (e.g., K=0) in the reference list. In one example, K may be predefined. In one example, K may be derived on the fly from reference picture information. In one example, K may be transmitted by signal. In one example, the extension tap may obtain information from a co-located frame. In one example, it may be determined which frame to utilize by decoding information. In one example, the frame to be utilized may be defined as the most frequently used reference picture of the first N (e.g., N=1) of the samples in the current slice / frame. In one example, the frames to be utilized may be defined as the top N (e.g., N=1) most frequently used reference pictures for each reference picture list (if available) for samples within the current slice / frame. In one example, the frames to be utilized may be defined as pictures with the top N (e.g., N=1) smallest POC distances / absolute POC distances relative to the current picture.
[0355] In one example, whether to obtain information from a previously coded frame may depend on decoded information (e.g., codec mode / statistics / characteristics) of at least one region of the block to be filtered. In one example, the encoder may signal to the decoder whether to obtain information from a previously coded frame of at least one region of the block to be filtered. In one example, whether to obtain information from a previously coded frame may depend on the slice / picture type. In one example, it may only apply to inter-coded slices / pictures (e.g., P or B slices / pictures). In one example, whether to obtain information from a previously coded frame may depend on the availability of a reference picture. In one example, whether to obtain information from a previously coded frame may depend on reference picture information or picture information in the DPB. In one example, if the minimum POC distance (e.g., the minimum POC distance between a picture in a reference picture / DPB and a current picture) is greater than a threshold, it is disabled. In one example, whether to obtain information from a previously coded frame may depend on a temporal layer index and / or a QP and / or the size of a picture. In one example, it may apply to blocks with a given temporal layer index (e.g., the highest temporal layer). In one example, if the block to be filtered contains a portion of samples encoded in a non-inter-frame mode, the extended tap cannot use information from a previously encoded frame to filter the block. In one example, the non-inter-frame mode can be defined as an intra-frame mode. In one example, the non-inter-frame mode can be defined as a set of encoding and decoding modes including but not limited to intra-frame / IBC / palette mode. In one example, the distortion between the current block and the matching block is calculated and used to decide whether to obtain information from the previously encoded frame to filter the current block. In one example, the distortion between the co-located block in the previously encoded frame and the current block can be used to decide whether to obtain information from the previously encoded frame to filter the current block. In one example, motion estimation can be first used to find a matching block from at least one previously encoded frame. In one example, when the distortion is greater than a predefined threshold, information from the previously encoded frame cannot be used.
[0356] In one example, the information includes two reference blocks and / or co-located blocks for the current block, one of which is from the first reference frame in list-0 and the other is from the first reference frame in list-1.
[0357] Fig.194000 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed format or an encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or a cellular interface.
[0358] System 4000 may include a codec component 4004 that may implement various codecs or coding methods described in this document. Codec component 4004 may reduce the average bit rate of the video from the input 4002 to the codec component 4004 to the output to generate the codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 may be stored or sent via the connected communication, as shown in component 4006. The stored or transmitted bitstream representation (or codec representation) of the video received at input 4002 may be used by component 4008 to generate pixel values or displayable video, which is sent to display interface 4010. The process of generating user-visible video from bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "coding" operations or tools, it should be understood that coding tools or operations are used in encoders, and decoders will perform corresponding decoding tools or operations of the reverse process of coding.
[0359] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Display Port, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The techniques described in this document may be embodied in various electronic devices such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0360] Fig. 204106 is a block diagram of an example video processing device 4100. Device 4100 may be used to implement one or more methods described herein. Device 4100 may be implemented as a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor 4102 may be configured to implement one or more methods described in this document. Memory (memory) 4104 may be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 may be used to implement some of the techniques described in this document in hardware circuits. In some embodiments, video processing circuitry 4106 may be at least partially included in processor 4102 (e.g., a graphics coprocessor).
[0361] Fig.21 4200 is a flow chart of an example method 4200 for video processing. The method 4200 includes determining to apply an extension tap in an ALF at step 4202. A conversion between visual media data and a bitstream is performed based on the ALF at step 4204. According to an example, the conversion of step 4204 may include encoding at an encoder or decoding at a decoder.
[0362] It should be noted that the method 4200 may be implemented in a device for processing video data, the device comprising a processor and a non-transitory memory having instructions thereon, such as the video encoder 4400, the video decoder 4500, and / or the encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform the method 4200. In addition, the method 4200 may be performed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec device performs the method 4200.
[0363] Fig. 22 43 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, which may be referred to as a video encoding device. The target device 4320 may decode the encoded video data generated by the source device 4310, which may be referred to as a video decoding device.
[0364] Source device 4310 may include video source 4312, video encoder 4314 and input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider and / or a computer graphics system for generating video data, or a combination of such sources. Video data may include one or more pictures. Video encoder 4314 encodes video data from video source 4312 to generate a bitstream. The bitstream may include a series of bits that form a codec representation of video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be directly sent to target device 4320 via network 4330 via I / O interface 4316. The encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0365] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320, or may be outside the target device 4320, and the target device may be configured to be connected to an external display device through an interface.
[0366] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the HEVC standard, the VVC standard, and other current and / or further standards.
[0367] Fig.23 is a block diagram showing an example of a video encoder 4400, which may be Fig. 22 Video encoder 4314 in system 4300 shown. Video encoder 4400 can be configured to perform any or all of the techniques of the present disclosure. Video encoder 4400 includes multiple functional components. The techniques described in the present disclosure can be shared between the various components of video encoder 4400. In some examples, the processor can be configured to perform any or all of the techniques described in the present disclosure.
[0368] The functional components of the video encoder 4400 may include a segmentation unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a cache 4413 and an entropy coding unit 4414.
[0369] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is a picture where the current video block is located.
[0370] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated but are shown separately in the example of the video encoder 4400 for explanation purposes.
[0371] The segmentation unit 4401 may segment the picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0372] The mode selection unit 4403 may select one of the codec modes, for example, based on the error result, and provide the resulting intra or inter codec block to the residual generation unit 4407 to generate residual block data and to the reconstruction unit 4412 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 4403 may select a combined intra and inter prediction (CIIP) mode, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 4403 may also select a resolution of motion vectors (e.g., sub-pixel or integer pixel precision) for the block in the case of inter prediction.
[0373] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 4413 (not the picture associated with the current video block).
[0374] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0375] In some examples, the motion estimation unit 4404 may perform unidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures of list 0 or list 1. The motion estimation unit 4404 may then generate a reference index indicating a reference picture in list 0 or list 1, the reference picture containing the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0376] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block, and the motion estimation unit 4404 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in the reference pictures in list 1. Then, the motion estimation unit 4404 may generate a reference index indicating a reference picture in list 0 and list 1, the reference picture containing a reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0377] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in a decoding process by a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Rather, motion estimation unit 4404 may reference motion information of another video block to signal motion information of the current video block. For example, motion estimation unit 4404 may determine that motion information of the current video block is sufficiently similar to motion information of a neighboring video block.
[0378] In one example, the motion estimation unit 4404 may indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0379] In another example, the motion estimation unit 4404 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0380] As discussed above, the video encoder 4400 can signal the motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0381] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0382] The residual generating unit 4407 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0383] In other examples, such as in skip mode, there may be no residual data for the current video block of the current video block, and the residual generation unit 4407 may not perform a subtraction operation.
[0384] The transform processing unit 4408 may generate a transform coefficient video block for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0385] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0386] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.
[0387] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0388] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.
[0389] Fig.24 is a block diagram showing an example of a video decoder 4500, which may be Fig. 22 Video decoder 4324 in the system 4300 shown. Video decoder 4500 can be configured to perform any or all of the techniques of the present disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in the present disclosure can be shared between the various components of video decoder 4500. In some examples, the processor can be configured to perform any or all of the techniques described in the present disclosure.
[0390] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, and a reconstruction unit 4506 and a buffer 4507. In some examples, the video decoder 4500 may perform a decoding process that is substantially the inverse of the encoding process described with respect to the video encoder 4400.
[0391] The entropy decoding unit 4501 may take out the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 may decode the entropy-encoded video data, and the motion compensation unit 4502 may determine motion information based on the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.
[0392] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. The identification of the interpolation filter used with sub-pixel precision may be included in a syntax element.
[0393] The motion compensation unit 4502 may calculate interpolation of sub-integer pixels of a reference block using an interpolation filter used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to received syntax information and use the interpolation filter to generate a prediction block.
[0394] The motion compensation unit 4502 can use some syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the encoded video sequence.
[0395] The intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes, i.e., dequantizes, the video block coefficients provided in the bitstream and decoded and quantized by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0396] The reconstruction unit 4506 may add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter may also be applied to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction and also generates a decoded video for presentation on a display device.
[0397] Fig.25 4600 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602 that uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, and utilizing the side information of the codec to signal the offset and filter coefficients. The ALF 4606 is located at the last processing stage of each picture and can be seen as a tool that attempts to capture and repair artifacts created by previous stages.
[0398] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using a reference picture obtained from a reference picture cache 4612. The residual block from the inter prediction or intra prediction is fed to the transform (T) component 4614 and the quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed to the entropy coding component 4618. The entropy coding component 4618 entropy codes and decodes the prediction result and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized component output from the quantization component 4616 can be fed to the inverse quantization (IQ) component 4620, the inverse transform component 4622, and the reconstruction (REC) component 4624. The REC component 4624 is capable of outputting images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these pictures are stored in the reference picture cache 4612 .
[0399] A list of solutions preferred by some examples is provided next.
[0400] The following solutions illustrate examples of the techniques discussed herein.
[0401] 1. A method for processing video data, comprising: determining to apply an extended tap in an adaptive loop filter (ALF); and performing conversion between visual media data and a bitstream based on the ALF.
[0402] 2. The method as described in solution 1, wherein the ALF adopts a spatial tap that uses information of spatially adjacent samples of the target component, and wherein the extended tap is different from the spatial tap.
[0403] 3. The method as described in any one of solutions 1 to 2, wherein the extension tap and at least one spatial domain tap coexist inside the ALF.
[0404] 4. A method as described in any of solutions 1 to 3, wherein the expansion tap uses one or more input sources.
[0405] 5. A method as described in any one of solutions 1 to 4, wherein the use of the extension tap is different for different color formats or different color components.
[0406] 6. A method as described in any one of solutions 1 to 5, wherein the coefficients of the extended taps correspond to one or more input sample points.
[0407] 7. The method as described in any one of solutions 1 to 6, wherein the ALF adopts different shapes or sizes.
[0408] 8. A method as described in any of solutions 1 to 7, wherein the ALF includes different shapes for spatial domain taps and the extension taps.
[0409] 9. The method as described in any one of solutions 1 to 8, wherein the shape and size of one or more spatial domain taps used inside the ALF are different from the shape and size of the extension taps used inside the ALF.
[0410] 10. The method as described in any one of solutions 1 to 9, wherein the shape and size of one or more spatial domain taps used inside the ALF are the same as the shape and size of the extension taps used inside the ALF.
[0411] 11. A method as described in any one of solutions 1 to 10, wherein the filter shape of the spatial tap inside the ALF includes a diamond shape, a square shape, a cross shape, a symmetrical shape, an asymmetrical shape, or a combination of the above shapes.
[0412] 12. A method as described in any one of solutions 1 to 11, wherein the filter shape of the extended tap inside the ALF includes a diamond shape, a square shape, a cross shape, a symmetrical shape, an asymmetrical shape, or a combination of the above shapes.
[0413] 13. A method as described in any one of solutions 1 to 12, wherein a transformation based on geometric information is applied to the spatial domain taps or the extended taps in the ALF.
[0414] 14. A method as described in any one of solutions 1 to 13, wherein the extended tap is input from an input source, and wherein the input source includes reconstruction before ALF, reconstruction before DBF, a filtering result generated by reconstruction before ALF and an offline trained filter of ALF, an intermediate filtering result generated by reconstruction before deblocking filter (DBF) and an offline trained filter of ALF, an intermediate filtering result generated by feeding reconstruction before ALF to classifier-0 and an offline trained filter group corresponding to a filter group indicator transmitted by a signal, an intermediate filtering result generated by feeding reconstruction before ALF to classifier-1 and an offline trained filter group corresponding to a filter transmitted by a signal, a reconstruction before a decoded picture buffer (DBF) of a current picture, an intermediate filtering result generated by feeding reconstruction before DBF to classifier-0 and an offline trained filter group-0 A filtering result, an intermediate filtering result generated by feeding the reconstruction before DBF to the classifier-0 and the filter group-1 trained offline, a codec picture named reference picture 0 inside the reference list 0, a codec picture named reference picture 0 inside the reference list 1, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group-0 trained offline, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group-1 trained offline, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group corresponding to the indicator opposite to the filter group indicator transmitted by the signal, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-1 and the filter group trained offline corresponding to the indicator opposite to the filter group indicator transmitted by the signal, or a combination thereof.
[0415] 15. The method as described in any one of solutions 1 to 14, wherein application of the extended tap in the ALF depends on classification in the ALF, and wherein the classification is based on gradient information of the input source or frequency band information of the input source.
[0416] 16. The method as described in any one of solutions 1 to 15, wherein the input samples of the extended tap are padded between picture, slice, slice or codec tree unit (CTU) boundaries.
[0417] 17. The method as described in any one of solutions 1 to 16, wherein a first syntax element is transmitted by a signal to indicate whether the filter with the extended tap is enabled.
[0418] 18. The method as described in any of solutions 1 to 17, wherein a first syntax element is transmitted by signaling to indicate an input source for an extended tap inside the ALF filter.
[0419] 19. The method as described in any of solutions 1 to 18, wherein the extended taps inside the ALF filter are signaled in an adaptive parameter set (APS).
[0420] 20. A method as in any one of solutions 1 to 19, wherein an intermediate filtering result of the filter is used as input to the expansion tap.
[0421] 21. A method as described in any one of solutions 1 to 20, wherein an intermediate filtering result of an offline trained ALF, an online trained ALF, a predefined filter or other online trained filters is used as an input of the expansion tap.
[0422] 22. A method as described in any one of solutions 1 to 21, wherein an intermediate filtering result of an online trained ALF filter is used as an input of the expansion tap.
[0423] 23. The method as described in any one of solutions 1 to 22, wherein the input for the median filtering includes reconstructed samples at different encoding and decoding stages.
[0424] 24. The method as described in any one of solutions 1 to 23, wherein reconstructed samples at different encoding and decoding stages of the current frame are used as input of the expansion tap.
[0425] 25. The method as described in any one of solutions 1 to 24, wherein samples inside the coded picture are used as input sources of the extension taps.
[0426] 26. A method as described in any of solutions 1 to 25, wherein samples inside a reference frame in a reference picture list (RPL) or a reference picture set (RPS) associated with a current block, a current slice or a current frame are used as input sources for the extended tap.
[0427] 27. The method as described in any one of solutions 1 to 26, wherein samples inside the frame in the decoded picture buffer are used as input sources of the expansion tap.
[0428] 28. The method of any one of solutions 1 to 27, wherein an indicator is signaled to indicate a codec picture containing samples used as input source for the extended tap.
[0429] 29. The method of any one of solutions 1 to 28, wherein a codec picture contains samples used as an input source for the extended taps, and wherein the codec picture is dynamically determined.
[0430] 30. A method as described in any one of solutions 1 to 29, wherein whether to obtain information from a previously encoded and decoded frame for extending the tap depends on decoded information of at least one area of the block to be filtered.
[0431] 31. A method as described in any of solutions 1 to 30, wherein the information contains two reference blocks or co-located blocks for the current block, one block is from the first reference frame in list-0, and the other block is from the first reference frame in list-1.
[0432] 32. A device for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method as described in any one of solutions 1 to 31.
[0433] 33. A non-temporary computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-temporary computer-readable medium, so that when executed by a processor, the video codec device performs a method as described in any one of Solutions 1 to 31.
[0434] 34. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining to apply an extended tap in an adaptive loop filter (ALF); and generating the bitstream based on the determination.
[0435] 35. A method for storing a bitstream of a video, comprising: determining to apply an extended tap in an adaptive loop filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0436] 36. A method, apparatus or system as described in this patent document.
[0437] In the solution described herein, an encoder can comply with the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can parse syntax elements in the codec representation using the format rules and generate decoded video based on the presence and absence of syntax elements known according to the format rules.
[0438] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation (or vice versa). For example, the bitstream representation of a current video block may correspond to bits that are co-located or distributed at different locations in the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on an error residual value from a transform and encoding, and also encoded using bits in a header and other fields in the bitstream. In addition, during conversion, the decoder may parse the bitstream based on a determination and knowing that certain fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include certain syntax fields and generate the encoded representation accordingly by including or excluding the syntax fields in the encoded representation.
[0439] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuits or computer software, firmware or hardware or a combination of one or more thereof, which includes the structures disclosed in this document and their equivalents. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions, encoded on a computer-readable medium to be executed by a data processing device or to control the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that affects a machine-readable propagation signal, or a combination of one or more thereof. The term "data processing device" covers all devices, equipment and machines for processing data, including, for example, a programmable processor, a computer or multiple processors or computers. In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagation signal is an artificially generated signal, for example, a machine-generated electrical, optical or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0440] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or any other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store portions of one or more modules, subroutines, or code). A computer program may be deployed to execute on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0441] The processes or logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and devices may also be implemented as, special purpose logic circuits, such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs).
[0442] Processors suitable for executing computer programs include, for example, both general and special purpose microprocessors and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by or incorporated in dedicated logic circuits.
[0443] Although the present disclosure contains many details, these details should not be interpreted as limitations on the scope of any subject matter or content that may be claimed, but should be interpreted as descriptions of features of specific embodiments that may be specific to a particular technology. Certain features described in the context of separate embodiments in the present disclosure may also be implemented in combination in a single embodiment. Conversely, the individual features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any appropriate sub-combination. In addition, although the above may describe features as acting in certain combinations and even initially claiming protection, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve a sub-combination or a deformation of a sub-combination.
[0444] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order shown or in a sequential order or to perform all of the operations shown to achieve the desired results. In addition, the separation of various system components described in this disclosure should not be understood as requiring such separation in all embodiments.
[0445] Only a few implementations and examples are described, and other implementations, enhancements, and modifications may be made based on what is described and illustrated in this disclosure.
[0446] A first component is directly coupled to a second component when there are no intermediate components other than a line, trace, or another medium between the first and second components. A first component is indirectly coupled to a second component when there are intermediate components other than a line, trace, or another medium between the first and second components. The term "coupled" and variations thereof include direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" is meant to include a range of ±10% of the subsequent value.
[0447] Although several embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are considered to be illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0448] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems and methods described and shown as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques or methods. Other items shown or discussed as coupled may be directly connected, or may be indirectly coupled or communicated through some interface, device or intermediate component, whether electrically, mechanically or otherwise coupled or communicated. Other examples of changes, substitutions and alterations may be identified by those skilled in the art, and these changes, substitutions and alterations may be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, wherein include: Determining to apply extended taps in an adaptive loop filter (ALF); as well as Conversion between visual media data and a bit stream is performed based on the ALF.
2. The method of claim 1, wherein the ALF employs a spatial tap, the spatial tap utilizing information of spatially adjacent samples of a target component, and wherein the extended tap is different from the spatial tap.
3. The method according to any one of claims 1 to 2, wherein the spatial neighboring samples are from reconstruction before or after a deblocking filter (DBF), a sample adaptive offset (SAO) filter and / or a bilateral filter (BF).
4. The method of any one of claims 1 to 3, wherein the spatial neighboring samples are spatial neighboring luma samples used by the ALF to filter a central luma sample, or wherein the spatial neighboring samples are spatial neighboring chroma samples used by the ALF to filter a central chroma sample.
5. The method according to any one of claims 1 to 4, wherein the extension tap and the at least one spatial domain tap coexist inside the ALF.
6. The method of any one of claims 1 to 5, wherein the ALF comprises M spatial domain taps and N extension taps, wherein both M and N are greater than zero.
7. The method of any one of claims 1 to 6, wherein the expansion taps use one or more input sources.
8. The method of any one of claims 1 to 7, wherein the input source is a reconstruction before DBF, or an intermediate filtering result of a predefined filter, or a reconstruction before SAO or BF.
9. The method according to any one of claims 1 to 8, wherein the input source is a reconstruction before DBF and an intermediate filtering result of a predefined filter.
10. The method of any one of claims 1 to 9, wherein the input source is also used by other taps inside the ALF.
11. The method of any one of claims 1 to 9, wherein the input source is not used by any other tap inside the ALF.
12. The method of any one of claims 1 to 11, wherein the input source comprises prediction samples, comprises filtered prediction samples, comprises a weighted sum or filtering of multiple sources, comprises the output of a function of multiple sources, is derived based on samples in different color components, or a combination of the above sources.
13. The method of any one of claims 1 to 12, wherein the use of the extension tap is different for different color formats or different color components.
14. The method of any one of claims 1 to 13, wherein the ALF having at least one extended tap is applied only to process a luminance component, or wherein the ALF having at least one extended tap is applied only to process one of two chrominance components, or wherein the ALF having at least one extended tap is applied to process the two chrominance components, or wherein the ALF having at least one extended tap is applied to process the luminance component and the two chrominance components.
15. The method of any one of claims 1 to 14, wherein the coefficients of the extended taps correspond to N input samples.
16. The method of any one of claims 1 to 15, wherein N=1.
17. The method of any one of claims 1 to 15, wherein N>1, and wherein the N input samples are designed in a symmetrical manner, or wherein the N input samples are designed in an asymmetrical manner.
18. The method of any one of claims 1 to 17, wherein the coefficients of the extended taps are shared by a plurality of input samples based on geometric information.
19. The method of any one of claims 1 to 18, wherein the plurality of input samples are located at symmetrical positions.
20. The method of any one of claims 1 to 19, wherein the ALFs with at least one extended tap use different shapes or sizes.
21. The method of any one of claims 1 to 20, wherein the ALF comprises different shapes for spatial domain taps and the extension taps.
22. The method of any one of claims 1 to 21, wherein a shape and size of one or more spatial domain taps used inside the ALF are different from a shape and size of the extension taps used inside the ALF.
23. The method of any one of claims 1 to 22, wherein a shape and size of one or more spatial domain taps used inside the ALF are the same as a shape and size of the extension tap used inside the ALF.
24. The method according to any one of claims 1 to 23, wherein the filter shape of the spatial domain tap inside the ALF comprises a diamond shape, a square shape, a cross shape, a symmetrical shape, an asymmetrical shape, or a combination of the above shapes.
25. The method according to any one of claims 1 to 24, wherein the filter shape of the extended tap inside the ALF comprises a diamond shape, a square shape, a cross shape, a symmetrical shape, an asymmetrical shape or a combination of the above shapes.
26. The method of any one of claims 1 to 25, wherein a transformation based on geometric information is applied to a spatial domain tap or the extended tap in the ALF.
27. A method as described in any one of claims 1 to 26, wherein the transformation based on geometric information is applied independently to one or more spatial domain taps, independently to one or more extended taps, or jointly to one or more spatial domain taps and one or more extended taps.
28. The method of any one of claims 1 to 27, wherein the input of the extended tap is from an input source, and wherein the input source comprises reconstruction before ALF, reconstruction before DBF, a filtering result generated by reconstruction before ALF and an offline trained filter of ALF, an intermediate filtering result generated by reconstruction before deblocking filter (DBF) and an offline trained filter of ALF, an intermediate filtering result generated by feeding reconstruction before ALF to classifier-0 and an offline trained filter bank corresponding to a filter bank indicator transmitted by a signal, an intermediate filtering result generated by feeding reconstruction before ALF to classifier-1 and an offline trained filter bank corresponding to a filter transmitted by a signal, a reconstruction before a decoded picture buffer (DBF) of a current picture, a reconstruction before DBF to classifier-0 and an offline trained filter bank-0 An intermediate filtering result, an intermediate filtering result generated by feeding the reconstruction before DBF to the classifier-0 and the filter group-1 trained offline, a codec picture named reference picture 0 inside the reference list 0, a codec picture named reference picture 0 inside the reference list 1, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group-0 trained offline, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group-1 trained offline, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-0 and the filter group trained offline corresponding to the opposite indicator of the filter group indicator transmitted by the signal, an intermediate filtering result generated by feeding the reconstruction before ALF to the classifier-1 and the filter group trained offline corresponding to the opposite indicator of the filter group indicator transmitted by the signal, or a combination thereof.
29. The method of any one of claims 1 to 28, wherein application of the extended tap in the ALF depends on classification in the ALF, and wherein classification is based on gradient information of an input source or frequency band information of an input source.
30. The method of any one of claims 1 to 29, wherein the input source comprises a reconstruction before ALF, a reconstruction before DBF, an intermediate filtering result generated by a reconstruction before ALF and an offline trained filter of ALF, an intermediate filtering result generated by a reconstruction before a deblocking filter (DBF) and an offline trained filter of ALF, or a combination of the above sources.
31. The method of any one of claims 1 to 30, wherein input samples of the extended taps are padded between picture, slice, slice or codec tree unit (CTU) boundaries.
32. The method of any one of claims 1 to 31, wherein padding is applied to a picture / sub-picture boundary, a slice / slice boundary, a CTU / codec tree block (CTB) boundary, a codec unit (CU) / transform unit (TU) / picture unit (PU) boundary, a block boundary, a unit boundary, or a virtual boundary.
33. The method of any one of claims 1 to 32, wherein different fills are applied to the borders, and wherein the fills include extended fills, mirror fills, repeated fills, or a combination of the above fill methods.
34. The method of any one of claims 1 to 33, wherein a first syntax element is signaled to indicate whether the filter with the extended taps is enabled.
35. The method of any one of claims 1 to 34, wherein a first syntax element is signaled to indicate an input source for the extended taps inside the ALF filter.
36. The method of any one of claims 1 to 35, wherein the first syntax element is encoded and decoded by arithmetic coding, encoded and decoded by bypass coding, encoded and decoded based on context information of a current block or a neighboring block, or a combination of the above coding and decoding methods.
37. The method of any one of claims 1 to 36, wherein the first syntax element is binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, or a truncated exponential Golomb code.
38. The method of any one of claims 1 to 37, wherein the first syntax element is signaled when the extended tap is available.
39. The method of any one of claims 1 to 38, wherein the first syntax element is predicted based on an on / off decision of an extended tap of at least one neighboring block.
40. The method of any one of claims 1 to 39, wherein the first syntax element is signaled independently for different color components, wherein the first syntax element is signaled and shared for different color components, or wherein the first syntax element is signaled for a first color component but not for a second color component.
41. The method of any one of claims 1 to 40, wherein the extended taps inside the ALF filter are signaled in an adaptive parameter set (APS).
42. The method of any one of claims 1 to 41, wherein the APS comprises coefficients of extended taps, clipping parameters of extended taps, class merging results of extended taps, or a combination thereof.
43. The method of any one of claims 1 to 42, wherein the coefficients of the extended taps are encoded in a predictive manner, encoded with an arithmetic codec and at least one context, encoded with a bypass codec, or jointly encoded with coefficients of the spatial domain taps.
44. A method as claimed in any one of claims 1 to 43, wherein an intermediate filtering result of a filter is used as an input to the expansion tap.
45. The method according to any one of claims 1 to 44, wherein an intermediate filtering result of an offline-trained ALF, an online-trained ALF, a predefined filter, or other online-trained filters is used as an input to the extended tap.
46. The method according to any one of claims 1 to 45, wherein the intermediate filtering result of the offline-trained filter of the ALF is generated by reconstruction before the ALF and the offline-trained filter of the ALF, or wherein the intermediate filtering result of the offline-trained ALF is generated by reconstruction before the DBF and the offline-trained filter of the ALF.
47. The method according to any one of claims 1 to 46, wherein an intermediate filtering result of an online-trained ALF filter is used as an input to the extended tap.
48. The method according to any one of claims 1 to 47, wherein the predefined filter includes a Gaussian filter.
49. The method according to any one of claims 1 to 48, wherein the predefined filter includes a bilateral filter, a guided filter, a median filter, a filter with low-pass property, or a filter with high-pass property.
50. The method according to any one of claims 1 to 49, wherein the input for intermediate filtering includes reconstructed samples at different coding and decoding stages.
51. The method according to any one of claims 1 to 50, wherein reconstructed samples before or after the ALF from the current frame or a reference frame are used to generate the intermediate filtering result, or reconstructed samples before or after the DBF from the current frame or a reference frame are used to generate the intermediate filtering result.
52. The method according to any one of claims 1 to 51, wherein reconstructed samples before or after the SAO / Cross-Component SAO (CCSAO) from the current frame or a reference frame are used to generate the intermediate filtering result, or reconstructed samples before or after the Bilateral Filtering (BIF) from the current frame or a reference frame are used to generate the intermediate filtering result.
53. The method according to any one of claims 1 to 52, wherein reconstructed samples at different coding and decoding stages of the current frame are used as an input to the extended tap.
54. The method according to any one of claims 1 to 53, wherein reconstructed samples before or after the DBF of the current frame are used as an input to the extended tap.
55. The method according to any one of claims 1 to 54, wherein reconstructed samples before or after the SAO / CCSAO of the current frame are used as an input to the extended tap, or wherein reconstructed samples before or after the BIF of the current frame are used as an input to the extended tap.
56. The method according to any one of claims 1 to 55, wherein samples inside the coded and decoded picture are used as an input source for the extended tap.
57. The method according to any one of claims 1 to 56, wherein samples inside a reference frame in a Reference Picture List (RPL) or a Reference Picture Set (RPS) associated with the current block, current stripe, or current frame are used as an input source for the extended tap.
58. The method of any one of claims 1 to 57, wherein samples inside a frame in a decoded picture buffer are used as input sources for the expansion taps.
59. The method of any one of claims 1 to 58, wherein an indicator is signaled to indicate a codec picture containing samples used as input source for the extended taps.
60. The method of any one of claims 1 to 59, wherein a codec picture contains samples used as an input source for the extended taps, and wherein the codec picture is dynamically determined.
61. The method of any one of claims 1 to 60, wherein the input to the expansion tap comprises information from one or more previously encoded and decoded frames in a decoded picture buffer (DPB).
62. The method of any one of claims 1 to 61, wherein whether to obtain information from a previously coded frame for extending the taps depends on decoded information of at least one region of the block to be filtered.
63. The method of any one of claims 1 to 62, wherein whether to obtain information from the previously encoded frame is based on a slice or picture type.
64. The method of any one of claims 1 to 63, wherein whether to obtain information from the previously coded frame is only applicable to inter-coded slices or pictures.
65. The method of any one of claims 1 to 64, wherein whether to obtain information from the previously encoded frame depends on the availability of one or more reference pictures.
66. The method of any one of claims 1 to 65, wherein the information contains two reference blocks or co-located blocks for the current block, wherein one block is from the first reference frame in list-0 and the other block is from the first reference frame in list-1.
67. The method of any one of claims 1 to 66, further comprising: include: Applying the extended tap in a loop filter is determined, and conversion between the visual media data and the bitstream is performed based on the loop filter.
68. The method of any one of claims 1 to 67, wherein the loop filter comprises a CCALF filter.
69. The method of any one of claims 1 to 68, wherein the converting comprises decoding the visual media data from the bitstream.
70. The method of any one of claims 1 to 68, wherein the converting comprises encoding the visual media data into the bitstream.
71. An apparatus for processing video data, include: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 70.
72. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method as described in any one of claims 1 to 70.
73. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein the method include: Determining to apply extended taps in an adaptive loop filter (ALF); as well as The bitstream is generated based on the determination.
74. A method for storing a bit stream of a video, include: Determining to apply extended taps in an adaptive loop filter (ALF); generating the bitstream based on the determination; as well as The bit stream is stored in a non-transitory computer-readable recording medium.