Conditional filter shape switching for adaptive loop filters in video coding
By introducing extended taps and conditional filter shape switching into the adaptive loop filter, the bandwidth usage and quality of video encoding and decoding are optimized, solving the problem of low efficiency in the prior art.
Patent Information
- Application Number
- CN202480026971.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-18
- Filing Date
- 2024-04-18
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to effectively utilize extended taps for adaptive loop filtering in video encoding and decoding, resulting in low bandwidth utilization efficiency.
An adaptive loop filter (ALF) is used with extended taps to filter samples based on spatial proximity information. This includes deblocking filter, sample adaptive compensation, cross-component SAO, and reconstructed samples after bilateral filtering. The filtering effect is optimized by switching the shape of the conditional filter.
It improves the bandwidth utilization efficiency of video encoding and decoding, enhances video quality, and reduces resource consumption during the encoding and decoding process.
Smart Images

Figure CN121002873A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2023 / 088914, filed April 18, 2023, which is incorporated by reference herein. TECHNICAL FIELD
[0003] The present disclosure relates to the generation, storage, and use of digital audio video media information in file formats. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use can continue to grow. SUMMARY
[0005] A first aspect relates to a method for processing video data, comprising: determining to use at least one extended tap in an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
[0006] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap is different from a spatial domain tap in the ALF.
[0007] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF uses only information of spatial domain neighboring samples of a target component, and wherein the spatial domain neighboring samples are neighboring a center sample to be filtered.
[0008] Optionally, in any of the above aspects, another embodiment of the aspect provides that the spatial domain neighboring samples are obtained from a reconstruction after a deblocking filter (DBF), a sample adaptive offset (SAO) filter, a cross component SAO (CCSAO), or a bilateral filter (BF).
[0009] Optionally, in any of the above aspects, another embodiment of the aspect provides that the spatial domain tap uses only spatial domain neighboring luma samples to filter a center luma sample within a single ALF.
[0010] Optionally, in any of the above aspects, another embodiment of the aspect provides that the spatial domain tap uses only spatial domain neighboring chroma samples to filter a center chroma sample within a single ALF.
[0011] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap is configured to coexist with at least one spatial tap in a single ALF.
[0012] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF consists of both spatial taps and extended taps.
[0013] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF consists of M spatial taps and N extended taps, where M and N are each positive integers.
[0014] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap uses one or more input sources.
[0015] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap uses a single input source.
[0016] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input source includes intermediate filtering results from one or more predefined filters.
[0017] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters include a Gaussian filter.
[0018] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters include a low-pass filter.
[0019] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters include a high-pass filter.
[0020] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input source includes a modification of the Gaussian filter, a modification of the low-pass filter, or a modification of the high-pass filter.
[0021] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input source includes a fusion or a weighted sum of the Gaussian filter, the low-pass filter, and the high-pass filter.
[0022] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap uses multiple input sources.
[0023] Optionally, in any of the above aspects, another embodiment of the aspect provides that the plurality of input sources comprises intermediate filtering results from one or more predefined filters.
[0024] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters comprises a Gaussian filter.
[0025] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters comprises a low-pass filter.
[0026] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more predefined filters comprises a high-pass filter.
[0027] Optionally, in any of the above aspects, another embodiment of the aspect provides that the plurality of input sources comprises a modification of the Gaussian filter, a modification of the low-pass filter, or a modification of the high-pass filter.
[0028] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input source comprises a fusion or a weighted sum of the Gaussian filter, the low-pass filter, and the high-pass filter.
[0029] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input source of the at least one extended tap is derived based on samples in different color components.
[0030] Optionally, in any of the above aspects, another embodiment of the aspect provides that whether and / or how the ALF with the at least one extended tap is applied is different for different color formats and / or different color components.
[0031] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF with the at least one extended tap is applied only for processing a luma component.
[0032] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF with the at least one extended tap is applied only for processing one chroma component, and wherein the one chroma component comprises a blue-difference chroma component (Cb) or a red-difference chroma component (Cr).
[0033] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF with the at least one extended tap is applied to process all chroma components, and wherein all chroma components include a blue-difference chroma component (Cb) and a red-difference chroma component (Cr).
[0034] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF with the at least one extended tap is applied to process a luma component and chroma components, and wherein the luma component and the chroma components include a luma component (Y), a blue-difference chroma component (Cb), and a red-difference chroma component (Cr).
[0035] Optionally, in any of the above aspects, another embodiment of the aspect provides that the ALF with the at least one extended tap uses a different shape or size than an ALF filter without an extended tap.
[0036] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one extended tap includes different shapes for spatial taps and extended taps.
[0037] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape and / or size used by the ALF for one or more spatial taps is different than the shape and / or size used by the ALF for an extended tap.
[0038] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape and / or size used by the ALF for one or more spatial taps is the same as the shape and / or size used by the ALF for an extended tap.
[0039] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is diamond-shaped.
[0040] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is square-shaped.
[0041] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is cross-shaped.
[0042] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a symmetric shape.
[0043] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is an asymmetric shape.
[0044] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a pre-designed shape.
[0045] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is determined in real-time, specified in the bitstream, or derived.
[0046] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for the one or more extended taps is a diamond shape.
[0047] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a square shape.
[0048] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a cross shape.
[0049] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a symmetric shape.
[0050] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is an asymmetric shape.
[0051] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is a pre-designed shape.
[0052] Optionally, in any of the above aspects, another embodiment of the aspect provides that the filter shape used by the ALF for one or more spatial taps is determined in real-time, specified in the bitstream, or derived.
[0053] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that a filter size used by the ALF for the one or more extended taps is MxN, where M and N are each positive integers.
[0054] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that M is equal to N.
[0055] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that M or N is equal to 1.
[0056] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that M or N is equal to 3.
[0057] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that M or N is equal to 5.
[0058] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that M is not equal to N.
[0059] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the bitstream includes a first syntax element to indicate whether the ALF with the at least one extended tap is enabled.
[0060] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the first syntax element is coded by arithmetic coding.
[0061] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the first syntax element is coded with at least one context.
[0062] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the at least one context depends on coding information of a current block or a neighboring block.
[0063] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the at least one context depends on a filter shape of at least one neighboring block.
[0064] Optionally, in any of the preceding aspects, another embodiment of the aspect provides that the first syntax element is coded by bypass coding.
[0065] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, or a truncated exponential Golomb code.
[0066] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is conditionally included in the bitstream.
[0067] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the bitstream only when the at least one extended tap is available.
[0068] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is coded in a predictive manner.
[0069] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is predicted based on whether an extended tap of at least one neighboring block is on or off.
[0070] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is independently included in the bitstream for different color components.
[0071] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is shared for different color components.
[0072] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the bitstream for a first color component but not for a second color component.
[0073] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the bitstream in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptation parameter set (APS), a coding tree unit (CTU), or a coding unit (CU).
[0074] Optionally, in any of the above aspects, another embodiment of the aspect provides that the bitstream includes a first syntax element to indicate which input sources are used for the at least one extended tap in the ALF.
[0075] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is coded by arithmetic coding.
[0076] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is coded with at least one context.
[0077] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one context depends on coding information of the current block or a neighboring block.
[0078] Optionally, in any of the above aspects, another embodiment of the aspect provides that the at least one context depends on a filter shape of at least one neighboring block.
[0079] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is coded by bypass coding.
[0080] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is binarized by unary code, truncated unary code, fixed length code, exponential Golomb code, or truncated exponential Golomb code.
[0081] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is conditionally included in the bitstream.
[0082] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the bitstream only when the at least one extended tap is available.
[0083] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is coded in a predictive manner.
[0084] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is predicted based on whether an extended tap of at least one neighboring block is on or off.
[0085] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is independently included in the bitstream for different color components.
[0086] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is shared for different color components.
[0087] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the bitstream for a first color component but not for a second color component.
[0088] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in a bitstream in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptation parameter set (APS), a coding tree unit (CTU), or a coding unit (CU).
[0089] Optionally, in any of the above aspects, another embodiment of the aspect provides that the first syntax element is included in the APS in the bitstream.
[0090] Optionally, in any of the above aspects, another embodiment of the aspect provides that one or more coefficients of the at least one extended tap of the ALF are included in a syntax element structure in the bitstream.
[0091] Optionally, in any of the above aspects, another embodiment of the aspect provides that the syntax element structure comprises an adaptation parameter set (APS).
[0092] Optionally, in any of the above aspects, another embodiment of the aspect provides that the APS comprises a clipping parameter of the at least one extended tap.
[0093] Optionally, in any of the above aspects, another embodiment of the aspect provides that the APS comprises a category merge result of the at least one extended tap.
[0094] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more coefficients are coded in a predictive manner.
[0095] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more coefficients are coded using arithmetic coding with at least one context.
[0096] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more coefficients are coded using bypass coding.
[0097] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more coefficients are jointly coded with coefficients of a spatial domain tap.
[0098] Optionally, in any of the above aspects, another embodiment of the aspect provides that the APS includes an additional parameter of the extended taps.
[0099] Optionally, in any of the above aspects, another embodiment of the aspect provides that the one or more coefficients include a pre-defined fixed value.
[0100] Optionally, in any of the above aspects, another embodiment of the aspect provides that an intermediate filtering result of at least one fixed filter or an intermediate filtering result of at least one adaptive filter is used as an input for the at least one extended tap.
[0101] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is from an offline trained filter.
[0102] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is generated from a reconstruction before the ALF and an offline trained filter of the ALF.
[0103] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is generated from a reconstruction before the ALF and a de-blocking filter (DBF).
[0104] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is from an online trained ALF filter.
[0105] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is from a pre-defined filter.
[0106] Optionally, in any of the above aspects, another embodiment of the aspect provides that the pre-defined filter is a Gaussian filter.
[0107] Optionally, in any of the above aspects, another embodiment of the aspect provides that the pre-defined filter is a bilateral filter.
[0108] Optionally, in any of the above aspects, another embodiment of the aspect provides that the pre-defined filter is a guided filter.
[0109] Optionally, in any of the above aspects, another embodiment of the aspect provides that the pre-defined filter is a median filter.
[0110] Optionally, in any of the above aspects, another embodiment of the aspect provides that the median filter is a local median filter.
[0111] Optionally, in any of the above aspects, another embodiment of the aspect provides that the median filter is a non-local median filter.
[0112] Optionally, in any of the above aspects, another embodiment of the aspect provides that the median filter is a filter with low-pass properties.
[0113] Optionally, in any of the above aspects, another embodiment of the aspect provides that the median filter is a filter with high-pass properties.
[0114] Optionally, in any of the above aspects, another embodiment of the aspect provides that the intermediate filtering result is from an online trained filter.
[0115] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input for generating the intermediate filtering result includes reconstructed samples at different coding stages.
[0116] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input for generating the intermediate filtering result includes reconstructed samples before the ALF for the current frame, reconstructed samples after the ALF for the current frame, reconstructed samples before the ALF for the reference frame, or reconstructed samples after the ALF for the reference frame.
[0117] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input for generating the intermediate filtering result includes reconstructed samples before a sample adaptive offset (SAO) filter for the current frame, reconstructed samples after the SAO filter for the current frame, reconstructed samples before a cross component SAO (CCSAO) filter for the reference frame, or reconstructed samples after the CCSAO filter for the reference frame.
[0118] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input for generating the intermediate filtering result includes reconstructed samples before a bilateral filter (BF) for the current frame, reconstructed samples after the BF for the current frame, reconstructed samples before the BF for the reference frame, or reconstructed samples after the BF for the reference frame.
[0119] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input used to generate the intermediate filter result comprises reconstructed samples before a deblocking filter (DBF) for the current frame, reconstructed samples after the DBF for the current frame, reconstructed samples before the DBF for the reference frame, or reconstructed samples after the DBF for the reference frame.
[0120] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input used to generate the intermediate filter result comprises reconstructed samples before a filter stage for the current frame, reconstructed samples after the filter stage for the current frame, reconstructed samples before the filter stage for the reference frame, or reconstructed samples after the filter stage for the reference frame.
[0121] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is applied to the ALF.
[0122] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is applied only to a spatial domain shape.
[0123] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is applied only to a shape of the at least one extended tap.
[0124] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 1x1 diamond.
[0125] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 3x3 diamond.
[0126] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 5x5 diamond.
[0127] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 5x5 cross.
[0128] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 7x7 cross.
[0129] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a 9x9 cross.
[0130] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape is a diamond or cross shape with a size of N, where N is a positive integer.
[0131] Optionally, in any of the above aspects, another embodiment of the aspect provides that N is 1, 3, 5, 7, 9, 11, or 13.
[0132] Optionally, in any of the above aspects, another embodiment of the aspect provides that N is 0 to indicate that the at least one extended tap is disabled.
[0133] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is jointly applied to a spatial domain shape and a shape of the at least one extended tap.
[0134] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is based on a configuration setting.
[0135] Optionally, in any of the above aspects, another embodiment of the aspect provides that an all-intra configuration setting is different from other configuration settings.
[0136] Optionally, in any of the above aspects, another embodiment of the aspect provides that a random access configuration setting uses a different spatial domain shape than other configuration settings.
[0137] Optionally, in any of the above aspects, another embodiment of the aspect provides that a low-delay bi-directional (B) / uni-directional (P) configuration setting uses a different spatial domain shape than other configuration settings.
[0138] Optionally, in any of the above aspects, another embodiment of the aspect provides that a random access configuration setting and a low-delay bi-directional (B) / uni-directional (P) configuration setting use the same shape.
[0139] Optionally, in any of the above aspects, another embodiment of the aspect provides that an all-intra configuration setting uses a different shape of the at least one extended tap than other configuration settings.
[0140] Optionally, in any of the above aspects, another embodiment of the aspect provides that a random access configuration setting uses a different shape of the at least one extended tap than other configuration settings.
[0141] Optionally, in any of the above aspects, another embodiment of the aspect provides that a low-delay bi-directional (B) / uni-directional (P) configuration setting uses a different shape of the at least one extended tap than other configuration settings.
[0142] Optionally, in any of the above aspects, another embodiment of the aspect provides that a random access configuration setting and a low delay bi-directional (B) / unidirectional (P) configuration setting use a same shape of the at least one extended tap.
[0143] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is based on a slice type.
[0144] Optionally, in any of the above aspects, another embodiment of the aspect provides that the slice type is an I slice, and wherein the I slice uses a different spatial domain shape than other slice types.
[0145] Optionally, in any of the above aspects, another embodiment of the aspect provides that the slice type is a B slice, and wherein the B slice uses a different spatial domain shape than other slice types.
[0146] Optionally, in any of the above aspects, another embodiment of the aspect provides that the slice type is a P slice, and wherein the P slice uses a different spatial domain shape than other slice types.
[0147] Optionally, in any of the above aspects, another embodiment of the aspect provides that a B slice and a P slice use a same spatial domain shape.
[0148] Optionally, in any of the above aspects, another embodiment of the aspect provides that an I slice uses a different shape of the at least one extended tap than other slice types.
[0149] Optionally, in any of the above aspects, another embodiment of the aspect provides that a B slice uses a different shape of the at least one extended tap than other slice types.
[0150] Optionally, in any of the above aspects, another embodiment of the aspect provides that a P slice uses a different shape of the at least one extended tap than other slice types.
[0151] Optionally, in any of the above aspects, another embodiment of the aspect provides that a B slice and a P slice use a same spatial domain shape of the at least one extended tap.
[0152] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is based only on a configuration setting or a slice type.
[0153] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape of the at least one extended tap in the full frame intra configuration setting is 5x5 cross shape and 1x1 diamond shape in the random access and low delay bi-directional (B) / uni-directional (P) configuration setting.
[0154] Optionally, in any of the above aspects, another embodiment of the aspect provides that the shape of the at least one extended tap is 5x5 cross shape for I slice and 1x1 diamond shape for B slice or P slice.
[0155] Optionally, in any of the above aspects, another embodiment of the aspect provides that the conditional filter shape switching is based on the configuration setting and the slice type jointly.
[0156] Optionally, in any of the above aspects, another embodiment of the aspect provides that the decision for the conditional filter shape switching is predefined, included in the bitstream or derived in real time.
[0157] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is used in post-processing and / or pre-processing.
[0158] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is used jointly.
[0159] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is used individually.
[0160] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is applied to any in-loop filtering tool, any pre-processing filtering method or post-processing filtering method in video coding, including the ALF or cross component ALF (CCALF).
[0161] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is applied to the in-loop filtering method including the ALF or CCALF.
[0162] Optionally, in any of the above aspects, another embodiment of the aspect provides that the method is applied to a video unit, and wherein the video unit comprises a sequence, a picture, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), any other region containing more than one luma or chroma sample or pixel.
[0163] Optionally, in any of the above aspects, another embodiment of the aspect provides that whether and / or how the method is applied is signaled in the bitstream.
[0164] Optionally, in any of the above aspects, another embodiment of the aspect provides that the signaling is performed at a sequence level, a picture group level, a picture level, a slice level, a tile group level, or in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), a decoder capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.
[0165] Optionally, in any of the above aspects, another embodiment of the aspect provides that the signaling is performed in a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a sub-picture, or other region containing more than one sample or pixel.
[0166] Optionally, in any of the above aspects, another embodiment of the aspect provides that whether and / or how the method is applied depends on coded information, and wherein the coded information comprises a block size, a color format, a single tree partitioning, a dual tree partitioning, a color component, a slice type, or a picture type.
[0167] Optionally, in any of the above aspects, another embodiment of the aspect provides that the converting comprises encoding the media data into a bitstream.
[0168] Optionally, in any of the above aspects, another embodiment of the aspect provides that the converting comprises decoding the media data from a bitstream.
[0169] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any disclosed embodiment.
[0170] A third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to perform the method of any disclosed embodiment.
[0171] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises: determining to use at least one extended tap in an adaptive loop filter (ALF); and generating the bitstream based on the determination.
[0172] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining to use at least one extended tap in an adaptive loop filter (ALF); generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0173] A sixth aspect relates to the methods, apparatuses, or systems described in this disclosure.
[0174] For the sake of clarity, any of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0175] These and other features will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF DRAWINGS
[0176] For a more complete understanding of the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and specific embodiments, wherein like reference numerals represent like parts.
[0177] Figure 1 An example of nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture is shown.
[0178] Figure 2 An example encoder block diagram is shown.
[0179] Figure 3 An example picture partitioned into raster scan slices is shown.
[0180] Figure 4 An example picture partitioned into rectangular scan slices is shown.
[0181] Figure 5 An example picture split into tiles is shown.
[0182] Figures 6A-6C An example of a CTB across a picture boundary is shown.
[0183] Figure 7 An example of an intra prediction mode is shown.
[0184] Figure 8 An example of a block boundary in a picture is shown.
[0185] Figure 9 An example of a pixel involved in filter usage is shown.
[0186] Figure 10 An example of a filter shape for ALF is shown.
[0187] Figure 11 An example of transform coefficients supported by a 5x5 diamond filter is shown.
[0188] Figure 12 An example of relative coordinates supported by a 5x5 diamond filter is shown.
[0189] Figure 13 is a block diagram illustrating an example video processing system.
[0190] Figure 14 is a block diagram of an example video processing device.
[0191] Figure 15 is a flowchart of an example method of video processing.
[0192] Figure 16 is a block diagram illustrating an example video coding system.
[0193] Figure 17 is a block diagram illustrating an example encoder.
[0194] Figure 18 is a block diagram illustrating an example decoder.
[0195] Figure 19 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0196] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or in existence. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations set forth herein, but can be modified in various ways within the scope of the appended claims along with their full scope of equivalents.
[0197] The use of section headings in this disclosure is for convenience only and is not to be construed as limiting the techniques and applicability of the techniques disclosed in each section to only the section in which it is disclosed. Further, the embodiments described herein are applicable to other video codec protocols and designs.
[0198] 1. Preliminary Discussion
[0199] The present disclosure relates to video coding techniques. In particular, it relates to loop filters and other coding tools in image / video coding. These ideas can be applied to video codecs such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video coding technologies, alone or in various combinations.
[0200] 2. Abbreviations
[0201] The disclosure includes the following acronyms. Advanced Video Coding (Recommendation ITU-T H.264 | ISO / IEC 14496-10) (AVC), coded picture buffer (CPB), clean random access (CRA), coding tree unit (CTU), coded video sequence (CVS), decoded picture buffer (DPB), decoding parameter set (DPS), general constraint information (GCI), High Efficiency Video Coding, also known as Recommendation ITU-T H.265 | ISO / IEC 23008-2, (HEVC), joint exploration model (JEM), motion constrained tile set (MCTS), network abstraction layer (NAL), output layer set (OLS), picture header (PH), picture parameter set (PPS), profile, tier, and level (PTL), picture unit (PU), reference picture resampling (RPR), raw byte sequence payload (RBSP), supplemental enhancement information (SEI), slice header (SH), sequence parameter set (SPS), video coding layer (VCL), video parameter set (VPS), Versatile Video Coding, also known as Recommendation ITU-T H.266 | ISO / IEC 23090-3, (VVC), VVC test model (VTM), video usability information (VUI), transform unit (TU), coding unit (CU), de-blocking filter (DF), sample adaptive offset (SAO), adaptive loop filter (ALF), coded block flag (CBF), quantization parameter (QP), rate-distortion optimization (RDO), and bilateral filter (BF).
[0202] 3. Video coding standard
[0203] Video coding standards have evolved primarily through the development of international telecommunication union (ITU) - telecommunication standardization sector (ITU-T) and international organization for standardization (ISO) / international electrotechnical commission (IEC) standards. The ITU-T produced H.261 and H.263 standards, the ISO / IEC produced the motion picture expert group (MPEG)-1 and MPEG-4 visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 advanced video coding (AVC) and H.265 / HEVC [1] standards. Starting with H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction is combined with transform coding. To explore future video coding technologies beyond HEVC, the joint video exploration team (JVET) was formed by the video coding experts group (VCEG) and MPEG jointly. The JVET adopted many methods and incorporated them into a reference software named joint exploration model (JEM). When the versatile video coding (VVC) project was officially launched, the JVET was renamed as joint video experts team (JVET). VVC is a coding standard, aiming to reduce 50% bit rate compared to HEVC. The working draft of VVC and VVC test model (VTM) are constantly updated.
[0204] An example version of the VVC draft, namely Versatile Video Coding (Draft 10) can be found at: https: / / jvet-experts.org / doc_end_user / documents / 19_Teleconference / wgl l / JVET-S2001-vl7.zip. An example version of the VVC reference software named VTM can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet-u-ee2 / VVCSoftware_VTM / - / tree / VTM-11.2.
[0205] The Joint Technical Committee (JTC) 1 / Subcommittee (SC) 29 / Working Group (WG) 11 of the Video Coding Experts Group (VCEG) of the Telecommunication Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization and the International Electrotechnical Commission (ISO / IEC) is studying the potential need for standardization of future video coding technology with compression capabilities significantly over current VVC standard. Such future standardization effort can take the form of extension(s) of VVC or a completely new standard. These groups are jointly conducting this study activity as a Joint Exploration Testbed (JET) to evaluate compression technology designs proposed by experts in the field. The JET established a first Exploration Experiment (EE) and uses a reference software named Enhanced Compression Mode (ECM). The test model ECM is continuously updated.
[0206] 3.1 Color Space and Chroma Downsampling
[0207] A color space, also called a color model (or color system), is a mathematical model that describes a range of colors as digital tuples, e.g. 3 or 4 values or color components (e.g. RGB). Generally, a color space is a refinement of a coordinate system and a subspace. For video compression, the most commonly used color spaces are luma, blue-difference chroma and red-difference chroma (YCbCr) and red, green, blue (RGB).
[0208] YCbCr, Y'CbCr or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, CB and CR are the blue-difference chroma component and the red-difference chroma component. Y' (with the prime) distinguishes from Y, Y is the luminance, meaning that the light intensity is non-linearly encoded based on the gamma-corrected RGB primaries.
[0209] Chroma downsampling is the practice of encoding an image with lower resolution for chroma information than for luma information, exploiting the fact that the human visual system is less acute in its discrimination of color differences than in its discrimination of luminance. 3.1.1 4:4:4
[0211] In 4:4:4, each of the Y'CbCr three components has the same sampling rate. There is thus no chroma downsampling. This scheme is sometimes used for high-end film scanners and movie post-production. 3.1.2 4:2:2
[0213] In 4:2:2, the two chroma components are sampled at half the sampling rate of the luma. The horizontal chroma resolution is halved, while the vertical chroma resolution remains the same. This reduces the bandwidth of the uncompressed video signal by a factor of three, with little visual difference. Figure 1 An example of the nominal vertical and horizontal positions of the 4:2:2 color format is shown. 3.1.3 4:2:0
[0215] In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved since in this scheme the Cb and Cr channels are only sampled on every alternate line. The data rate is therefore the same. Cb and Cr are downsampled by a factor of 2 in both the horizontal and vertical directions. There are three variants of the 4:2:0 scheme, with different horizontal and vertical positions.
[0216] In MPEG-2, Cb and Cr are co-located in the horizontal direction. Cb and Cr are located between pixels in the vertical direction (in the gap position). In Joint Photographic Experts Group (JPEG) / JPEG File Interchange Format (JFIF), H.261, and MPEG-1, Cb and Cr are located in the gap position, in the middle of alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternate lines.
[0217] chroma_format_idc separate_colour_plane_flag chroma format SubWidthC SubHeightC 0 0 monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1
[0218] Table 1. SubWidthC and SubHeightC values derived from chroma format idc and separate colour plane flag
[0219] 3.2 Example coding process of a video codec
[0220] Figure 2 An example of the encoder block diagram of VVC is shown, which contains three in-loop filtering blocks: the Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF, which uses a pre-defined filter, SAO and ALF exploit the original samples of the current picture to reduce the mean square error between the original and reconstructed samples by adding an offset and by applying a Finite Impulse Response (FIR) filter, respectively, and exploit the coded side information to signal the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be seen as a tool trying to capture and fix artifacts caused by previous stages.
[0221] 3.3 Definition of a video / coding unit
[0222] A picture is partitioned into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular region of a picture. A tile can be partitioned into one or more bricks, each brick including multiple CTU rows within the tile. A tile that is not partitioned into multiple bricks can also be referred to as a brick. However, a brick that is a proper subset of a tile cannot be referred to as a tile. A slice contains multiple tiles of a picture, or contains multiple bricks of a tile.
[0223] Two modes of slices are supported, i.e., raster-scan slice mode and rectangular slice mode. In the raster-scan slice mode, a slice contains a sequence of tiles in a raster scan of the picture. In the rectangular slice mode, a slice contains multiple bricks of a picture that collectively form a rectangular region of the picture. The bricks within a rectangular slice are arranged in the order of a brick raster scan of the slice. Figure 3 An example of raster-scan slice partitioning of a picture is shown, where the picture is partitioned into 12 tiles and 3 raster-scan slices.
[0224] Figure 4 An example of rectangular slice partitioning of a picture is shown, where the picture is partitioned into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0225] Figure 5 An example of a picture partitioned into tiles, bricks, and rectangular slices is shown, where the picture is partitioned into 4 tiles (2 tile columns and 2 tile rows), 11 bricks (the top-left tile contains 1 brick, the top-right tile contains 5 bricks, the bottom-left tile contains 2 bricks, and the bottom-right tile contains 3 bricks), and 4 rectangular slices.
[0226] 3.3.1 CTU / CTB size
[0227] In VVC, the CTU size, which is signaled in the sequence parameter set (SPS) by the syntax element log2_ctu_size_minus2, can be as small as 4x4.
[0228] 7.3.2.3 Sequence parameter set RBSP syntax
[0229] seq_parameter_set_rbsp( ) { descriptor sps_decoding_parameter_set_id u(4) sps_video_parameter_set_id u(4) sps_max_sub_layers_minus1 u(3) sps_reserved_zero_5bits u(5) profile_tier_level( sps_max_sub_layers_minus1 ) gra_enabled_flag u(1) sps_seq_parameter_set_id ue(v) chroma_format_idc ue(v) if( chroma_format_idc = = 3 ) separate_colour_plane_flag u(1) pic_width_in_luma_samples ue(v) pic_height_in_luma_samples ue(v) conformance_window_flag u(1) if( conformance_window_flag ) { conf_win_left_offset ue(v) conf_win_right_offset ue(v) conf_win_top_offset ue(v) conf_win_bottom_offset ue(v) } bit_depth_luma_minus8 ue(v) bit_depth_chroma_minus8 ue(v) log2_max_pic_order_cnt_lsb_minus4 ue(v) sps_sub_layer_ordering_info_present_flag u(1) for( i = ( sps_sub_layer_ordering_info_present_flag? 0 : sps_max_sub_layers_minus1 ); i <= sps_max_sub_layers_minus1; i++ ) { sps_max_dec_pic_buffering_minus1[ i ] ue(v) sps_max_num_reorder_pics[ i ] ue(v) sps_max_latency_increase_plus1[ i ] ue(v) } long_term_ref_pics_flag u(1) sps_idr_rpl_present_flag u(1) rpl1_same_as_rpl0_flag u(1) for( i = 0; i <!rpl1_same_as_rpl0_flag? 2 : 1; i++ ) { num_ref_pic_lists_in_sps[ i ] ue(v) for( j = 0; j < num_ref_pic_lists_in_sps[ i ]; j++) ref_pic_list_struct( i, j ) } qtbtt_dual_tree_intra_flag u(1) log2_ctu_size_minus2 ue(v) log2_min_luma_coding_block_size_minus2 ue(v) partition_constraints_override_enabled_flag u(1) sps_log2_diff_min_qt_min_cb_intra_slice_luma ue(v) sps_log2_diff_min_qt_min_cb_inter_slice ue(v) sps_max_mtt_hierarchy_depth_inter_slice ue(v) sps_max_mtt_hierarchy_depth_intra_slice_luma ue(v) if( sps_max_mtt_hierarchy_depth_intra_slice_luma!= 0 ) { sps_log2_diff_max_bt_min_qt_intra_slice_luma ue(v) sps_log2_diff_max_tt_min_qt_intra_slice_luma ue(v) } if( sps_max_mtt_hierarchy_depth_inter_slices!= 0 ) { sps_log2_diff_max_bt_min_qt_inter_slice ue(v) sps_log2_diff_max_tt_min_qt_inter_slice ue(v) } if( qtbtt_dual_tree_intra_flag ) { sps_log2_diff_min_qt_min_cb_intra_slice_chroma ue(v) sps_max_mtt_hierarchy_depth_intra_slice_chroma ue(v) if ( sps_max_mtt_hierarchy_depth_intra_slice_chroma!= 0 ) { sps_log2_diff_max_bt_min_qt_intra_slice_chroma ue(v) sps_log2_diff_max_tt_min_qt_intra_slice_chroma ue(v) } } … rbsp_trailing_bits( ) }
[0230] log2_ctu_size_minus2 plus 2 specifies the luma coding tree block size of each CTU. log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size. The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, MaxTbSizeY, PicWidthInCtbsY, PicHeightInCtbsY, PicSizeInCtbsY, PicWidthInMinCbsY, PicHeightInMinCbsY, PicSizeInMinCbsY, PicSizeInSamplesY, PicWidthInSamplesC and PicHeightInSamplesC are derived as follows:
[0231] CtbLog2SizeY = log2_ctu_size_minus2 + 2 (7-9)
[0232] CtbSizeY = 1 << CtbLog2SizeY (7-10)
[0233] MinCbLog2SizeY = log2_min_luma_coding_block_size_minus2 + 2 (7-11)
[0234] MinCbSizeY = 1 << MinCbLog2SizeY (7-12)
[0235] MinTbLog2SizeY = 2 (7-13)
[0236] MaxTbLog2SizeY = 6 (7-14)
[0237] MinTbSizeY = 1 << MinTbLog2SizeY (7-15)
[0238] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-16)
[0239] PicWidthInCtbsY = Ceil( pic_width_in_luma_samples ÷ CtbSizeY ) (7-17)
[0240] PicHeightInCtbsY = Ceil( pic_height_in_luma_samples ÷ CtbSizeY ) (7-18)
[0241] PicSizeInCtbsY = PicWidthInCtbsY × PicHeightInCtbsY (7-19)
[0242] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY (7-20)
[0243] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY (7-21)
[0244] PicSizeInMinCbsY = PicWidthInMinCbsY × PicHeightInMinCbsY (7-22)
[0245] PicSizeInSamplesY = pic_width_in_luma_samples × pic_height_in_luma_samples (7-23)
[0246] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC (7-24)
[0247] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC (7-25)
[0248] 3.3.2 CTU in a picture
[0249] Let the CTB / largest coding unit (LCU) LCU size be denoted by M x N (typically M is equal to N), and for a CTB located at a picture boundary (or slice or tile or other type of boundary, taking the picture boundary as an example), K x L samples are inside the picture boundary, where K < M or L < N. For those CTBs as shown in FIG. 3.3.2-1, the CTB size is still equal to M x N. However, the lower boundary / right boundary of the CTB is outside the picture. FIGS. 6A-6C
[0250] 3.4 Intra prediction
[0251] To capture arbitrary edge directions presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65. The extended directional modes are shown in FIG. 7
[0252] As shown in FIG. 7
[0253] In HEVC, each intra coded block has a square shape and the length of each side of the block is a power of 2. Therefore, no division operation is needed to generate the intra prediction value using the DC mode. In VVC, the block can have a rectangular shape and in general case, a division operation is needed to be used for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value for non-square blocks.
[0254] 3.5 Inter prediction
[0255] For each inter prediction coded unit (CU), the motion parameters include the motion vector, the reference picture index, the reference picture list usage index, and the extension information for the new coding features of VVC that will be used for the inter prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded with the skip mode, the CU is associated with one prediction unit (PU) and does not have significant residual coefficients, coded motion vector delta and / or reference picture index. The merge mode is specified that the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates and the extension scheduling introduced in VVC. The merge mode can be applied to any inter prediction CU, not only for the skip mode. An alternative to the merge mode is the explicit transmission of the motion parameters, where the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list usage flag and other useful information are explicitly signaled for each CU.
[0256] 3.6 Deblocking filter
[0257] Deblocking filtering is an example in-loop filter in video codecs. In VVC, the deblocking filtering process is applied to CU boundaries, transform subblock boundaries, and prediction subblock boundaries. Prediction subblock boundaries include prediction unit boundaries introduced by subblock-based temporal motion vector prediction (SbTMVP) and affine mode. Transform subblock boundaries include transform unit boundaries introduced by subblock transform (SBT) and intra subpartition (ISP) mode and transforms due to implicit partitioning of large CUs. The processing order of the deblocking filter is defined as horizontal filtering on vertical edges of the whole picture first, and then vertical filtering on horizontal edges. This particular order enables multiple horizontal filtering or vertical filtering processes to be applied in parallel threads. The filtering process can also be implemented on a CTB-by-CTB basis with only a small processing delay.
[0258] Vertical edges in a picture are filtered first. Then, horizontal edges in the picture are filtered with the samples modified by the vertical edge filtering process as input. The vertical and horizontal edges in the CTBs of each CTU are processed individually on a coding unit basis. The vertical edges of the coding blocks in a coding unit are filtered, starting from the edges on the left-hand side of the coding blocks, proceeding through the edges in their geometric order towards the right-hand side of the coding blocks. The horizontal edges of the coding blocks in a coding unit are filtered, starting from the edges on the top of the coding blocks, proceeding through the edges in their geometric order towards the bottom of the coding blocks.
[0259] 3.6.1 Boundary decision
[0260] Filtering is applied to 8x8 block boundaries. In addition, such a boundary must be a transform block boundary or a coding subblock boundary, e.g., a boundary resulting from the use of affine motion prediction (ATMVP). For other boundaries, deblocking filtering is disabled.
[0261] 3.6.2 Boundary strength calculation
[0262] For transform block boundaries / coding subblock boundaries, if the boundary is located in the 8x8 grid, the boundary can be filtered and the setting of bS[ xDi ][ yDj ] (where [ xDi ][ yDj ] denotes the coordinates) of this edge is defined as Table 2 and Table 3, respectively.
[0263] priority condition Y U V 5 at least one of the neighboring blocks is intra coded 2 2 2 4 TU boundary and at least one neighboring block have non-zero transform coefficients 1 1 1 3 The number of reference pictures or MVs (1 for uni-prediction, 2 for bi-prediction) of neighboring blocks are different 1 N / A N / A 2 The absolute difference between motion vectors belonging to neighboring blocks of the same reference picture is greater than or equal to one integer luma sample 1 N / A N / A 1 Others 0 0 0
[0264] Table 2. Boundary strength (when SPS IBC is disabled)
[0265] Priority Condition Y U V 8 At least one of the neighboring blocks is intra coded 2 2 2 7 TU boundary and at least one neighboring block have non-zero transform coefficients 1 1 1 6 The prediction modes of neighboring blocks are different (e.g., one is IBC and one is inter coded) 1 5 The absolute difference between IBC and motion vectors belonging to neighboring blocks are both greater than or equal to one integer luma sample 1 N / A N / A 4 The number of reference pictures or MVs (1 for uni-prediction, 2 for bi-prediction) of neighboring blocks are different 1 N / A N / A 3 The absolute difference between motion vectors belonging to neighboring blocks of the same reference picture is greater than or equal to one integer luma sample 1 N / A N / A 1 Others 0 0 0
[0266] Table 3. Boundary strength (when SPS IBC is enabled)
[0267] 3.6.3 Deblocking decision for luma component
[0268] The wider and stronger luma filter is used only when condition 1, condition 2 and condition 3 are all true. Condition 1 is the "large block condition". This condition checks whether the samples on the P-side and the Q-side belong to a large block, which are denoted by variables bSidePisLargeBlk and bSideQisLargeBlk, respectively. bSidePisLargeBlk and bSideQisLargeBlk are defined as follows.
[0269] bSidePisLargeBlk = ((edgeType is vertical and p0 belongs to a CU with width >= 32) || (edgeType is horizontal and p0 belongs to a CU with height >= 32))? true : false
[0270] bSideQisLargeBlk = ((edgeType is vertical and q0 belongs to a CU with width >= 32) || (edgeType is horizontal and q0 belongs to a CU with height >= 32))? true : false
[0271] Based on bSidePisLargeBlk and bSideQisLargeBlk, condition 1 is defined as follows:
[0272] condition 1 = (bSidePisLargeBlk || bSidePisLargeBlk)? true : false
[0273] Next, if condition 1 is true, condition 2 will be further checked. First, the following variables are derived:
[0274] First, dp0, dp3, dq0, dq3 are derived in the way of HEVC
[0275] if (p-side is greater than or equal to 32)
[0276] dp0 = ( dp0 + Abs( p50 - 2 x p40 + p30 ) + 1 ) » 1
[0277] dp3 = ( dp3 + Abs( p53 - 2 x p43 + p33 ) + 1 ) » 1
[0278] if (q-side is greater than or equal to 32)
[0279] dq0 = ( dq0 + Abs( q50 - 2 x q40 + q30 ) + 1 ) » 1
[0280] dq3 = ( dq3 + Abs( q53 - 2 x q43 + q33 ) + 1 ) » 1
[0281] Condition2 = ( d < β )? True : False
[0282] where d = dp0 + dq0 + dp3 + dq3.
[0283] If Conditionl and Condition2 are valid, then further check whether either of the blocks uses sub-blocks:
[0284] If ( bSidePisLargeBlk )
[0285] {
[0286] If ( BlockP'sMode == SUBBLOCKMODE )
[0287] Sp = 5
[0288] else
[0289] Sp = 7
[0290] }
[0291] else
[0292] Sp = 3
[0293] If ( bSideQisLargeBlk )
[0294] {
[0295] If ( BlockQ'sMode == SUBBLOCKMODE )
[0296] Sq = 5
[0297] else
[0298] Sq = 7
[0299] }
[0300] else
[0301] Sq = 3
[0302] Finally, if both Conditionl and Condition2 are valid, then the deblocking method will check Condition3 (large block strong filter condition), which is defined as follows. In Condition3 StrongFilterCondition, the following variables are derived:
[0303] dpq is derived as in HEVC.
[0304] sp3 = Abs( p3 - p0 ) is derived as in HEVC
[0305] if (pSide is greater than or equal to 32)
[0306] if (Sp == 5)
[0307] sp3 = ( sp3 + Abs( p5 - p3 ) + 1) » 1
[0308] else
[0309] sp3 = ( sp3 + Abs( p7 - p3 ) + 1) » 1
[0310] sq3 = Abs( q0 - q3 ) is derived as in HEVC
[0311] if (qSide is greater than or equal to 32)
[0312] if (Sq == 5)
[0313] sq3 = ( sq3 + Abs( q5 - q3 ) + 1) » 1
[0314] else
[0315] sq3 = ( sq3 + Abs( q7 - q3 ) + 1) » 1
[0316] StrongFilterCondition = ( dpq is less than ( β » 2 ), sp3 + sq3 is less than ( 3 x β » 5 ), and Abs( p0 - q0 ) is less than ( 5 x tC + 1 ) » 1)? TRUE : FALSE, as in HEVC.
[0317] 3.6.4 Stronger Deblocking Filter for Luma
[0318] When the samples on either side of the boundary belong to a large block, the bilinear filter is used. A sample is defined to belong to a large block when the width of the vertical edge >= 32, and when the height of the horizontal edge >= 32. The bilinear filter is listed below. Then, the block boundary samples pi (i = 0 to Sp-1) and qi (j = 0 to Sq-1) in the above HEVC Deblocking (pi and qi are the i-th sample within the row for filtering the vertical edge, or within the column for filtering the horizontal edge) are replaced by the following linear interpolation:
[0319]
[0320]
[0321] wherein and term is the above position dependent clipping, and , , , and are given as follows.
[0322] 3.6.5 Deblocking decision for chroma
[0323] Chroma strong filter is used on both sides of the block boundary. Here, when the two sides of the chroma edge are greater than or equal to 8 (chroma position), the chroma filter is selected and the following decision with three conditions is satisfied: The first decision is the decision for the boundary strength and the large block. When the block width or height that crosses the block boundary orthogonally in the chroma sample domain is equal to or greater than 8, the filter can be applied. The second and third decisions are basically the same as the HEVC luma deblocking decisions, which are the on / off decision and the strong filter decision, respectively.
[0324] In the first decision, the boundary strength (bS) is modified for chroma filtering and the conditions are checked in order. If the condition is satisfied, the remaining conditions with lower priority are skipped. When bS is equal to 2, or in the case of bS equal to 1 when a large block boundary is detected, chroma deblocking is performed. The second and third conditions are basically the same as the HEVC luma strong filter decision as follows.
[0325] In the second condition, d is derived in the way of HEVC luma deblocking. The second condition will be true when d is less than β. In the third condition, StrongFilterCondition is derived as follows:
[0326] dpq is derived in the way of HEVC;
[0327] sp3 = Abs( p3 - p0 ) is derived in the way of HEVC;
[0328] sq3 = Abs( q0 - q3 ) is derived in the way of HEVC.
[0329] According to the HEVC design, StrongFilterCondition = (dpq is less than ( β » 2 ), sp3 + sq3 is less than ( β » 3 ), and Abs( p0 - q0 ) is less than ( 5 x tC + 1 ) » 1).
[0330] 3.6.6 Strong deblocking filter for chroma
[0331] The strong deblocking filter for chroma is defined as follows:
[0332] p2' = (3 x p3 + 2 x p2 + p1 + p0 + q0 + 4) » 3
[0333] p1' = (2 x p3 + p2 + 2 x p1 + p0 + q0 + q1 + 4) » 3
[0334] p0' = (p3 + p2 + p1 + 2 x p0 + q0 + q1 + q2 + 4) » 3
[0335] The example chroma filter performs deblocking on a 4x4 grid of chroma samples.
[0336] 3.6.7 Position-dependent clipping
[0337] The position-dependent clipping tcPD is applied to the output samples of the luma filtering process that involves the strong filter and the long filter modifying 7, 5 and 3 samples at the boundary. Assuming a quantization error distribution, the clipping value can be increased for samples that are expected to have higher quantization noise, thus expected to have a higher deviation from the true sample value.
[0338] For each P or Q boundary filtered with an asymmetric filter, a position-dependent threshold table is selected from two tables provided as side information to the decoder (e.g., Tc7 and Tc3 as listed below) according to the result of the decision-making process:
[0339] Tc7 = { 6, 5, 4, 3, 2, 1, 1}; Tc3 = { 6, 4, 2};
[0340] tcPD = (Sp == 3)? Tc3 : Tc7;
[0341] tcQD = (Sq == 3)? Tc3 : Tc7;
[0342] For P or Q boundaries filtered with a short symmetric filter, a lower amplitude position-dependent threshold is applied:
[0343] Tc3 = { 3, 2, 1};
[0344] After the threshold is defined, the filtered p’i and q’i sample values are clipped according to the tcP and tcQ clipping values:
[0345] p''i = Clip3(p’i + tcPi, p’i – tcPi, p’i );
[0346] q''j = Clip3(q’j + tcQj, q’j – tcQ j, q’j );
[0347] where p’i and q’i are the filtered sample values, p''i and q''j are the clipped output sample values, and tcPi is the clipping threshold derived from the VVC tc parameters and tcPD and tcQD. The function Clip3 is the clipping function as specified in VVC.
[0348] 3.6.8 Subblock deblocking adjustment
[0349] To enable parallel friendly deblocking using both long filter and subblock deblocking, the long filter is restricted to modify at most 5 samples on the side using subblock deblocking (AFFINE or ATMVP or decoder side motion vector refinement (DMVR)), as shown in the luma control of long filter. Extending, the subblock deblocking is adjusted such that the subblock boundaries close to the CU or implicit TU boundaries on the 8x8 grid are restricted to modify at most two samples on each side.
[0350] The following applies to subblock boundaries that are not aligned with the CU boundaries.
[0351] If (block Q's mode == SUBBLOCKMODE && edge!=0) {
[0352] if (!(implicitTU && (edge == (64 / 4))))
[0353] if (edge == 2 || edge == (orthogonalLength-2) || edge == (56 / 4) || edge == (72 / 4))
[0354] Sp = Sq = 2;
[0355] else
[0356] Sp = Sq = 3;
[0357] else
[0358] Sp = Sq = bSideQisLargeBlk? 5:3
[0359] }
[0360] where edge equal to 0 corresponds to a CU boundary, edge equal to 2 or equal to orthogonalLength - 2 corresponds to a sub-block boundary 8 samples away from a CU boundary, etc. If implicit partitioning of TUs is used, implicit TU is true.
[0361] 3.7 Sample Adaptive Offset
[0362] Sample Adaptive Offset (SAO) is applied to the reconstructed signal after the deblocking filter by using an offset specified by the encoder for each CTB. The video encoder first decides whether to apply the SAO process to the current slice. If SAO is applied to the slice, each CTB is classified into one of five SAO types as shown in Table 4. The concept of SAO is to classify pixels into multiple categories and reduce distortion by adding an offset to the pixels of each category. The SAO operation includes Edge Offset (EO) and Band Offset (BO), where EO uses edge properties for pixel classification in SAO types 1 to 4, and BO uses pixel intensity for pixel classification in SAO type 5. Each applicable CTB has SAO parameters including sao_merge_left_flag, sao_merge_up_flag, SAO type, and four offsets. If sao_merge_left_flag is equal to 1, the current CTB will reuse the SAO type and offsets of the left CTB. If sao_merge_up_flag is equal to 1, the current CTB will reuse the SAO type and offsets of the above CTB.
[0363] SAO type The type of sample adaptive offset to be used Number of categories 0 None 0 1 One-dimensional 0-degree pattern edge offset 4 2 One-dimensional 90-degree pattern edge offset 4 3 One-dimensional 135-degree pattern edge offset 4 4 One-dimensional 45-degree pattern edge offset 4 5 Band offset 4
[0364] Table 4. SAO type specification
[0365] 3.8 Adaptive Loop Filter
[0366] Adaptive loop filtering for video coding is to minimize the mean square error between original and decoded samples by using a Wiener-based adaptive filter. ALF is located at the last processing stage of each picture and can be regarded as a tool to capture and fix artifacts from previous stages. Suitable filter coefficients are determined by the encoder and explicitly signaled to the decoder. To achieve better coding efficiency, especially for high resolution videos, local adaptation is used for the luma signal by applying different filters to different regions or blocks in a picture. In addition to filter adaptation, filter on / off control at the coding tree unit (CTU) level also helps to improve coding efficiency. In terms of syntax, filter coefficients are sent in a picture-level header called adaptive parameter set, and filter on / off flags of a CTU are interleaved at the CTU level in slice data. This syntax design not only supports picture-level optimization, but also enables low encoding delay.
[0367] 3.8.1 Signaling of parameters
[0368] According to the ALF design in VTM, the filter coefficients and clipping indices are carried in the ALF APS. The ALF APS can include up to 8 chroma filters and one luma filter set with up to 25 filters. An index is also included for each of the 25 luma categories. Categories with the same index share the same filter. By merging different categories, the number of bits needed to represent the filter coefficients is reduced. The absolute values of the filter coefficients are represented using a 0thorder Exp-Golomb code followed by a sign bit for non-zero coefficients. When clipping is enabled, a two-bit fixed length code is also used to signal the clipping index for each filter coefficient. The decoder can use up to 8 ALF APSs at the same time.
[0369] The filter control syntax elements for ALF in VTM include two types of information. First, the ALF on / off flag is signaled at the sequence, picture, slice and CTB level. Chroma ALF can be enabled at the picture and slice level only if luma ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice and CTB level, the filter usage information is signaled at that level. If all slices within a picture use the same APS, the referenced ALF APS ID is coded at the slice or picture level. The luma component can refer to up to 7 ALF APSs and the chroma component can refer to 1 ALF APS. For luma CTBs, an index is signaled which indicates which ALF APS or offline trained luma filter set is used. For chroma CTBs, the index indicates which filter in the referred APS is used.
[0370] The data syntax elements for ALF associated with the luma component in VTM are listed as follows:
[0371] alf_data( ) { descriptor alf_luma_filter_signal_flag u(1) if( alf_luma_filter_signal_flag ) { alf_luma_clip_flag u(1) alf_luma_num_filters_signalled_minus1 ue(v) if( alf_luma_num_filters_signalled_minus1 > 0 ) for( filtIdx = 0; filtIdx < NumAlfFilters; filtIdx++ ) alf_luma_coeff_delta_idx[ filtIdx ] u(v) for( sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1; sfldx++ ) for( j = 0; j < 12; j++ ) { alf_luma_coeff_abs[ sfIdx ][ j ] ue(v) if( alf_luma_coeff_abs[ sfIdx ][ j ] ) alf_luma_coeff_sign[ sfIdx ][ j ] u(1) } if( alf_luma_clip_flag ) for( sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1; sfldx++ ) for( j = 0; j < 12; j++ ) alf_luma_clip_idx[ sfIdx ][ j ] u(2) }
[0372] alf_luma_filter_signal_flag equal to 1 specifies that the luma filter set is signaled. alf_luma_filter_signal_flag equal to 0 specifies that the luma filter set is not signaled. alf_luma_clip_flag equal to 0 specifies that linear adaptive loop filtering is applied to the luma component. alf_luma_clip_flag equal to 1 specifies that non-linear adaptive loop filtering can be applied to the luma component. alf_luma_num_filters_signalled_minus1 plus 1 specifies the number of adaptive loop filter classes whose luma coefficients can be signaled. The value of alf_luma_num_filters_signalled_minus1 shall be in the range of 0 to NumAlfFilters - 1, inclusive. alf_luma_coeff_delta_idx[ filtIdx ] specifies the index of the signaled adaptive loop filter luma coefficient delta for the filter class indicated by filtIdx in the range of 0 to NumAlfFilters - 1. When alf_luma_coeff_delta_idx[ filtIdx ] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[ filtIdx ] is Ceil( Log2( alf_luma_num_filters_signalled_minus1 + 1 ) ) bits. The value of alf_luma_coeff_delta_idx[ filtIdx ] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1, inclusive.
[0373] alf_luma_coeff_abs[ sfIdx ][ j ] specifies the absolute value of the j-th coefficient of the signaled luma filter indicated by sfIdx. When alf_luma_coeff_abs[ sfIdx ][ j ] is not present, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[ sfIdx ][ j ] shall be in the range of 0 to 128, inclusive. alf_luma_coeff_sign[ sfIdx ][ j ] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx, as follows:
[0374] If alf_luma_coeff_sign[ sfldx ][ j ] is equal to 0, the corresponding luma filter coefficient has a positive value.
[0375] Otherwise (alf_luma_coeff_sign[ sfldx ][ j ] is equal to 1 ), the corresponding luma filter coefficient has a negative value.
[0376] When alf_luma_coeff_sign[ sfldx ][ j ] is not present, it is inferred to be equal to 0.
[0377] alf_luma_clip_idx[ sfldx ][ j ] specifies the clipping index of the clipping value to be used before multiplying the j-th coefficient of the luma filter indicated by sfldx that is signaled. When alf_luma_clip_idx[ sfldx ][ j ] is not present, it is inferred to be equal to 0. The coding tree unit syntax elements of ALF associated with the luma component in VTM are listed as follows:
[0378] coding_tree_unit( ) { descriptor xCtb = CtbAddrX « CtbLog2SizeY yCtb = CtbAddrY « CtbLog2SizeY if( sh_alf_enabled_flag ){ alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ae(v) if( alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ) { if( sh_num_alf_aps_ids_luma > 0 ) alf_use_aps_flag ae(v) if( alf_use_aps_flag ) { if( sh_num_alf_aps_ids_luma > 1 ) alf_luma_prev_filter_idx ae(v) } else alf_luma_fixed_filter_idx ae(v) } }
[0379] alf_ctb_flag[ cldx ][ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] equal to 1 specifies that the adaptive loop filter is applied to the coding tree block of the color component indicated by cldx of the coding tree unit at luma location ( xCtb, yCtb ). alf_ctb_flag[ cldx ][ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] equal to 0 specifies that the adaptive loop filter is not applied to the coding tree block of the color component indicated by cldx of the coding tree unit at luma location ( xCtb, yCtb ).
[0380] When alf_ctb_flag[ cIdx ][ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] is not present, it is inferred to be equal to 0. alf_use_aps_flag equal to 0 specifies that one of the fixed filter sets in the fixed filter set is applied to the luma CTB. alf_use_aps_flag equal to 1 specifies that a filter set from the APS is applied to the luma CTB. When alf_use_aps_flag is not present, it is inferred to be equal to 0. alf_luma_prev_filter_idx specifies the previous filter applied to the luma CTB. The value of alf_luma_prev_filter_idx shall be in the range of 0 to sh_num_alf_aps_ids_luma - 1, inclusive. When alf_luma_prev_filter_idx is not present, it is inferred to be equal to 0.
[0381] The variable AlfCtbFiltSetIdxY[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ], which specifies the filter set index of the luma CTB at position ( xCtb, yCtb ), is derived as follows:
[0382] If alf_use_aps_flag is equal to 0, AlfCtbFiltSetIdxY[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] is set equal to alf_luma_fixed_filter_idx.
[0383] Otherwise, AlfCtbFiltSetIdxY[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] is set equal to 16 + alf_luma_prev_filter_idx.
[0384] alf_luma_fixed_filter_idx specifies the fixed filter applied to the luma CTB. The value of alf_luma_fixed_filter_idx shall be in the range of 0 to 15, inclusive.
[0385] Based on the VTM-based ALF design, the ALF design in ECM further introduces the concept of alternative filter set into the luma filter. The luma filter conducts multiple alternative / rounds of training based on the updated luma CTU ALF on / off decision for each alternative / round. In this way, there will be multiple filter sets associated with each training alternative, and the category merge result for each filter set can be different. The best filter set for each CTU can be selected by RDO, and the related alternative information will be signaled. The data syntax elements for ALF associated with luma component in ECM are listed as follows:
[0386] alf_data( ) { descriptor alf_luma_filter_signal_flag u(1) if( alf_luma_filter_signal_flag ) { alf_luma_num_alts_minus1 ue(v) for (altIdx = 0; altIdx < alf_luma_num_alts_minus1 +1; altIdx++){ alf_luma_clip_flag[altIdx] u(1) alf_luma_num_filters_signalled_minus1[altIdx] ue(v) if(alf_luma_num_filters_signalled_minus1[altIdx] > 0){ for( filtIdx = 0; filtIdx < NumAlfFilters; filtIdx++ ) alf_luma_coeff_delta_idx[altIdx][filtIdx] u(v) } for (sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1[altldx]; sfldx++) { for(j = 0; j < 19; j++){ alf_luma_coeff_abs[altIdx][ sfIdx ][ j ] ue(v) if( alf_luma_coeff_abs[altIdx][ sfIdx ][ j ] ) alf_luma_coeff_sign[altIdx][ sfIdx ][ j ] u(1) } } if( alf_luma_clip_flag [altIdx]) for (sfldx = 0; sfldx <= alf_luma_num_filters_signalled_minus1[altldx]; sfldx++ ) for( j = 0; j <19; j++ ) alf_luma_clip_idx[altIdx][ sfIdx ][ j ] u(2) } }
[0387] alf_luma_num_alts_minus1 plus 1 specifies the number of alternative filter sets for luma component. The value of alf_luma_num_alts_minus1 shall be in the range of 0 to 3, inclusive. alf_luma_clip_flag[altIdx] equal to 0 specifies that linear adaptive loop filtering is applied to the luma component of the alternative luma filter set with index altIdx. alf_luma_clip_flag[altIdx] equal to 1 specifies that non-linear adaptive loop filtering can be applied to the luma component of the alternative luma filter set with index altIdx. alf_luma_num_filters_signalled_minus1[altIdx] plus 1 specifies the number of adaptive loop filter categories whose luma coefficients can be signaled in the alternative luma filter set with index altIdx. The value of alf_luma_num_filters_signalled_minus1[altIdx] shall be in the range of 0 to NumAlfFilters - 1, inclusive.
[0388] alf_luma_coeff_delta_idx[ altIdx ][ filtIdx ] specifies the index of the signalled adaptive loop filter luma coefficient delta for the filter class indicated by filtIdx in the range of 0 to NumAlfFilters - 1 for the alternative luma filter set with index altIdx. When alf_luma_coeff_delta_idx[ filtIdx ][ altIdx ] is not present, it is inferred to be equal to 0. The length of alf_luma_coeff_delta_idx[ altIdx ][ filtIdx ] is Ceil( Log2( alf_luma_num_filters_signalled_minus1[ altIdx ] + 1 ) ) bits. The value of alf_luma_coeff_delta_idx[ altIdx ][ filtIdx ] shall be in the range of 0 to alf_luma_num_filters_signalled_minus1[ altIdx ], inclusive. alf_luma_coeff_abs[ altIdx ][ sfIdx ][ j ] specifies the absolute value of the j-th coefficient of the signalled luma filter indicated by sfIdx for the alternative luma filter set with index altIdx. When alf_luma_coeff_abs[ altIdx ][ sfIdx ][ j ] is not present, it is inferred to be equal to 0. The value of alf_luma_coeff_abs[ altIdx ][ sfIdx ][ j ] shall be in the range of 0 to 128, inclusive.
[0389] alf_luma_coeff_sign[ altIdx ][ sfIdx ][ j ] specifies the sign of the j-th luma coefficient of the filter indicated by sfIdx for the alternative luma filter set with index altIdx as follows:
[0390] If alf_luma_coeff_sign[ altIdx ][ sfIdx ][ j ] is equal to 0, the corresponding luma filter coefficient has a positive value.
[0391] Otherwise (alf_luma_coeff_sign[ altIdx ][ sfIdx ][ j ] is equal to 1), the corresponding luma filter coefficient has a negative value.
[0392] When alf_luma_coeff_sign[ altIdx ][ sfldx ][ j ] is not present, it is inferred to be equal to 0.
[0393] alf_luma_clip_idx[ altIdx ][ sfldx ][ j ] specifies the clipping index of the clipping value to be used before multiplying by the j-th coefficient of the luma filter signaled by sfldx of the alternative luma filter set with index altIdx. When alf_luma_clip_idx[ altIdx ][ sfldx ][ j ] is not present, it is inferred to be equal to 0. The coding tree unit syntax elements of ALF associated with the luma component in the ECM are listed as follows:
[0394] coding_tree_unit( ) { descriptor xCtb = CtbAddrX « CtbLog2SizeY yCtb = CtbAddrY « CtbLog2SizeY if( sh_alf_enabled_flag ){ alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ae(v) if( alf_ctb_flag[ 0 ][ CtbAddrX ][ CtbAddrY ] ) { if( sh_num_alf_aps_ids_luma > 0 ) alf_use_aps_flag ae(v) if( alf_use_aps_flag ) { if( sh_num_alf_aps_ids_luma > 1 ) alt_ctb_luma_filter_alt_idx[CtbAddrX][CtbAddrY] ae(v) alf_luma_prev_filter_idx ae(v) } else alf_luma_fixed_filter_idx ae(v) } }
[0395] alf_ctb_luma_filter_alt_idx[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] specifies the index of the alternative luma filter applied to the coding tree block of the luma component of the coding tree unit at location (xCtb, yCtb). When alf_ctb_luma_filter_alt_idx[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] is not present, it is inferred to be equal to 0.
[0396] 3.8.2 Filter shape
[0397] In JEM, up to three diamond filter shapes (as shown in Figure 10 ) can be selected for the luma component. An index is signaled at the picture level to indicate the filter shape used for the luma component. Each square represents a sample, and Ci (i is 0-6 (left), 0-12 (middle), 0-20 (right)) represents the coefficient to be applied to the sample. For the chroma components in a picture, the 5x5 diamond shape is always used. In VVC, the 7x7 diamond shape is always used for luma, while the 5x5 diamond shape is always used for chroma.
[0398] 3.8.3 Classification of ALF
[0399] Each 2x2 (or 4x4) block is classified into one of the 25 categories. The classification index C is derived based on the quantized values of its directionality and activity as follows:
[0400] .
[0401] To compute and , first the gradients in the horizontal, vertical and two diagonal directions are computed using one-dimensional Laplacians:
[0402]
[0403]
[0404]
[0405]
[0406] The indices and refer to the coordinates of the top-left sample in the 2x2 block, and indicates the reconstructed sample at coordinates . Then the maximum and minimum of the gradients in the horizontal and vertical directions are set to:
[0407]
[0408] and the maximum and minimum of the gradients in the two diagonal directions are set to:
[0409]
[0410] To derive the values of the directionality , these values are compared to each other and to two thresholds and :
[0411] Step 1. If and are both true, then is set to .
[0412] Step 2. If , continue from step 3; otherwise continue from step 4.
[0413] Step 3. If , then is set to ; otherwise is set to .
[0414] Step 4. If , then is set to ; otherwise Set as .
[0415] Activity value Calculated as:
[0416]
[0417] It is further quantized to the range of 0 to 4 (inclusive), and the quantized value is represented as For the two chromaticity components in the image, no classification method is applied; that is, a single set of ALF coefficients is applied for each chromaticity component.
[0418] 3.8.4 Geometric Transformation of Filter Coefficients
[0419] Before filtering each 2×2 block, geometric transformations (such as rotation or diagonal and vertical flips) are applied to the filter coefficients associated with coordinates (k, l), based on the gradient values calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make different blocks to which ALF is applied more similar by aligning the directionality of the different blocks.
[0420] Three geometric transformations are introduced: diagonal, vertical flip, and rotation.
[0421] diagonal: ,
[0422] Vertical Flip: ,
[0423] Rotation:
[0424] in It is the size of the filter, and These are coefficient coordinates, which make the position... In the top left corner, and in position In the bottom right corner. Based on the gradient values calculated for this block, the transform is applied to the filter coefficients f(k, l). The relationship between the transform and the four gradients in the four directions is summarized in Table 5. Figure 11 The transformation coefficients for each position based on a 5x5 rhombus are shown.
[0425] Gradient value Transform g d2 g d1 and g h g v ]]> No transform g d2 g d1 and g v g h ]]> Diagonal line g d1 g d2 and g h g v ]]> Vertical flip g d1 g d2 and g v g h ]]> Rotation
[0426] Table 5. Mapping of gradients to transformations computed for a block
[0427] 3.8.5 Filtering Process
[0428] On the decoder side, when ALF is enabled for a block, each sample within the block... are filtered, resulting in sample values as shown below where L denotes the filter length, denotes the filter coefficients, and denotes the decoded filter coefficients.
[0429]
[0430] Figure 12 An example of relative coordinates used for a 5x5 diamond filter support is shown assuming the coordinates (i, j) of the current sample are (0, 0). The samples in different coordinates that are filled with the same color are multiplied with the same filter coefficient.
[0431] 3.8.6 Non-linear filter reformulation
[0432] Linear filtering can be reformulated to the following expression without affecting the coding efficiency:
[0433]
[0434] where are the same filter coefficients.
[0435] VVC introduces non-linearity by using a simple clipping function to reduce the impact of neighboring sample values ( ) when they differ too much from the current sample value being filtered ( ), thus making the ALF more efficient. More specifically, the ALF filter is modified as follows:
[0436]
[0437] where is a clipping function, and is a clipping parameter, which depends on the filter coefficients. The encoder performs an optimization to find the best .
[0438] The clipping parameter is specified for each ALF filter, and each filter coefficient is signaled with a clipping value. This means that up to 12 clipping values can be signaled in the bitstream for each luma filter, and up to 6 clipping values can be signaled in the bitstream for each chroma filter. To limit the signaling cost and encoder complexity, only 4 fixed values that are the same for inter and intra slices are used.
[0439] Because the variance of local differences in luminance is typically higher than that of chrominance, two different sets of filters are applied for luminance and chrominance. A maximum sample value in each set (here, 1024 for a 10-bit bit-depth) is also introduced so that clipping can be disabled when not necessary. The 4 values are chosen by roughly evenly dividing the full range of sample values (coded on 10 bits) for luminance and the range from 4 to 1024 for chrominance in the log domain. More precisely, the luminance clipping value table has been derived by:
[0440] AlfClip L where M = 2 10 and N = 4
[0441] Similarly, the chrominance clipping value table is derived according to:
[0442] AlfClip C where M = 2 10 , N = 4, A = 4
[0443] 3.9 Bilateral loop filter
[0444] 3.9.1 Bilateral image filter
[0445] The bilateral image filter is a non-linear filter that smooths noise while preserving edge structures. Bilateral filtering is a technique that makes the filter weights decrease not only with the distance between samples but also with the increase of intensity difference. In this way, over-smoothing of edges can be improved. The weight is defined as
[0446]
[0447] where and are the distances in the vertical and horizontal directions, and is the intensity difference between samples.
[0448] The edge-preserving denoising bilateral filter employs a low-pass Gaussian filter for both the domain filter and the range filter. The domain low-pass Gaussian filter gives higher weights to pixels that are close to the center pixel in the spatial domain. The range low-pass Gaussian filter gives higher weights to pixels that are similar to the center pixel. In combination with the range filter and the domain filter, the bilateral filter at an edge pixel becomes an elongated Gaussian filter that is oriented along the edge and greatly reduced in the gradient direction. This is the reason why the bilateral filter can smooth noise while preserving edge structures.
[0449] 3.9.2 Bilateral filter in video coding
[0450] The bilateral filter in video coding is a coding tool for VVC [2]. The filter is applied as an in-loop filter in parallel with the sample adaptive offset (SAO) filter. Both the bilateral filter and the SAO act on the same input samples, each filter produces a compensation, and these compensations are then added to the input samples to produce output samples that go to the next stage after clipping. The spatial filter strength is determined by the block size, where smaller blocks are filtered more strongly, and the strength filter strength is determined by the quantization parameter, where the higher the quantization parameter (QP), the higher the filter strength. Only the four nearest samples are used, so the filtered sample strength can be calculated as
[0451]
[0452] where denotes the strength of the center sample, denotes the strength difference between the center sample and the above sample. denotes the strength difference between the center sample and the below, left and right samples, respectively.
[0453] 4. Technical problems solved by the disclosed technology
[0454] Example designs of the adaptive loop filter (ALF) in video coding have the following problems.
[0455] In example ALF designs, only the spatially reconstructed samples after other filtering tools such as deblocking filter / SAO / bilateral filter (BF) are used for filter training and filtering. However, there are other valuable information that can be potentially exploited, such as the samples filtered / generated by one or more pre-defined filters.
[0456] 5. List of solutions and embodiments
[0457] To solve the above problems, the methods as outlined below are disclosed. The embodiments should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Furthermore, the embodiments can be applied individually or combined in any way.
[0458] It should be noted that the disclosed methods can be used as a loop filter or post-processing.
[0459] In this disclosure, a video unit can refer to a sequence, a picture, a subpicture, a slice, a CTU, a block, and / or a region. A video unit can include one color component or multiple color components.
[0460] In this disclosure, an ALF processing unit can refer to a sequence, a picture, a subpicture, a slice, a CTU, a block, a region, or a sample. An ALF processing unit can include one color component or multiple color components.
[0461] In this disclosure, an extended tap can refer to a filter tap that is not used by ALF in VVC.
[0462] 1) It is proposed to use at least one extended tap for ALF filter to further improve the efficiency of ALF.
[0463] a. In one example, the at least one extended tap can be different from the spatial taps in the ALF filter that only utilize the information of the spatial neighboring samples (adjacent to the center sample to be filtered) of the target component.
[0464] a) In one example, the spatial neighboring samples can come from the reconstruction after DBF / SAO / CCSAO / BF.
[0465] b) In one example, the spatial tap only uses the spatial neighboring luma samples to filter the center luma sample inside one ALF filter.
[0466] c) Optionally, the spatial tap only uses the spatial neighboring chroma samples to filter the center chroma sample inside one ALF filter.
[0467] b. In one example, the at least one extended tap and the at least one spatial tap can coexist inside one ALF filter.
[0468] a) In one example, the ALF filter can be composed of both spatial taps and extended taps.
[0469] 1. In one example, the ALF filter can be composed of M (e.g., M > 0) spatial taps and N (e.g., N > 0) extended taps.
[0470] c. In one example, one or more extended taps of the ALF filter use one or more input sources.
[0471] a) In one example, one or more extended taps of the ALF filter use one input source.
[0472] 1. For example, the input source can be based on one of the following:
[0473] a. The intermediate filtering results of one or more predefined filters.
[0474] a) In one example, the predefined filter can be a Gaussian filter.
[0475] b) In one example, the predefined filter can be a low-pass filter.
[0476] c) In one example, the predefined filter can be a high-pass filter.
[0477] b. Any modification of the above sources
[0478] c. Any combination or fusion (e.g. weighted sum) of at least two sources.
[0479] b) In one example, one or more extended taps of the ALF filter use multiple input sources.
[0480] 1. For example, the multiple input sources can include more than one of the following:
[0481] a. An intermediate filtering result of one or more predefined filters.
[0482] a) In one example, the predefined filter can be a Gaussian filter.
[0483] b) In one example, the predefined filter can be a low-pass filter.
[0484] c) In one example, the predefined filter can be a high-pass filter.
[0485] b. Any modification of the above sources
[0486] c. Any combination or fusion (e.g. weighted sum) of at least two sources.
[0487] c) The input sources of the extended taps can be derived based on samples in different color components.
[0488] d. In one example, whether and / or how the filter with at least one extended tap is applied can be different for different color formats and / or different color components.
[0489] a) In one example, the ALF filter with at least one extended tap can be applied only to process the luma component.
[0490] b) Optionally, the ALF filter with at least one extended tap can be applied only to process one of the chroma components (e.g. the Cb or Cr component).
[0491] c) Optionally, the ALF filter with at least one extended tap can be applied to process all chroma components. (e.g. the Cb and Cr components).
[0492] d) Optionally, an ALF filter with at least one extended tap can be applied to filter the luma and chroma components. (e.g., Y, Cb and / or Cr components).
[0493] e. In one example, the ALF filter with at least one extended tap can use a shape or size that is different from the shape or size used by an ALF without extended taps, such as the ALF in VVC.
[0494] a) In one example, the ALF filter can contain different shapes for spatial taps and extended taps.
[0495] b) In one example, within the ALF filter, the shape / size used for one or more spatial taps can be different from the shape / size used for one or more extended taps.
[0496] c) Optionally, within the ALF filter, the shape / size used for one or more spatial taps can be the same as the shape / size used for one or more extended taps.
[0497] d) In one example, within the ALF filter with at least one extended tap, the filter shape used for one or more spatial taps can be as follows.
[0498] 1. In one example, the filter shape used for spatial taps can be diamond shape.
[0499] 2. In one example, the filter shape used for spatial taps can be square shape.
[0500] 3. In one example, the filter shape used for spatial taps can be cross shape.
[0501] 4. Optionally, the filter shape used for spatial taps can be symmetric shape.
[0502] 5. Optionally, the filter shape used for spatial taps can be asymmetric shape.
[0503] 6. Optionally, the filter shape used for spatial taps can be other designed shape.
[0504] 7. In one example, the filter shape used for spatial taps can be determined / signaled / derived on the fly.
[0505] e) In one example, within the ALF filter with at least one extended tap, the filter shape used for one or more extended taps can be as follows.
[0506] 1. In one example, the filter shape used for the extension taps can be diamond shape.
[0507] 2. In one example, the filter shape used for the extension taps can be square shape.
[0508] 3. In one example, the filter shape used for the extension taps can be cross shape.
[0509] 4. Optionally, the filter shape used for the extension taps can be symmetric shape.
[0510] 5. Optionally, the filter shape used for the extension taps can be asymmetric shape.
[0511] 6. Optionally, the filter shape used for the extension taps can be other design shape.
[0512] 7. In one example, the filter shape used for the extension taps can be determined / signaled / derived on the fly.
[0513] f) In one example, within the ALF filter with at least one extension tap, the filter size used for one or more extension taps can be as follows:
[0514] 1. In one example, the filter size used for the extension taps can be MxN.
[0515] a. In one example, M can be equal to N.
[0516] a) In one example, M or N can be equal to 1.
[0517] b) In one example, M or N can be equal to 3.
[0518] c) In one example, M or N can be equal to 5.
[0519] b. In one example, M can not be equal to N.
[0520] f. In one example, a first syntax element can be signaled to indicate whether the filter with at least one extension tap is enabled or not.
[0521] a) In one example, the first syntax element can be coded by arithmetic coding.
[0522] 1. In one example, the first syntax element can be coded with at least one context.
[0523] a. The context can depend on the coding information of the current block or neighboring blocks.
[0524] b. The context can depend on the filter shape of at least one neighboring block.
[0525] 2. In one example, the first syntax element can be coded with bypass coding.
[0526] b) In one example, the first syntax element can be binarized by unary code, or truncated unary code, or fixed length code, or exponential Golomb code, truncated exponential Golomb code, etc.
[0527] c) In one example, the first syntax element can be conditionally signaled.
[0528] 1. For example, the first syntax element can be signaled only when extended taps are available.
[0529] d) The first syntax element can be coded in a predictive manner.
[0530] 1. The first syntax element can be predicted by the on / off decision of the extended taps of at least one neighboring block.
[0531] e) The first syntax element can be independently signaled for different color components.
[0532] 1. Optionally, the first syntax element can be signaled and shared for different color components.
[0533] 2. Optionally, the first syntax element can be signaled for the first color component but not for the second color component.
[0534] f) The syntax element can be signaled in SPS / PPS / picture header / slice header / APS / CTU / CU / etc.
[0535] g. In one example, the first syntax element can be signaled to indicate which / what input source is used for the extended taps inside the ALF filter.
[0536] a) In one example, the first syntax element can be coded by arithmetic coding.
[0537] 1. In one example, the first syntax element can be coded with at least one context.
[0538] a. The context can depend on the coding information of the current block or neighboring blocks.
[0539] b. The context can depend on the filter shape of at least one neighboring block.
[0540] 2. In one example, the first syntax element can be coded with bypass coding.
[0541] b) In one example, the first syntax element can be binarized by unary code, or truncated unary code, or fixed length code, or exponential Golomb code, truncated exponential Golomb code, etc.
[0542] c) In one example, the first syntax element can be conditionally signaled.
[0543] 1. For example, the first syntax element can be signaled only when the extended tap is available.
[0544] d) The first syntax element can be coded in a predictive manner.
[0545] 1. The first syntax element can be predicted by the on / off decision of the extended tap of at least one neighboring block.
[0546] e) The first syntax element can be independently signaled for different color components.
[0547] 1. Optionally, the first syntax element can be signaled and shared for different color components.
[0548] 2. Optionally, the first syntax element can be signaled for the first color component but not for the second color component.
[0549] f) The syntax element can be signaled in SPS / PPS / picture header / tile header / APS / CTU / CU / etc.
[0550] 1. In one example, the syntax element can be signaled in APS for the signaled ALF filter.
[0551] h. The coefficients of at least one extended tap inside the ALF filter can be signaled in a syntax element structure such as APS.
[0552] a) In one example, the coefficients of the extended tap can be included in APS.
[0553] b) In one example, the clipping parameters of the extended tap can be included in APS.
[0554] c) In one example, the category merge results of the extended tap can be included in APS.
[0555] d) In one example, the coefficients of the extended tap can be coded in a predictive manner.
[0556] e) In one example, the coefficients of the extension tap can be coded by using arithmetic coding with at least one context.
[0557] f) In one example, the coefficients of the extension tap can be coded using bypass coding.
[0558] g) In one example, the coefficients of the extension tap can be jointly coded with the coefficients of the spatial tap.
[0559] h) Optionally, other parameters of the extension tap can be included in the APS.
[0560] i) Optionally, the coefficients of the extension tap can be predefined fixed values.
[0561] 2) It is proposed to use the intermediate filtering result of at least one fixed or adaptive filter as input for one or more extension taps.
[0562] a. In one example, the intermediate filtering result of an offline trained filter can be used as input for one or more extension taps.
[0563] a) In one example, the intermediate filtering result can be generated from the reconstruction before the ALF and the offline trained filter of the ALF.
[0564] b) Optionally, the intermediate filtering result can be generated from the reconstruction before the DBF and the offline trained filter of the ALF.
[0565] b. In one example, the intermediate filtering result of an online trained ALF filter can be used as input for one or more extension taps.
[0566] c. In one example, the intermediate filtering result of other predefined filters can be used as input for one or more extension taps.
[0567] a) In one example, a Gaussian filter can be applied.
[0568] b) In one example, a bilateral filter can be applied.
[0569] c) In one example, a guided filter can be applied.
[0570] d) In one example, a median filter can be applied.
[0571] 1. In one example, a local median filter can be applied.
[0572] 2. In one example, a non-local median filter can be applied.
[0573] e) In one example, a filter with low-pass property can be applied.
[0574] f) In one example, a filter with high-pass property can be applied.
[0575] d. In one example, the intermediate filtering result of other online trained filter(s) can be used as input for one or more extended taps.
[0576] e. In one example, the input for generating the intermediate filtering result can be the reconstructed samples at different coding stages.
[0577] a) In one example, the reconstruction before / after ALF of the current frame / reference frame can be used to generate the intermediate filtering result.
[0578] b) In one example, the reconstruction before / after SAO / CCSAO of the current frame / reference frame can be used to generate the intermediate filtering result.
[0579] c) In one example, the reconstruction before / after BIF of the current frame / reference frame can be used to generate the intermediate filtering result.
[0580] d) In one example, the reconstruction before / after DBF of the current frame / reference frame can be used to generate the intermediate filtering result.
[0581] e) In one example, the reconstruction before / after any other stage of the current frame / reference frame can be used to generate the intermediate filtering result.
[0582] 3) It is proposed to apply conditional filter shape switching for ALF.
[0583] a. In one example, shape switching is applied only for spatial shape.
[0584] b. In one example, shape switching is applied only for shape of extended taps.
[0585] a) In one example, one possible shape of the extended taps can be diamond 1x1.
[0586] b) In one example, one possible shape of the extended taps can be diamond 3x3.
[0587] c) In one example, one possible shape of the extended taps can be diamond 5x5.
[0588] d) In one example, one possible shape of the extended taps can be cross 5x5.
[0589] e) In one example, one possible shape of the extended tap can be cross 7x7.
[0590] f) In one example, one possible shape of the extended tap can be cross 9x9.
[0591] g) In one example, one possible shape of the extended tap can be diamond / cross / square of size N.
[0592] 1. In one example, N can be equal to 1 / 3 / 5 / 7 / 9 / 11 / 13.
[0593] 2. In one example, N can be equal to 0. In this case, it results in disabling the extended tap.
[0594] c. In one example, shape switching can be jointly applied to the spatial domain shape and the shape of the extended tap.
[0595] d. In one example, shape switching can be based on configuration setting.
[0596] a) In one example, the spatial domain shape can be different for all-intra than for other configurations.
[0597] b) In one example, the spatial domain shape can be different for random access than for other configurations.
[0598] c) In one example, the spatial domain shape can be different for low-delay B / P than for other configurations.
[0599] d) In one example, the spatial domain shape can be the same for random access and low-delay B / P.
[0600] e) In one example, the shape of the extended tap can be different for all-intra than for other configurations.
[0601] f) In one example, the shape of the extended tap can be different for random access than for other configurations.
[0602] g) In one example, the shape of the extended tap can be different for low-delay B / P than for other configurations.
[0603] h) In one example, the shape of the extended tap can be the same for random access and low-delay B / P.
[0604] e. In one example, shape switching can be based on slice type.
[0605] a) In one example, the spatial domain shape can be different for I slice than for other slice types.
[0606] b) In one example, B slices can use a different spatial shape than other slice types.
[0607] c) In one example, P slices can use a different spatial shape than other slice types.
[0608] d) In one example, B slices and P slices can use the same spatial shape.
[0609] e) In one example, I slices can use a different shape of extended tap than other slice types.
[0610] f) In one example, B slices can use a different shape of extended tap than other slice types.
[0611] g) In one example, P slices can use a different shape of extended tap than other slice types.
[0612] h) In one example, B slices and P slices can use the same shape of extended tap.
[0613] f. In one example, shape switching can be based on configuration setting or slice type only.
[0614] a) In one example, shape of extended tap can be cross 5x5 in full intra and diamond 1x1 in random access and low delay BP.
[0615] b) In one example, shape of extended tap can be cross 5x5 in I slice and diamond 1x1 in B slice / P slice.
[0616] g. In one example, shape switching can be based on configuration setting and slice type jointly.
[0617] h. In one example, shape switching decision can be predefined / signaled / derived on the fly.
[0618] 4) In one example, the disclosed method can be used for post-processing and / or pre-processing.
[0619] 5) In one example, the above methods can be used jointly.
[0620] 6) Optionally, the above methods can be used individually.
[0621] 7) In one example, the proposed / described extended tap for ALF method can be applied to any in-loop filtering tool, pre-processing or post-processing filtering method in video coding including but not limited to ALF / CCALF or any other filtering method.
[0622] a. In one example, the proposed extended-tap method can be applied to in-loop filtering methods.
[0623] a) In one example, the proposed extended-tap method can be applied to ALF.
[0624] b) In one example, the proposed extended-tap method can be applied to CCALF.
[0625] c) Optionally, the proposed extended-tap method can be applied to other in-loop filtering methods.
[0626] b. In one example, the proposed extended-tap method can be applied to pre-processing filtering methods.
[0627] c. In one example, the proposed extended-tap method can be applied to post-processing filtering methods.
[0628] 8) In the above examples, the video unit can refer to sequence / picture / sub-picture / slice / tile / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transform unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transform block (TB) / any other region containing more than one luma or chroma sample / pixel.
[0629] 9) Whether and / or how to apply the above disclosed methods can be signaled in the bitstream.
[0630] a. In one example, whether and / or how to apply the above disclosed methods can be signaled at sequence level / picture group level / picture level / slice level / tile group level, such as in sequence header / picture header / S PS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.
[0631] b. In one example, whether and / or how to apply the above disclosed methods can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / tile / sub-picture / other kind of region containing more than one sample or pixel.
[0632] 10) Whether and / or how to apply the above disclosed methods can depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type.
[0633] 6. References
[0634] [1] J. Strom, P. Wennersten, J. Enhorn, D. Liu, K. Andersson and R. Sjoberg, “Bilateral Loop Filter in Combination with SAO,” in proceeding of IEEE Picture Coding Symposium (PCS), Nov. 2019.
[0635] Figure 13 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 4000. The system 4000 can include an input 4002 to receive video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be received in a compressed or encoded format. The input 4002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as wireless-fidelity (Wi-Fi) or cellular interfaces.
[0636] The system 4000 can include a codec component 4004 that can implement various coding or encoding methods described in this disclosure. The codec component 4004 can reduce the average bitrate of video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. The coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 can be stored, or transmitted via a communication connection as represented by component 4006. The bitstream (or coded) representation of the video received at the input 4002, or stored or communicated transmission, can be used by component 4008 to generate pixel values or displayable video that is transmitted to a display interface 4010. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “coding” operations or tools, it can be understood that the coding tools or operations are used by an encoder, and that the decoding tools or operations that reverse the coding results will be performed by a decoder.
[0637] Examples of a peripheral bus interface or display interface can include a universal serial bus (USB) or a high-definition multimedia interface (HDMI) or a display interface, etc. Examples of a storage interface can include a serial advanced technology attachment (SATA), a peripheral component interconnect (PCI), an integrated drive electronics (IDE) interface, etc. The techniques described in this disclosure can be embodied in various electronic devices, such as a mobile phone, a laptop, a smart phone, or other devices capable of performing digital data processing and / or video display.
[0638] Figure 14 is a block diagram of an example video processing apparatus 4100. The apparatus 4100 can be used to implement one or more methods described herein. The apparatus 4100 can be embodied in a smart phone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor(s) 4102 can be configured to implement one or more methods described in this disclosure. The memory(ies) 4104 can be used for storing data and code used during operation of the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be included at least in part in the processor 4102, e.g., a graphics co-processor.
[0639] Figure 15 is a flowchart of an example method 4200 of video processing. The method 4200 includes determining, at step 4202, to use at least one extended tap in an adaptive loop filter (ALF). A conversion between video data and a bitstream is performed based on the ALF, at step 4204. According to an example, the conversion of step 4204 can include encoding at an encoder or decoding at a decoder.
[0640] It is noted that the method 4200 can be implemented in an apparatus that processes video data, the apparatus including a processor and a non-transitory memory having instructions thereon, such as the video encoder 4400, the video decoder 4500, and / or the encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform the method 4200. Furthermore, the method 4200 can be performed by a non-transitory computer readable medium including a computer program product for use by a video codec device. The computer program product includes computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause a video codec device to perform the method 4200.
[0641] Figure 16is a block diagram illustrating an example video coding system 4300 that can utilize the techniques of this disclosure. Video coding system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where the source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where the destination device 4320 can be referred to as a video decoding device.
[0642] Source device 4310 can include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to destination device 4320 via I / O interface 4316 and network 4330. The encoded video data can also be stored onto a storage medium / server 4340 for access by destination device 4320.
[0643] Destination device 4320 can include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 can include a receiver and / or a modem. I / O interface 4326 can acquire, from source device 4310 or storage medium / server 4340, the encoded video data. Video decoder 4324 can decode the encoded video data. Display device 4322 can display the decoded video data to a user. Display device 4322 can be integrated with destination device 4320, or can be external to destination device 4320, which can be configured to interface with an external display device.
[0644] Video encoder 4314 and video decoder 4324 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or further standards.
[0645] Figure 17 is a block diagram illustrating an example of a video encoder 4400, which can beFigure 16 The video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0646] The functional components of the video encoder 4400 can include a partitioning unit 4401, a prediction unit 4402 (which can include a mode select unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414.
[0647] In other examples, the video encoder 4400 can include more, less, or different functional components. In one example, the prediction unit 4402 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0648] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405, can be highly integrated, but are represented separately in the example of the video encoder 4400 for explanatory purposes.
[0649] The partitioning unit 4401 can partition a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0650] The mode select unit 4403 can select one of a plurality of coding modes (intra- coded or inter-coded) based on, for example, error results, and provide a resulting intra- coded block or inter-coded block to the residual generation unit 4407 for generation of residual block data and to the reconstruction unit 4412 for reconstruction of the coded block for use as a reference picture. In some examples, the mode select unit 4403 can select a combined intra-inter prediction (CIIP) mode, in which prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode select unit 4403 can also select a resolution for a motion vector (e.g., sub-pixel precision or integer pixel precision) for the block.
[0651] To perform inter prediction on a current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 to the current video block. Motion compensation unit 4405 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 4413 other than the picture associated with the current video block.
[0652] Motion estimation unit 4404 and motion compensation unit 4405 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0653] In some examples, motion estimation unit 4404 can perform uni-prediction on a current video block, and motion estimation unit 4404 can search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 4404 can then generate a reference index that indicates a reference picture in list 0 or list 1 that contains the reference video block and a motion vector that indicates a spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0654] In other examples, motion estimation unit 4404 can perform bi-prediction on a current video block, motion estimation unit 4404 can search reference pictures in list 0 for a reference video block for the current video block and also search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 4404 can then generate a reference index that indicates reference pictures in list 0 and list 1 that contain the reference video blocks and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference indices and the motion vectors for the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0655] In some examples, motion estimation unit 4404 can output a full set of motion information for a current video block for decoding processing at a decoder. In some examples, motion estimation unit 4404 can not output a full set of motion information for a current video block. Instead, motion estimation unit 4404 can reference motion information for another video block to signal motion information for the current video block. For example, motion estimation unit 4404 can determine that the motion information for a current video block is sufficiently similar to motion information for a neighboring video block.
[0656] In one example, the motion estimation unit 4404 can indicate, in a syntax structure associated with the current video block, a value that indicates that the current video block has the same motion information as another video block to the video decoder 4500.
[0657] In another example, the motion estimation unit 4404 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0658] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0659] The intra prediction unit 4406 can perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0660] The residual generation unit 4407 can generate residual data for the current video block by subtracting the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0661] In other examples, such as in skip mode, the current video block can not have residual data, and the residual generation unit 4407 can not perform the subtraction operation.
[0662] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0663] After the transform processing unit 4408 generates the transform coefficient video blocks associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0664] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.
[0665] After the reconstruction unit 4412 reconstructs the video block, an in-loop filtering operation can be performed to reduce video block artifacts in the video block.
[0666] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0667] Figure 18 is a block diagram illustrating an example of a video decoder 4500 that can be Figure 16 the system 4300 shown. The video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, the video decoder 4500 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0668] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 4400.
[0669] The entropy decoding unit 4501 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy encoded video data and, from the entropy decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. The motion compensation unit 4502 may, for example, determine this information by performing AMVP and Merge modes.
[0670] Motion compensation unit 4502 can generate the motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel accuracy can be included in the syntax elements.
[0671] Motion compensation unit 4502 can calculate the interpolation for sub-integer pixels of the reference block using the interpolation filter as used by video encoder 4400 during encoding of the video block. Motion compensation unit 4502 can determine the interpolation filter used by video encoder 4400 from the received syntax information and can use the interpolation filter to generate the prediction block.
[0672] Motion compensation unit 4502 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information used to decode the encoded video sequence.
[0673] Intra prediction unit 4503 can use, for example, intra prediction modes received in the bitstream to form the prediction block from spatially neighboring blocks. Dequantization unit 4504 dequantizes (i.e., inverse quantizes) quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.
[0674] Reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by motion compensation unit 4502 or intra prediction unit 4503 to form a decoded block. If desired, a deblocking filter can also be used to filter the decoded block to remove blockiness artifacts. The decoded video blocks are then stored in buffer 4507, which provides reference blocks for subsequent motion compensation / intra prediction, and also generates decoded video for presentation on a display device.
[0675] Figure 19is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques of VVC. The encoder 4600 includes three in-loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a pre-defined filter, the SAO 4604 and the ALF 4606 utilize original samples of a current picture, respectively by adding an offset and by applying a finite impulse response (FIR) filter, and utilize coded side information to signal the offset and the filter coefficients to reduce the mean square error between the original samples and the reconstructed samples. The ALF 4606 is located at the last processing stage of each picture and can be viewed as a tool that attempts to capture and fix artifacts caused by previous stages.
[0676] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction with reference pictures obtained from a reference picture cache 4612. Residual blocks from either inter or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding component 4618. The entropy coding component 4618 entropy codes the prediction results and the quantized transform coefficients and transmits them toward a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting pictures to the DF 4602, the SAO 4604, and the ALF 4606 to filter them before they are stored in the reference picture cache 4612.
[0677] Next, a list of some example preferred solutions is provided.
[0678] The following solutions show examples of embodiments discussed herein.
[0679] 1. A method for processing video data, comprising: determining to use at least one extended tap in an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
[0680] 2. The method of solution 1, wherein the extended taps receive input from: spatially neighboring samples of reconstructed after a deblocking filter (DBF), a sample adaptive offset (SAO) filter, a cross-component SAO (CCSAO), or a bilateral filter (BF); spatially neighboring luma samples for filtering a center luma sample; or spatially neighboring chroma samples for filtering a center chroma sample.
[0681] 3. The method of any of solutions 1-2, wherein the ALF comprises M spatial taps and N extended taps, where M and N are both greater than 0.
[0682] 4. The method of any of solutions 1-3, wherein the extended taps receive input from one or more of a Gaussian filter, a low pass filter, a high pass filter, or a combination thereof.
[0683] 5. The method of any of solutions 1-4, wherein the input source for the extended taps is derived based on samples in different color components.
[0684] 6. The method of any of solutions 1-5, wherein the ALF is applied to filter only a luma component, only one chroma component, all chroma components, or all luma and chroma components.
[0685] 7. The method of any of solutions 1-6, wherein the filter shapes for spatial taps and the extended taps can be the same or different.
[0686] 8. The method of any of solutions 1-7, wherein the filter shape for spatial taps is a diamond, a square, a cross, a symmetric shape, an asymmetric shape, a designed shape, or a combination thereof, and wherein the filter shape for the spatial taps is signaled, derived, or determined.
[0687] 9. The method of any of solutions 1-8, wherein the filter shape for the extended taps is a diamond, a square, a cross, a symmetric shape, an asymmetric shape, a designed shape, or a combination thereof, and wherein the filter shape for the extended taps is signaled, derived, or determined.
[0688] 10. The method of any of solutions 1-9, wherein the filter size for the extended taps is MxN, where M is 1, 3, or 5, and N is 1, 3, or 5.
[0689] 11. The method according to any of solutions 1-10, wherein the syntax element indicates whether the ALF is enabled or not, and wherein the syntax element is coded by arithmetic coding, context coding, bypass coding, unary code, truncated unary code, fixed length code, exponential Golomb code, truncated exponential Golomb code, conditional coding, predictive coding, independent signaling for different color components, coding in a header or parameter set, or a combination thereof.
[0690] 12. The method according to any of solutions 1-11, wherein the syntax element indicates the input source for the extended tap, and wherein the syntax element is coded by arithmetic coding, context coding, bypass coding, unary code, truncated unary code, fixed length code, exponential Golomb code, truncated exponential Golomb code, conditional coding, predictive coding, independent signaling for different color components, coding in a header or parameter set, or a combination thereof.
[0691] 13. The method according to any of solutions 1-12, wherein the coefficients, clipping parameters, category merge results or other parameters of the extended tap are signaled in an adaptive parameter set (APS).
[0692] 14. The method according to any of solutions 1-13, wherein the coefficients of the extended tap are predefined fixed values or coded according to predictive coding, arithmetic coding with context, bypass coding, joint coding with coefficients of spatial domain taps, or a combination thereof.
[0693] 15. The method according to any of solutions 1-14, wherein an intermediate filtering result of a fixed filter or adaptive filter is used as input for the extended tap.
[0694] 16. The method according to any of solutions 1-15, wherein the intermediate filtering result is a result of an offline trained filter of ALF and a reconstruction before ALF, a result of an offline trained filter of ALF and a reconstruction before DBF, a result of an online trained ALF filter, a Gaussian filter result, a BF result, a directional filter result, a median filter result, a local median filter result, a non-local median filter result, a low pass filter result, a high pass filter result, a result of an online trained filter, a reconstructed sample before applying ALF, a reconstructed sample after applying ALF, a reconstructed sample before applying SAO, a reconstructed sample after applying SAO, a reconstructed sample before applying BF, a reconstructed sample after applying BF, a reconstructed sample before applying DBF, a reconstructed sample after applying DBF, or a combination thereof.
[0695] 17. The method of any of solutions 1-16, wherein a conditional filter shape is applied to the ALF.
[0696] 18. The method of any of solutions 1-17, wherein the conditional filter shape is applied to a spatial domain shape.
[0697] 19. The method of any of solutions 1-18, wherein the conditional filter shape is applied to a shape of the extended tap.
[0698] 20. The method of any of solutions 1-19, wherein the conditional filter shape comprises a 1x1 diamond, a 3x3 diamond, a 5x5 diamond, a 5x5 cross, a 7x7 cross, a 9x9 cross, or a NxN shape, where N equals 0, 1, 3, 5, 7, 9, 11, or 13.
[0699] 21. The method of any of solutions 1-20, wherein the conditional filter shape is jointly applied to a spatial domain shape and a shape of the extended tap.
[0700] 22. The method of any of solutions 1-21, wherein the conditional filter shape is based on a configured shape switch.
[0701] 23. The method of any of solutions 1-22, wherein an all-intra, random access, or low delay B / P employs a different spatial domain shape than other configurations or the same spatial domain shape as a spatial domain shape of the extended tap.
[0702] 24. The method of any of solutions 1-23, wherein an intra (I) slice, a predictive (P) slice, a bi-predictive (B) slice, or an extended tap thereof, uses different shapes, the same shape, or a combination thereof.
[0703] 25. The method of any of solutions 1-24, wherein the shape switch is based on a configuration setting or a slice type, or a shape switch decision is predefined, signaled, or derived.
[0704] 26. The method of any of solutions 1-25, wherein the shape of the extended tap comprises a 5x5 cross for all-intra, a 1x1 diamond for random access, a 1x1 diamond for low delay BP, a 5x5 cross for I slice, a 1x1 diamond for B slice, a 1x1 diamond for P slice, or a combination thereof.
[0705] 27. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of solutions 1-26.
[0706] 28. A non-transitory computer readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to perform the method of any of solutions 1-26.
[0707] 29. A non-transitory computer readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises determining to use at least one extended tap in an adaptive loop filter (ALF), and generating the bitstream based on the determination.
[0708] 30. A method of storing a bitstream of a video, comprising determining to use at least one extended tap in an adaptive loop filter (ALF), generating the bitstream based on the determination, and storing the bitstream in a non-transitory computer readable recording medium.
[0709] 31. A method, apparatus or system described in the present disclosure.
[0710] In the described solutions, an encoder can conform to a format rule by producing a coded representation according to the format rule. In the described solutions, a decoder can parse syntax elements in a coded representation according to the format rule with known information of presence and absence of syntax elements to produce a decoded video.
[0711] In the present disclosure, the term “video processing” can refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, a bitstream representation of a current video block can correspond to a bit position in a bitstream defined by a syntax or bits that are spread over different positions. For example, a macroblock can be encoded according to transformed and coded error residual values, and also using bits in headers and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination that some fields can be present or absent, as described in the above solutions. Similarly, an encoder can determine to include or not include a particular syntax field, and generate a coded representation accordingly by including the syntax field or excluding the syntax field from the coded representation.
[0712] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this disclosure can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include, in addition to a hardware component, code that creates an execution environment for a program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0713] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or code portions). A computer program can be deployed for execution on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0714] The processes and logic flows described in this disclosure can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0715] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or input to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or tape. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0716] Although the present disclosure includes a number of details, these should not be construed as limitations on any subject matter or scope of protection that can be required to practice the subject technology. In the present disclosure, certain features that are described in the context of separate embodiments can also be implemented in combination with each other. Conversely, various features that are described in the context of a single embodiment can also be implemented separately from each other. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0717] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such an order, nor that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this disclosure should not be understood as requiring such separation in all embodiments.
[0718] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this disclosure.
[0719] A first component is directly coupled to a second component when there are no intervening components between the first component and the second component other than a wire, trace, or other medium between the first component and the second component. A first component is indirectly coupled to a second component when there are intervening components between the first component and the second component other than a wire, trace, or other medium between the first component and the second component. The term “coupled” and variations thereof include both direct and indirect couplings. The use of the term “about” means a range of ±10% of the subsequent number, unless otherwise indicated.
[0720] While multiple embodiments are provided in the present disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention is not to limit the details given herein. For example, various elements or components can be combined or integrated within another system, or certain features can be omitted or not implemented.
[0721] Also, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as separate from the other items or hierarchical levels can be implemented as integrated components in accordance with the principles of the present disclosure. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: determining to use at least one extended tap in an adaptive loop filter (ALF); and performing a conversion between visual media data and a bitstream based on the ALF.
2. The method of claim 1, wherein, The at least one extended tap is different from spatial domain taps in the ALF.
3. The method of claim 2, wherein, The ALF uses only information of spatial neighboring samples of a target component, and wherein the spatial neighboring samples are neighboring a center sample to be filtered.
4. The method of any one of claims 1-3, wherein, The spatial neighboring samples are obtained from a de-blocking filter (DBF), a sample adaptive offset (SAO) filter, a cross-component SAO (CCSAO), or a bilateral filter (BF) after reconstruction.
5. The method of claim 2, wherein, The spatial domain taps use only spatial neighboring luma samples to filter a center luma sample within a single ALF.
6. The method of claim 2, wherein, The spatial domain taps use only spatial neighboring chroma samples to filter a center chroma sample within a single ALF.
7. The method of claim 1, wherein, The at least one extended tap is configured to co-exist with at least one spatial domain tap in a single ALF.
8. The method of claim 1, wherein, The ALF consists of both spatial domain taps and extended taps.
9. The method of claim 1, wherein, The ALF consists of M spatial domain taps and N extended taps, where M and N are each positive integers.
10. The method of claim 1, wherein, The at least one extended tap uses one or more input sources.
11. The method of claim 1, wherein, The at least one extended tap uses a single input source.
12. The method of claim 11, wherein, The input source includes intermediate filtering results from one or more predefined filters.
13. The method of claim 12, wherein, The one or more predefined filters include a Gaussian filter.
14. The method of claim 12, wherein, The one or more predefined filters include a low-pass filter.
15. The method of claim 12, wherein, The one or more predefined filters include a high-pass filter.
16. The method of any one of claims 12-15, wherein, The input source includes a modification of the Gaussian filter, a modification of the low-pass filter, or a modification of the high-pass filter.
17. The method of any one of claims 12-15, wherein, The input source includes a fusion or a weighted sum of the Gaussian filter, the low-pass filter, and the high-pass filter.
18. The method of claim 1, wherein, The at least one extended tap uses multiple input sources.
19. The method of claim 18, wherein, The multiple input sources include intermediate filtering results from one or more predefined filters.
20. The method of claim 19, wherein, The one or more predefined filters include a Gaussian filter.
21. The method of claim 19, wherein, The one or more predefined filters include a low-pass filter.
22. The method of claim 19, wherein, The one or more predefined filters include a high-pass filter.
23. The method of any one of claims 20-22, wherein, The multiple input sources include a modification of the Gaussian filter, a modification of the low-pass filter, or a modification of the high-pass filter.
24. The method of any one of claims 20-22, wherein, The input source includes a fusion or a weighted sum of the Gaussian filter, the low-pass filter, and the high-pass filter.
25. The method of claim 1, wherein, The input source of the at least one extended tap is derived based on samples in different color components.
26. The method of claim 1, wherein, Whether and / or how to apply the ALF with the at least one extended tap is different for different color formats and / or different color components.
27. The method of any one of claims 1-26, wherein, The ALF with the at least one extended tap is only applied to process a luma component.
28. The method of any one of claims 1-26, wherein, The ALF with the at least one extended tap is only applied to process one chroma component, and wherein the one chroma component includes a blue-difference chroma component (Cb) or a red-difference chroma component (Cr).
29. The method of any one of claims 1-26, wherein, The ALF with the at least one extended tap is applied to process all chroma components, and wherein all chroma components include a blue-difference chroma component (Cb) and a red-difference chroma component (Cr).
30. The method of any one of claims 1-26, wherein, The ALF with the at least one extended tap is applied to process a luma component and chroma components, and wherein the luma component and the chroma components include a luma component (Y), a blue-difference chroma component (Cb), and a red-difference chroma component (Cr).
31. The method of any one of claims 1-26, wherein, The ALF with the at least one extended tap uses a different shape or size than an ALF filter without an extended tap uses.
32. The method of claim 31, wherein, The ALF with the at least one extended tap includes different shapes for spatial taps and extended taps.
33. The method of claim 31, wherein, The shape and / or size used by the ALF for one or more spatial taps is different than the shape and / or size used by the ALF for an extended tap.
34. The method of claim 31, wherein, The shape and / or size used by the ALF for one or more spatial taps is the same as the shape and / or size used by the ALF for an extended tap.
35. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is diamond-shaped.
36. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is square-shaped.
37. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is cross-shaped.
38. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is a symmetric shape.
39. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is an asymmetric shape.
40. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is a pre-designed shape.
41. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is determined in real-time, specified in the bitstream, or derived.
42. The method of claim 31, wherein, The filter shape used by the ALF for one or more extended taps is diamond-shaped.
43. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is square-shaped.
44. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is cross-shaped.
45. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is a symmetric shape.
46. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is an asymmetric shape.
47. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is a pre-designed shape.
48. The method of claim 31, wherein, The filter shape used by the ALF for one or more spatial taps is determined in real-time, specified in the bitstream, or derived.
49. The method of any one of claims 1-48, wherein, The filter size used by the ALF for the one or more extended taps is MxN, where M and N are each positive integers.
50. The method of claim 49, wherein, M is equal to N.
51. The method of claim 49, wherein, M or N is equal to 1.
52. The method of claim 49, wherein, M or N is equal to 3.
53. The method of claim 49, wherein, M or N is equal to 5.
54. The method of claim 49, wherein, M is not equal to N.
55. The method of any one of claims 1-54, wherein, The bitstream includes a first syntax element to indicate whether the ALF with the at least one extended tap is enabled.
56. The method of claim 55, wherein, The first syntax element is coded by arithmetic coding.
57. The method of claim 55, wherein, The first syntax element is coded with at least one context.
58. The method of claim 57, wherein, The at least one context depends on coding information of the current block or a neighboring block.
59. The method of claim 57, wherein, The at least one context depends on a filter shape of at least one neighboring block.
60. The method of claim 55, wherein, The first syntax element is coded by bypass coding.
61. The method of claim 55, wherein, The first syntax element is binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, or a truncated exponential Golomb code.
62. The method of claim 55, wherein, The first syntax element is conditionally included in the bitstream.
63. The method of claim 62, wherein, The first syntax element is included in the bitstream only if the at least one extended tap is available.
64. The method of claim 55, wherein, The first syntax element is coded in a predictive manner.
65. The method of claim 64, wherein, The first syntax element is predicted based on whether an extended tap of at least one neighboring block is on or off.
66. The method of claim 55, wherein, The first syntax element is independently included in the bitstream for different color components.
67. The method of claim 55, wherein, The first syntax element is shared for different color components.
68. The method of claim 55, wherein, The first syntax element is included in the bitstream for a first color component but not for a second color component.
69. The method of claim 55, wherein, The first syntax element is included in a bitstream in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptive parameter set (APS), a coding tree unit (CTU), or a coding unit (CU).
70. The method of any one of claims 1-54, wherein, The bitstream includes a first syntax element to indicate which input sources are used for the at least one extended tap in the ALF.
71. The method of claim 70, wherein, The first syntax element is coded by arithmetic coding.
72. The method of claim 70, wherein, The first syntax element is coded with at least one context.
73. The method of claim 72, wherein, The at least one context depends on coding information of the current block or a neighboring block.
74. The method of claim 72, wherein, The at least one context depends on a filter shape of at least one neighboring block.
75. The method of claim 70, wherein, The first syntax element is coded by bypass coding.
76. The method of claim 70, wherein, The first syntax element is binarized by a unary code, a truncated unary code, a fixed length code, an exponential Golomb code, or a truncated exponential Golomb code.
77. The method of claim 70, wherein, The first syntax element is conditionally included in the bitstream.
78. The method of claim 77, wherein, The first syntax element is included in the bitstream only if the at least one extended tap is available.
79. The method of claim 70, wherein, The first syntax element is coded in a predictive manner.
80. The method of claim 79, wherein, The first syntax element is predicted based on whether an extended tap of at least one neighboring block is on or off.
81. The method of claim 70, wherein, The first syntax element is independently included in the bitstream for different color components.
82. The method of claim 70, wherein, The first syntax element is shared for different color components.
83. The method of claim 70, wherein, The first syntax element is included in the bitstream for a first color component but not for a second color component.
84. The method of claim 70, wherein, The first syntax element is included in a bitstream in a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice header, an adaptive parameter set (APS), a coding tree unit (CTU), or a coding unit (CU).
85. The method of claim 84, wherein, The first syntax element is included in the APS in the bitstream.
86. The method of any one of claims 1-54, wherein, One or more coefficients of the at least one extended tap of the ALF are included in a syntax element structure in the bitstream.
87. The method of claim 86, wherein, The syntax element structure includes an adaptive parameter set (APS).
88. The method of claim 87, wherein, The APS includes a clipping parameter for the at least one extended tap.
89. The method of claim 87, wherein, The APS includes a category merge result for the at least one extended tap.
90. The method of claim 86, wherein, The one or more coefficients are coded in a predictive manner.
91. The method of claim 86, wherein, The one or more coefficients are coded by arithmetic coding with at least one context.
92. The method of claim 86, wherein, The one or more coefficients are coded by bypass coding.
93. The method of claim 86, wherein, The one or more coefficients are jointly coded with coefficients of a spatial domain tap.
94. The method of any one of claims 86-93, wherein, The APS includes an additional parameter for the extended tap.
95. The method of claim 86, wherein, The one or more coefficients include a predefined fixed value.
96. The method of any one of claims 1-95, wherein, An intermediate filtering result of at least one fixed filter or an intermediate filtering result of at least one adaptive filter is used as input for the at least one extended tap.
97. The method of claim 96, wherein, The intermediate filtering result is from an offline trained filter.
98. The method of claim 96, wherein, The intermediate filtering result is generated from a reconstruction before the ALF and an offline trained filter of the ALF.
99. The method of claim 96, wherein, The intermediate filtering result is generated from a reconstruction before the ALF and a deblocking filter (DBF) of the ALF.
100. The method of claim 96, wherein, The intermediate filtering result is from an online trained ALF filter.
101. The method of claim 96, wherein, The intermediate filtering result is from a predefined filter.
102. The method of claim 101, wherein, The predefined filter is a Gaussian filter.
103. The method of claim 101, wherein, The predefined filter is a bilateral filter.
104. The method of claim 101, wherein, The predefined filter is a directional filter.
105. The method of claim 101, wherein, The predefined filter is a median filter.
106. The method of claim 105, wherein, The median filter is a local median filter.
107. The method of claim 105, wherein, The median filter is a non-local median filter.
108. The method of claim 105, wherein, The median filter is a filter with low pass properties.
109. The method of claim 105, wherein, The median filter is a filter with high pass properties.
110. The method of claim 96, wherein, The intermediate filtering result is from an online trained filter.
111. The method of claim 96, wherein, The input for generating the intermediate filtering result includes reconstructed samples at different coding stages.
112. The method of claim 96, wherein, The input for generating the intermediate filtering result includes reconstructed samples before the ALF for a current frame, reconstructed samples after the ALF for a current frame, reconstructed samples before the ALF for a reference frame, or reconstructed samples after the ALF for a reference frame.
113. The method of claim 96, wherein, The input for generating the intermediate filtering result includes reconstructed samples before a sample adaptive offset (SAO) filter for a current frame, reconstructed samples after the SAO filter for a current frame, reconstructed samples before a cross component SAO (CCSAO) filter for a reference frame, or reconstructed samples after the CCSAO filter for a reference frame.
114. The method of claim 96, wherein, The input for generating the intermediate filtering result includes reconstructed samples before a bilateral filter (BF) for a current frame, reconstructed samples after the BF for a current frame, reconstructed samples before the BF for a reference frame, or reconstructed samples after the BF for a reference frame.
115. The method of claim 96, wherein, The input for generating the intermediate filtering result includes reconstructed samples before a deblocking filter (DBF) for a current frame, reconstructed samples after the DBF for a current frame, reconstructed samples before the DBF for a reference frame, or reconstructed samples after the DBF for a reference frame.
116. The method of claim 96, wherein, The input for generating the intermediate filter result comprises reconstructed samples before the filter stage for the current frame, reconstructed samples after the filter stage for the current frame, reconstructed samples before the filter stage for a reference frame, or reconstructed samples after the filter stage for a reference frame.
117. The method of any of claims 1-116, further comprising determining to apply a conditional filter shape switch to the ALF.
118. The method of claim 117, wherein, The conditional filter shape switch is only applied to a spatial domain shape.
119. The method of claim 117, wherein, The conditional filter shape switch is only applied to a shape of the at least one extended tap.
120. The method of claim 119, wherein, The shape is a 1x1 diamond.
121. The method of claim 119, wherein, The shape is a 3x3 diamond.
122. The method of claim 119, wherein, The shape is a 5x5 diamond.
123. The method of claim 119, wherein, The shape is a 5x5 cross.
124. The method of claim 119, wherein, The shape is a 7x7 cross.
125. The method of claim 119, wherein, The shape is a 9x9 cross.
126. The method of claim 119, wherein, The shape is a diamond or cross of size N, where N is a positive integer.
127. The method of claim 126, wherein, N is 1, 3, 5, 7, 9, 11, or 13.
128. The method of claim 126, wherein, N is 0 to indicate that the at least one extended tap is disabled.
129. The method of claim 119, wherein, The conditional filter shape switch is jointly applied to a spatial domain shape and a shape of the at least one extended tap.
130. The method of claim 119, wherein, The conditional filter shape switch is based on a configuration setting.
131. The method of claim 130, wherein, The all-intra configuration setting uses a different spatial domain shape than other configuration settings.
132. The method of claim 130, wherein, The random access configuration setting uses a different spatial domain shape than other configuration settings.
133. The method of claim 130, wherein, The low-delay bi-directional (B) / uni-directional (P) configuration setting uses a different spatial domain shape than other configuration settings.
134. The method of claim 130, wherein, The random access configuration setting and the low-delay bi-directional (B) / uni-directional (P) configuration setting use the same shape.
135. The method of claim 130, wherein, The all-intra configuration setting uses a different shape of the at least one extended tap than other configuration settings.
136. The method of claim 130, wherein, The random access configuration setting uses a different shape of the at least one extended tap than other configuration settings.
137. The method of claim 130, wherein, The low-delay bi-directional (B) / uni-directional (P) configuration setting uses a different shape of the at least one extended tap than other configuration settings.
138. The method of claim 130, wherein, The random access configuration setting and the low-delay bi-directional (B) / uni-directional (P) configuration setting use the same shape of the at least one extended tap.
139. The method of claim 117, wherein, The conditional filter shape switch is based on a slice type.
140. The method of claim 139, wherein, The slice type is an I slice, and wherein the I slice uses a different spatial domain shape than other slice types.
141. The method of claim 139, wherein, The slice type is a B slice, and wherein the B slice uses a different spatial domain shape than other slice types.
142. The method of claim 139, wherein, The slice type is a P slice, and wherein the P slice uses a different spatial domain shape than other slice types.
143. The method of claim 139, wherein, The B slice and the P slice use the same spatial domain shape.
144. The method of claim 139, wherein, The I slice uses a different shape of the at least one extended tap than other slice types.
145. The method of claim 139, wherein, The B slice uses a different shape of the at least one extended tap than other slice types.
146. The method of claim 139, wherein, The P slice uses a different shape of the at least one extended tap than other slice types.
147. The method of claim 139, wherein, The B slice and the P slice use the same shape of the at least one extended tap.
148. The method of claim 117, wherein, The conditional filter shape switch is only based on a configuration setting or a slice type.
149. The method of claim 148, wherein, The shape of the at least one extended tap is 5x5 cross in full frame intra configuration setting and 1x1 diamond in random access and low delay bi-directional (B) / uni-directional (P) configuration setting.
150. The method of claim 148, wherein, The shape of the at least one extended tap is 5x5 cross for I slice and 1x1 diamond for B slice or P slice.
151. The method of claim 117, wherein, The conditional filter shape switching is based on the configuration setting and the slice type jointly.
152. The method of claim 117, wherein, The decision for the conditional filter shape switching is predefined, included in the bitstream or derived in real time.
153. The method of any one of claims 1-152, wherein, The method is used in post-processing and / or pre-processing.
154. The method of any one of claims 1-152, wherein, The method is used jointly.
155. The method of any one of claims 1-152, wherein, The method is used individually.
156. The method of any one of claims 1-152, wherein, The method is applied to any in-loop filtering tool, any pre-processing filtering method or post-processing filtering method in video coding, including the ALF or cross component ALF (CCALF).
157. The method of any one of claims 1-152, wherein, The method is applied to in-loop filtering method including the ALF or CCALF.
158. The method of any one of claims 1-157, wherein, The method is applied to video units, and wherein the video units include sequence, picture, sub-picture, slice, tile, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), any other region containing more than one luma or chroma sample or pixel.
159. The method of any one of claims 1-158, wherein, Whether and / or how to apply the method is signaled in the bitstream.
160. The method of claim 159, wherein, The signaling is performed at sequence level, group of pictures level, picture level, slice level, tile group level, or in sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoder capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header or tile group header.
161. The method of claim 159, wherein, The signaling is performed in prediction block (PB), transform block (TB), coding block (CB), prediction unit (PU), transform unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), coding tree unit (CTU), CTU row, slice, tile, sub-picture or other region containing more than one sample or pixel.
162. The method of claim 159, wherein, Whether and / or how to apply the method depends on coded information, and wherein the coded information includes block size, color format, single tree partitioning, dual tree partitioning, color component, slice type or picture type.
163. The method of any one of claims 1-162, wherein, The conversion includes encoding the media data into a bitstream.
164. The method of any one of claims 1-162, wherein, The conversion includes decoding the media data from a bitstream.
165. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of claims 1-164. a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of claims 1-164.
166. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to perform the method of any of claims 1-164.
167. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing apparatus, wherein, The method comprises: determining to use at least one extended tap in an adaptive loop filter (ALF); and generating the bitstream based on the determining.
168. A method for storing a bitstream of a video, comprising: determining to use at least one extended tap in an adaptive loop filter (ALF); generating the bitstream based on the determining; and storing the bitstream in a non-transitory computer-readable recording medium.
169. A method, apparatus or system described in the present disclosure.