Method and apparatus for adaptive loop filter selection for position tapping in video coding

By introducing position taps into the adaptive loop filter (ALF) of the video codec system and signaling the target ALF according to horizontal and vertical periods, the problem of position tap selection and signalization of ALF filters in the prior art is solved, and video quality is improved.

CN120077666APending Publication Date: 2025-05-30MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073743.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-24
Filing Date
2023-09-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing video codec systems use adaptive loop filters (ALFs), it is difficult to effectively select and signal the position taps of the ALF filter, resulting in poor video quality.

Method used

By introducing a position tap into the ALF and signaling the target ALF according to horizontal and vertical periods, the position tap and position function of the filter are dynamically adjusted to meet the needs of different codec areas.

Benefits of technology

The video quality of the video encoding and codec system is improved, and the adaptability and effectiveness of the filtered output is enhanced by more fine control of the position tap of the ALF filter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077666A_ABST
    Figure CN120077666A_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding using an adaptive loop filter (ALF). According to the method, a target ALF is derived, where the target ALF comprises one or more position taps, and a position function associated with at least one position tap outputs a variable. A current filtered output is obtained by applying the target ALF to the current block. A filtered reconstructed pixel including the current filtered output is provided. According to another method, a target horizontal cycle and a target vertical cycle are explicitly or implicitly determined, where the target horizontal cycle is determined in a set of horizontal cycles and the target vertical cycle is determined in a set of vertical cycles. A target ALF comprising one or more position taps is determined, wherein a total number of the one or more position taps and one or more respective position functions depends on a target horizontal period and a target vertical period.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This invention is a non - provisional application and claims priority from U.S. Provisional Patent Application No. 63 / 379,923, filed on October 18, 2022, and U.S. Provisional Patent Application No. 63 / 380,590, filed on October 24, 2022. The entire contents of the above - mentioned U.S. provisional patent applications are hereby incorporated herein by reference.

Technical Field

[0003] This invention relates to video coding and decoding systems using an ALF (Adaptive Loop Filter). In particular, this invention relates to the selection of the ALF filter and the signaling of position taps.

Background Art

[0004] Versatile Video Coding (VVC) is the latest international video coding and decoding standard developed by the Joint Video Team (JVET) of the Video Coding Experts Group (VCEG) of the International Telecommunication Union - Telecommunication Standardization Sector (ITU - T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission. This standard has been published as an international standard: ISO / IEC 23090 - 3:2021, Information technology - Coding representation of immersive media - Part 3: Versatile Video Coding, published in February 2021. VVC is developed based on its predecessor, High Efficiency Video Coding (HEVC), by adding more coding and decoding tools to improve coding and decoding efficiency and handle various types of video sources, including three - dimensional (3D) video signals.

[0005] Figure 1AShows an exemplary adaptive intra / inter video codec system that includes loop processing. For intra prediction, the prediction data is derived based on previously encoded video data in the current picture. For inter prediction 112, the encoder performs motion estimation (ME) and performs motion compensation (MC) based on the result of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects either intra prediction 110 or inter prediction 112, and the selected prediction data is sent to adder 116 to form a prediction error, also known as a residual. The prediction error then undergoes processing by a transform (T) 118 and quantization (Q) 120. The transformed and quantized residuals are then encoded by entropy encoder 122 to be included in the video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed together with side information, such as motion and codec modes related to intra and inter prediction, and other information such as parameters related to the loop filter applied to the underlying image regions. As Figure 1A shown, side information related to intra prediction 110, inter prediction 112, and loop filter 130 is provided to entropy encoder 122. When using the inter prediction mode, the reference picture or pictures must also be reconstructed at the encoder side. Therefore, the transformed and quantized residuals undergo inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residuals. The residuals are then added back to the prediction data 136 to reconstruct the video data at reconstruction (REC) 128. The reconstructed video data may be stored in the reference picture buffer 134 and used for prediction of other frames.

[0006] As Figure 1A shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to the series of processes. Therefore, the loop filter 130 is typically applied to the reconstructed video data before it is stored in the reference picture buffer 134 to improve video quality. For example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF) may be used. Loop filter information may need to be incorporated into the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to entropy encoder 122 to be incorporated into the bitstream. InFigure 1A In this case, before the loop filter 130 is applied to the reconstructed video, the reconstructed samples are stored in the reference picture buffer 134. Figure 1A The system in [description] is designed to illustrate the exemplary structure of a typical video encoder. It may correspond to an efficient video coding (HEVC) system, VP8, VP9, H.264, or VVC.

[0007] As Figure 1B shown, the decoder can use the same or partially the same functional modules as the encoder, except for the transform 118 and quantization 120, because the decoder only needs the inverse quantization 124 and inverse transform 126. The decoder uses an entropy decoder 140 instead of the entropy encoder 122 to decode the video bitstream into quantized transform coefficients and the required coding and decoding information (e.g., In-Loop Prediction Filter (ILPF) information, intra prediction information, and inter prediction information). The intra prediction 150 at the decoder side does not need to perform a mode search. Instead, the decoder only needs to generate an intra prediction according to the intra prediction information received from the entropy decoder 140. In addition, for inter prediction, the decoder only needs to perform motion compensation (MC 152) according to the inter prediction information received from the entropy decoder 140 without performing motion estimation.

[0008] According to VVC, the input picture is divided into non-overlapping square block regions called Coding Tree Units (CTUs), similar to HEVC. Each CTU can be divided into one or more smaller-sized Coding Units (CUs). The resulting CU partition can be square or rectangular in shape. In addition, VVC divides the CTU into Prediction Units (PUs) as the units to which the prediction process, such as inter prediction, intra prediction, etc., is applied.

[0009] Adaptive Loop Filter in VVC

[0010] In VVC, an Adaptive Loop Filter (ALF) based on block-based filter adaptation is applied. For the luminance component, one filter is selected from 25 filters according to the direction and activity of the local gradient in each 4×4 block.

[0011] Filter Shape

[0012] Two diamond filter shapes are used (as Figure 2 shown). The 7×7 diamond shape 220 is used for the luminance component, and the 5×5 diamond shape 210 is used for the chrominance component.

[0013] Block Classification

[0014] For the luminance component, each 4×4 block is classified into one of 25 categories. The classification index C is derived based on its directivity D and the quantization value of activity as follows:

[0015]

[0016] To calculate D and first, the gradients in the horizontal, vertical, and two diagonal directions are calculated using a one-dimensional Laplacian operator:

[0017]

[0018] where the indices i and j refer to the coordinates of the top-left sample within the 4×4 block, and R(i, j) represents the reconstructed sample at the coordinates (i, j).

[0019] To reduce the complexity of block classification, downsampled one-dimensional Laplacian calculations are applied to the vertical direction ( Figure 3A ) and the horizontal direction ( Figure 3B ). As Figure 3C -D shows, the same downsampling positions ( Figure 3C g in d1 and Figure 3D g in d2 ) are used for the gradient calculations in all directions.

[0020] Then, the maximum and minimum values of the horizontal and vertical direction gradients are set to:

[0021]

[0022] The maximum and minimum values of the two diagonal direction gradients are set to:

[0023]

[0024] To derive the value of directivity D, these values are compared with each other and with two thresholds t 1 and t 2 as follows:

[0025] Step 1. If and are both true, D is set to 0.

[0026] Step 2. If proceed to Step 3; otherwise, proceed to Step 4.

[0027] Step 3. If D is set to 2; otherwise, D is set to 1.

[0028] Step 4. If D is set to 4; otherwise, D is set to 3.

[0029] The active value A is calculated as:

[0030]

[0031] A is further quantized to the range from 0 to 4 (including 0 and 4), and the quantized value is denoted as

[0032] For the chrominance component in the image, no classification is applied.

[0033] Geometric transformation of filter coefficients and shear values

[0034] Before filtering each 4×4 luminance block, a geometric transformation such as rotation or diagonal and vertical flipping is applied to the filter coefficients f(k, l) and the corresponding filter shear values c(k, l) according to the gradient value calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The purpose of doing this is to make the different blocks to which the ALF is applied more similar by aligning their directivities.

[0035] Three geometric transformations are introduced, including diagonal, vertical flipping, and rotation:

[0036] Diagonal: f D (k, l) = f(l, k), c D (k, l) = c(l, k),

[0037] Vertical flipping: f V (k, l) = f(k, K - l - 1), c V (k, l) = c(k, K - l - 1),

[0038] Rotation: f R (k, l) = f(K - l - 1, k), c R (k, l) = c(K - l - 1, k),

[0039] where K is the size of the filter, 0 ≤ k, l ≤ K - 1 are the coefficient coordinates such that the position (0, 0) is at the upper left corner and the position (K - 1, K - 1) is at the lower right corner. These transformations are applied to the filter coefficients f(k, l) and the shear values c(k, l) of the gradient value calculated for that block. The relationship between the transformations and the four gradients in the four directions is summarized in the following table.

[0040] Table 1. Mapping of gradients calculated for a block and the transformations

[0041] Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 and g v <g h > Diagonal <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation

[0042] Filtering process

[0043] At the decoder side, when ALF is enabled for a CTB, each sample R(i,j) within a CU is filtered to produce a sample value R′(i,j) as follows:

[0044]

[0045] where f(k,l) represents the decoded filter coefficients, K(x,y) is the clipping function, and c(k,l) represents the decoded clipping parameters. The variables k and l vary between -L / 2 and L / 2, where L represents the filter length. The clipping function K(x,y) = min(y, max(-y,x)) corresponds to the function Clip3(-y,y,x). The clipping operation introduces non-linearity, making ALF more effective by reducing the influence of neighboring sample values that are too different from the current sample value.

[0046] Cross Component Adaptive Loop Filter

[0047] CC-ALF uses the luminance sample values to refine each chrominance component by applying an adaptive linear filter to the luminance channel and then using the output of this filtering operation for chrominance refinement. Figure 4A A system-level diagram of the CC-ALF process related to the SAO, luminance ALF, and chrominance ALF processes is provided. As Figure 4A shown, each color component (i.e., Y, Cb, and Cr) is processed by its corresponding SAO (i.e., SAO luminance 410, SAO Cb 412, and SAO Cr 414). After SAO, ALF luminance 420 is applied to the SAO-processed luminance, and ALF chrominance 430 is applied to the SAO-processed Cb and Cr. However, there is a cross-component term from the luminance to the chrominance components (i.e., CC-ALF Cb 422 and CC-ALF Cr 424). The output of the cross-component ALF (using adders 432 and 434 respectively) is added to the output of ALF chrominance 430.

[0048] The filtering in CC-ALF is done by applying a linear diamond filter (e.g., Figure 4B the filters 440 and 442 in Figure 4B ) to the luminance channel. In

[0049]

[0050] where (x,y) is the position of the chrominance component i being refined, (x Y ,y Y ) is the luminance position based on (x,y), S i is the filter support region in the luminance component, and c i(x 0 , y 0 ) represents the filter coefficients.

[0051] As Figure 4B shown, the luma filter support region is the region that coincides with the current chroma sample, taking into account the spatial scaling factor between the luma and chroma planes.

[0052] In the VVC reference software, the CC-ALF filter coefficients are calculated by minimizing the mean squared error of each chroma channel with respect to the original chroma content. To this end, the VTM (VVC Test Model) algorithm uses a coefficient derivation process similar to that of chroma ALF. Specifically, a correlation matrix is derived and the coefficients are calculated using a Cholesky decomposition solver to attempt to minimize the mean squared error metric. When designing the filter, up to 8 CC-ALF filters can be designed and transmitted per picture. These resulting filters are then indicated for use on a CTU basis for the two chroma channels.

[0053] Other features of CC-ALF include:

[0054] ● Designed to use a 3x4 diamond with 8 taps.

[0055] ● Seven filter coefficients are transmitted in the APS.

[0056] ● Each transmitted coefficient has a 6-bit dynamic range and is limited to a power-of-two value.

[0057] ● The eighth filter coefficient is derived at the decoder such that the sum of the filter coefficients equals 0.

[0058] ● The APS can be referenced in the picture header.

[0059] ● CC-ALF filter selection is controlled for each chroma component at the CTU level.

[0060] ● The boundary filling of the horizontal virtual boundary uses the same memory access pattern as luma ALF.

[0061] As an additional feature, the reference encoder can be configured to enable some basic subjective adjustments via a configuration file. When enabled, VTM weakens the application of CC-ALF in regions encoded with a high QP, which are either close to mid-gray or contain a large amount of luma high frequency. Algorithmically, this is achieved by disabling the application of CC-ALF in CTUs that meet any of the following conditions:

[0062] ● The slice QP value minus 1 is less than or equal to the base QP value.

[0063] ● The number of chroma samples with local contrast greater than (1<<(bitDepth–2))–1 exceeds the CTU height, where the local contrast is the difference between the maximum and minimum luminance sample values within the filter support region.

[0064] ● More than a quarter of the chroma samples are between (1<<(bitDepth–1))–16 and (1<<(bitDepth–1))+16

[0065] The motivation for this feature is to provide some assurance that CC-ALF does not amplify artifacts introduced earlier in the decoding path (this is mainly because VTM does not currently explicitly optimize chroma subjective quality). It is expected that alternative encoder implementations may not use this feature or adopt alternative strategies suitable for their coding characteristics.

[0066] Filter parameter signaling

[0067] The ALF filter parameters are signaled in the Adaptation Parameter Set (APS). In one APS, up to 25 sets of luminance filter coefficients and shear value indices, and up to eight sets of chroma filter coefficients and shear value indices can be signaled. To reduce the bit overhead, the filter coefficients of different classifications of the luminance component can be merged. In the slice header, the index of the APS for the current slice is signaled.

[0068] The shear value index, decoded from the APS, allows the use of the shear value tables for the luminance and chroma components to determine the shear value. These shear values depend on the internal bit depth. More precisely, the shear value is obtained by the following formula:

[0069] AlfClip = {round(2 B-α*n ) for n ∈ [0..N-1]}

[0070] where B is the internal bit depth, ( is a predefined constant value equal to 2.35, and N is equal to 4, which is the number of shear values allowed in VVC. Then AlfClip is rounded to the format of the nearest power of 2.

[0071] In the slice header, up to 7 APS indices can be signaled to specify the luminance filter set for the current slice. The filtering process can be further controlled at the CTB level. A flag is always signaled to indicate whether ALF is applied to the luminance CTB. The luminance CTB can select a filter set between 16 fixed filter sets and the filter set in the APS. A filter set index is signaled for the luminance CTB to indicate which filter set is applied. The 16 fixed filter sets are predefined and hard-coded in the encoder and decoder.

[0072] For the chrominance component, an APS index is signaled in the slice header to indicate the chrominance filter set used for the current slice. At the CTB level, if there are multiple chrominance filter sets in the APS, a filter index is signaled for each chrominance CTB.

[0073] The filter coefficients are quantized with a standard of 128. To limit the multiplication complexity, a bitstream consistency is applied such that the coefficient values at non-central positions should be in the range of -2 7 to 2 7 - 1, including the boundaries. The coefficient at the central position is not signaled in the bitstream and is considered equal to 128.

[0074] The adaptive loop filter is in the ECM

[0075] ALF Simplification

[0076] The ALF gradient subsampling and the ALF virtual boundary processing are removed. The classified block size is reduced from 4x4 to 2x2. The filter sizes for luminance and chrominance, and the ALF coefficients are signaled and increased to 9x9.

[0077] ALF with Fixed Filters

[0078] To filter a luminance sample, three different classifiers (C 0 , C 1 and C 2 ) and three different filter sets (F 0 , F 1 and F 2 ) are used. The sets F 0 and F 1 contain fixed filters whose coefficients are trained by the classifiers C 0 and C 1 . The filter coefficients in F 2 are signaled. Which filter set's filter to use is determined by the class C i assigned to the sample, and the sample is filtered using the classifier C i[equationId74]

[0079] Filtering

[0080] First, two 13x13 diamond-shaped fixed filters F 0 and F 1 are applied to obtain two intermediate samples R 0 (x,y) and R 1 (x,y). After that, F 2 is applied to R 0 (x,y), R 1 (x,y) and neighboring samples to obtain the filtered sample

[0081]

[0082] where f i,j is the shear difference between the neighboring sample and the current sample R(x,y), and g i is the shear difference between R i-20 (x,y) and the current sample. The filter coefficients c i , i = 0, … 21, are signaled.

[0083] Classification

[0084] Assigns a class C to each 2x2 block based on the directionality D i and the activity : i

[0085]

[0086] where M D,i represents the total number of directionality D i .

[0087] As in the Versatile Video Coding (VVC), the horizontal, vertical, and two diagonal gradient values of each sample are calculated using a one-dimensional Laplacian operator. The sum of the sample gradients within a 4x4 window covering the target 2x2 block is used for the classifier C 0 , while the sum of the sample gradients within a 12x12 window is used for the classifiers C 1 and C 2 . The sums of the horizontal, vertical, and two diagonal gradients are denoted as and respectively. The directionality D i is determined by comparing

[0088]

[0089] with a set of thresholds. The directionality D 2 is derived using the thresholds 2 and 4.5 as in the VVC. For D 0 and D 1 , the horizontal / vertical edge strength and the diagonal edge strength are first calculated using the thresholds Th = [1.25, 1.5, 2, 3, 4.5, 8]. The edge strength is 0 if Otherwise, is the largest integer such that The edge strength is 0 if Otherwise, is the largest integer such that When ​That is, when horizontal / vertical edges are dominant, D i is derived by using Table 2A; otherwise, diagonal edges are dominant, and D i is derived by using Table 2B.

[0090] Table 2A. and the mapping from i to D

[0091]

[0092]

[0093] Table 2B. and the mapping from i to D

[0094]

[0095] To obtain the sum A of vertical and horizontal gradients is mapped to the range from 0 to n, where n is equal to 4 for i and 15 for and and

[0096] In an ALF_APS, up to 4 sets of luminance filter sets can be signaled, and each set may have up to 25 filters.

[0097] In the present invention, an ALF with position taps is disclosed.

Summary of the Invention

[0098] A method and apparatus for video coding and decoding using an Adaptive Loop Filter (ALF). According to the method, reconstructed pixels related to a current block are received. A target ALF is derived, where the ALF includes one or more position taps and a position function associated with at least one position tap outputs a variable. By applying the target ALF to the current block, a current filtered output is obtained. Filtered reconstructed pixels are provided, where the filtered reconstructed pixels include the current filtered output.

[0099] In one embodiment, the variable is related to the current sample value of the current sample, the one or more neighboring sample values of one or more neighboring samples of the current sample, or both.

[0100] In one embodiment, the position function outputs a variable or a constant according to conditions related to pixel positions in the horizontal and vertical directions. In one embodiment, the variable includes a predetermined function that takes the current sample value, the one or more neighboring sample values, or both as input data. In one embodiment, the predetermined function includes a clipping function for clipping a target input value. In one embodiment, the target input value corresponds to a scaled current sample value. In another embodiment, the target input value corresponds to the difference between the current sample value and one of the one or more neighboring sample values.

[0101] In one embodiment, the variable includes a first clipping function applied to a first difference between the current sample value and a first neighboring sample value of a first neighboring sample, and a second clipping function applied to a second difference between the current sample value and a second neighboring sample value of a second neighboring sample, and wherein the first neighboring sample and the second neighboring sample are located at symmetric positions relative to the current sample.

[0102] In one embodiment, the variable includes a first clipping function applied to the current sample value multiplied by a first scaled current sample value, and a second clipping function applied to a second scaled current sample value.

[0103] In one embodiment, the variable is related to one or more source values of one or more corresponding existing taps. In one embodiment, each source value corresponds to a clipped neighboring difference, a first correction value from another filter, or a second correction value from another loop filtering stage. In another embodiment, the variable corresponds to a linear function or a quadratic function. In one embodiment, the position function outputs a variable or a constant according to conditions related to pixel positions in the horizontal and vertical directions.

[0104] According to another method, reconstructed pixels related to the current block are received. A target horizontal period and a target vertical period are determined explicitly or implicitly, where the target horizontal period is determined from a set of horizontal periods, and the target vertical period is determined from a set of vertical periods. A target ALF including one or more position taps is determined, where the total number of the one or more position taps and the one or more corresponding position functions depends on the target horizontal period and the target vertical period. By applying the target ALF to the current block, a current filtered output is obtained. Filtered reconstructed pixels are provided, where the filtered reconstructed pixels include the current filtered output.

[0105] In one embodiment, the target horizontal period and the target vertical period are signaled or parsed explicitly or implicitly in the bitstream. In one embodiment, the target horizontal period and the target vertical period are signaled or parsed using respective indices. In another embodiment, the target horizontal period and the target vertical period are signaled or parsed jointly using an index to select the target horizontal period and the target vertical period from a set of predetermined period pairs.

[0106] In one embodiment, up to MxN coefficients and shear indices associated with one or more position taps are signaled or parsed in each filter at the adaptive parameter set (APS) level, where M and N are positive integers representing the target horizontal period and the target vertical period, respectively.

[0107] In one embodiment, up to MxN coefficients and shear indices associated with one or more position taps are signaled or parsed in each filter of a filter set.

[0108] In one embodiment, one or more coefficients and shear indices associated with one or more position taps are signaled or parsed at the filter level.

[0109] In one embodiment, the target horizontal period and the target vertical period are signaled or parsed at the adaptive parameter set (APS) level, and up to MxN coefficients and shear indices associated with one or more position taps are signaled for all filters in each filter set at the APS level.

[0110] In one embodiment, the target horizontal period and the target vertical period are signaled or parsed at the filter set level, and up to MxN coefficients and shear indices associated with one or more position taps are signaled or parsed for all filters at the filter set level.

[0111] In one embodiment, one or more coefficients and shear indices associated with one or more position taps are signaled or parsed at a first level different from a second level used for signaling or parsing non-position taps.

[0112] In one embodiment, one or more coefficients and shear indices associated with one or more position taps are signaled or parsed at the slice level, and information of non-position taps is signaled or parsed at the adaptive parameter set (APS) level.

[0113] In one embodiment, the target horizontal period and the target vertical period are implicitly derived based on a scaling factor. In one embodiment, the scaling factor depends on the picture resolution.

BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1ADescribes an exemplary adaptive Inter / Intra video codec system incorporating loop processing.

[0115] Figure 1B Describes Figure 1A the corresponding decoder of the encoder in

[0116] Figure 2 Describes the ALF filter shapes for the chrominance (left) and luminance (right) components.

[0117] Figure 3A -D describes the downsampled Laplacian calculation for g v (3A), g h (3B), g d1 (3C) and g d2 (3D).

[0118] Figure 4A Describes the placement of CC-ALF relative to other loop filters.

[0119] Figure 4B Describes the diamond filter for chrominance samples.

[0120] Figure 5 Describes a flowchart of an exemplary video codec system utilizing diverse location ALF, according to one embodiment of the present invention.

[0121] Figure 6 Describes a flowchart of an exemplary video encoding system for emitting horizontal and vertical periodic signals of diverse location ALF, according to one embodiment of the present invention.

DETAILED DESCRIPTION

[0122] The components of the present invention, as generally described and depicted in the figures herein, can be arranged and designed in a variety of different configurations. Accordingly, the following more detailed description of embodiments of the systems and methods of the present invention is not intended to limit the scope of the present invention, as claimed, but is merely representative of selected embodiments of the present invention. The phrase "in one embodiment" or "in an embodiment" as used throughout the specification means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" throughout the specification are not necessarily all referring to the same embodiment.

[0123] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be practiced without one or more of the specific details, or using other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the present invention. The embodiments of the present invention can be best understood by reference to the drawings, in which like parts are designated by like numerals. The following description is only by way of example and simply illustrates certain selected device and method embodiments consistent with the invention claimed herein.

[0124] In the following, a scheme for deriving, signaling, and utilizing a diverse position tap ALF is disclosed.

[0125] ALF Filter Shape Selection for Position Taps

[0126] Generally, the ALF reconstruction process can be expressed as:

[0127]

[0128] where R(x,y) is the sample value before ALF filtering, is the sample value after ALF filtering, c i is the i-th filter coefficient, n i is the i-th filter tap input. In particular, n i can be a clipped neighboring difference, a correction value from another filter, or a correction value from another loop filtering stage. In some cases, a position tap can be added to the reconstruction equation:

[0129]

[0130] where the additional P term is the position tap, f i (x,y) is a position embedding function that takes the current sample position (x,y) as input.

[0131] However, for different sequences, the position attributes may be different. In the present invention, a filter shape selection mechanism for position taps is illustrated to adaptively change the position taps used in the ALF.

[0132] In one embodiment, the horizontal period M and the vertical period N are explicitly signaled. The number of position taps P and the position embedding function f i (x,y) are determined according to M and N. For example, there is one position tap for the sample at a specific position in each M×N block. Specifically, P = M×N and f i (x,y) is defined as follows:

[0133] f0 (x, y) = (x mod M == 0 && y mod N == 0)? C : 0,

[0134] f 1 (x, y) = (x mod M == 1 && y mod N == 0)? C : 0,

[0135] …

[0136] f Mn+m (x, y) = (x mod N == m && y mod N == n)? C : 0,

[0137] …

[0138] f MN-1 (x, y) = (x mod M == (M - 1) && y mod N == (N - 1))? C : 0,

[0139] where "mod" represents the modulo operation, and C is a predetermined constant value or a value determined by the shear index. Note that for different numbers of M and N, the number of position taps P and the position embedding function f i (x, y) may also be different.

[0140] For the periodic signalization (i.e., M and N) in the above embodiments, they can be signalized separately or jointly. In the latter case, one index is signalized to select a period pair (M, N) from several predetermined period pairs. Additionally, the period information can be signalized at the APS level, filter set level, or filter level. In the above example (P = M × N), if the period is signalized at the APS level, the maximum M × N coefficients and shear index of the position taps of each filter in the APS are signalized; if the period is signalized at the filter set level, the maximum M × N coefficients and shear index of the position taps of each filter in the filter set are signalized; if the period is signalized at the filter level, the position taps of the M × N coefficients and shear index are signalized in the filter.

[0141] In the above embodiments, the coefficients and shear index of the position taps can be signalized at a higher level than those of other taps. For example, if the period is signalized at the APS level, the maximum M × N coefficients and shear index of the position taps of each filter set in the APS are signalized instead of each filter, and the coefficients and shear index of these position taps are shared by all filters in the filter set. If the period is signalized at the filter set level, the position taps of the M × N coefficients and shear index are signalized, and the coefficients and shear index of these position taps are shared by all filters in the filter set.

[0142] In the above embodiments, the coefficients of the position taps and the shear indices may be signaled at a different level from other taps. For example, the position taps are signaled at the slice level. In this design, the position taps signaled at the slice level are combined with other taps signaled in the APS to form a filter for ALF reconstruction.

[0143] In another embodiment, the horizontal period M and the vertical period N are implicitly derived based on a scaling factor. The number P of position taps and the position embedding function f i (x,y) are determined according to M and N. For example, if M = 4 and N = 2 are used for the original resolution (i.e., scaling factor = 1), M = 4 * 0.5 = 2 and N = 2 * 0.5 = 1 are used for the half resolution (i.e., scaling factor = 0.5). Note that this method is useful when codec tools that change the picture resolution, such as reference picture resampling (RPR), are enabled.

[0144] In the above embodiments, if a historical APS is from a frame with a different codec resolution from the current frame, the position taps of the filter of that APS will be disabled to prevent period mismatches.

[0145] ALF with diverse position taps

[0146] In one embodiment, each position tap is activated only for a subset of samples in the current codec region, and whether a sample belongs to the subset is determined by the position of the sample. If a position tap is not activated for a sample, the corresponding position embedding function outputs 0. If a position tap is activated for a sample, the output of the corresponding position embedding function can be a constant offset (e.g., Examples 1 and 2 below), a variable related to the current and / or neighboring sample values (e.g., Example 3 below), or a variable related to the source values of existing taps (n i ), e.g., Example 4 below).

[0147] Example 1. There are 4 position taps, and the position embedding function is defined as:

[0148] f 0 (x,y) = (x mod 2 == 0 && y mod 2 == 0)? C : 0,

[0149] f 1 (x,y) = (x mod 2 == 1 && y mod 2 == 0)? C : 0,

[0150] f 2 (x,y) = (x mod 2 == 0 && y mod 2 == 1)? C : 0,

[0151] f 3(x, y) = (x mod 2 == 1 && y mod 2 == 1)? C : 0,

[0152] where "mod" represents the modulo operation, and C can be a predefined constant value or a clipping index c based on the corresponding coefficient i+K selected value.

[0153] Example 2. There are 3 position taps, and the position embedding function is defined as:

[0154] f 0 (x, y) = (x mod 2 == 0 && y mod 2 == 0)? C : 0,

[0155] f 1 (x, y) = (x mod 2 == 1 && y mod 2 == 0) || (x mod 2 == 0 && y mod 2 ==

[0156] 1)? C : 0,

[0157] f 2 (x, y) = (x mod 2 == 1 && y mod 2 == 1)? C : 0.

[0158] Example 3. The position taps almost follow the same design as in Example 1, but C is modified.

[0159] g 0 (x, y) = (x mod 2 == 0 && y mod 2 == 0)? g(R, x, y) : 0,

[0160] f 1 (x, y) = (x mod 2 == 1 && y mod 2 == 0)? g(R, x, y) : 0,

[0161] f 2 (x, y) = (x mod 2 == 0 && y mod 2 == 1)? g(R, x, y) : 0,

[0162] f 3 (x, y) = (x mod 2 == 1 && y mod 2 == 1)? g(R, x, y) : 0,

[0163] where g(R, x, y) is a predefined function that takes the current processed sample value R(x, y) and / or its neighboring sample values R(x + p, y + q) as input, where p and q are integers. Some example function forms of g(R, x, y) are as follows:

[0164] g(R, x, y) = Clip(a * R(x, y)) + b,

[0165] g(R, x, y) = Clip((R(x + p, y + q) - R(x, y))),

[0166] g(R, x, y)

[0167] = Clip((R(x + p, y + q) - R(x, y)))

[0168] + Clip((R(x - p, y - q) - R(x, x))),

[0169] where a and b are predetermined integers, and Clip() represents the same clipping function used with the existing ALF taps. Example 4. The position taps almost follow the same design as in Example 3, but with a modification to g(R, x, y):

[0170] f 0 (x, y) = (x mod 2 == 0 && y mod 2 == 0)? h(n i ): 0,

[0171] f 1 (x, y) = (x mod 2 == 1 && y mod 2 == 0)? h(n i ): 0,

[0172] f 2 (x, y) = (x mod 2 == 0 && y mod 2 == 1)? h(n i ): 0,

[0173] f 3 (x, y) = (x mod 2 == 1 && y mod 2 == 1)? h(n i ): 0,

[0174] where h(n i ) is a predetermined function that takes as input the source of an existing tap n i . An example functional form of h(n i ) is:

[0175] h(n i ) = a * n i + b,

[0176] where a and b are predetermined integers. Note that the clipping operation for this tap has already been in the derivation of n i . In such a design, if n t is used for the position tap, the original existing tap can be removed, resulting in the following filtering equation:

[0177]

[0178] In the above embodiments, the position taps in each example can be combined.

[0179] Example 5. This example shows a combination of Example 1 and Example 4. There are a total of 8 position taps.

[0180] f 0 (x,y) = (x mod 2 == 0 && y mod 2 == 0)? C : 0,

[0181] f 1 (x,y) = (x mod 2 == 0 && y mod 2 == 0)? h(n i ) : 0,

[0182] f 2 (x,y) = (x mod 2 == 1 && y mod 2 == 0)? C : 0,

[0183] f 3 (x,y) = (x mod 2 == 1 && y mod 2 == 0)? h(n i ) : 0,

[0184] f 4 (x,y) = (x mod 2 == 0 && y mod 2 == 1)? C : 0,

[0185] f 5 (x,y) = (x mod 2 == 0 && y mod 2 == 1)? h(n i ) : 0,

[0186] f 6 (x,y) = (x mod 2 == 1 && y mod 2 == 1)? C : 0,

[0187] f 7 (x,y) = (x mod 2 == 1 && y mod 2 == 1)? h(n i ) : 0,

[0188] In this example, for each type of position, there are 2 coefficients to form a linear model to refine the samples. Note that h(n i ) can be replaced by g(R,x,y) in Example 3, which results in a combination of Example 1 and Example 3.

[0189] In the above embodiments, more non - linearity can be introduced into g(R,x,y) and h(n i ) For example, quadratic terms can be used:

[0190] g(R, x, y) = Clip(a * (R(x, y)) 2 ) + Clip(b * R(x, y)) + c

[0191]

[0192] Any of the ALFs described above can be implemented in the encoder and / or decoder. For example, any proposed method can be implemented in the loop filter module of the encoder or decoder (e.g., Figure 1A and Figure 1B the ILPF 130 therein). Alternatively, any proposed method can be implemented as a circuit that is connected to the internal encoding / decoding module and / or the motion compensation module of the encoder, and the merge candidate derivation module of the decoder. The ALF method can also be implemented using executable software or firmware code stored on a medium, such as a hard disk or flash memory, for a CPU (Central Processing Unit) or a programmable device (e.g., a DSP (Digital Signal Processor) or an FPGA (Field Programmable Gate Array)).

[0193] Figure 5 Shows a flowchart of an exemplary video encoding / decoding system that utilizes diverse location ALF according to an embodiment of the present invention. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart can also be implemented based on hardware, such as one or more electronic devices or processors arranged to execute the steps in the flowchart. According to the method, reconstructed pixels related to the current block are received in step 510. A target ALF is derived in step 520, where the ALF includes one or more position taps and a position function output variable related to at least one position tap. By applying the target ALF to the current block, the current filtered output is obtained in step 530. Filtered reconstructed pixels are provided in step 540, where the filtered reconstructed pixels include the current filtered output.

[0194] Figure 6A flowchart showing an exemplary video coding and decoding system signals horizontal and vertical periods for diverse location ALF according to an embodiment of the present invention. According to this method, reconstructed pixels related to a current block are received in step 610. A target horizontal period and a target vertical period are explicitly or implicitly determined in step 620, where the target horizontal period is determined among a set of horizontal periods and the target vertical period is determined among a set of vertical periods. A target ALF including one or more location taps is determined in step 630, where the total number of the one or more location taps and one or more corresponding location functions depends on the target horizontal period and the target vertical period. By applying the target ALF to the current block, a current filtering output is obtained in step 640. Filtered reconstructed pixels are provided in step 650, where the filtered reconstructed pixels include the current filtering output.

[0195] The flowchart is intended to illustrate an example of video coding and decoding according to the present invention. Those skilled in the art can modify each step, rearrange steps, split steps, or combine steps to practice the present invention without departing from the spirit of the present invention. In this disclosure, specific syntax and semantics have been used to illustrate examples of practicing the present invention. Those skilled in the art can practice the present invention by replacing the said syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0196] The above description is intended to enable a person of ordinary skill in the art to practice the present invention in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details have been set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art will understand that the present invention can be practiced.

[0197] Embodiments of the present invention as described above can be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuits integrated into a video compression chip, or program codes integrated into video compression software to perform the processing described herein. An embodiment of the present invention can also be program codes to be executed on a digital signal processor (DSP) to perform the processing described herein. The present invention may also relate to multiple functions executed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors can be configured according to the invention to perform specific tasks by executing machine-readable software codes or firmware codes that define the specific methods embodied by the invention. The software codes or firmware codes can be developed in different programming languages and different formats or styles. The software codes can also be compiled for different target platforms. However, different code formats, styles, and languages of the software codes, as well as other means of configuring the codes to perform tasks according to the invention, do not deviate from the spirit and scope of the invention.

[0198] The present invention can be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples should be considered illustrative in all respects and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes that come within the meaning and scope of equivalence of the claims are to be embraced within their scope.

Claims

1. A method for adaptive loop filter (ALF) processing for video reconstruction, the method comprises: Receiving reconstructed pixels related to a current block; Deriving a target ALF, wherein the target ALF includes one or more position taps and a position function outputting a variable associated with at least one position tap; Obtaining a current filtered output by applying the target ALF to the current block; and Providing filtered reconstructed pixels, wherein the filtered reconstructed pixels include the current filtered output.

2. The method according to claim 1, wherein the variable is related to the current sample value of the current sample, one or more neighboring sample values of one or more neighboring samples of the current sample, or both.

3. The method according to claim 2, wherein the variable includes a predetermined function that takes the current sample value, the one or more neighboring sample values, or both as input data.

4. The method according to claim 3, wherein the predetermined function includes a clipping function for clipping a target input value.

5. The method according to claim 4, wherein the target input value corresponds to a scaled current sample value.

6. The method according to claim 4, wherein the target input value corresponds to the difference between the current sample value and one of the one or more neighboring sample values.

7. The method according to claim 2, wherein the variable includes a first clipping function applied to a first difference between the current sample value and a first neighboring sample value of a first neighboring sample, and a second clipping function applied to a second difference between the current sample value and a second neighboring sample value of a second neighboring sample, and wherein the first neighboring sample and the second neighboring sample are located at symmetric positions relative to the current sample.

8. The method according to claim 2, wherein the variable includes a first clipping function applied to a first scaled current sample value multiplied by the current sample value, and a second clipping function applied to a second scaled current sample value.

9. The method according to claim 1, wherein the output of the position function depends on conditions related to pixel positions in horizontal and vertical directions.

10. The method according to claim 1, wherein the position function further outputs a constant under conditions related to pixel positions in horizontal and vertical directions.

11. The method according to claim 1, wherein the variable is related to one or more source values of one or more corresponding existing taps.

12. The method according to claim 11, wherein each source value corresponds to a clipped neighboring difference, a first correction value from another filter, or a second correction value from another loop filtering stage.

13. The method according to claim 11, wherein the variable corresponds to a linear function or a quadratic function.

14. The method according to claim 11, wherein the output of the position function depends on conditions related to pixel positions in horizontal and vertical directions.

15. An apparatus for adaptive loop filtering (ALF) processing for video reconstruction, the apparatus includes one or more electronic devices or processors configured to: Receive reconstructed pixels related to a current block; Derive a target ALF, where the target ALF includes one or more position taps and a variable output by a position function associated with at least one position tap; Obtain a current filtering output by applying the target ALF to the current block; and Provide filtered reconstructed pixels, where the filtered reconstructed pixels include the current filtering output.

16. A method for adaptive loop filtering (ALF) processing for video reconstruction, the method comprises: Receive reconstructed pixels related to a current block; Explicitly or implicitly determine a target horizontal period and a target vertical period, where the target horizontal period is determined from a set of horizontal periods, and the target vertical period is determined from a set of vertical periods; Determine a target ALF including one or more position taps, where the total number of the one or more position taps and one or more corresponding position functions depends on the target horizontal period and the target vertical period; Obtain a current filtering output by applying the target ALF to the current block; and Provide filtered reconstructed pixels, where the filtered reconstructed pixels include the current filtering output.

17. The method according to claim 16, wherein the target horizontal period and the target vertical period are signaled or parsed explicitly or implicitly in a bitstream.

18. The method according to claim 17, wherein the target horizontal period and the target vertical period are signaled or parsed using respective indices.

19. The method according to claim 17, wherein the target horizontal period and the target vertical period are signaled or parsed jointly using an index to select the target horizontal period and the target vertical period from a set of predetermined period pairs.

20. The method according to claim 16, wherein at most MxN coefficients and shear indices associated with the one or more position taps are signaled or parsed in each filter at the adaptive parameter set (APS) level, where M and N are positive integers representing the target horizontal period and the target vertical period respectively.

21. The method according to claim 16, wherein at most MxN coefficients and shear indices associated with the one or more position taps are signaled or parsed in each filter of a filter set, and where M and N are positive integers representing the target horizontal period and the target vertical period respectively.

22. The method according to claim 16, wherein one or more coefficients and shear indices related to the one or more position taps are signaled or parsed at the filter level.

23. The method according to claim 16, wherein for all filters in a filter set, at most MxN coefficients and shear indices related to the one or more position taps are signaled or parsed in each of the filter sets at the APS level, where M and N are positive integers representing the target horizontal period and the target vertical period respectively.

24. The method according to claim 16, wherein for all filters of the filter set, at most MxN coefficients and shear indices associated with the one or more position taps are signaled or resolved at the filter set level.

25. The method according to claim 16, wherein the target horizontal period and the target vertical period are signaled or resolved at the adaptive parameter set (APS) level.

26. The method according to claim 16, wherein the target horizontal period and the target vertical period are signaled or resolved at the filter set level.

27. The method according to claim 16, wherein one or more coefficients and shear indices associated with the one or more position taps are signaled or resolved at a first level different from a second level used for signaling or resolving non-position taps.

28. The method according to claim 16, wherein one or more coefficients and shear indices associated with the one or more position taps are signaled or resolved at the slice level, and information on non-position taps is signaled or resolved at the adaptive parameter set (APS) level.

29. The method according to claim 16, wherein the target horizontal period and the target vertical period are implicitly derived based on a scaling factor.

30. The method according to claim 29, wherein the scaling factor depends on the picture resolution.

31. An apparatus for adaptive loop filter (ALF) processing for reconstructing video, the apparatus comprising one or more electronic devices or processors configured to: receive reconstructed pixels associated with a current block; explicitly or implicitly determine a target horizontal period and a target vertical period, wherein the target horizontal period is determined among a set of horizontal periods and the target vertical period is determined among a set of vertical periods; determine a target ALF including one or more position taps, wherein the total number of the one or more position taps and one or more corresponding position functions depend on the target horizontal period and the target vertical period; derive a current filtered output by applying the target ALF to the current block; provide filtered reconstructed pixels, wherein the filtered reconstructed pixels include the current filtered output.