Method and apparatus for adaptive loop filter with non-sample tapping for video coding

By introducing adaptive loop filters with non-sample value filter taps into the video encoding and decoding system, the problem that existing ALFs are difficult to utilize non-sample value information is solved, and more efficient video data processing and quality improvement is achieved.

CN119948879APending Publication Date: 2025-05-06MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066768.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-08-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing adaptive loop filter (ALF) is difficult to effectively utilize non-sample value information in video encoding and decoding systems, limiting its performance and adaptability.

Method used

The adaptive loop filter using a non-sample value filter tap is improved by introducing non-sample value terms derived from target information independent of the sample value to improve its performance and adaptability.

Benefits of technology

By using ALF with non-sample value filter taps, video data can be processed more efficiently, video quality can be improved, and the adaptability and performance of the encoding and codec system can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948879A_ABST
    Figure CN119948879A_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding using an adaptive loop filter (ALF) with non-sample value filter taps. According to the method, a reconstructed pixel is received, where the reconstructed pixel includes a current block. A current filtered output is derived from an ALF of a current sample in a current block, where the ALF comprises at least one non-sample value term derived using target information independent of a sample value of the current block, and the target information is derived on the encoder side, or derived or received on the decoder side. A filtered reconstructed pixel is provided, wherein the filtered reconstructed pixel includes the current filtered output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Cross-reference to related applications]

[0002] This application is a non-provisional application and claims priority from U.S. Provisional Patent Application No. 63 / 375,882, filed on September 16, 2022, and U.S. Provisional Patent Application No. 63 / 377,731, filed on September 30, 2022. The entire contents of the above-mentioned U.S. Provisional Patent Applications are hereby incorporated by reference. [Technical field]

[0003] The present invention relates to a video coding and decoding system using an adaptive loop filter (ALF), and in particular to an ALF using non-sample taps. [Background technology]

[0004] Versatile Video Coding (VVC) is the latest international video codec standard developed by the Joint Video Experts Team (JVET) of the Video Coding Experts Group (VCEG) of the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) and the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Coded representation of immersive media - Part 3: Versatile video codec, published in February 2021. VVC is developed on the basis of its predecessor High Efficiency Video Coding (HEVC), by adding more codec tools to improve codec efficiency and handle various types of video sources including three-dimensional (3D) video signals.

[0005] Figure 1AAn exemplary adaptive intra / inter video codec system incorporating in-loop processing is shown. For intra prediction, prediction data is derived based on previously encoded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed on the encoder side, and motion compensation (MC) is then performed based on the results of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects intra prediction 110 or inter prediction 112, and the selected prediction data is provided to adder 116 to form a prediction error, also referred to as a residual. The prediction error is then processed by transform (T) 118 and quantization (Q) 120. The transformed and quantized residual is then encoded by an entropy encoder 122 for inclusion in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged with side information such as motion and codec modes associated with intra and inter predictions, as well as other information such as parameters associated with the loop filter applied to the underlying image area. As Figure 1A As shown, side information related to intra prediction 110, inter prediction 112 and loop filter 130 is provided to entropy encoder 122. When inter prediction mode is used, the reference picture or pictures must also be reconstructed at the encoder end. Therefore, the transformed and quantized residual is processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to recover the residual. The residual is then added back to the prediction data 136 and the video data is reconstructed at reconstruction (REC) 128. The reconstructed video data may be stored in a reference picture buffer 134 and used for prediction of other frames.

[0006] like Figure 1A As shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subjected to various impairments due to the series of processes. Therefore, a loop filter 130 is typically applied to the reconstructed video data before it is stored in a reference picture buffer 134 to improve the video quality. For example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) may be used. The loop filter information may need to be included in the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 for inclusion in the bitstream. Figure 1A In , a loop filter 130 is applied to the reconstructed video and the reconstructed samples are then stored in a reference picture buffer 134 . Figure 1AThe system in is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Codec (HEVC) system, VP8, VP9, ​​H.264, or VVC.

[0007] like Figure 1B The decoder shown may use the same or partially the same functional blocks as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses an entropy decoder 140 instead of the entropy encoder 122 to decode the video bitstream into quantized transform coefficients and required codec information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). The intra-frame prediction 150 on the decoder side does not require a pattern search. Instead, the decoder only needs to generate an intra-frame prediction based on the intra-frame prediction information received from the entropy decoder 140. In addition, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from the entropy decoder 140, without the need for motion estimation.

[0008] According to VVC, the input picture is divided into non-overlapping square block areas, called CodingTree Units (CTUs), similar to HEVC. Each CTU can be divided into one or more smaller-sized Coding Units (CUs). The resulting CU division can be square or rectangular. In addition, VVC divides CTU into Prediction Units (PUs) as units for applying prediction processes (e.g., inter-frame prediction, intra-frame prediction, etc.).

[0009] Adaptive loop filter in VVC

[0010] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luminance component, a filter is selected from 25 filters in each 4×4 block based on the direction and activity of the local gradient.

[0011] Filter shape

[0012] Using two diamond filter shapes (such as Figure 2 ). The 7×7 diamond shape 220 is applied to the luma component and the 5×5 diamond shape 210 is applied to the chroma components.

[0013] Block Classification

[0014] For the luma component, each 4×4 block is classified into one of 25 classes. The class index C is based on the quantized values ​​of its directionality D and activity The results are as follows:

[0015]

[0016] To calculate D and First, the gradients in the horizontal, vertical and two diagonal directions are calculated using a one-dimensional Laplacian operator.

[0017]

[0018] where the indices i and refer to the coordinates of the top-left sample within the 4×4 block, and R(i, j) represents the reconstructed sample at coordinate (i, j).

[0019] In order to reduce the complexity of block classification, the vertical direction ( Figure 3A ) and horizontal direction ( Figure 3B ) applies a downsampled one-dimensional Laplace calculation. Figure 3C -D, the same downsampling position is used for gradient calculation in all directions ( Figure 3C g in d1 and Figure 3D g in d2 ).

[0020] Then set the maximum and minimum values ​​of the horizontal and vertical gradients to:

[0021]

[0022] The maximum and minimum values ​​of the two diagonal gradients are set as:

[0023]

[0024] To derive the value of the directionality D, these values ​​are compared with each other and with two thresholds t1 and t2:

[0025] Step 1. If and are both true, then D is set to 0.

[0026] Step 2. If Then continue from step 3; otherwise continue from step 4.

[0027] Step 3. If Then D is set to 2; otherwise D is set to 1.

[0028] Step 4. If Then D is set to 4; otherwise D is set to 3.

[0029] The activity value A is calculated as follows:

[0030]

[0031] A is further quantized to the range of 0 to 4 (inclusive), and the quantized value is expressed as

[0032] For the chrominance components in the image, no classification is applied.

[0033] Geometric transformation of filter coefficients and clipping values

[0034] Before filtering each 4×4 luma block, a geometric transformation such as rotation or diagonal and vertical flipping is applied to the filter coefficients f(k, l) and the corresponding filter shear values ​​c(k, l) according to the gradient values ​​calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The purpose of this is to make different blocks to which the ALF is applied more similar by aligning their directionality.

[0035] Three geometric transformations are introduced, including diagonal, vertical flip and rotation:

[0036] Diagonal: f D (k, l) = f(l, k), c D (k, l) = c(l, k),

[0037] Flip vertically: f V (k, l) = f(k, Kl-1), c V (k, l) = c(k, Kl-1),

[0038] Rotation: f R (k, l) = f(Kl-1, k), c R (k, l) = c(Kl-1, k),

[0039] Where K is the size of the filter and 0≤k,l≤K-1 are the coefficient coordinates such that position (0,0) is at the top left corner and position (K-1,K-1) is at the bottom right corner. These transformations are applied to the filter coefficients f(k,1) and the shear value c(k,l) depending on the gradient values ​​calculated for the block. The relationship between the transformations and the four gradients in the four directions is summarized in the following table.

[0040] Table 1. Mapping of gradients computed for a block to transformations

[0041] Gradient Value Transform <![CDATA[g d2 <g d1 and g h <g v ]]> No change <![CDATA[g d2 <g d1 and g v <g h ]]> Diagonal <![CDATA[g d1 <g d2 and g h <g v ]]> Flip Vertically <![CDATA[g d1 <g d2 and g v <g h ]]> Rotation

[0042] Filtration process

[0043] At the decoder side, when ALF is enabled for a CTB, each sample R(i, j) within the CU is filtered, resulting in a sample value R′(i, j) as shown below,

[0044] R′(i, j)=R(i, j)+((∑ k≠0∑ l≠0 f(k, l)×K(R(i+k, j+l)-R(i, j), c(k, l))+64)>>7.)where f(k, l) represents the decoded filter coefficient, K(x, y) is the clipping function, and c(k, l) represents the decoded clipping parameter. The variables k and 1 vary between -L / 2 and L / 2, where L represents the filter length. The clipping function K(x, y)=min(y, max(-y, x)) corresponds to the function Clip3(-y, y, x). The clipping operation introduces nonlinearity, making the ALF more effective by reducing the influence of neighboring sample values ​​that differ too much from the current sample value.

[0045] Cross-component adaptive loop filter

[0046] CC-ALF uses luma sample values ​​to refine the chroma components by applying an adaptive, linear filter to the luma channel and then using the output of this filtering operation. Figure 4A A system-level diagram of the CC-ALF process relative to SAO, luma ALF, and chroma ALF processing is provided. Figure 4A As shown, each color component (i.e., Y, Cb, and Cr) is processed by its corresponding SAO (i.e., SAO Luma 410, SAO Cb 412, and SAO Cr 414). After SAO, ALF Luma 420 is applied to the SAO processed Luma, and ALF Chroma 430 is applied to the SAO processed Cb and Cr. However, there is a cross-component term from Luma to Chroma components (i.e., CC-ALF Cb 422 and CC-ALF Cr 424). The output of the cross-component ALFs (using adders 432 and 434, respectively) is added to the output of ALF Chroma 430.

[0047] Filtering in CC-ALF is achieved by applying a linear, diamond filter (e.g. Figure 4B This is accomplished by using filters 440 and 442 in FIG. Figure 4B In the figure, a hollow circle represents a brightness sample, and a dot-filled circle represents a chrominance sample. Each chrominance channel uses a filter, and the operation is expressed as:

[0048]

[0049] where (x, y) is the position of the chrominance component i being refined, and (x Y ,y Y ) is the brightness position based on (x, y), S i is the filter support area in the luminance component, c i (x0, y0) represents the filter coefficients.

[0050] like Figure 4BAs shown, the luma filter support is the area that coincides with the current chroma sample, taking into account the spatial scaling factor between the luma and chroma planes.

[0051] In the VVC reference software, the CC-ALF filter coefficients are calculated by minimizing the mean squared error of each chroma channel relative to the original chroma content. To achieve this, the VTM (VVC Test Model) algorithm uses a coefficient derivation process similar to that used for the chroma ALF. Specifically, a correlation matrix is ​​derived and the coefficients are calculated using a Cholesky decomposition solver in an attempt to minimize the mean squared error metric. When designing the filters, up to 8 CC-ALF filters can be designed and transmitted per image. The resulting filters are then indicated for both chroma channels according to the CTU.

[0052] Additional features of CC-ALF include:

[0053] The design uses a 3x4 diamond with 8 taps.

[0054] Seven filter coefficients are transmitted in APS.

[0055] Each transmitted coefficient has a 6-bit dynamic range and is limited to power-of-two values.

[0056] The eighth filter coefficient is derived in the decoder such that the sum of the filter coefficients is equal to zero.

[0057] The APS may be referenced in the slice header.

[0058] CC-ALF filter selection is controlled at CTU level for each chroma component.

[0059] Border filling for horizontal virtual borders uses the same memory access pattern as luma ALF.

[0060] As an additional feature, the reference encoder can be configured to enable some basic subjective adjustments via a configuration file. When enabled, VTM reduces the application of CC-ALF in areas that are encoded at high QP and are close to mid-grey or contain a lot of luminance high frequencies. Algorithmically, this is achieved by disabling the application of CC-ALF in CTUs that meet any of the following conditions:

[0061] The slice QP value minus 1 is less than or equal to the base QP value.

[0062] The number of chroma samples with local contrast greater than (1<<(bitDepth-2))-1 exceeds the CTU height, where local contrast is the difference between the maximum and minimum luma sample values ​​within the filter support area.

[0063] More than a quarter of the chroma samples are in the range (1<<(bitDepth-1))-16 and (1<<(bitDepth-1))+16

[0064] The motivation for this feature is to provide some assurance that CC-ALF does not amplify artifacts introduced earlier in the decoding path (primarily because VTM does not currently explicitly optimize for chroma subjective quality). It is expected that alternative encoder implementations may not use this feature or adopt alternative strategies appropriate to their encoding characteristics.

[0065] Filter parameter signal

[0066] ALF filter parameters are signaled in an Adaptive Parameter Set (APS). In one APS, up to 25 groups of luma filter coefficients and shear value indices can be signaled, as well as up to eight groups of chroma filter coefficients and shear value indices. To reduce bit overhead, filter coefficients for luma components of different classifications can be merged. In the slice header, the index of the APS used by the current slice is signaled.

[0067] The clipping value index, decoded from the Adaptive Parameter Set (APS), allows to determine the clipping value using the clipping value tables for the luma and chroma components. These clipping values ​​depend on the internal bit depth. More precisely, the clipping value is obtained by the following formula:

[0068] AlfClip = {round(2 B-α*n )for n∈[0..N-1]}

[0069] Where B is equal to the internal bit depth, (is a predefined constant value equal to 2.35, and N is equal to 4, which is the number of clipping values ​​allowed in VVC. AlfClip is then rounded to the nearest power of 2 format value.

[0070] In the slice header, up to 7 APS indices can be signaled to specify the luma filter set used for the current slice. The filtering process can be further controlled at the coding tree block (CTB) level. There is always a flag signaled to indicate whether the adaptive loop filter (ALF) is applied to the luma CTB. The luma CTB can select a filter set from 16 fixed filter sets and the filter set from the APS. A filter set index is signaled for the luma CTB to indicate which filter set is applied. These 16 fixed filter sets are predefined and hard-coded in the encoder and decoder.

[0071] For chroma components, one APS index is signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there are multiple chroma filter sets in the APS, one filter index is signaled for each chroma CTB.

[0072] The filter coefficients are quantized to a standard value of 128. To limit the multiplication complexity, a bitstream consistency is applied so that coefficient values ​​in non-central positions are within -2 7 To 2 7 The coefficients in the center are not signaled in the bitstream and are considered equal to 128.

[0073] Adaptive loop filter in ECM

[0074] ALF Simplification

[0075] ALF gradient subsampling and ALF virtual boundary processing are removed. The block size for classification is reduced from 4x4 to 2x2. The filter size for luma and chroma, where ALF coefficients are signaled, is increased to 9x9.

[0076] ALF with fixed filter

[0077] To filter a brightness sample, three different classifiers (C0, C1 and C2) and three different sets of filters (F0, F1 and F2) are used. Sets F0 and F1 contain fixed filters whose coefficients are trained for classifiers C0 and C1. The coefficients in the filters are signaled in F2. Which one from set F is used? i The filter is determined for a given sample by using the classifier C i[equationId74] The class C assigned to this sample i Decide.

[0078] filter

[0079] First, two 13x13 diamond shape fixed filters F0 and F1 are applied to obtain two intermediate samples R0(x, y) and R1(x, y). Afterwards, F2 is applied to R0(x, y), R1(x, y) and neighboring samples to obtain filtered samples such as

[0080]

[0081] where f i,j is the shear difference between the neighboring sample and the current sample R(x, y), g i YesR i-20 The shear difference between (x, y) and the current sample. Filter coefficient c i , i=0,...21, are transmitted by the signal.

[0082] Classification

[0083] Based on the directionality D i and activity Assign a class C to each 2x2 block i :

[0084]

[0085] Among them, M D,i Represents the total number of directions D i As in VVC, the values ​​of horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within the 4×4 window covered by the target 2×2 block for classifier C0, and the sum of the sample gradients within the 12×12 window for classifiers C1 and C2. The sum of the horizontal, vertical, and two diagonal gradients are expressed as and Directionality D i Determined by comparison.

[0086]

[0087] A set of thresholds are used. Directivity D2 is derived using thresholds 2 and 4.5 as per VVC. For D0 and D1, first calculate the horizontal / vertical edge strength and diagonal edge strength Use threshold Th = [1.25, 1.5, 2, 3, 4.5, 8]. If Edge Strength is 0; otherwise, is the largest integer such that if Edge Strength is 0; otherwise, is the largest integer such that when When horizontal / vertical edges dominate, use Table 2A to obtain D i Otherwise, when diagonal edges dominate, use Table 2B to obtain D i .

[0088] Table 2A. and To D i Mapping

[0089]

[0090]

[0091] Table 2B. and To D i Mapping

[0092]

[0093] To obtain The sum of the vertical and horizontal gradients Ai Mapped to the range 0 to n, where n is is equal to 4, for and Equals 15.

[0094] In ALF_APS, up to 4 luminance filter sets can be signaled, with up to 25 filters in each set.

[0095] In the present invention, a novel adaptive loop filter (ALF) whose input corresponds to a non-sample value is disclosed to improve the performance of the ALF. [Summary of the invention]

[0096] Brief Summary of the Invention

[0097] The present invention discloses a method and apparatus for video encoding and decoding using an adaptive loop filter (ALF). According to the method, a reconstructed pixel is received, wherein the reconstructed pixel includes a current block. A current filtering output is obtained from the ALF of a current sample in the current block, wherein the ALF includes at least one non-sample value item, and the non-sample value item is obtained using target information that is independent of the sample value of the current block, and the target information is obtained on the encoder side or obtained or received on the decoder side. A filtered reconstructed pixel is provided, wherein the filtered reconstructed pixel includes the current filtering output.

[0098] In one embodiment, at least one non-sample value term based on the target information is derived as a sum of one or more non-sample value filter taps that are independent of sample values ​​of the current block. In one embodiment, each of the one or more non-sample value filter taps corresponds to a target function of the target information.

[0099] In one embodiment, the target function corresponds to a position function that takes as input the position information of one or more current samples of the current block. In one embodiment, the position function corresponds to a periodic function of one or more positions associated with the one or more current samples. In one embodiment, the periodic function of one or more positions corresponds to a sine function, a square wave function, a triangle wave function, or a sawtooth wave function.

[0100] In one embodiment, the target function corresponds to a binary function of the target information, wherein the binary function outputs a first value when the target information satisfies a condition and outputs a second value when the target information does not satisfy the condition. In one embodiment, the first value corresponds to a predefined offset or a first value selected according to a shear value index associated with the current block. In one embodiment, the second value corresponds to 0.

[0101] In one embodiment, the target information corresponds to position information of one or more current samples of the current block as input. In one embodiment, the condition is determined based on the position of the target sample relative to a repeating pattern in horizontal direction, vertical direction or both, or one or more derived positions.

[0102] In one embodiment, the target information corresponds to coding unit (Coding Unit, referred to as CU) encoding information. In one embodiment, the CU encoding information includes CU mode, prediction mode, CU boundary, CU residual, motion vector (Motion Vector, referred to as MV) information or a combination thereof. In one embodiment, the condition is determined by comparing the horizontal component, the vertical component or both of the motion vector of the current block with one or more thresholds. In another embodiment, the condition is determined based on the proximity of the target sample to the boundary of the current block. In another embodiment, the condition is determined based on the prediction direction of the target sample of the current block.

[0103] In one embodiment, the target information corresponds to picture information. In one embodiment, the picture information includes a picture order count (POC), a time ID, a layer ID or a combination thereof.

[0104] In one embodiment, the target information corresponds to ALF classification information. In one embodiment, the ALF classification information includes a transposition index, an activity, a directionality, a quantization activity, a quantization directionality, or a combination thereof derived from an ALF classifier. In another embodiment, the condition is determined based on a transposition index of the current block. In another embodiment, the condition is determined based on a quantization activity of the current block.

[0105] In one embodiment, the target information corresponds to a joint correlation computed from luma and chroma samples.

Brief Description of the Drawings

[0106] Figure 1A An exemplary adaptive intra / inter video coding and decoding system with integrated in-loop processing is described.

[0107] Figure 1B Describes Figure 1A The corresponding decoder of the encoder in .

[0108] Figure 2 Depicted are the ALF filter shapes for the chroma (left) and luma (right) components.

[0109] Figure 3A -D describes the v (3A), g h (3B), g d1 (3C) and g d2(3D) Downsampled Laplacian computation.

[0110] Figure 4A Describes the position of the CC-ALF relative to other loop filters.

[0111] Figure 4B Describes a diamond filter for chroma samples.

[0112] Figure 5A -D shows various periodic functions used to derive the ALF tap signal: (A) sine wave, (B) square wave, (C) triangle wave, and (D) sawtooth wave.

[0113] Figure 6 A flow chart of an exemplary video codec system utilizing an ALF with non-sample valued filter taps according to one embodiment of the present invention is described. [Specific implementation method]

[0114] The components of the present invention, as generally described and depicted in the figures, can be arranged and designed into a variety of different configurations. Therefore, the following more detailed description of the embodiments of the systems and methods of the present invention, as shown in the figures, is not intended to limit the scope of the present invention, as claimed, but is merely representative of selected embodiments of the present invention. References to "one embodiment", "certain embodiment" or similar language in this specification mean that at least one embodiment of the present invention may be included in the description of a particular feature, structure or characteristic associated with that embodiment. Therefore, the phrases "in one embodiment" or "in a certain embodiment" appearing throughout this specification do not necessarily all refer to the same embodiment.

[0115] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be practiced without one or more specific details, or using other methods, components, etc. In other cases, in order to avoid obscuring aspects of the present invention, well-known structures or operations are not shown or described in detail. Embodiments of the present invention may be best understood by reference to the drawings, in which like parts are represented by like numbers throughout. The following description is by way of example only, and simply illustrates embodiments of certain selected devices and methods consistent with the invention claimed herein.

[0116] In the following, a new type of input for an adaptive loop filter (ALF) is disclosed. In a conventional ALF, the filtering operation is applied to a signal related to a sample value (e.g., the current sample value, a neighboring sample value, or the difference between two sample values, etc.). The new type of input is derived using information that is independent of the sample value.

[0117] Adaptive loop filter with non-sample valued taps

[0118] In general, the ALF reconstruction process can be expressed as

[0119]

[0120] Where R(x, y) is the sample value before ALF filtering, is the sample value after ALF filtering, c i is the i-th filter coefficient, n i is the i-th filter tap input. Specifically, n i It can be a clipped neighbor difference, a correction value from another filter, or a correction value from another loop filtering stage.In this disclosure, several additional tap generation methods related to information other than sample values ​​are described to increase the capability and adaptability of the ALF filter.

[0121] According to one embodiment of the present invention, the reconstruction equation of ALF is modified as follows:

[0122]

[0123] where f j is a function related to information other than the sample value at (x, y), and M is the total number of additional taps.

[0124] In one embodiment, f j is a function that is related to the transposed index (the index used to determine how to perform geometric transformations) at (x, y). For example:

[0125] f0(x,y)=(transposeIndex==0)? C:0,

[0126] f1(x,y)=(transposeIndex==1)? C:0,

[0127] f2(x,y)=(transposeIndex==2)? C:0,

[0128] f3(x,y)=(transposeIndex==3)? C:0,

[0129] Where C is a predefined offset value or a value selected by cutting the index. In the above expression, "x? y: z" will output y or z depending on x. If x is true (for example, x is equal to 1), then output y; otherwise output z. In other words, "x" can be regarded as a test condition, and the output depends on the condition.

[0130] In another embodiment, f jis a function related to the activity value computed for the gradient classifier at (x, y). For example,

[0131]

[0132] in is the quantized activity value at (x, y), and C is a predefined offset value or a value selected by the clipping index.

[0133] Let me give you another example.

[0134] f0(x, y) = A,

[0135] Where A is the active value at (x, y). In this example, the active value is used directly as the source and clipping can be applied to it like any other tap.

[0136] In the above embodiment, the activity value A (or the quantized value ) can be replaced by the directionality D (or quantized value) computed for the gradient classifier at (x, y) ).

[0137] In another embodiment, f j is a function related to the mean of the sample values ​​in the block computed for the band classifier at (x, y). For example,

[0138] f0(x,y)=(M%4==0)? C:0

[0139] f1(x,y)=(M%4==1)? C:0

[0140] f2(x,y)=(M%4==2)? C:0

[0141] f3(x,y)=(M%4==3)? C:0,

[0142] Where M%4 is the 2 least significant bits (LSBs) of the average of the sample values ​​in the current block, and C is a predefined offset value or a value selected by the clipping index.

[0143] ALF with position tap

[0144] In general, the ALF reconstruction process can be expressed as

[0145]

[0146] Where R(x, y) is the sample value before ALF filtering, is the sample value after ALF filtering, c i is the i-th filter coefficient, and n i is the i-th filter tap input. Specifically, n iIt can be a clipped neighboring difference, a correction value from another filter, or a correction value from another loop filtering stage. In this proposal, several additional tap generation methods related to sample positions are shown to increase the capability and adaptability of the ALF filter.

[0147] Specifically, the sample position (x, y) is modeled with a periodic function and added as an extra tap to the ALF reconstruction equation:

[0148]

[0149] where f j is a periodic function and M is the total number of position taps.

[0150] In one embodiment, two position taps are introduced in the ALF (M=2), where the periodic function f j is the following sine function:

[0151]

[0152] where P0 and P1 are the periods of the function, which are chosen according to the shear index. For example, if the corresponding coefficient c K+0 The shear index is 0 / 1 / 2 / 3, and the period P0 is 32 / 16 / 8 / 4 respectively.

[0153] In another embodiment, four position taps are included in the ALF (M=4), where the periodic function f j is the following sine function.

[0154]

[0155] Where P0, P1, P2 and P3 are the periods of the functions, which are selected according to the shear index. Since there is a phase difference between the sine function and the cosine function, more types of refinement can be performed by combining the two functions.

[0156] In the above embodiment, each periodic function may include an amplitude term A. For example, instead of sin(2πx / P0), A*sin(2πx / P0) is used, where A may be a predefined value or vary with the shear index.

[0157] In the above embodiment, the sine function can be Figure 5A-5D Other non-sinusoidal periodic functions are substituted as shown, where Figure 5A shows a sine function, Figure 5B A square wave is shown. Figure 5C A triangle wave is shown. Figure 5D A sawtooth wave is shown.

[0158] In another embodiment, a periodic function such as f is used. j (x, y) = sin(2π(a j x+b j y) / P0) makes positions x and y co-embedded, where (a j , b j ) is a predetermined pair of integers.

[0159] In another embodiment, a binary function is used as a source to generate the additional taps. Some examples are shown below.

[0160] Example 1:

[0161] f0(x,y)=(x mod 2==0)? 1:0,

[0162] f1(x,y)=(x mod 2==1)? 1:0,

[0163] f2(x,y)=(y mod 2==0)? 1:0,

[0164] f3(x,y)=(y mod 2==1)? 1:0.

[0165] In Example 1, x represents the horizontal position of the sample and y represents the vertical position of the sample. The condition "(x mod 2 == 0)" corresponds to x being divisible by 2. On the other hand, the condition "(x mod 2 == 1)" corresponds to x being divisible by 2 with a remainder of 1. Instead of "mod 2", other values ​​(such as 4, or 8) can be used for the test condition. The test condition for "mod n" is equivalent to checking the position of the x position relative to the repeating pattern (for example, for n = 2, 2, 4, 6, 8, etc.). A similar situation applies to y. Although the position of the target sample (i.e., x, y or (x, y)) is checked as a test condition in Example 1, other position information based on the position (referred to as a derived position in this article) can also be used for the test condition. Some examples of using one or more derived positions are shown in Examples 2 and 3.

[0166] Example 2:

[0167] f0(x,y)=((x+y)mod 2==0)? 1:0,

[0168] f1(x,y)=((x+y)mod 2==1)? 1:0,

[0169] f2(x,y)=(|xy|mod 2==0)? 1:0,

[0170] f3(x,y)=(|xy|mod 2==1)? 1:0.

[0171] Example 3:

[0172] f0(x,y)=((ax+y)mod 2==0)? 1:0,

[0173] f1(x,y)=((ax+y)mod 2==1)? 1:0,

[0174] f2(x,y)=(|x-ay|mod 2==0)? 1:0,

[0175] f3(x,y)=(|x-ay|mod 2==1)? 1:0,

[0176] Where a is the slope, which can be a fixed predetermined value or vary with the shear index selection.

[0177] Example 4:

[0178] f0(x,y)=(x mod 2==0&&y mod 2==0)? 1:0,

[0179] f1(x,y)=(x mod 2==1&&y mod 2==0)? 1:0,

[0180] f2(x,y)=(x mod 2==0&&y mod 2==1)? 1:0,

[0181] f3(x,y)=(x mod 2==1&&y mod 2==1)? 1:0.

[0182] In the above-mentioned Example 4, the positions in both the horizontal and vertical directions are checked.

[0183] In the above method, the number of "2" in "mod 2" is an example, indicating that the pattern is repeated in a 2x2 block pattern. The number of "2" can be replaced by other numbers if the repeating pattern is larger, such as an MxN pattern, where M and N are non-zero integers. The shape of the repeating pattern can be square or non-square.

[0184] Example 5:

[0185] f0(x,y)=((x,y)is near a CU boundary)? 1:0.

[0186] In Example 5, the proximity of the target sample to the CU boundary is checked. The proximity can be measured by the distance to the CU boundary.

[0187] Example 6:

[0188] f0(x,y)=(Residual at(x,y)is smaller than a threshold T)? 1:0,

[0189] Where the threshold T is a fixed predetermined value or varies with the shear index selection. Note that it is also possible to use multiple taps to cover more different T values.

[0190] Example 7:

[0191] f0(x,y)=(Mvx at(x,y)is smaller than a threshold T)? 1:0,

[0192] f1(x,y)=(Mvy at(x,y)is smaller than a threshold T)? 1:0,

[0193] f2(x,y)=(Mvx at(x,y)is larger than a threshold T)? 1:0,

[0194] f3(x,y)=(MVy at(x,y)is larger than a threshold T)? 1:0,

[0195] Where the threshold T is a fixed predetermined value or varies with the shear index selection, and Mvx and Mvy correspond to the horizontal component and vertical component of the motion vector of the current block. Note that it is also possible to use multiple taps to cover more different T values.

[0196] Example 8:

[0197] f0(x,y)=(inter dir at(x,y)is Bi)? 1:0,

[0198] f1(x,y)=(inter dir at(x,y)is L0)? 1:0,

[0199] f2(x,y)=(inter dir at(x,y)is L1)? 1:0.

[0200] In Example 8, the inter prediction direction (ie, Bi, L0, or 11) is checked as a test condition.

[0201] According to an embodiment of the present invention, various encoding information other than sample intensity can also be used as filter input. The encoding information can be derived or received at the decoder. For example, sample position, value derived from sample position, codec unit mode, prediction mode, codec unit boundary, residual, motion vector information, chroma sampling position / phase, adaptive loop filter / loop filter / post-filter information, sample clipping value, picture order count (POC), time ID, layer ID or joint correlation calculated from luminance and chroma samples.

[0202] In the above embodiment, the binary function

[0203] f j (x,y)=(some conditions)? 1:0

[0204] Can be replaced by

[0205] f j (x,y)=(some conditions)? C:0,

[0206] Where C is the magnitude based on the clip index selection. For example, C = clipValue >> 1, the magnitude of the function varies with the clip index selection.

[0207] Any of the adaptive loop filter methods with non-sample value taps described above can be implemented in an encoder and / or decoder. For example, any of the proposed methods can be implemented in a loop filter module of an encoder or decoder (e.g., Figure 1A and Figure 1B ILPF 130 in ). Alternatively, any of the proposed methods can be implemented as a circuit connected to the inter-coding module of the encoder and / or the motion compensation module, merged candidate derivation module of the decoder. The adaptive loop filter method can also be implemented using executable software or firmware code stored on a medium, such as a hard disk or flash memory, for a CPU (central processing unit) or a programmable device (e.g., a DSP (digital signal processor) or an FPGA (field programmable gate array)).

[0208] Figure 6A flow chart of a video encoding and decoding system using an adaptive loop filter with non-sample value filter taps according to one embodiment of the present invention is shown. The steps shown in the flow chart can be implemented as executable program code on one or more processors (e.g., one or more CPUs) at the encoder end. The steps shown in the flow chart can also be implemented based on hardware, such as one or more electronic devices or processors arranged to perform the steps in the flow chart. According to this method, a reconstructed pixel associated with a current block is received in step 610. In step 620, a current filtered output of an adaptive loop filter for a current sample in the current block is determined, wherein the adaptive loop filter includes at least one non-sample value item derived using target information that is independent of the sample value of the current block, and the target information is derived on the encoder side, or derived or received on the decoder side. In step 630, a filtered reconstructed pixel is provided, wherein the filtered reconstructed pixel includes the current filtered output.

[0209] The flowchart shown is intended to illustrate an example of implementing video encoding and decoding according to the present invention. A skilled person may modify each step, rearrange the steps, split the steps, or combine the steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics are used to illustrate examples of implementing embodiments of the present invention. A skilled person may practice the present invention by replacing syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0210] The above description is intended to enable persons with ordinary skill to practice the present invention in the context of specific applications and their requirements. For the described embodiments, persons skilled in the art will see various modifications, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but rather to be given the widest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details are shown in order to provide a deeper understanding of the present invention. However, persons skilled in the art will appreciate that the present invention may be practiced.

[0211] Embodiments of the present invention are described above and can be implemented in various hardware, software codes, or a combination of the two. For example, one embodiment of the present invention may be one or more circuits integrated into a video compression chip, or program code integrated into video compression software to perform the processing described herein. One embodiment of the present invention may also be a program code executed on a digital signal processor (DSP) to perform the processing described herein. The present invention may also involve multiple functions performed by a computer processor, a digital signal processor, a microprocessor, or a field programmable gate array (FPGA). These processors may be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines the specific method embodied in the present invention. The software code or firmware code can be developed in different programming languages ​​and different formats or styles. The software code can also be compiled for different target platforms. However, the different code formats, styles, and languages ​​of the software code and other configuration codes do not deviate from the spirit and scope of the present invention in the manner in which the tasks are performed according to the present invention.

[0212] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples should be considered in all respects as illustrative only and not restrictive. Therefore, the scope of the present invention should be indicated by the appended claims rather than the above description. All changes that fall within the meaning and equivalent scope of the claims should be included within their scope.

Claims

1. A method for adaptive loop filter processing for reconstructing video, the method comprising: receiving reconstructed pixels associated with a current block; deriving a current filter output from an adaptive loop filter of a current sample in the current block, wherein the adaptive loop filter includes at least one non-sample value term, the non-sample value term being derived using target information that is independent of sample values ​​of the current block, and the target information being derived at an encoder side or derived or received at a decoder side; and A filtered reconstructed pixel is provided, wherein the filtered reconstructed pixel comprises the current filter output.

2. The method according to claim 1, characterized in that The at least one non-sample value term based on the target information is derived as a sum of one or more non-sample value filter taps that are independent of sample values ​​of the current block.

3. The method according to claim 2, characterized in that Each of the one or more non-sample valued filter taps corresponds to a target function of the target information.

4. The method according to claim 3, characterized in that The objective function corresponds to a position function, which takes position information of one or more current samples of the current block as input.

5. The method according to claim 4, characterized in that The position function corresponds to a periodic function of one or more positions associated with one or more current samples.

6. The method according to claim 3, characterized in that The periodic function of the position corresponds to a sine wave function, a square wave function, a triangle wave function or a sawtooth wave function.

7. The method according to claim 3, characterized in that The target function corresponds to a binary function of the target information, wherein the binary function outputs a first value when the target information satisfies a condition, and outputs a second value when the target information does not satisfy the condition.

8. The method according to claim 7, characterized in that The first value corresponds to a predefined offset or the first value is selected according to a clipping value index associated with the current block, and the second value corresponds to zero.

9. The method according to claim 7, characterized in that The second value corresponds to zero.

10. The method according to claim 7, characterized in that The target information corresponds to position information of one or more current samples of the current block as input.

11. The method according to claim 10, characterized in that The condition is determined based on the position or one or more derived positions of the target sample relative to a repeating pattern in the horizontal direction, the vertical direction, or both.

12. The method according to claim 7, characterized in that The target information corresponds to coding unit (CU) encoding information.

13. The method according to claim 12, characterized in that The CU encoding information includes CU mode, prediction mode, CU boundary, CU residual, motion vector (MV) information, or a combination thereof.

14. The method according to claim 12, characterized in that The condition is determined by comparing the horizontal component, the vertical component, or both of the motion vector of the current block with one or more thresholds.

15. The method according to claim 12, characterized in that The condition is determined based on how close the target sample is to the boundary of the current block.

16. The method according to claim 12, characterized in that The condition is determined based on the prediction direction of the target sample of the current block.

17. The method according to claim 7, characterized in that The object information corresponds to the picture information.

18. The method according to claim 17, characterized in that The picture information includes a picture order count (POC), a time ID, a layer ID, or a combination thereof.

19. The method according to claim 7, characterized in that The target information corresponds to adaptive loop filter classification information.

20. The method of claim 19, wherein: The adaptive loop filter classification information includes a transposition index, an activity, a directionality, a quantization activity, a quantization directionality, or a combination thereof derived from an ALF classifier.

21. The method of claim 19, wherein: The condition is determined based on the transposed index of the current block.

22. The method of claim 19, wherein: The condition is determined based on the quantization activity of the current block.

23. The method of claim 7, wherein: This target information corresponds to the joint correlation computed from the luma and chroma samples.

24. An apparatus for adaptive loop filter processing for reconstructing video, the apparatus comprising one or more electronic circuits or processors configured to: receiving reconstructed pixels associated with a current block; deriving a current filter output from an adaptive loop filter of a current sample in the current block, wherein the adaptive loop filter includes at least one non-sample value term, the non-sample value term being derived using target information that is independent of sample values ​​of the current block, and the target information being derived at an encoder side, or derived or received at a decoder side; and A filtered reconstructed pixel is provided, wherein the filtered reconstructed pixel comprises the current filter output.