Method and device of alternative clipping in adaptive loop filter
Patent Information
- Application Number
- PCT/CN2025/081080
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video coding standards like VVC face challenges in efficiently handling video impairments due to processing, particularly in the application of Adaptive Loop Filters (ALF) for improving video quality, as they often require complex block classification and filtering operations that can be computationally intensive and inefficient.
The introduction of alternative clipping functions, such as ReLU, LeakyReLU, inverse ReLU, and hyperbolic tangent functions, along with geometric transformations and adaptive filtering techniques, to enhance the ALF process, allowing for more efficient and flexible video decoding and encoding.
These alternative clipping functions and adaptive filtering techniques improve video quality by reducing computational complexity and enhancing the efficiency of ALF operations, leading to better video reconstruction and coding efficiency.
Smart Images

Figure CN2025081080_02102025_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE OF ALTERNATIVE CLIPPING IN ADAPTIVE LOOP FILTERCROSS-REFERENCE TO RELATED INVENTION
[0001] This patent application claims the benefit of United States Provisional Application No. 63 / 562,328 filed March 7, 2024 and the disclosure is incorporated herein by reference in its entirety.FIELD OF INVENTION
[0002] The present description relates generally to video coding. In particular, the present disclosure relates to a method and a device for Adaptive Loop Filter (ALF) .BACKGROUND
[0003] Unless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.
[0004] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0005] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A CTU consists of one luma CTB, two chroma CTBs.SUMMARY
[0006] In some embodiments, a decoder for decoding of a video bitstream associated with video data comprises: a storage module for receiving the video bitstream associated with a picture of the video data wherein the picture comprises a current video region; a processing module for performing the following steps: determining from a syntax of the video stream a filtering operation is applied for the current video region; determining, for the current video region, a clipping index element set; and a filter for performing the filtering operation, according to the clipping index element set, to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions, wherein the clipping index element set is used for indicating the selected clipping function, a segment coordinate point element and a slope element for formulating the selected clipping function.
[0007] In some embodiments, the clipping index element set is determined from an Adaptation Parameter Set (APS) of the syntax of the video bitstream.
[0008] In some embodiments, the clipping functions correspond to a piecewise clipping function group comprising a ReLU function, a LeakyReLU function, an inverse ReLU function, inverse a LeakyReLU function, a ReLU function with clipping, a leakyReLU function with clipping, a shifted ReLU function with / without clipping, a shifted leakyReLU function with / without clipping, a shifted ReLU function with non-constant clipping, a shifted LeakyReLU function with non-constant clipping, a Sawtooth wave function with clipping and a Periodic Sawtooth Wave function.
[0009] In some embodiments, the clipping functions correspond to a non-piecewise clipping function group comprising a sigmoid function and a hyperbolic tangent (tanh) function.
[0010] In some embodiments, the clipping index element set is further used for indicating a shift vector element for shifting the segment coordinate point element toward a direction and a distance according to the shift vector.
[0011] In some embodiments, the current video region comprises a current sample, and wherein the filtering operation is performed using the clipping function for clipping a difference value between a value of a neighboring sample of the current sample and a value of the current sample.
[0012] In some embodiments, the decoder further comprises: performing a classification process for the current region to categorize the current region into a class indicated by a class index; and determining the clipping index element set according to the class index.
[0013] In some embodiments, a method for processing video data comprises: receiving a picture of the video data wherein the picture comprises a current video region; determining a filtering operation is applied for the current video region; determining, for the current video region, a clipping index element set; and performing the filtering operation to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions, wherein the clipping index element set is used for indicating the selected clipping function, a segment coordinate point element and a slope element for formulating the selected clipping function.
[0014] In some embodiments, the method further comprises: singing the clipping index element set into a syntax of a video bitstream.
[0015] In some embodiments, the clipping index element set is signaled into an Adaptation Parameter Set (APS) , Filter-set or filter level of the syntax of the video bitstream.
[0016] In some embodiments, the clipping functions correspond to a piecewise clipping function group comprising a ReLU function, a LeakyReLU function, an inverse ReLU function, inverse a LeakyReLU function, a ReLU function with clipping, a leakyReLU function with clipping, a shifted ReLU function with / without clipping, a shifted leakyReLU function with / without clipping, a shifted ReLU function with non-constant clipping, a shifted LeakyReLU function with non-constant clipping, a Sawtooth wave function with clipping and a Periodic Sawtooth Wave function.
[0017] In some embodiments, the clipping functions correspond to a non-piecewise clipping function group comprising a sigmoid function and a hyperbolic tangent (tanh) function.
[0018] In some embodiments, the clipping index element set is further used for indicating a shift vector element for shifting the segment coordinate point element toward a direction and distance according to the shift vector.
[0019] In some embodiments, the current video region comprises a current sample, and the filtering operation is performed using the clipping function for clipping a difference value between a value of a neighboring sample of the current sample and a value of the current sample.
[0020] In some embodiments, the method further comprises: performing a classification process for the current region to categorize the current region into a class indicated by a class index; and determining the clipping index element set according to the class index.
[0021] In some embodiments, the clipping index element set is determined according to a filter coefficient precision.
[0022] In some embodiments, the clipping index element set is determined according to a source of filtering sample.
[0023] In some embodiments, the clipping index element set is determined according to an ALF filter coefficient bit-depth.
[0024] In some embodiments, a method for processing video data comprises: receiving a picture of the video data wherein the picture comprises a current video region; determining a filtering operation is applied for the current video region; determining, for the current video region, a clipping index element set; and performing the filtering operation to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions, wherein the clipping index element set is used to indicating the selected clipping function and is implicitly determined by an existing flag.
[0025] In some embodiments, the existing flag associates with a class index, a filter coefficient precision, a source of filtering sample, or an ALF filter coefficient bit-depth.BRIEF DESCRIPTION OF DRAWINGS
[0026] Aspects of the present disclosure are best understood from the following detailed description when read with the accompanying figures. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several exemplary embodiments and, together with the corresponding descriptions, provide examples for explaining the disclosed embodiment consistent with the present disclosure and related principles. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.
[0027] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.
[0028] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0029] Figs. 2A to 2D show example subsampled Laplacian calculations for adaptive loop filter (ALF) classification.
[0030] Fig. 3 illustrates the ALF filter shapes for the chroma (left) and luma (right) components.
[0031] Fig. 4A illustrates the placement of CC-ALF with respect to other loop filters.
[0032] Fig. 4B illustrates a diamond shaped filter for the chroma samples.
[0033] Fig. 5 illustrates the 25-tap large filter used in CCALF process.
[0034] Fig. 6 illustrates the diamond shaped ALF in ECM-5.0.
[0035] Fig. 7 illustrates a longer ALF as an alternative to the diamond shaped ALF in Fig. 6.
[0036] Fig. 8 illustrates the filter shape of ALF in ECM-7.0.
[0037] Fig. 9 illustrates an example of ALF with additional fixed filter.
[0038] Fig. 10 illustrates an example of filter shape for CCALF.
[0039] Figs. 11A-11R illustrate the functions for ALF clipping according to some embodiments of the present disclosure.
[0040] Fig. 12 illustrates a flowchart of an exemplary method of ALF processing for reconstructed video according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0041] The following disclosure provides many different embodiments, or examples, for implementing different features of the provided subject matter. Specific examples of components and arrangements are described below to simplify the present disclosure. These are, of course, merely examples and are not intended to be limiting. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure.
[0042] In this specification, "signaling" and "signaled" may refer to either embedding, inserting, receiving or retrieving information within a bitstream about controlling a filter, such as enabling or disabling modes or other control parameters. It is also understood that a "class" denotes a category of elements, and an "active" flag signals that the corresponding filter / tool is in use.
[0043] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system 1000A incorporating loop processing. As illustrated, the adaptive Inter / Intra video encoding system 1000A receives input video signal 1002 and encodes the signal into video bitstream 1004. The adaptive Inter / Intra video encoding system 1000A has several components or modules for encoding the video signal 1002, at least including some components selected from Intra Prediction Module 110, Inter-Prediction Module 112, Switch Modules 114, Adder Modules 116, Transform (T) Module 118, Quantization (Q) Module 120, Entropy Encoder Modules 122, Inverse Quantization (IQ) Module 124, Inverse Transformation (IT) Module 126, Reconstruction (REC) Module 128, In-Loop Filter (ILPF) Module 130, Reference Picture Buffer 134.
[0044] In some embodiments, the modules 112 –130 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 112-130 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 112-130 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0045] For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction Module 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch Modules 114 selects Intra Prediction 110 or Inter-Prediction Module 112 and the selected prediction data is supplied to Adder Modules 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) Module 118 followed by Quantization (Q) Module 120. The transformed and quantized residues are then coded by Entropy Encoder Modules 122 to be included in a video bitstream 1004 corresponding to the compressed video data. The video bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction Module 112 and in-loop filter Modules 130, are provided to Entropy Encoder Modules 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) Module 124 and Inverse Transformation (IT) Module 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) Module 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other pictures.
[0046] As shown in Fig. 1A, incoming video signal 1002 undergoes a series of processing in the encoding system 1000A. The reconstructed video data from REC Module 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter Modules 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the video bitstream 1004 so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder Modules 122 for incorporation into the bitstream. In Fig. 1A, Loop filter Modules 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0047] Fig. 1B illustrates an exemplary adaptive Inter / Intra video decoding system 1000B incorporating loop processing. As illustrated, the adaptive Inter / Intra video decoding system 1000B receives video bitstream 1004 and decodes it into video signal 1006. The adaptive Inter / Intra video decoding system 1000B has several components or modules for decoding the video bitstream 1004, at least including some components selected from Switch Modules 114, Inverse Quantization (IQ) Module 124, Inverse Transformation (IT) Module 126, Reconstruction (REC) Module 128, In-Loop Filter (ILPF) Module 130, Reference Picture Buffer 134, Entropy Decoder Module 140, Intra prediction Module 150, Motion Compensation (MC) Module 152.
[0048] In some embodiments, the modules 114 –152 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 114-152 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 114-152 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0049] The decoder 1000B, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform Module 118 and Quantization Module 120 since the decoder only needs Inverse Quantization Module 124 and Inverse Transform Module 126. Instead of Entropy Encoder Modules 122, the decoder uses an Entropy Decoder Module 140 to decode the video bitstream 1004 into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction Module 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder Module 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC Module 152) according to Inter prediction information received from the Entropy Decoder Module 140 without the need for motion estimation. The reconstructed video data from REC Module 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter Modules 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used.
[0050] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc. Further, A CTU consists of one luma CTB, two chroma CTBs, and a CU also consists of one luma CU, two chroma CUs. To minimize artifacts, DBF, SAO, or ALF processing is individually applied to Luma and Chroma CUs, potentially with different parameters (e.g. index, mode, function, and / or coefficients) for each. In some cases, one CU's filter parameters may be adopted by another. In VVC ■ Adaptive Loop Filter
[0051] If the ALF tool is active, it is active at least for Luma component, and an index in the bitstream further indicates if the ALF is active for each Chroma component. The Adaptive Loop Filter (ALF) is applied with block-based filter adaption. For the luma component, one among 25 filters is selected for each sample block with N×N samples (e.g. N=4) in the luma CU, based on the direction and activity of local gradients. ■ Block classification
[0052] For luma component, each 4×4 block is locally categorized into one out of 25 classes. A class index C is derived based on its directionality D and a quantized value of activity as follows:
[0053] To calculate D and gradients of the horizontal, vertical and two diagonal directions are first calculated using 1-D Laplacian:
[0054] Where indices i and j refer to the coordinates of the upper left sample within the 4×4 block and R (i, j) indicates a reconstructed sample at coordinate (i, j) .
[0055] Please refer to Figs. 2A to 2D, Figs. 2A to 2D show example subsampled Laplacian calculations for adaptive loop filter (ALF) classification. To reduce the complexity of block classification, the subsampled 1-D Laplacian calculation is applied to the vertical direction (Fig. 2A) and the horizontal direction (Fig. 2B) . As shown in Figs. 2C-2D, the same subsampled positions are used for gradient calculation of all directions (gd1 in Fig. 2C and gd2 in Fig. 2D) .
[0056] Then D maximum and minimum values of the gradients of horizontal and vertical directions are set as:
[0057] The maximum and minimum values of the gradient of two diagonal directions are set as:
[0058] To derive the value of the directionality D, these values are compared against each other and with two thresholds t1 and t2: Step 1: If both and are true, D is set to 0. Step 2: If continue from Step 3; otherwise continue from Step 4. Step 3: If D is set to 2; otherwise D is set to 1. Step 4: If D is set to 4; otherwise D is set to 3.
[0059] The activity value A is calculated as:
[0060] A is further quantized to the range of 0 to 4, inclusively, and the quantized value is denoted as
[0061] For chroma components in a picture, no classification method is applied. In such application, a single set of ALF coefficients is applied for each chroma CUs. ■ Filter shape
[0062] Please refer to Fig. 3, Fig. 3 illustrates the ALF filter shapes for the chroma (left) and luma (right) components. The 7×7 diamond shape (with filter length of 12, e.g. C0-C11) is applied for luma component and the 5×5 diamond shape (with filter length of 6, e.g. C0-C5) is applied for chroma components. After the local classification, a filter is selected based on class index C to filter the sample block ■ Geometric transformations of filter coefficients and clipping values
[0063] In another embodiment, before filtering each N×N luma block, geometric transformations such as rotation or diagonal and vertical flipping may be applied to the filter coefficients f (k, l) and to the corresponding filter clipping values c (k, l) depending on gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning their directionality.
[0064] Three geometric transformations, including diagonal, vertical flip and rotation are introduced: Diagonal: fD (k, l) =f (l, k) , cD (k, l) =c (l, k) , Vertical flip: fV (k, l) =f (k, K-l-1) , cV (k, l) =c (k, K-l-1) Rotation: fR (k, l) =f (K-l-1, k) , cR (k, l) =c (K-l-1, k) where K is the size of the filter and 0≤k≤K-1 & 0≤l≤K-1 are coefficients coordinates, such that location (0, 0) is at the upper left corner and location (K-1, K-1) is at the lower right corner. The transformations are applied to the filter coefficients f (k, l) and to the clipping values c (k, l) depending on gradient values calculated for that block. The relationship between the transformation and the four gradients of the four directions are summarized in the following table. Table 1 -Mapping of the gradient calculated for one block and the transformations. ■ Filtering process
[0065] At decoder side, in VVC, when ALF is enabled for a CTB, each sample R (i, j) within the CU is filtered, resulting in sample value R′ (i, j) as shown below, where f (k, l) denotes the decoded filter coefficients, K (x, y) is the clipping function and c (k, l) denotes the decoded clipping parameters. The variable k and l vary between and where L denotes the filter length. The clipping function K (x, y) =min (y, max (-y, x) ) which corresponds to the function Clip3 (-y, y, x) . The clipping operation introduces non-linearity to make ALF more efficient by reducing the impact of neighbor sample values that are too different with the current sample value. ■ Cross component adaptive loop filter
[0066] Fig. 4A illustrates the placement of CC-ALF with respect to other loop filters. In another embodiment, CC-ALF refines chroma components using luma sample values by applying an adaptive linear filter to the luma channel and utilizing the output for chroma refinement. Fig. 4A shows a system-level diagram of the CC-ALF process alongside SAO, luma ALF, and chroma ALF processes. As shown in Fig. 4A, each colour component (i.e., Y, Cb and Cr) is processed by its respective SAO (i.e., SAO Luma 410, SAO Cb 412 and SAO Cr 414) . After SAO, ALF Luma 420 is applied to the SAO-processed luma and ALF Chroma 430 is applied to SAO-processed Cb and Cr. However, there is a cross-component term from luma to a chroma component (i.e., CC-ALF Cb 422 and CC-ALF Cr 424) . The outputs from the cross-component ALF are added (using adders 432 and 434 respectively) to the outputs from ALF Chroma 430.
[0067] Please refer to Fig. 4B. Fig. 4B illustrates a diamond shaped filter for the chroma samples. Filtering in CC-ALF is accomplished by applying a linear, diamond shaped filter (e.g. filters 440 and 442 in Fig. 4B) to the luma channel. In Fig. 4B, a blank circle indicates a luma sample and a dot-filled circle indicate a chroma sample. One filter is used for each chroma channel, and the operation is expressed as: where (x, y) is chroma component i location being refined (xY, yY) is the luma location based on (x, y) , Si is filter support area in luma component, ci (x0, y0) represents the filter coefficients.
[0068] As shown in Fig, 4B, the luma filter support is the region collocated with the current chroma sample after accounting for the spatial scaling factor between the luma and chroma planes.
[0069] The VVC reference software calculates CC-ALF filter coefficients by minimizing the mean square error for each chroma channel to match the original content, using a VTM algorithm process akin to chroma ALF. Specifically, a correlation matrix is derived, and the coefficients are computed using a Cholesky decomposition solver in an attempt to minimize a mean square error metric. In designing the filters, a maximum of 8 CC-ALF filters can be designed and transmitted per picture. The resulting filters are then indicated for each of the two chroma channels on a CTU basis.
[0070] Additional characteristics of CC-ALF include: ● The design uses a 3x4 diamond shape with 8 taps. ● Seven filter coefficients are transmitted in the APS. ● Each of the transmitted coefficients has a 6-bit dynamic range and is restricted to power-of-2 values. ● The eighth filter coefficient is derived at the decoder such that the sum of the filter coefficients is equal to 0. ● An APS may be referenced in the slice header. ● CC-ALF filter selection is controlled at CTU-level for each chroma component ● Boundary padding for the horizontal virtual boundaries uses the same memory access pattern as luma ALF.
[0071] As an additional feature, the reference encoder can be configured to enable some basic subjective tuning through the configuration file. When enabled, the VTM attenuates the application of CC-ALF in regions that are coded with high QP and are either near mid-grey or contain a large amount of luma high frequencies. Algorithmically, this is accomplished by disabling the application of CC-ALF in CTUs where any of the following conditions are true: ● The slice QP value minus 1 is less than or equal to the base QP value ● The number of chroma samples for which the local contrast is greater than (1 << (bitDepth –2) ) –1 exceeds the CTU height, where the local contrast is the difference between the maximum and minimum luma sample values within the filter support region. ● More than a quarter of chroma samples are in the range between (1 << (bitDepth –1) ) –16 and (1 << (bitDepth –1) ) + 16
[0072] The motivation for this functionality is to provide some assurance that CC-ALF does not amplify artifacts introduced earlier in the decoding path (This is largely due the fact that the VTM currently does not explicitly optimize for chroma subjective quality) . It is anticipated that alternative encoder implementations would either not use this functionality or incorporate alternative strategies suitable for their encoding characteristics. ■ Filter parameters signaling
[0073] ALF filter parameters are signaled in Adaptation Parameter Set (APS) . In one APS, up to 25 sets of luma filter coefficients and clipping value indexes, and up to eight sets of chroma filter coefficients and clipping value indexes could be signaled. To reduce bits overhead, filter coefficients of different classification for luma component can be merged. In slice header, the indices of the APSs used for the current slice are signaled.
[0074] Clipping value indexes, which are decoded from the APS, allow determining clipping values using a table of clipping values for both luma and Chroma components. These clipping values are dependent of the internal bitdepth. More precisely, the clipping values are obtained by the following formula: AlfClip= {round (2B-a*n) for n∈ [0.. N-1] } with B equal to the internal bitdepth, a is a pre-defined constant value equal to 2.35, and N equal to 4 which is the number of allowed clipping values in VVC. The AlfClip is then rounded to the nearest value with the format of power of 2.
[0075] In slice header, up to 7 APS indices can be signaled to specify the luma filter sets that are used for the current slice. The filtering process can be further controlled at CTB level. A flag is always signaled to indicate whether ALF is applied to a luma CTB. A luma CTB can choose a filter set among 16 fixed filter sets and the filter sets from APSs. A filter set index is signaled for a luma CTB to indicate which filter set is applied. The 16 fixed filter sets are pre-defined and hard-coded in both the encoder and the decoder.
[0076] For chroma component, an APS index is signaled in slice header to indicate the chroma filter sets being used for the current slice. At CTB level, a filter index is signaled for each chroma CTB if there is more than one chroma filter set in the APS.
[0077] The filter coefficients are quantized with norm equal to 128. In order to restrict the multiplication complexity, a bitstream conformance is applied so that the coefficient value of the non-central position shall be in the range of -27 to 27 -1, inclusive. The central position coefficient is not signaled in the bitstream and is considered as equal to 128. Adaptive Loop Filter in ECM (Enhanced Compression Model) ■ ALF simplification removal
[0078] In ECM, ALF gradient subsampling and ALF virtual boundary processing are removed. Sample block size (NxN) for classification is reduced from 4x4 to 2x2. Filter size for both luma and chroma, for which ALF coefficients are signaled, is increased to 9x9. ■ Block classification
[0079] In ECM, to filter a luma sample, three different classifiers (C0, C1 and C2) and three different sets of filters (F0, F1 and F2) are used. Sets F0 and F1 contain fixed filters, with coefficients trained for classifiers C0 and C1. Coefficients of filters in F2 are signaled. Which filter from a set Fi is used for a given sample is decided by a class Ci assigned to this sample using classifier Ci.
[0080] Further, based on directionality Di and activity aclass Ci is assigned to each 2x2 block: where MD, i represents the total number of directionalities Di.
[0081] As in VVC, values of the horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within a 4×4 window that covers the target 2×2 block is used for classifier C0 and the sum of sample gradients within a 12×12 window is used for classifiers C1 and C2. The sums of horizontal, vertical and two diagonal gradients are denoted, respectively, as and The directionality Di is determined by comparing with a set of thresholds. The directionality D2 is derived as in VVC using thresholds 2 and 4.5. For D0 and D1, horizontal / vertical edge strength and diagonal edge strength are calculated first. Thresholds Th= [1.25, 1.5, 2, 3, 4.5, 8] are used. Edge strength is 0 if Th [0] ; otherwise, is the maximum integer such that Edge strength is 0 if otherwise, is the maximum integer such that When i.e., horizontal / vertical edges are dominant, the Di is derived by using Table 2A; otherwise, diagonal edges are dominant, the Di is derived by using Table 2B. Table 2A. Mapping of and to Di Table 2B. Mapping of and to Di
[0082] To obtain the sum of vertical and horizontal gradients Ai is mapped to the range of 0 to n, where n is equal to 4 for and 15 for and
[0083] In an ALF_APS, up to 4 luma filter sets are signaled, each set may have up to 25 filters. ■ Alternative 2x2 ALF classifier
[0084] Classification in ALF is extended with an additional alternative classifier. For a signaled luma filter set, a flag is signaled to indicate whether the alternative classifier is applied. Geometrical transformation is not applied to the alternative band classifier. When the band-based classifier is applied, the sum of sample values of a 2x2 luma block is calculated at first. Then the class index is calculated as below, class_index = (sum *25) >> (sample bit depth + 2) . ■ Residual based classifier
[0085] A third classifier based on luma residual sample values. For each 2x2 luma block, the sum of absolute values of the residual samples in a neighboring 8x8 window is calculated, and the class index is derived as: classIdx = sum >> (sample bit depth –4) .
[0086] The value of classIdx is in the range of 0 to 24, same as in ECM-8.0. The classifier usage is signaled for each luma filter set in APS. ■ Filtering process
[0087] At first, two 13x13 diamond shape fixed filters F0 and F1 are applied to derive two intermediate samples R_0 (x, y) and R_1 (x, y) . After that, F2 is applied to R_0 (x, y) , R_1 (x, y) , and neighboring samples to derive a filtered sample as where fi, j is the clipped difference between a neighboring sample and current sample R (x, y) and gi is the clipped difference between Ri-20 (x, y) and current sample. The filter coefficients ci, i=0, …21, are signaled. ■ CCALF with long tap filter
[0088] The CCALF process uses a linear filter to filter luma sample values and generate a residual correction for the chroma samples. A 25-tap large filter is used in CCALF process, which is illustrated in Fig. 4. For a given slice, the encoder can collect the statistics of the slice, analyze them and can signal up to 16 filters through APS.
[0089] The CCALF process uses a linear filter to filter luma sample values and generate a residual correction for the chroma samples. A 25-tap large filter is used in CCALF process, which is illustrated in Fig. 5. In Fig. 5, taps for luma samples are shown in grey dots and the location of the corresponding chroma sample is shown as a small dash-lined circle. For a given slice, the encoder can collect the statistics of the slice, analyze them and can signal up to 16 filters through APS. ■ Adaptive filter shape switch and using samples before deblocking filter for adaptive loop filter
[0090] Two candidate filter shapes: a diamond shape as shown in Fig. 6 and a new cross shape as shown in Fig. 7, can be adaptively selected by the luma filters in ALF. The number of coefficients of a luma filter is 22 for both the filter shapes. Please note that these 22 taps are constituted with 20 spatial taps (610 and 710 in Fig. 6 and Fig. 7 respectively) and 2 fixed filters based taps (620 and 720 in Fig. 6 and Fig. 7 respectively) in both shapes.
[0091] In each adaptation parameter set (APS) , a shape index for the derived luma filters is signaled to decoder. Each APS contains the luma filters that are associated with the filter shape index.
[0092] For each CTB, an APS index is signaled to indicate which luma filter shape is used to filter the current CTB. When filtering a luma sample, the coefficients and clip indices are also rearranged according to the corresponding filter shape.
[0093] The diamond shape luma ALF is replaced by the longer filter shown in Fig. 7.
[0094] The samples before deblocking filters are used as additional inputs for ALF. A final ALF sample is derived by weighting the regular ALF and the filter applied to the samples before the deblocking filter. Specifically, a filtered sample is derived as where fi, j is the clipped difference between a neighboring sample and current sample R (x, y) , gi is the clipped difference between an intermediate sample and current sample R (x, y) and hi, j is the clipped difference between a neighboring sample before DBF and current sample R (x, y) . The filter coefficients ci, i=0, …24 are signalled. In this test, 3x3 diamond shape is applied to samples before deblocking filter. In an APS, a flag is signalled to indicate whether samples before DBF are used for ALF which is always set as true at encoder. ■ Extended Fixed-Filter-Output based Taps for ALF
[0095] In ALF online-trained filters consist of 4 kinds of filter taps: spatial taps (810) , reconstruction-before-DBF based taps (840) , residual based taps (850) and fixed-filter-output based taps (820 and 830) as shown in Fig. 8. ■ ALF with residual samples
[0096] The residual samples are used as additional inputs to the ALF. A filtered sample is derived as where ri is the clipped neighboring residual sample value and rFilteredi is the clipped residual sample filtered by the fixed-filter. For residual samples, the fixed filter reuses the offline fixed filter trained for reconstruction after SAO. ■ Additional fixed filter for ALF
[0097] Additional fixed filter with a shape of diamond 7x7 is introduced, the filter parameters are stored at both encoder and decoder. There is no classification for the newly added fixed filter.
[0098] An online filter or online-trained filter of the proposed method is shown in Fig. 9, where spatial taps 910 (i.e., tap #0 ~ #19) , reconstruction-before-DBF-based taps 940 (i.e., tap #26, #27, #36) , residual-based taps 950 (i.e., #37 ~ #38) and fixed-filter-output-based taps 920 and 930 (i.e., tap #20 ~ #25, #34, #35) are kept the same as the ECM-8.0, and several extended taps 960 (i.e., tap #28 ~ #33, #39) are introduced into luma online-trained filters. The reconstruction before DBF is fed into the additional fixed filter to produce the filter outputs, then these filter outputs are used as input for newly extended taps. The online filter or online-trained filter refers to a filter specified in APS (Adaptation Parameter Set) , where the filter is trained at the encoder and signaled to decoder. The online filter or online-trained filter is in contrast to fixed filters, which are offline-trained and pre-defined in the specification.
[0099] This filter is always enabled without any filter shape switching. ■ Improved fixed filters for ALF
[0100] Two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to C′by a scaling factor. The value of C′is an integer between 0 and 7, inclusively. With i=0, 1, let Ci denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index Ci′is derived as C′i= C′*896+Ci.
[0101] The total number of the fixed filters is not changed.
[0102] Then a class index is determined based on the activity and directionality values. Two diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in the table below. Table 3. Mapping of and to Di
[0103] Fixed filter f1 is applied to outputs of f0 (instead of ALF input) and samples before DBF.
[0104] Finally, a signaled filter is applied to the ALF input samples, samples before the deblocking filter (DBF) , outputs of the two fixed filters, output of a gaussian filter and the residual data. ■ Luma residual taps in CCALF (JVET-AF0197)
[0105] Fig. 10 illustrates an example of filter shape for CCALF. For CCALF, five luma residual taps in a cross 3x3 shape are added. The extended taps take the co-located and neighboring luma residual values as input. ALF with Alternative Clipping Functions
[0106] In VVC and ECM ALF, if non-linear filtering is enabled for a filter set, one clipping index would be signaled and enable a clipping operation during the filter process. When applying the filter set, one filter coefficient would be multiplied by one clipped sample difference, where the clipping threshold is determined by one corresponding clipping index. Such non-linear mechanism provides a wide variety of ways to model the relationship between neighboring differences of a reconstructed sample and the coding distortion of the sample. In this disclosure, more clipping operations and related syntax design are provided to further improve coding efficiency.
[0107] In some embodiments, a piecewise linear function is used for ALF clipping. For example, sample differences are split into more than one interval, and for each interval, there is a corresponding linear function may be applied to derive the clipped value. Embodiment A:
[0108] A ReLU (rectified linear unit) function is used for ALF clipping. If the sample difference is larger than or equal to zero, the linear function is an identical function (slope equal to 1) ; while if the sample difference is smaller than zero, the linear function is a zero function, as shown in Fig. 11A. The ReLU function consists of two linear segments intersecting at the origin (0, 0) , and can be formulated as Embodiment B:
[0109] A leakyReLU function is used for ALF clipping. If the sample difference is larger than or equal to zero, the linear function is an identical function (slope equal to 1) ; while if the sample difference is smaller than zero, the slope of the linear function is smaller than 1, as shown in Fig. 11B. The leakyReLU function consists of two linear segments intersecting at the origin (0, 0) , and can be formulated as Embodiment C:
[0110] A more complicated piecewise linear function is used for ALF clipping, as shown in Fig. 11C. There can be multiple different slopes inside linear function, for instance, slope greater than 1, slope equal to 1 or slope smaller than 1. The multiple slope values inside linear function can be the same in some intervals. The piecewise linear function consists of three linear segments segmented at x= ki &x=-ki, can be formulated as Embodiment D:
[0111] An inverse ReLU function is used for ALF clipping. If the sample difference is smaller than or equal to zero, the linear function is an identical function (slope equal to 1) ; while if the sample difference is larger than zero, the linear function is a zero function, as shown in Fig. 11D. The inverse ReLU function consists of two linear segments intersecting at the origin (0, 0) , and can be formulated as Embodiment E:
[0112] An inverse LeakyReLU function is used for ALF clipping. If the sample difference is smaller than or equal to zero, the linear function is an identical function (slope equal to 1) ; while if the sample difference is larger than zero, the slope of the linear function is smaller than 1, as shown in Fig. 11E. The inverse leakyReLU function consists of two linear segments intersecting at the origin (0, 0) , and can be formulated as Embodiment F:
[0113] A ReLU function with clipping is used for ALF clipping. If the sample difference is larger than or equal to zero and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to a threshold, the sample difference is clipped to the threshold value. If the sample difference is smaller than zero, the linear function is a zero function, as shown in Fig. 11F. The ReLU function with clipping consists of three linear segments segmented at x= ki &the origin (0, 0) , and can be formulated as Embodiment G:
[0114] A leakyReLU function with clipping is used for ALF clipping. If the sample difference is larger than or equal to zero and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to a threshold, the sample difference is clipped to the threshold value. If the sample difference is smaller than zero, the slope of the linear function is smaller, as shown in Fig. 11G. The leakyReLU function with clipping consists of three linear segments segmented at x= ki &the origin (0, 0) , and can be formulated as Embodiment H:
[0115] Absolute clipping is used for ALF clipping. If the sample difference is larger than or equal to zero and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to a threshold, the sample difference is clipped to the threshold value. If the sample difference is negative, the clipped value will be the absolute value of the sample difference, as shown in Fig. 11H. The Absolute clipping function consists of four linear segments segmented at x= ki, x=-ki &the origin (0, 0) , and can be formulated as Embodiment I:
[0116] Inverse clipping is used for ALF clipping. If the sample difference is smaller than a positive threshold or larger than a negative threshold, the sample difference will be clipped to the threshold value. For the sample difference larger than a positive threshold or smaller than a negative threshold, an identical function (slope equal to 1) is applied, as shown in Fig. 11I. The Inverse clipping function consists of four linear segments segmented at x= ki, x=-ki &the origin (0, 0) , and can be formulated as Embodiment J:
[0117] Min-max clipping is used for ALF clipping. If the sample difference lies within a positive threshold1 and positive threshold2, an identical function (slope equal to 1) is applied. For the sample difference larger than positive threshold2 / smaller than position threshold1, the sample difference will be clipped to threshold2 value / threshold1 value. Similar concept can be extended to negative part, as shown in Fig. 11J. The Inverse clipping function consists of six linear segments segmented at x= ±ki, x=±kj &the origin (0, 0) , and can be formulated as Embodiment K:
[0118] A shifted ReLU function is used for ALF clipping. If the sample difference is larger than or equal to an offset or a vector, the linear function is an identical function (slope equal to 1) . If the sample difference is smaller than the offset, the linear function is a zero function, as shown in Fig. 11K. The shifted ReLU function consists of two linear segments intersecting at the coordinate point shifted from the origin (0, 0) and offset value, and can be formulated as Embodiment L:
[0119] A shifted ReLU function with clipping is used for ALF clipping. If the sample difference is larger than or equal to an offset and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to the threshold, the sample difference is clipped to the threshold value. If the sample difference is smaller than the offset, the linear function is a zero function, as shown in Fig. 11L. The shifted ReLU function consists of three linear segments segmented at x= ki &x= offset, and can be formulated as Embodiment M:
[0120] A shifted LeakyReLU function is used for ALF clipping. If the sample difference is larger than or equal to an offset, the linear function is an identical function (slope equal to 1) ; while if the sample difference is smaller than the offset, the slope of the linear function is smaller than 1, as shown in Fig. 11M. The shifted LeakyReLU function consists of two linear segments segmented at x= offset, and can be formulated as Embodiment N:
[0121] A shifted LeakyReLU function with clipping is used for ALF clipping. If the sample difference is larger than or equal to an offset and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to the threshold, the sample difference is clipped to the threshold value. If the sample difference is smaller than the offset, the slope of the linear function is smaller than 1, as shown in Fig. 11N. The shifted LeakyReLU function with clipping consists of three linear segments segmented at x= ki &x= offset, and can be formulated as Embodiment O:
[0122] A shifted ReLU function with non-constant clipping is used for ALF clipping. If the sample difference is larger than or equal to an offset and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to the threshold, the slope of linear function is smaller than 1. If the sample difference is smaller than the offset, the linear function is a zero function, as shown in Fig. 11O. The shifted ReLU function with non-constant clipping consists of three linear segments segmented at x= ki &x= offset, and can be formulated as Embodiment P:
[0123] A shifted LeakyReLU function with non-constant clipping is used for ALF clipping. If the sample difference is larger than or equal to an offset and smaller than a threshold, the linear function is an identical function (slope equal to 1) . For the sample difference is larger than or equal to the threshold, the slope of linear function is smaller than 1. If the sample difference is smaller than the offset, the slope of the linear function is smaller than 1, as shown in Fig. 11P. The shifted LeakyReLU function with non-constant clipping consists of three linear segments segmented at x= ki &x= offset, and can be formulated as Embodiment Q:
[0124] A sawtooth function with clipping is used for ALF clipping. If the sample difference is within an interval, the linear function is an identical function (slope equal to 1) . If the sample difference is larger than an offset or smaller than an offset, the linear function is an inverse identical function (slope equal to -1) . For sample difference larger than 2*offset or smaller than -2*offset, the sample difference is clipped to be zero, as shown in Fig. 11Q. The sawtooth function with clipping consists of three linear segments segmented at x= ki &x=-ki, and can be formulated as Embodiment R:
[0125] A periodic sawtooth function is used for ALF clipping. Linear functions with slope equal to 1 and -1 are periodically used to clip the sample difference, as shown in Fig. 11R. The periodic sawtooth function consists of multiple linear segments.
[0126] In another embodiments, non-piecewise function may be used for ALF clipping. Embodiment S :
[0127] A sigmoid function is used for ALF clipping. Embodiment T: A hyperbolic tangent (tanh) function is used for ALF clipping. Embodiment U:
[0128] VVC or ECM clipping operations (min-max clipping) could be combined with the other examples.
[0129] In one embodiment, for syntax design, in some embodiments, different coefficients have different clipping operation sets. The clipping operation set used for a coefficient could be determined by the source which the coefficient would be applied to, determined by the index of the coefficient, or determined by the classification results.
[0130] Yet in other embodiment, an ALF is composed of pre-ALF taps, pre-DBF taps, fixed-filter-output-based taps, and cross-component taps, where cross-component taps use chroma / luma sample values with signaled coefficients for luma / chroma component filtering. In such ALF, cross-component taps use alternative clipping functions (e.g. embodiments A~T, for example ReLU / leakyReLU clipping functions) , for filtering operations, while other taps use VVC / ECM-like clipping (e.g. embodiment U) for filtering operations.
[0131] In ALF, two index sets are pre-defined. For coefficients with index in the first index set, the clipping thresholds are {1024, 128, 32, 8} , while for coefficients with index in the second index set, the clipping thresholds are {1024, 64, 16, 4} .
[0132] Further, in another embodiment, for classification operation of a sample block, each NxN block will be classified into different classes. For some classes (indicated by the class_Index) the alternative clipping functions (e.g. embodiments A~T, for example ReLU / leakyReLU clipping functions) are used, while for the other classes, VVC / ECM-like clipping operations (e.g. embodiment U) are used.
[0133] In another embodiments, whether to use alternative clipping and / or which alternative clipping operation set (s) to use are signaled at APS / filter-set / filter level or at picture / sub-picture / slice / tile / CTB / block level of the syntax of the video bitstream.
[0134] In a further embodiment, as shown in the following table, the first bin for encoding and decoding indicates the usage of ReLU-like clipping operation. If the first bin is 1, current coefficient uses ReLU-like clipping operation. If the first bin is 0, two additional bins are signaled to indicate the threshold of VVC / ECM-like clipping operations with corresponding threshold. Table 4A
[0135] In addition, the following table shows another example, where truncated binary is used for syntax design. Table 4B
[0136] Further, as shown in the following tables, at APS / filter set level, one flag is signaled to indicate the VVC / ECM-like clipping operation is used or the ReLU-like clipping operation is used. If the VVC / ECM-like clipping operation is used, clipping syntax same as VVC / ECM is signaled at filter level; if the ReLU-like clipping operation is used, an alternative clipping syntax signaling is used or no clipping syntax signaling is required. Table 5A
[0137] Another table example: Table 5B
[0138] In another embodiments, whether to use alternative clipping and / or which alternative clipping operation set (s) to use are implicitly determined by some existing flags in APS, such as classifier flags, ALF filter coefficient bit-depth or coefficient precision flags either alone or in combination.
[0139] In more embodiment, if one filter set selects gradient-based classifier, one clipping operation set is used for filters in the filter set; if one filter set selects band-based classifier, another clipping operation set is used for filters in the filter set; if one filter set selects residual-based classifier, yet another clipping operation set is used for filters in the filter set.
[0140] If one filter set selects to signal coefficients in 8-bit, VVC / ECM-like clipping operation with clipping threshold set {1024, 128, 32, 8} is used for filters in the filter set; if one filter set selects to signal coefficients in less than 8-bit, VVC / ECM-like clipping operation with clipping threshold set {1024, 64, 16, 4} is used for filters in the filter set. Otherwise, is will select the alternative clipping functions (e.g. embodiments A~T, for example ReLU / leakyReLU clipping functions) .
[0141] Fig. 12 illustrates a flowchart of an exemplary method for processing video data according to an embodiment of the present disclosure. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. The method includes: receiving a picture of the video data wherein the picture comprises a current video region (1210) ; determining a filtering operation is applied for the current video region (1220) ; determining, for the current video region, a clipping index element set (1230) ; and performing the filtering operation, according to the clipping index element set, to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions, wherein the clipping index element set comprises a first clipping index element for indicating the selected clipping function and a second clipping index element comprising a segment coordinate point element and a slope element corresponding to function parameters for formulating the selected clipping function (1240) .
[0142] In some embodiments, the current video region comprises a current sample, and the filtering operation is performed using the clipping function for clipping a difference value between a value of a neighboring sample of the current sample and a value of the current sample.
[0143] In some embodiments, the clipping functions correspond to a piecewise clipping function group or a non-piecewise clipping function group; wherein the piecewise clipping function group comprises a ReLU function, a LeakyReLU function, an inverse ReLU function, inverse a LeakyReLU function, a ReLU function with clipping, a leakyReLU function with clipping, a shifted ReLU function with / without clipping, a shifted leakyReLU function with / without clipping, a shifted ReLU function with non-constant clipping, a shifted LeakyReLU function with non-constant clipping, a Sawtooth wave function with clipping and a Periodic sawtooth wave function, and wherein the non-piecewise clipping function group comprises a sigmoid function and a hyperbolic tangent (tanh) function.
[0144] In some embodiments, the method may further include: singing the clipping index element set into a syntax of a video bitstream. The clipping index element set is signaled into an Adaptation Parameter Set (APS) , Filter-set or filter level of the syntax of the video bitstream.
[0145] In some embodiments, the first clipping index element may be determined based on a filter coefficient precision. Alternatively or additionally, the first clipping index element may be determined based on a source of filtering sample.
[0146] In some embodiments, the second clipping index element further comprises a shift vector element for shifting the segment coordinate point element toward a direction and distance according to the shift vector.
[0147] In some embodiments, the method may further includes: performing a classification process for the current region to categorize the current region into a class indicated by a class index; and determining the first clipping index element according to the class index.
[0148] The foregoing proposed methods can be implemented in encoders and / or decoders. For example, the proposed method can be implemented in an in-loop filtering module of an encoder, and / or an in-loop filtering module of a decoder.
[0149] Any of the methods of ALF processing described above can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in the in-loop filter module (e.g. ILPF 130 in Fig. 1A and Fig. 1B) of an encoder or a decoder. Alternatively, any of the proposed methods can be implemented as circuits coupled to the inter coding module of an encoder and / or motion compensation module, a merge candidate derivation module of the decoder. The simplified ALF methods may also be implemented using executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
[0150] The flowchart shown is intended to illustrate an example of video coding according to the present disclosure. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present disclosure without departing from the spirit of the present disclosure. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present disclosure. A skilled person may practice the present disclosure by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present disclosure.
[0151] The above description is presented to enable a person of ordinary skill in the art to practice the present disclosure as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present disclosure is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present disclosure. Nevertheless, it will be understood by those skilled in the art that the present disclosure may be practiced.
[0152] Embodiment of the present disclosure as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present disclosure can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present disclosure may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The disclosure may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the disclosure, by executing machine-readable software code or firmware code that defines the particular methods embodied by the disclosure. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the disclosure will not depart from the spirit and scope of the disclosure.
[0153] The foregoing outlines features of several embodiments or examples so that those skilled in the art may better understand the aspects of the present disclosure. Those skilled in the art should appreciate that they may readily use the present disclosure as a basis for designing or modifying other processes and structures for carrying out the same purposes and / or achieving the same advantages of the embodiments or examples introduced herein. Those skilled in the art should also realize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they may make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.
Claims
1.A decoder for decoding of a video bitstream associated with video data comprising:a storage module for receiving the video bitstream associated with a picture of the video data wherein the picture comprises a current video region;a processing module for performing the following steps:determining from a syntax of the video stream a filtering operation is applied for the current video region;determining, for the current video region, a clipping index element set; anda filter for performing the filtering operation, according to the clipping index element set, to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions,wherein the clipping index element set is used for indicating the selected clipping function, a segment coordinate point element and a slope element for formulating the selected clipping function.2.The decoder of claim 1, wherein the clipping index element set is determined from an Adaptation Parameter Set (APS) of the syntax of the video bitstream.3.The decoder of claim 1, wherein the clipping functions correspond to a piecewise clipping function group comprising a ReLU function, a LeakyReLU function, an inverse ReLU function, inverse a LeakyReLU function, a ReLU function with clipping, a leakyReLU function with clipping, a shifted ReLU function with / without clipping, a shifted leakyReLU function with / without clipping, a shifted ReLU function with non-constant clipping, a shifted LeakyReLU function with non-constant clipping, a Sawtooth wave function with clipping and a Periodic Sawtooth Wave function.4.The decoder of claim 1, wherein the clipping functions correspond to a non-piecewise clipping function group comprising a sigmoid function and a hyperbolic tangent (tanh) function.5.The decoder of claim 1, wherein the clipping index element set is further used for indicating a shift vector element for shifting the segment coordinate point element toward a direction and a distance according to the shift vector.6.The decoder of claim 1, wherein the current video region comprises a current sample, and wherein the filtering operation is performed using the clipping function for clipping a difference value between a value of a neighboring sample of the current sample and a value of the current sample.7.The decoder of claim 1, further comprises:performing a classification process for the current region to categorize the current region into a class indicated by a class index; anddetermining the clipping index element set according to the class index.8.A method for processing video data, comprising:receiving a picture of the video data wherein the picture comprises a current video region;determining a filtering operation is applied for the current video region;determining, for the current video region, a clipping index element set; andperforming the filtering operation to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions,wherein the clipping index element set is used for indicating the selected clipping function, a segment coordinate point element and a slope element for formulating the selected clipping function.9.The method of claim 8, further comprising:singing the clipping index element set into a syntax of a video bitstream.10.The method of claim 9, wherein the clipping index element set is signaled into an Adaptation Parameter Set (APS) , Filter-set or filter level of the syntax of the video bitstream.11.The method of claim 8, wherein the clipping functions correspond to a piecewise clipping function group comprising a ReLU function, a LeakyReLU function, an inverse ReLU function, inverse a LeakyReLU function, a ReLU function with clipping, a leakyReLU function with clipping, a shifted ReLU function with / without clipping, a shifted leakyReLU function with / without clipping, a shifted ReLU function with non-constant clipping, a shifted LeakyReLU function with non-constant clipping, a Sawtooth wave function with clipping and a Periodic Sawtooth wave function.12.The method of claim 8, wherein the clipping functions correspond to a non-piecewise clipping function group comprising a sigmoid function and a hyperbolic tangent (tanh) function.13.The method of claim 8, wherein the clipping index element set is further used for indicating a shift vector element for shifting the segment coordinate point element toward a direction and distance according to the shift vector.14.The method of claim 8, wherein the current video region comprises a current sample, and wherein the filtering operation is performed using the clipping function for clipping a difference value between a value of a neighboring sample of the current sample and a value of the current sample.15.The method of claim 8, further comprises:performing a classification process for the current region to categorize the current region into a class indicated by a class index; anddetermining the clipping index element set according to the class index.16.The method of claim 8, wherein the clipping index element set is determined according to a filter coefficient precision.17.The method of claim 8, wherein the clipping index element set is determined according to a source of filtering sample.18.The method of claim 8, wherein the clipping index element set is determined according to an ALF filter coefficient bit-depth.19.A method for processing video data, comprising:receiving a picture of the video data wherein the picture comprises a current video region;determining a filtering operation is applied for the current video region;determining, for the current video region, a clipping index element set; andperforming the filtering operation to the current video region for generating a filtered current video region using a selected clipping function among a plurality of clipping functions,wherein the clipping index element set is used to indicating the selected clipping function and is implicitly determined by an existing flag.20.The method of claim 19, wherein the existing flag associates with a class index, a filter coefficient precision, a source of filtering sample, or an ALF filter coefficient bit-depth.