Method and apparatus of syntax design for prediction filtering in video coding
Patent Information
- Application Number
- PCT/CN2026/083835
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-17
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026083835_01102026_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS OF SYNTAX DESIGN FOR PREDICTION FILTERING IN VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 777,045, filed on March 25, 2025. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding. In particular, the present invention discloses signalling schemes of filter information for a smoothing filter used to filter reference pictures to generate filtered predictor in video coding. Also, conditions for using the filtered prediction are also disclosed. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, are provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.
[0008] I. 1 Adaptive Loop Filter in VVC
[0009] In VVC, an Adaptive Loop Filter (ALF) with block-based filter adaption is applied. For the luma component, one filter is selected among 25 filters for each 4×4 block, based on the direction and activity of local gradients.
[0010] I. 1.1. Filter shape
[0011] Two diamond filter shapes (as shown in Fig. 2) are used. The 7×7 diamond shape 220 is applied for luma component and the 5×5 diamond shape 210 is applied for the chroma components.
[0012] I. 1.2. Block classification and geometric transformation
[0013] For luma component, each 4×4 block is categorized into one out of 25 classes. Before filtering each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f (k, l) and to the corresponding filter clipping values c (k, l) depending on gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning their directionality.
[0014] For chroma components in a picture, no classification method is applied.
[0015] More details could be found in VVC specification section 8.8.5.3.
[0016] I. 1.3. Filtering process
[0017] At decoder side, when ALF is enabled for a CTB, each sample R (i, j) within the CU is filtered, resulting in sample value R′ (i, j) as shown below, where f (k, l) denotes the decoded filter coefficients, K (x, y) is the clipping function and c (k, l) denotes the decoded clipping parameters. The variable k and l varies between –L / 2 and L / 2, where L denotes the filter length. The clipping function K (x, y) =min (y, max (-y, x) ) which corresponds to the function Clip3 (-y, y, x) . The clipping operation introduces non-linearity to make ALF more efficient by reducing the impact of neighbour sample values that are too different with the current sample value.
[0018] I. 1.4. Cross component adaptive loop filter
[0019] CC-ALF uses luma sample values to refine each chroma component by applying an adaptive, linear filter to the luma channel and then using the output of this filtering operation for chroma refinement. Fig. 3A provides a system level diagram of the CC-ALF process with respect to the SAO, luma ALF and chroma ALF processes. As shown in Fig. 3A, each colour component (i.e., Y, Cb and Cr) is processed by its respective SAO (i.e., SAO Luma 310, SAO Cb 312 and SAO Cr 314) . After SAO, ALF Luma 320 is applied to the SAO-processed luma and ALF Chroma 330 is applied to SAO-processed Cb and Cr. However, there is a cross-component term from luma to a chroma component (i.e., CC-ALF Cb 322 and CC-ALF Cr 324) . The outputs from the cross-component ALF are added (using adders 332 and 334 respectively) to the outputs from ALF Chroma 330.
[0020] Filtering in CC-ALF is accomplished by applying a linear, diamond shaped filter (e.g. filters 340 and 342 in Fig. 3B) to the luma channel. In Fig. 3B, a blank circle indicates a luma sample and a dot-filled circle indicate a chroma sample. One filter is used for each chroma channel, and the operation is expressed as: where (x, y) is chroma component i location being refined, (xY, yY) is the luma location based on (x, y) , Si is filter support area in luma component, and ci (x0, y0) represents the filter coefficients.
[0021] As shown in Fig, 3B, the luma filter support is the region collocated with the current chroma sample after accounting for the spatial scaling factor between the luma and chroma planes. In Fig. 3B, circles represent luma samples, while dotted circles represent chroma samples being refined.
[0022] I. 1.5. Filter parameters signalling
[0023] ALF filter parameters are signalled in Adaptation Parameter Set (APS) . In one APS, up to 25 sets of luma filter coefficients and clipping value indexes, and up to eight sets of chroma filter coefficients and clipping value indexes can be signalled. To reduce bits overhead, filter coefficients of different classification for luma component can be merged. In slice header, the indices of the APSs used for the current slice are signalled.
[0024] Clipping value indexes, which are decoded from the APS, allow determining clipping values using a table of clipping values for both the luma and chroma components. These clipping values are dependent of the internal bit-depth. More precisely, the clipping values are obtained by the following formula: AlfClip= {round (2B-α*n ) for n∈ [0.. N-1] } with B equal to the internal bit-depth, α is a pre-defined constant value equal to 2.35, and N equal to 4 which is the number of allowed clipping values in VVC. The AlfClip is then rounded to the nearest value with the format of power of 2.
[0025] In slice header, up to 7 APS indices can be signalled to specify the luma filter sets that are used for the current slice. The filtering process can be further controlled at CTB level. A flag is always signalled to indicate whether ALF is applied to a luma CTB. A luma CTB can choose a filter set among 16 fixed filter sets and the filter sets from APSs. A filter set index is signalled for a luma CTB to indicate which filter set is applied. The 16 fixed filter sets are pre-defined and hard-coded in both the encoder and the decoder.
[0026] For the chroma component, an APS index is signalled in slice header to indicate the chroma filter sets being used for the current slice. At CTB level, a filter index is signalled for each chroma CTB if there is more than one chroma filter set in the APS.
[0027] The filter coefficients are quantized with norm equal to 128. In order to restrict the multiplication complexity, a bitstream conformance is applied so that the coefficient value of the non-central position shall be in the range of -27 to 27 -1, inclusive. The central position coefficient is not signalled in the bitstream and is considered as equal to 128.
[0028] I. 2 Adaptive Loop Filter in ECM
[0029] In ECM8 (Muhammed Coban, et al., “Algorithm description of Enhanced Compression Model 8 (ECM 8) ” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29) , 29th Meeting, by teleconference, 11–20 January 2023, Document: JVET-AC2025) , some changes from the VVC ALF are disclosed. A brief overview is shown below.
[0030] I. 2.1. ALF simplification removal
[0031] ALF gradient subsampling and ALF virtual boundary processing are removed. Block size for classification is reduced from 4x4 to 2x2. Filter size for both the luma and chroma, for which ALF coefficients are signalled, is increased.
[0032] I. 2.2. ALF with fixed filters
[0033] To filter a luma sample, three different classifiers (C0, C1 and C2) and three different sets of filters (F0, F1 and F2) are used. Sets F0 and F1 contain fixed filters, with coefficients trained for classifiers C0 and C1. Coefficients of filters in F2 are signalled. Which filter from a set Fi is used for a given sample is decided by a class Ci assigned to this sample using classifier Ci.
[0034] The number of bits used to represent the fractional part of a luma coefficient is adaptive from 5 to 8, inclusively. For each luma filter set, which contains up to 25 filters, a 2-bit syntax element is signalled in APS to indicate the number of bits used for the coefficients in this set. The value range of a coefficient is not changed.
[0035] I. 2.3. Filtering
[0036] First, two 13x13 diamond shape fixed filters F0 and F1 are applied to derive two intermediate samples R0 (x, y) and R1 (x, y) . After that, F2 is applied to R0 (x, y) , R1 (x, y) , and neighbouring samples to derive a filtered sample as where fi, j is the clipped difference between a neighbouring sample and current sample R (x, y) and gi is the clipped difference between Ri-20 (x, y) and current sample. The filter coefficients ci, i=0, …21, are signalled.
[0037] I. 2.4. Classification
[0038] Based on directionality Di and activity aclass Ci is assigned to each 2x2 block: where MD, i represents the total number of directionalities Di.
[0039] As in VVC, values of the horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within a 4×4 window that covers the target 2×2 block is used for classifier C0 and the sum of sample gradients within a 12×12 window is used for classifiers C1 and C2. The sums of horizontal, vertical and two diagonal gradients are denoted, respectively, as The directionality Di is determined by comparing with a set of thresholds. The directionality D2 is derived as in VVC using thresholds 2 and 4.5. For D0 and D1, horizontal / vertical edge strength and diagonal edge strength are calculated first. Thresholds Th= [1.25, 1.5, 2, 3, 4.5, 8] are used. Edge strength is 0 if otherwise, is the maximum integer such that Edge strength is 0 if otherwise, is the maximum integer such that When i.e., horizontal / vertical edges are dominant, the Di is derived by using Table 1A; otherwise, diagonal edges are dominant, the Di is derived by using Table 1B. Table 1A. Mapping of to Di Table 1B. Mapping of to Di
[0040] To obtain the sum of vertical and horizontal gradients Ai is mapped to the range of 0 to n, where n is equal to 4 for and 15 for
[0041] In an ALF_APS, up to 4 luma filter sets are signalled, each set may have up to 24 filters.
[0042] I. 2.5. Alternative 2x2 ALF classifier
[0043] Classification in ALF is extended with an additional alternative classifier. For a signalled luma filter set, a flag is signalled to indicate whether the alternative classifier is applied. Geometrical transformation is not applied to the alternative band classifier. When the band-based classifier is applied, the sum of sample values of a 2x2 luma block is calculated at first. Then the class index is calculated as below, class_index = (sum *12) >> (sample bit depth + 2) .
[0044] I. 2.6. Residual based classifier
[0045] A third classifier is based on luma residual sample values. For each 2x2 luma block, the sum of absolute values of the residual samples in a neighbouring 8x8 window is calculated, and the class index is derived as: classIdx = sum >> (sample bit depth - 3) .
[0046] The value of classIdx is in the range of 0 to 24, same as in ECM-8.0. The classifier usage is signalled for each luma filter set in APS.
[0047] I. 2.7. Coding information-based classifier
[0048] For the online filters (signalled filters) , each 2×2 unit is classified into 2 noise levels according to the partitioning information. For a 2×2 unit, it is classified into noise level-1 if it is located at a CU or TU boundary; and into noise level-0 if it is not. The class number of existing classifiers is reduced from 25 to 12. For the texture-based classifier, the 25 classes are mapped to 12 classes with a pre-defined LUT. For the band-based and residual based classifiers, the 25 classes are decreased to 12 classes by enlarging the band width. These 12 classes are further combined with the proposed 2 noise levels to generate final 12×2 = 24 classes in total. The number of classifiers for online filters is kept as 3 and no additional encoder selection is introduced.
[0049] For the offline filters (fixed filters) , each 2×2 unit is classified into 2 noise levels in the same way as the classification of online filters. Besides, each 2×2 unit is further classified into 2 residual levels based on a predefined threshold also in the same way as the classifier of online filters. Furthermore, the generated offset of offline filters is adjusted based on the boundary level and residual level accordingly, where a stronger offset is applied on the positions at boundaries or with higher residuals.
[0050] I. 2.8. CCALF with long tap filter
[0051] The CCALF process uses a linear filter to filter luma sample values, luma residual samples and generate a residual correction for the chroma samples. In addition, the CCALF filter shape is constructed by 23 luma spatial taps (410) and 5 luma residual taps (420) , which is illustrated in Fig. 4. For a given slice, the encoder can collect the statistics of the slice, analyse them and signal up to 16 filters through APS. The number of bits used to represent the fractional part of a CCALF coefficient can vary from 7 to 10 adaptively.
[0052] Chroma SAO output samples applied to 4 taps (430) in a 3x3 asymmetric cross shape are added as additional inputs to CCALF, as illustrated in Fig. 4.
[0053] I. 2.9. New luma ALF filter shape
[0054] In ALF online-trained filters consist of 4 kinds of filter taps: spatial taps (510) , reconstruction-before-DBF based taps (540) , residual based taps (550) and fixed-filter-output based taps (520 and 530) as shown in Fig. 5, where the fixed-filter-output based taps are extended, the residual based taps contain the clipped residual sample and the clipped residual sample filtered by the fixed filters, and the reconstruction-before-DBF (pre-DBF) based taps contain the clipped pre-DBF samples and the pre-DBF samples filtered by a Gaussian fixed filter. The shape of Gaussian fixed filter is a diamond 7x7 shape, and the filter parameters are stored at both encoder and decoder. There is no classification for this fixed filter.
[0055] I. 2.10. ALF residuals scaling
[0056] A scaling factor is signalled in slice header, the scaling factors is applied to the difference between the ALF input and ALF output, and the scaled residual is added to the ALF input (it produces a scaled ALF filtering) . A similar scaling process is applied to NN filtering in NNVC.
[0057] Different luma scaling factors may be associated with different group of class indexes, and the ALF output is derived as follows: rec’ (s) = rec (s) + (corr (s) *tab [sfi [class (s) ] ] + 4) >> 3, where ALF residual correction ‘corr (s) ’ is scaled using the scaling factor associated to the class index of the sample, ‘tab’ is the predefined LUT for mapping scaling index ‘sfi’ to scaling factor.
[0058] I. 2.11. Improved fixed filters for ALF
[0059] Two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to C′ by a scaling factor. The value of C′ is an integer between 0 and 7, inclusively. With i=0, 1, let Ci denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index Ci′ is derived as C′i=C′ *896+Ci.
[0060] The total number of the fixed filters is not changed.
[0061] Then a class index is determined based on the activity and directionality values. Two diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in Table 2. Table 2. Comparison of fixed filters between ECM-9.0 and JVET-AE0139
[0062] Fixed filter f1 is applied to outputs of f0 (instead of ALF input) and samples before DBF.
[0063] Finally, a signalled filter is applied to the ALF input samples, samples before the deblocking filter (DBF) , outputs of the two fixed filters, output of a Gaussian filter and the residual data.
[0064] I. 2.12. Chroma ALF Fixed Filter
[0065] A classifier based on Laplacian values and variance is applied to a 2x2 chroma block. Compared to the luma classifier of a fixed filter, when calculating the activity value, the sum of the chroma vertical and horizontal Laplacian values is multiplied by 2 before scaling. Similarly, the chroma variance is multiplied by 2 before scaling. The derived class index is then used to select a fixed filter from a chroma filter set. A chroma fixed filter is applied to chroma ALF input samples in a 13x13 diamond shape and DBF input samples in a 7x7 diamond shape. The first luma classifier is applied to each 2x2 chroma block. The derived class index is then used to select a fixed filter from the luma fixed filter set related to this classifier. A fixed filter is applied to chroma ALF input sample in a 9x9 diamond shape and DBF input samples in a 9x9 diamond shape. In a signalled chroma filter, 5x5 crossing extra taps are introduced, which are applied to the fixed filter output.
[0066] I. 3. OBMC
[0067] When OBMC is applied, top and left boundary pixels of a CU are refined using the motion information of neighbouring blocks with a weighted prediction as described in JVET-L0101.
[0068] Conditions of not applying OBMC are as follows: ● When OBMC is disabled at the SPS level ● When the current block is coded in intra mode or IBC mode ● When the current luma block area is smaller or equal to 32
[0069] Additionally, OBMC is adaptively controlled at a block level as follows: ● OBMC flag is inherited from a neighbouring affine block for affine merge mode. ● OBMC is not applied to a block if there is a neighbour block coded with IBC, palette, or BDPCM mode. ● When applying OBMC to a block, block boundary, check whether OBMC is applied to the boundary is further made based on the reference samples of the current block. If any absolute difference between the prediction sample and non-interpolated (integer pel) reference sample is greater than a threshold, the OBMC is not applied to that boundary.
[0070] A subblock-boundary OBMC is performed by applying the same blending to the top, left, bottom, and right subblock boundary pixels using motion information of neighbouring subblocks. It is enabled for the subblock based coding tools: ● Affine AMVP modes; ● Affine merge modes and subblock-based temporal motion vector prediction (SbTMVP) ; ● Subblock-based bilateral matching.
[0071] When OBMC mode is used in CIIP mode with LMCS, inter blending is performed prior to LMCS mapping of inter samples. LMCS is applied to blended inter samples which are combined with LMCS applied intra samples in CIIP mode, where InterpredY represents the samples predicted by the motion of current block in the original domain, IntrapredY represents the samples predicted in the mapped domain, OBMCpredY represents the samples predicted by the motion of neighboring blocks in the original domain, and w0 and w1 are the weights.
[0072] When OBMC mode is used in an LIC coded block, the LIC parameters are applied to generate the corresponding prediction samples for the OBMC of the LIC coded block. Besides, to reduce the complexity, the OBMC is only applied to the top and left CU boundaries while being always disabled for the boundaries of the internal sub-blocks of the LIC coded block.
[0073] I. 4. Multi-Pass Decoder-Side Motion Vector Refinement (MP-DMVR)
[0074] A multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 subblock within the coding block. In the third pass, MV in each 8x8 subblock is refined by applying bi-directional optical flow (BDOF) . The refined MVs are stored for both spatial and temporal motion vector prediction.
[0075] I. 4.1 First pass - Block based bilateral matching MV refinement
[0076] In the first pass, a refined MV is derived by applying BM to a coding block. Similar to decoder-side motion vector refinement (DMVR) , in the bi-prediction operation, a refined MV is searched around the two initial MVs (i.e., MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (i.e., MV0_pass1 and MV1_pass1) are derived around the initiate MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0077] BM performs local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to loop through the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0078] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW *cbH is greater than 64, MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. When the bilCost at the centre point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new centre point of the 3×3 search pattern and continue to search for the minimum cost, until it reaches the end of the search range.
[0079] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MVs after the first pass are then derived as: MV0_pass1 = MV0 + deltaMV MV1_pass1 = MV1 - deltaMV
[0080] I. 4.2 Second pass - Subblock based bilateral matching MV refinement
[0081] In the second pass, a refined MV is derived by applying BM to a 16×16 grid subblock. For each subblock, a refined MV is searched around the two MVs (e.g. MV0_pass1 and MV1_pass1) , obtained during the first pass, in the reference picture list L0 and L1. The refined MVs (i.e., MV0_pass2 (sbIdx2) and MV1_pass2 (sbIdx2) ) are derived based on the minimum bilateral matching cost between the two reference subblocks in L0 and L1.
[0082] For each subblock, BM performs full search to derive integer sample precision intDeltaMV. The full search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.
[0083] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference subblocks, as: bilCost = satdCost *costFactor. The search area (2*sHor + 1) * (2*sVer + 1) is divided up to 5 diamond shape search regions shown on Fig. 6, where the 5 search regions are shown in 5 different shades. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in the order starting from the centre of the search area. In each region, the search points are processed in the raster scan order starting from the top left going to the bottom right corner of the region. When the minimum bilCost within the current search region is less than a threshold equal to sbW *sbH, the int-pel full search is terminated; otherwise, the int-pel full search continues to the next search region until all search points are examined. Additionally, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold that is equal to the area of the block, the search process terminates.
[0084] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2) . The refined MVs at second pass is then derived as: ● MV0_pass2 (sbIdx2) = MV0_pass1 + deltaMV (sbIdx2) ● MV1_pass2 (sbIdx2) = MV1_pass1 –deltaMV (sbIdx2)
[0085] I. 4.3 Third pass - Subblock based bi-directional optical flow MV refinement
[0086] In the third pass, a refined MV is derived by applying BDOF to an 8×8 grid subblock. For each 8×8 subblock, BDOF refinement is applied to derive scaled Vx and Vy without clipping starting from the refined MV of the parent subblock of the second pass. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0087] The refined MVs (e.g. MV0_pass3 (sbIdx3) and MV1_pass3 (sbIdx3) ) at third pass are derived as: ● MV0_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) + bioMv ● MV1_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) –bioMv
[0088] I. 4.4 Fourth pass - Adaptive subblock based bi-directional optical flow MV refinement
[0089] In the fourth pass, a refined MV is derived by applying BDOF to a 4×4 or 8×8 or 16x16 grid subblock. When a block is smaller than 1024 pixels, the 4×4 grid subblock is used. Otherwise, 8×8 grid subblock is used. The MV of each subblock is refined in the same way as that used in third pass.
[0090] In all aforementioned sub-clauses, when wrap around motion compensation is enabled, the motion vectors shall be clipped with wrap around offset taken into consideration. It is noted that in ECM, the DMVR is extended to non-equal POC distance cases, and the mean removed equations are utilized to derive the BDOF MV refinement parameters as: (∑Gx·Gx+R1) *vx + ∑Gx·Gy *vy = ∑dI ·Gx→ (∑Gx·Gx+R1) *vx + ∑Gx·Gy *vy = ∑dI ·Gx -dM ·∑Gx ∑Gx·Gy *vx + (∑Gy·Gy+R1) *vy= ∑dI ·Gy→∑Gx·Gy *vx + (∑Gy·Gy+R1) *vy = ∑dI ·Gy -dM ·∑Gy
[0091] I. 5 Affine subblock BDOF refinement
[0092] BDOF subblock MV refinement and sample adjustment is applied to an affine or SbTMVP coded block with subblock MC when BDOF condition is satisfied.
[0093] An affine coded block, e.g. affine regular merge mode, affine BM merge mode, affine AMVP mode, derives MVs for each 4×4 subblock from the affine model. The BDOF process starts with the 4×4 subblocks grouping with identical MVs. The first iteration of BDOF MV refinement is processed in 8x8 subblock grid as in ECM-10.0. When the grouped subblock size is less than 256, the second iteration of BDOF MV refinement is processed in 4×4 subblock grid, and otherwise in 8×8 subblock grid. When the grouped subblock size is 4xN or Nx4, the first iteration of BDOF MV refinement is bypassed.
[0094] I. 6 Sample-based BDOF
[0095] In the sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, it is performed per sample. In addition, the high accuracy sample adjustment method, is applied to derive the sample adjustment.
[0096] The coding block is divided into 8×8 subblocks. For each subblock, whether to apply BDOF or not is determined by checking the SAD between the two reference subblocks against a threshold. If decided to apply BDOF to a subblock, for every sample in the subblock, a sliding 5×5 window is used and the existing BDOF process is applied for every sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predicted sample value for the centre sample of the window.
[0097] I. 7 Group of pictures (GOP) structure
[0098] A closed GOP prediction structure can be realized in VVC streams as shown in Fig. 7. In Fig. 7, a GOP structure with size 16 is illustrated. In a GOP structure, the frames are coded in a hierarchical layered structure with a re-ordered frame coding order. In Figure6, the term “Temp. id” indicates temporal index and the arrows indicate the frame reference directions (e.g., the intra frame is referenced by the frame with temporal index equal to 0) .
[0099] I. 8 Motion Compensated Temporal Filter (MCTF)
[0100] Motion Compensated Temporal Filter has been proposed as an effective pre-processing tool. More specifically, before the current input frame is sent to the encoder, MCTF uses temporal adjacent frames to remove noise for the current frame with a bilateral filter. A low-complexity hierarchical motion estimation is used in MCTF to obtain motion compensated blocks from temporal adjacent frames. These temporally adjacent motion compensation blocks and the current block are jointly treated as inputs to the temporal filter, to obtain the filtered block in the to-be-coded frame. The filter weights are determined by frame-level parameters including the relative Picture Order Count (POC) position in the GOP, the encoding parameter QP. The default MCTF with bilateral filtering can be described as follows, where I0 is the sample value of current block and Ir is the co-located sample value of the motion compensated block from temporal adjacent frames. b is the number of available neighbouring frames which are used for the filtering and i is the POC distance between current frames and neighbouring frames. In VVC, four previous and four following available neighbouring frames are used for referencing under the random-access configuration. Thus, N is set to -4 and M is set to 4. wr (b, i) is a weighting factor for the co-located sample.
[0101] I. 9 Bi-predictive temporal MVP (Bi-TMVP)
[0102] TMVP candidates based on motion trajectory crossing the current block are introduced as shown in Fig. 8.
[0103] It adds a bi-TMVP candidate (i.e., MV0 and MV1) after HMVP if a motion trajectory between a block (e.g. RefBlk0) in a reference picture (e.g. Reference picture 0) and its reference block (e.g. RefBlk1) in a reference picture (e.g. Reference picture 1) crosses the current block. For a given MV shown in a dashed line, which is a reference block MV, a pair of MV0 and MV1 is constructed and if it crosses the current block then those MVs are used as bi-TMVP.
[0104] I. 10 Local illumination compensation (LIC)
[0105] LIC is an inter prediction technique to model local illumination variation between the current block and its prediction block as a function of that between the current block template and the reference block template. The parameters of the function can be denoted by a scale α and an offset β, which forms a linear equation, that is, α*p [x] +β to compensate illumination changes, where p [x] is a reference sample pointed to by MV at a location x on reference picture. When wrap around motion compensation is enabled, the MV shall be clipped with wrap around offset taken into consideration. Since α and β can be derived based on current block template and reference block template, no signalling overhead is required for them.
[0106] The local illumination compensation proposed in JVET-O0066 is used for inter-coded CUs with the following modifications. ● Intra neighbour samples can be used in LIC parameter derivation. ● LIC is disabled for blocks with less than 32 luma samples. ● Samples of the reference block template are generated by using MC with the block MV without rounding it to integer-pel precision.
[0107] I. 11 Non-local LIC
[0108] Non-local illumination compensation (NLIC) is applied in ECM wherein the linear model is derived from the previously coded inter CUs by minimizing the difference between their reconstruction and prediction samples. When constructing the merge lists, up to 16 and 6 NLIC candidates (obtained from both spatial adjacent and non-adjacent positions) are inserted to the lists of regular merge and subblock merge respectively, and reordered with the existing merge candidates. The lengths of the output merge lists are kept unchanged. The same pattern used for non-adjacent merge mode is reused to locate the non-adjacent positions in the scheme.
[0109] I. 12 Affine Motion Compensated Prediction
[0110] In HEVC, only translational motion model is applied for motion compensation prediction (MCP) . While in the real world, there are many kinds of motion, e.g. zoom in / out, rotation, perspective motions and the other irregular motions. In VVC, a block-based affine transform motion compensation prediction is applied. As shown Figs. 9A-B, the affine motion field of the blocks 910 and 920 is described by motion information of two control point (4-parameter) in Fig. 9A or three control point motion vectors (6-parameter) in Fig. 9B.
[0111] For 4-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:
[0112] For 6-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:
[0113] Where (mv0x, mv0y) is motion vector of the top-left corner control point, (mv1x, mv1y) is motion vector of the top-right corner control point, and (mv2x, mv2y) is motion vector of the bottom-left corner control point.
[0114] In order to simplify the motion compensation prediction, block based affine transform prediction is applied. To derive motion vector of each 4×4 luma subblock, the motion vector of the centre sample of each subblock, as shown in Fig. 10, is calculated according to above equations, and rounded to 1 / 16 fraction accuracy. Then, the motion compensation interpolation filters are applied to generate the prediction of each subblock with the derived motion vector. The subblock size of chroma-components is also set to be 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luma subblocks in the collocated 8x8 luma region.
[0115] In the present invention, signalling schemes of filter information for a smoothing filter used to filter reference pictures to generate filtered predictor in video coding is disclosed. Also, conditions for using the filtered prediction are also disclosed. BRIEF SUMMARY OF THE INVENTION
[0116] A method and apparatus for video coding using filtered reference pictures to improve coding performance are disclosed. According to one method, input data associated with a current block is received, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. At least one set of filter parameters for one or more reference frames is signalled or parsed. At least one filtered reference block is generated by applying a target filter, configured according to said at least one set of filter parameters, to a target reference frame of said one or more reference frames associated with the current block. A filtered predictor for the current block is derived based on said at least one filtered reference block. The current block is encoded or decoded using the filtered predictor.
[0117] In one embodiment, said at least one set of filter parameters is coded in a bitstream using at least one coding scheme selected from a group consisting of: a fixed-length code, an exponential-Golomb code, and a Huffman code.
[0118] In one embodiment, a plurality of sets of filter parameters are signalled or parsed for a specific reference frame, and a target set of filter parameters is selected from the plurality of sets of filter parameters for filtering the specific reference frame. In one embodiment, the target set of filter parameters is selected from the plurality of sets of filter parameters based on a classification of the specific reference frame. In one embodiment, one or more flags are signalled or parsed in a bitstream to indicate a specific classification method used to select the target set of filter parameters.
[0119] In one embodiment, only one single set of filter parameters is signalled or parsed in a bitstream only for a leading reference frame in each reference frame list.
[0120] In one embodiment, a first flag is signalled or parsed in a bitstream to indicate whether the target filter is applied to the target reference frame associated with the current block to generate said at least one filtered reference block. In one embodiment, if the first flag indicates that the target filter is applied to the target reference frame, a second flag is signalled or parsed in the bitstream to indicate the target filter selected for the current block.
[0121] A method of applying the smoother filter only when certain conditions are satisfied is disclosed. According to this method, input data associated with a current block is received, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. Whether one or more conditions associated with the current block are satisfied is determined. Only in response to said one or more conditions being satisfied: at least one filtered reference block is generated by applying a target filter to one or more current reference frames associated with the current block; a filtered predictor for the current block is derived based on said at least one filtered reference block; and the current block is encoded or decoded using the filtered predictor.
[0122] In one embodiment, said one or more conditions comprise a requirement that the current block is coded in one or more specific inter modes.
[0123] In one embodiment, said one or more conditions comprise a requirement that a reference index or a temporal index associated with one current reference frame of the current block is smaller than or equal to a threshold, and wherein the threshold is a non-negative integer.
[0124] In one embodiment, said one or more conditions comprise a requirement that a predictor of the current block belongs to a specific colour component of a video signal.
[0125] In one embodiment, said one or more conditions comprise a requirement that samples with an absolute difference between filtered and un-filtered sample values are smaller than or equal to a threshold, and wherein the threshold is a non-negative integer.
[0126] In one embodiment, generating said at least one filtered reference block by applying the target filter to said one or more current reference frames associated with the current block comprises applying the target filter to one or more predictors of the current block and one or more corresponding reference template regions.BRIEF DESCRIPTION OF THE D RAWINGS
[0127] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.
[0128] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0129] Fig. 2 illustrates the ALF filter shapes for the chroma (left) and luma (right) components.
[0130] Fig. 3A illustrates the placement of CC-ALF with respect to other loop filters.
[0131] Fig. 3B illustrates a diamond shaped filter for the chroma samples.
[0132] Fig. 4 illustrates the cross 9x9 CCALF filter shape, luma residual based taps and 4 taps in a 3x3 asymmetric cross shape as additional inputs to CCALF.
[0133] Fig. 5 illustrates the filter shape of ALF in ECM-7.0.
[0134] Fig. 6 illustrates the diamond regions in the search area for multi-pass decoder-side motion vector refinement.
[0135] Fig. 7 illustrates an example of group of pictures (GOP) structure with 16 pictures in the group.
[0136] Fig. 8 illustrates an example of bi-predictive temporal MVP derivation.
[0137] Fig. 9A illustrates an example of the affine motion field of a block described by motion information of two control point (4-parameter) .
[0138] Fig. 9B illustrates an example of the affine motion field of a block described by motion information of three control point motion vectors (6-parameter) .
[0139] Fig. 10 illustrates an example of block based affine transform prediction, where the motion vector of each 4×4 luma subblock is derived from the control-point MVs.
[0140] Fig. 11 illustrates an example of prediction filtering scheme according to one embodiment of the present invention.
[0141] Fig. 12 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention, where the filter information is signalled.
[0142] Fig. 13 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention, where the filter is applied only when certain conditions are satisfied.DETAILED DESCRIPTION OF THE INVENTION
[0143] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0144] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0145] PROPOSED METHOD
[0146] The following methods are proposed to improve the luma / chroma inter coding prediction accuracy or coding performance. In one embodiment, whether to allow or apply the following proposed methods can depend on SPS / PPS / SH / PH syntax, or CTU or CU / PU / TU level syntax or semantic. In another embodiment, whether to allow or apply the following proposed methods can also depend on implicit conditions. For example, the proposed methods described herein can be applied based on one or more of block width, block height, block area, QP, POC, temporal index, reference index, or prediction mode. In another embodiment, the following proposed prediction filtering methods can be conditionally applied to partial positions / sub-block of the current block. In another embodiment, the term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB. Any combination of the following proposed methods in this invention can be applied.
[0147] II. 1 Prediction Filtering
[0148] In one invention, a prediction filtering scheme is proposed to improve the prediction quality or reduce the residual signalling overhead to improve the coding performance. In the proposed scheme, for a current coding frame, a filter is derived for each reference frame using regression that minimizes differences (e.g., mean squared error (MSE) ) between samples in the current original frame and corresponding samples in the reference frame. For an example, two filters are derived for the current frame if there are two reference frames for the current frame. The predictor of an inter CU can be refined by applying the filter derived from the corresponding reference frame. In Fig. 11, an example of applying prediction filters to an inter block is illustrated. Two filters (filter 0 and filter 1) are applied on the reference blocks (Refblk0 and Refblk1) from two reference pictures (reference picture 0 and reference picture 1) respectively to improve the prediction quality for the current block (CurrBlk) in the current picture before bi-predictive blending.
[0149] II. 1.1 Filter Derivation
[0150] To derive a filter, a regression process is performed to minimize the difference between the samples on the original frame and the samples on the reference frames as shown below, minf∑p (orgp- (refp+∑c (ncfc) ) ) 2 where org indicates samples on original frame or original frame after motion compensation, ref indicates samples of reference frames after motion compensation or samples of reference frames, nc indicates neighbouring information (i.e., neighbouring reference samples adjacent to refp) , fc indicates filter coefficients, and c indicates index of filter coefficients.
[0151] To simplify the description, a regression pair is defined as (orgp, refp) for the following embodiments.
[0152] In one embodiment, a regression pair is formed by an original sample (i.e., orgp) of the original frame and a reference sample (i.e., refp) of one reference frame after motion compensation. In another embodiment, a regression pair is formed by an original sample (i.e., orgp) of the original frame after motion compensation and a reference sample (i.e., refp) of one reference frame. In another embodiment, a regression pair is formed by an original sample (i.e., orgp) of the original frame and / or a sample of the original frame after motion compensation and a reference sample (i.e., refp) of one reference frame after motion compensation and / or a reference sample of one reference frame. The motion compensation process is performed to align the structure of the original and reference frames. To determine the MVs for the motion compensation process, any inter prediction encoding algorithm (or any combination of inter prediction encoding algorithms) , any motion estimation algorithms, TZ (Test Zone) search, diamond search, MCTF ME algorithm (hieratical search algorithm) or Bi-TMVP (use motion information of previously coded reference pictures) can be performed on the original frame to align with reference frames or on reference frames to align with original frame.
[0153] For example, motion information of previously coded reference pictures can be used as initial motions when performing the motion estimation.
[0154] For example, MVs determined in previously blocks can be used as initial motions when performing the motion estimation.
[0155] In one embodiment, a two-pass encoding scheme is proposed to derive the prediction filters for each coding frame. For a coding frame (except for the intra frame) , the inter prediction encoding algorithms (or any combination of inter prediction encoding algorithms) or motion estimation algorithms are performed in the first encoding pass to determine the MVs of each coding block. After performing the motion compensation of each block according to the MVs, multiple regression pairs can be collected for each reference frame and used to derive prediction filters for the corresponding reference frames.
[0156] II. 1.2 Signalling
[0157] In one embodiment, a filter is signalled / parsed for each reference frame for prediction filtering.
[0158] In one embodiment, a filter is signalled / parsed only for the first reference frame in each reference frame list (i.e., L0 and L1) for prediction filtering.
[0159] In one embodiment, a flag is signalled / parsed for a reference frame to indicate whether prediction filtering is used for the reference frame or not. If the prediction filter is used, a filter is further signalled / parsed.
[0160] In one embodiment, filter coefficients are signalled / parsed with fixed-length codes, exponential-Golomb codes, or Huffman codes.
[0161] In one embodiment, multiple filters are signalled / parsed for a reference frame, and classification is performed before applying prediction filters to select one filter among the multiple filters. A set of flags are signalled / parsed to indicate the classification method.
[0162] In one embodiment, multiple filters are signalled / parsed for a reference frame, and classification is performed before applying prediction filters to select one filter among the multiple filters. A first set of flags are signalled / parsed to indicate the number of filters, and a second set of flags are signalled / parsed to indicate how to select a filter based on the classification result.
[0163] II. 1.3 Condition to apply prediction filtering
[0164] In one embodiment, the filters are only applied to the coding blocks with LIC (including non-local LIC) disabled. In another embodiment, the filters are only applied to the coding blocks with LIC (including non-local LIC) enabled.
[0165] In another embodiment, the filters are only applied to the coding blocks with MP-DMVR disabled. In another embodiment, the filters are only applied to the coding blocks with BDOF disabled. In another embodiment, the filters are only applied to the predictors before or after a certain pass of MP-DMVR. In another embodiment, the filters are only applied to the predictors before or after sample-based BDOF. In another embodiment, the filters are applied to the predictors before or after several passes of MP-DMVR and / or sample-based BDOF.
[0166] In another embodiment, the filters are only applied to the coding blocks coded by certain inter modes. In one example, the filters are only applied to uni-predictive inter blocks. In one example, the filters are only applied to bi-predictive inter blocks. In one example, the filters are only applied to non-affine inter blocks. In one example, the filters are only applied to affine inter blocks.
[0167] In another embodiment, the filters are only applied to the predictors before or after the OBMC refinement.
[0168] In another embodiment, the filters are only applied to the predictors of coding blocks. In another embodiment, the filters are applied to the predictors of coding blocks and the corresponding reference template regions (i.e., samples adjacent to collocated blocks of reference frames) .
[0169] In another embodiment, the filters are only applied to the reference frames with a reference index equal to and / or smaller than T1 (T1 >= 0) . In another embodiment, the filters are only applied to the reference frames with a reference index equal to and / or larger than T2 (T2 >= 0) .
[0170] In another embodiment, the filters are only applied to the reference frames of the coding frames with temporal index equal to and / or smaller than T3 (T3 >= 0) . In another embodiment, the filters are only applied to the reference frames of the coding frames with temporal index equal to and / or larger than T4 (T4 >= 0) .
[0171] In another embodiment, the filters are only applied to certain colour components of predictors. In one example, the filters are only applied to the first colour component.
[0172] In another embodiment, the filters are only applied to the samples with the absolute difference between filtered and un-filtered sample values smaller than or equal to a threshold G1, where G1 can be any value. In another embodiment, the filters are only applied to the samples with the absolute difference between filtered and un-filtered sample values larger than or equal to a threshold G2, where G2 can be any value.
[0173] In one embodiment, the filters are only applied to the picture which LMCS is disabled.
[0174] In one embodiment, the filters are only applied to the pictures which include N percentage of samples in the filter derivation process. N can be any integer larger than 0.
[0175] In one embodiment, the filters are only applied to the samples in reference pictures including in the filter derivation process.
[0176] II. 2 Determine Prediction Filters from ALF Fixed Filters
[0177] In one embodiment, fixed filters used in ALF are also used for prediction filtering.
[0178] In one embodiment, when applying prediction filtering with ALF fixed filters, only partial coefficients of ALF fixed filters are used.
[0179] In one embodiment, a set of flags are signalled / parsed to indicate whether ALF fixed filters are used for prediction filtering or not.
[0180] In one embodiment, different ALF fixed filters are used to perform prediction filtering on different reference frames. A set of flags are signalled / parsed to indicate which ALF fixed filters are used for a reference frame.
[0181] II. 3 Combine Interpolation Filtering in MC with Prediction Filtering
[0182] In one embodiment, the prediction filters can be combined with interpolation filters and used to replace the current interpolation filters, or used as an additional interpolation filtering option. In one example, to replace a set of 12-tap interpolation filters, 64 prediction filters are derived and can be selected according to MV of the current coding block.
[0183] In one embodiment, a flag is signalled in a sequence parameter set (SPS) , a picture parameter set (PPS) , a slice header (SH) , and a picture header (PH) or CTU or CU / PU / TU level to indicate whether a block is filtered by the original interpolation filter or the prediction filter.
[0184] Any of the foregoing proposed methods of filtered prediction generated by applying smoothing filters to reference frames can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module.
[0185] With reference to the exemplary encoder and decoder in Fig. 1A and Fig 1B, the proposed methods can be implemented in the intra / inter prediction modules. For example, in the encoder side, the required processing can be implemented as part of the Inter-Pred. unit 112 or Intra Pred. unit 110 as shown in Fig. 1A. However, the encoder may also use additional processing unit to implement the required processing. For the decoder side, the required processing can be implemented as part of the MC unit 152 or Intra Pred. 150 as shown in Fig. 1B. However, the decoder may also use additional processing unit to implement the required processing. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module. While the Inter-Pred. 112 and Intra Pred. 110 in the encoder side and MC 152 and Intra Pred. 150 in the decoder side are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array)) .
[0186] Fig. 12 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention, where the filter information is signalled. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g. one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data associated with a current block in a current frame is received in step 1210, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. At least one set of filter parameters for one or more reference frames is signalled or parsed in step 1220. At least one filtered reference block is generated by applying a target filter in step 1230, configured according to said at least one set of filter parameters, to a target reference frame of said one or more reference frames associated with the current block. A filtered predictor for the current block is derived based on said at least one filtered reference block in step 1240. The current block is encoded or decoded using the filtered predictor in step 1250.
[0187] Fig. 13 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention, where the filter is applied only when certain conditions are satisfied. According to this method, input data associated with a current block is received in step 1310, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. Whether one or more conditions associated with the current block are satisfied is determined in step 1320. Only when said one or more conditions are satisfied (i.e., the “Yes” pass from step 1320) , steps 1330 to 1350 are performed. Otherwise (i.e., the “No” pass from step 1320) , steps 1330 to 1350 are skipped. In step 1330, at least one filtered reference block is generated by applying a target filter to one or more current reference frames associated with the current block; in step 1340, a filtered predictor for the current block is derived based on said at least one filtered reference block; and in step 1350, the current block is encoded or decoded using the filtered predictor.
[0188] The flowcharts shown are intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0189] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0190] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0191] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;signalling or parsing at least one set of filter parameters for one or more reference frames;generating at least one filtered reference block by applying a target filter, configured according to said at least one set of filter parameters, to a target reference frame of said one or more reference frames associated with the current block;deriving a filtered predictor for the current block based on said at least one filtered reference block; andencoding or decoding the current block using the filtered predictor.2.The method of Claim 1, wherein said at least one set of filter parameters is coded in a bitstream using at least one coding scheme selected from a group consisting of: a fixed-length code, an exponential-Golomb code, and a Huffman code.3.The method of Claim 1, wherein a plurality of sets of filter parameters are signalled or parsed for a specific reference frame, and a target set of filter parameters is selected from the plurality of sets of filter parameters for filtering the specific reference frame.4.The method of Claim 3, wherein the target set of filter parameters is selected from the plurality of sets of filter parameters based on a classification of the specific reference frame.5.The method of Claim 3, wherein one or more flags are signalled or parsed in a bitstream to indicate a specific classification method used to select the target set of filter parameters.6.The method of Claim 1, wherein only one single set of filter parameters is signalled or parsed in a bitstream only for a leading reference frame in each reference frame list.7.The method of Claim 1, wherein a first flag is signalled or parsed in a bitstream to indicate whether the target filter is applied to the target reference frame associated with the current block to generate said at least one filtered reference block.8.The method of Claim 7, wherein if the first flag indicates that the target filter is applied to the target reference frame, a second flag is signalled or parsed in the bitstream to indicate the target filter selected for the current block.9.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;signal or parse at least one set of filter parameters for one or more reference frames;generate at least one filtered reference block by applying a target filter, configured according to said at least one set of filter parameters, to a target reference frame of said one or more reference frames associated with the current block;derive a filtered predictor for the current block based on said at least one filtered reference block; andencode or decode the current block using the filtered predictor.10.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;determining whether one or more conditions associated with the current block are satisfied; andonly in response to said one or more conditions being satisfied:generating at least one filtered reference block by applying a target filter to one or more current reference frames associated with the current block;deriving a filtered predictor for the current block based on said at least one filtered reference block; andencoding or decoding the current block using the filtered predictor.11.The method of Claim 10, wherein said one or more conditions comprise a requirement that the current block is coded in one or more specific inter modes.12.The method of Claim 10, wherein said one or more conditions comprise a requirement that a reference index or a temporal index associated with one current reference frame of the current block is smaller than or equal to a threshold, and wherein the threshold is a non-negative integer.13.The method of Claim 10, wherein said one or more conditions comprise a requirement that a predictor of the current block belongs to a specific colour component of a video signal.14.The method of Claim 10, wherein said one or more conditions comprise a requirement that samples with an absolute difference between filtered and un-filtered sample values are smaller than or equal to a threshold, and wherein the threshold is a non-negative integer.15.The method of Claim 10, wherein generating said at least one filtered reference block by applying the target filter to said one or more current reference frames associated with the current block comprises applying the target filter to one or more predictors of the current block and one or more corresponding reference template regions.16.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;determine whether one or more conditions associated with the current block are satisfied; andonly in response to said one or more conditions being satisfied:generate at least one filtered reference block by applying a target filter to one or more current reference frames associated with the current block;derive a filtered predictor for the current block based on said at least one filtered reference block; andencode or decode the current block using the filtered predictor.