Method and apparatus of prediction filtering in video coding

WO2026200522A1PCT designated stage Publication Date: 2026-10-01MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/082497
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-10
Publication Date
2026-10-01

Smart Images

  • Figure CN2026082497_01102026_PF_FP_ABST
    Figure CN2026082497_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for video coding using filtered reference pictures to improve coding performance are disclosed. According to the method, input data associated with a current block in a current frame is received, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. One or more filters to be applied to one or more reference frames associated with the current block are determined. At least one filtered reference block is generated by applying said one or more filters to at least one of said one or more reference frames. A filtered predictor for the current block is derived based on said at least one filtered reference block. The current block is encoded or decoded using the filtered predictor.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS OF PREDICTION FILTERING IN VIDEO CODINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 777,040, filed on March 25, 2025. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to video coding. In particular, the present invention discloses a new scheme to improve the coding performance by introducing a refined prediction generated by applying a smoothing filter to reference pictures in video coding. BACKGROUND AND RELATED ART

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, are provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.

[0008] I. 1 Adaptive Loop Filter in VVC

[0009] In VVC, an Adaptive Loop Filter (ALF) with block-based filter adaption is applied. For the luma component, one filter is selected among 25 filters for each 4×4 block, based on the direction and activity of local gradients.

[0010] I. 1.1. Filter shape

[0011] Two diamond filter shapes (as shown in Fig. 2) are used. The 7×7 diamond shape 220 is applied for luma component and the 5×5 diamond shape 210 is applied for the chroma components.

[0012] I. 1.2. Block classification and geometric transformation

[0013] For luma component, each 4×4 block is categorized into one out of 25 classes. Before filtering each 4×4 luma block, geometric transformations such as rotation or diagonal and vertical flipping are applied to the filter coefficients f (k, l) and to the corresponding filter clipping values c (k, l) depending on gradient values calculated for that block. This is equivalent to applying these transformations to the samples in the filter support region. The idea is to make different blocks to which ALF is applied more similar by aligning their directionality.

[0014] For chroma components in a picture, no classification method is applied.

[0015] More details could be found in VVC specification section 8.8.5.3.

[0016] I. 1.3. Filtering process

[0017] At decoder side, when ALF is enabled for a CTB, each sample R (i, j) within the CU is filtered, resulting in sample value R′ (i, j) as shown below, where f (k, l) denotes the decoded filter coefficients, K (x, y) is the clipping function and c (k, l) denotes the decoded clipping parameters. The variable k and l varies between –L / 2 and L / 2, where L denotes the filter length. The clipping function K (x, y) =min (y, max (-y, x) ) which corresponds to the function Clip3 (-y, y, x) . The clipping operation introduces non-linearity to make ALF more efficient by reducing the impact of neighbour sample values that are too different with the current sample value.

[0018] I. 1.4. Cross component adaptive loop filter

[0019] CC-ALF uses luma sample values to refine each chroma component by applying an adaptive, linear filter to the luma channel and then using the output of this filtering operation for chroma refinement. Fig. 3A provides a system level diagram of the CC-ALF process with respect to the SAO, luma ALF and chroma ALF processes. As shown in Fig. 3A, each colour component (i.e., Y, Cb and Cr) is processed by its respective SAO (i.e., SAO Luma 310, SAO Cb 312 and SAO Cr 314) . After SAO, ALF Luma 320 is applied to the SAO-processed luma and ALF Chroma 330 is applied to SAO-processed Cb and Cr. However, there is a cross-component term from luma to a chroma component (i.e., CC-ALF Cb 322 and CC-ALF Cr 324) . The outputs from the cross-component ALF are added (using adders 332 and 334 respectively) to the outputs from ALF Chroma 330.

[0020] Filtering in CC-ALF is accomplished by applying a linear, diamond shaped filter (e.g. filters 340 and 342 in Fig. 3B) to the luma channel. In Fig. 3B, a blank circle indicates a luma sample and a dot-filled circle indicate a chroma sample. One filter is used for each chroma channel, and the operation is expressed as: where (x, y) is chroma component i location being refined, (xY, yY) is the luma location based on (x, y) , Si is filter support area in luma component, and ci (x0, y0) represents the filter coefficients.

[0021] As shown in Fig, 3B, the luma filter support is the region collocated with the current chroma sample after accounting for the spatial scaling factor between the luma and chroma planes. In Fig. 3B, circles represent luma samples, while dotted circles represent chroma samples being refined.

[0022] I. 1.5. Filter parameters signalling

[0023] ALF filter parameters are signalled in Adaptation Parameter Set (APS) . In one APS, up to 25 sets of luma filter coefficients and clipping value indexes, and up to eight sets of chroma filter coefficients and clipping value indexes can be signalled. To reduce bits overhead, filter coefficients of different classification for luma component can be merged. In slice header, the indices of the APSs used for the current slice are signalled.

[0024] Clipping value indexes, which are decoded from the APS, allow determining clipping values using a table of clipping values for both the luma and chroma components. These clipping values are dependent of the internal bit-depth. More precisely, the clipping values are obtained by the following formula: AlfClip= {round (2B-α*n) for n∈ [0.. N-1] } with B equal to the internal bit-depth, α is a pre-defined constant value equal to 2.35, and N equal to 4 which is the number of allowed clipping values in VVC. The AlfClip is then rounded to the nearest value with the format of power of 2.

[0025] In slice header, up to 7 APS indices can be signalled to specify the luma filter sets that are used for the current slice. The filtering process can be further controlled at CTB level. A flag is always signalled to indicate whether ALF is applied to a luma CTB. A luma CTB can choose a filter set among 16 fixed filter sets and the filter sets from APSs. A filter set index is signalled for a luma CTB to indicate which filter set is applied. The 16 fixed filter sets are pre-defined and hard-coded in both the encoder and the decoder.

[0026] For the chroma component, an APS index is signalled in slice header to indicate the chroma filter sets being used for the current slice. At CTB level, a filter index is signalled for each chroma CTB if there is more than one chroma filter set in the APS.

[0027] The filter coefficients are quantized with norm equal to 128. In order to restrict the multiplication complexity, a bitstream conformance is applied so that the coefficient value of the non-central position shall be in the range of -27 to 27 -1, inclusive. The central position coefficient is not signalled in the bitstream and is considered as equal to 128.

[0028] I. 2 Adaptive Loop Filter in ECM

[0029] In ECM8 (Muhammed Coban, et al., “Algorithm description of Enhanced Compression Model 8 (ECM 8) ” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29) , 29th Meeting, by teleconference, 11–20 January 2023, Document: JVET-AC2025) , some changes from the VVC ALF are disclosed. A brief overview is shown below.

[0030] I. 2.1. ALF simplification removal

[0031] ALF gradient subsampling and ALF virtual boundary processing are removed. Block size for classification is reduced from 4x4 to 2x2. Filter size for both the luma and chroma, for which ALF coefficients are signalled, is increased.

[0032] I. 2.2. ALF with fixed filters

[0033] To filter a luma sample, three different classifiers (C0, C1 and C2) and three different sets of filters (F0, F1 and F2) are used. Sets F0 and F1 contain fixed filters, with coefficients trained for classifiers C0 and C1. Coefficients of filters in F2 are signalled. Which filter from a set Fi is used for a given sample is decided by a class Ci assigned to this sample using classifier Ci.

[0034] The number of bits used to represent the fractional part of a luma coefficient is adaptive from 5 to 8, inclusively. For each luma filter set, which contains up to 25 filters, a 2-bit syntax element is signalled in APS to indicate the number of bits used for the coefficients in this set. The value range of a coefficient is not changed.

[0035] I. 2.3. Filtering

[0036] First, two 13x13 diamond shape fixed filters F0 and F1 are applied to derive two intermediate samples R0 (x, y) and R1 (x, y) . After that, F2 is applied to R0 (x, y) , R1 (x, y) , and neighbouring samples to derive a filtered sample as where fi, j is the clipped difference between a neighbouring sample and current sample R (x, y) and gi is the clipped difference between Ri-20 (x, y) and current sample. The filter coefficients ci, i=0, …21, are signalled.

[0037] I. 2.4. Classification

[0038] Based on directionality Di and activity aclass Ci is assigned to each 2x2 block: where MD, i represents the total number of directionalities Di.

[0039] As in VVC, values of the horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within a 4×4 window that covers the target 2×2 block is used for classifier C0 and the sum of sample gradients within a 12×12 window is used for classifiers C1 and C2. The sums of horizontal, vertical and two diagonal gradients are denoted, respectively, as and The directionality Di is determined by comparing with a set of thresholds. The directionality D2 is derived as in VVC using thresholds 2 and 4.5. For D0 and D1, horizontal / vertical edge strength and diagonal edge strength are calculated first. Thresholds Th= [1.25, 1.5, 2, 3, 4.5, 8] are used. Edge strength is 0 if otherwise,  is the maximum integer such that Edge strength is 0 if otherwise,  is the maximum integer such that When i.e., horizontal / vertical edges are dominant, the Di is derived by using Table 1A; otherwise, diagonal edges are dominant, the Di is derived by using Table 1B. Table 1A. Mapping of and to Di Table 1B. Mapping of   and   to Di

[0040] To obtain the sum of vertical and horizontal gradients Ai is mapped to the range of 0 to n, where n is equal to 4 for and 15 for and

[0041] In an ALF_APS, up to 4 luma filter sets are signalled, each set may have up to 24 filters.

[0042] I. 2.5. Alternative 2x2 ALF classifier

[0043] Classification in ALF is extended with an additional alternative classifier. For a signalled luma filter set, a flag is signalled to indicate whether the alternative classifier is applied. Geometrical transformation is not applied to the alternative band classifier. When the band-based classifier is applied, the sum of sample values of a 2x2 luma block is calculated at first. Then the class index is calculated as below, class_index = (sum *12) >> (sample bit depth + 2) .

[0044] I. 2.6. Residual based classifier

[0045] A third classifier is based on luma residual sample values. For each 2x2 luma block, the sum of absolute values of the residual samples in a neighbouring 8x8 window is calculated, and the class index is derived as: classIdx = sum >> (sample bit depth –3) .

[0046] The value of classIdx is in the range of 0 to 24, same as in ECM-8.0. The classifier usage is signalled for each luma filter set in APS.

[0047] I. 2.7. Coding information-based classifier

[0048] For the online filters (signalled filters) , each 2×2 unit is classified into 2 noise levels according to the partitioning information. For a 2×2 unit, it is classified into noise level-1 if it is located at a CU or TU boundary; and into noise level-0 if it is not. The class number of existing classifiers is reduced from 25 to 12. For the texture-based classifier, the 25 classes are mapped to 12 classes with a pre-defined LUT. For the band-based and residual based classifiers, the 25 classes are decreased to 12 classes by enlarging the band width. These 12 classes are further combined with the proposed 2 noise levels to generate final 12×2 = 24 classes in total. The number of classifiers for online filters is kept as 3 and no additional encoder selection is introduced.

[0049] For the offline filters (fixed filters) , each 2×2 unit is classified into 2 noise levels in the same way as the classification of online filters. Besides, each 2×2 unit is further classified into 2 residual levels based on a predefined threshold also in the same way as the classifier of online filters. Furthermore, the generated offset of offline filters is adjusted based on the boundary level and residual level accordingly, where a stronger offset is applied on the positions at boundaries or with higher residuals.

[0050] I. 2.8. CCALF with long tap filter

[0051] The CCALF process uses a linear filter to filter luma sample values, luma residual samples and generate a residual correction for the chroma samples. In addition, the CCALF filter shape is constructed by 23 luma spatial taps (410) and 5 luma residual taps (420) , which is illustrated in Fig. 4. For a given slice, the encoder can collect the statistics of the slice, analyse them and signal up to 16 filters through APS. The number of bits used to represent the fractional part of a CCALF coefficient can vary from 7 to 10 adaptively.

[0052] Chroma SAO output samples applied to 4 taps (430) in a 3x3 asymmetric cross shape are added as additional inputs to CCALF, as illustrated in Fig. 4.

[0053] I. 2.9. New luma ALF filter shape

[0054] In ALF online-trained filters consist of 4 kinds of filter taps: spatial taps (510) , reconstruction-before-DBF based taps (540) , residual based taps (550) and fixed-filter-output based taps (520 and 530) as shown in Fig. 5, where the fixed-filter-output based taps are extended, the residual based taps contain the clipped residual sample and the clipped residual sample filtered by the fixed filters, and the reconstruction-before-DBF (pre-DBF) based taps contain the clipped pre-DBF samples and the pre-DBF samples filtered by a Gaussian fixed filter. The shape of Gaussian fixed filter is a diamond 7x7 shape, and the filter parameters are stored at both encoder and decoder. There is no classification for this fixed filter.

[0055] I. 2.10. ALF residuals scaling

[0056] A scaling factor is signalled in slice header, the scaling factors is applied to the difference between the ALF input and ALF output, and the scaled residual is added to the ALF input (it produces a scaled ALF filtering) . A similar scaling process is applied to NN filtering in NNVC.

[0057] Different luma scaling factors may be associated with different group of class indexes, and the ALF output is derived as follows: rec’ (s) = rec (s) + (corr (s) *tab [sfi [class (s) ] ] + 4 ) >> 3, where ALF residual correction ‘corr (s) ’ is scaled using the scaling factor associated to the class index of the sample, ‘tab’ is the predefined LUT for mapping scaling index ‘sfi’ to scaling factor.

[0058] I. 2.11. Improved fixed filters for ALF

[0059] Two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to C′ by a scaling factor. The value of C′is an integer between 0 and 7, inclusively. With i=0, 1, let Ci denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index Ci′is derived as Ci′= C′*896+Ci.

[0060] The total number of the fixed filters is not changed.

[0061] Then a class index is determined based on the activity and directionality values. Two diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in Table 2. Table 2. Comparison of fixed filters between ECM-9.0 and JVET-AE0139

[0062] Fixed filter f1 is applied to outputs of f0 (instead of ALF input) and samples before DBF.

[0063] Finally, a signalled filter is applied to the ALF input samples, samples before the deblocking filter (DBF) , outputs of the two fixed filters, output of a Gaussian filter and the residual data.

[0064] I. 2.12. Chroma ALF Fixed Filter

[0065] A classifier based on Laplacian values and variance is applied to a 2x2 chroma block. Compared to the luma classifier of a fixed filter, when calculating the activity value, the sum of the chroma vertical and horizontal Laplacian values is multiplied by 2 before scaling. Similarly, the chroma variance is multiplied by 2 before scaling. The derived class index is then used to select a fixed filter from a chroma filter set. A chroma fixed filter is applied to chroma ALF input samples in a 13x13 diamond shape and DBF input samples in a 7x7 diamond shape. The first luma classifier is applied to each 2x2 chroma block. The derived class index is then used to select a fixed filter from the luma fixed filter set related to this classifier. A fixed filter is applied to chroma ALF input sample in a 9x9 diamond shape and DBF input samples in a 9x9 diamond shape. In a signalled chroma filter, 5x5 crossing extra taps are introduced, which are applied to the fixed filter output.

[0066] I. 3. OBMC

[0067] When OBMC is applied, top and left boundary pixels of a CU are refined using the motion information of neighbouring blocks with a weighted prediction as described in JVET-L0101.

[0068] Conditions of not applying OBMC are as follows: ● When OBMC is disabled at the SPS level ● When the current block is coded in intra mode or IBC mode ● When the current luma block area is smaller or equal to 32

[0069] Additionally, OBMC is adaptively controlled at a block level as follows: ● OBMC flag is inherited from a neighbouring affine block for affine merge mode. ● OBMC is not applied to a block if there is a neighbour block coded with IBC, palette, or BDPCM mode. ● When applying OBMC to a block, block boundary, check whether OBMC is applied to the boundary is further made based on the reference samples of the current block. If any absolute difference between the prediction sample and non-interpolated (integer pel) reference sample is greater than a threshold, the OBMC is not applied to that boundary.

[0070] A subblock-boundary OBMC is performed by applying the same blending to the top, left, bottom, and right subblock boundary pixels using motion information of neighbouring subblocks. It is enabled for the subblock based coding tools: ● Affine AMVP modes; ● Affine merge modes and subblock-based temporal motion vector prediction (SbTMVP) ; ● Subblock-based bilateral matching.

[0071] When OBMC mode is used in CIIP mode with LMCS, inter blending is performed prior to LMCS mapping of inter samples. LMCS is applied to blended inter samples which are combined with LMCS applied intra samples in CIIP mode, where InterpredY represents the samples predicted by the motion of current block in the original domain, IntrapredY represents the samples predicted in the mapped domain, OBMCpredY represents the samples predicted by the motion of neighboring blocks in the original domain, and w0 and w1 are the weights.

[0072] When OBMC mode is used in an LIC coded block, the LIC parameters are applied to generate the corresponding prediction samples for the OBMC of the LIC coded block. Besides, to reduce the complexity, the OBMC is only applied to the top and left CU boundaries while being always disabled for the boundaries of the internal sub-blocks of the LIC coded block.

[0073] I. 4. Multi-Pass Decoder-Side Motion Vector Refinement (MP-DMVR)

[0074] A multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coding block. In the second pass, BM is applied to each 16x16 subblock within the coding block. In the third pass, MV in each 8x8 subblock is refined by applying bi-directional optical flow (BDOF) . The refined MVs are stored for both spatial and temporal motion vector prediction.

[0075] I. 4.1 First pass - Block based bilateral matching MV refinement

[0076] In the first pass, a refined MV is derived by applying BM to a coding block. Similar to decoder-side motion vector refinement (DMVR) , in the bi-prediction operation, a refined MV is searched around the two initial MVs (i.e., MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (i.e., MV0_pass1 and MV1_pass1) are derived around the initiate MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.

[0077] BM performs local search to derive integer sample precision intDeltaMV. The local search applies a 3×3 square search pattern to loop through the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0078] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW *cbH is greater than 64, MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. When the bilCost at the centre point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new centre point of the 3×3 search pattern and continue to search for the minimum cost, until it reaches the end of the search range.

[0079] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MVs after the first pass are then derived as: MV0_pass1 = MV0 + deltaMV MV1_pass1 = MV1 –deltaMV

[0080] I. 4.2 Second pass - Subblock based bilateral matching MV refinement

[0081] In the second pass, a refined MV is derived by applying BM to a 16×16 grid subblock. For each subblock, a refined MV is searched around the two MVs (e.g. MV0_pass1 and MV1_pass1) , obtained during the first pass, in the reference picture list L0 and L1. The refined MVs (i.e., MV0_pass2 (sbIdx2) and MV1_pass2 (sbIdx2) ) are derived based on the minimum bilateral matching cost between the two reference subblocks in L0 and L1.

[0082] For each subblock, BM performs full search to derive integer sample precision intDeltaMV. The full search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, wherein, the values of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0083] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference subblocks, as: bilCost = satdCost *costFactor. The search area (2*sHor + 1) * (2*sVer + 1) is divided up to 5 diamond shape search regions shown on Fig. 6, where the 5 search regions are shown in 5 different shades. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in the order starting from the centre of the search area. In each region, the search points are processed in the raster scan order starting from the top left going to the bottom right corner of the region. When the minimum bilCost within the current search region is less than a threshold equal to sbW *sbH, the int-pel full search is terminated; otherwise, the int-pel full search continues to the next search region until all search points are examined. Additionally, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold that is equal to the area of the block, the search process terminates.

[0084] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2) . The refined MVs at second pass is then derived as: ● MV0_pass2 (sbIdx2) = MV0_pass1 + deltaMV (sbIdx2) ● MV1_pass2 (sbIdx2) = MV1_pass1 –deltaMV (sbIdx2)

[0085] I. 4.3 Third pass - Subblock based bi-directional optical flow MV refinement

[0086] In the third pass, a refined MV is derived by applying BDOF to an 8×8 grid subblock. For each 8×8 subblock, BDOF refinement is applied to derive scaled Vx and Vy without clipping starting from the refined MV of the parent subblock of the second pass. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.

[0087] The refined MVs (e.g. MV0_pass3 (sbIdx3) and MV1_pass3 (sbIdx3) ) at third pass are derived as: ● MV0_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) + bioMv ● MV1_pass3 (sbIdx3) = MV0_pass2 (sbIdx2) –bioMv

[0088] I. 4.4 Fourth pass - Adaptive subblock based bi-directional optical flow MV refinement

[0089] In the fourth pass, a refined MV is derived by applying BDOF to a 4×4 or 8×8 or 16x16 grid subblock. When a block is smaller than 1024 pixels, the 4×4 grid subblock is used. Otherwise, 8×8 grid subblock is used. The MV of each subblock is refined in the same way as that used in third pass.

[0090] In all aforementioned sub-clauses, when wrap around motion compensation is enabled, the motion vectors shall be clipped with wrap around offset taken into consideration. It is noted that in ECM, the DMVR is extended to non-equal POC distance cases, and the mean removed equations are utilized to derive the BDOF MV refinement parameters as: (∑Gx·Gx+R1) *vx + ∑Gx·Gy *vy = ∑dI ·Gx → (∑Gx·Gx+R1) *vx + ∑Gx·Gy *vy = ∑dI ·Gx -dM ·∑Gx ∑Gx·Gy *vx + (∑Gy·Gy+R1) *vy= ∑dI ·Gy →∑Gx·Gy *vx + (∑Gy·Gy+R1) *vy = ∑dI ·Gy -dM ·∑Gy

[0091] I. 5 Affine subblock BDOF refinement

[0092] BDOF subblock MV refinement and sample adjustment is applied to an affine or SbTMVP coded block with subblock MC when BDOF condition is satisfied.

[0093] An affine coded block, e.g. affine regular merge mode, affine BM merge mode, affine AMVP mode, derives MVs for each 4×4 subblock from the affine model. The BDOF process starts with the 4×4 subblocks grouping with identical MVs. The first iteration of BDOF MV refinement is processed in 8x8 subblock grid as in ECM-10.0. When the grouped subblock size is less than 256, the second iteration of BDOF MV refinement is processed in 4×4 subblock grid, and otherwise in 8×8 subblock grid. When the grouped subblock size is 4xN or Nx4, the first iteration of BDOF MV refinement is bypassed.

[0094] I. 6 Sample-based BDOF

[0095] In the sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, it is performed per sample. In addition, the high accuracy sample adjustment method, is applied to derive the sample adjustment.

[0096] The coding block is divided into 8×8 subblocks. For each subblock, whether to apply BDOF or not is determined by checking the SAD between the two reference subblocks against a threshold. If decided to apply BDOF to a subblock, for every sample in the subblock, a sliding 5×5 window is used and the existing BDOF process is applied for every sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predicted sample value for the centre sample of the window.

[0097] I. 7 Group of pictures (GOP) structure

[0098] A closed GOP prediction structure can be realized in VVC streams as shown in Fig. 7. In Fig. 7, a GOP structure with size 16 is illustrated. In a GOP structure, the frames are coded in a hierarchical layered structure with a re-ordered frame coding order. In Figure6, the term “Temp. id” indicates temporal index and the arrows indicate the frame reference directions (e.g., the intra frame is referenced by the frame with temporal index equal to 0) .

[0099] I. 8 Motion Compensated Temporal Filter (MCTF)

[0100] Motion Compensated Temporal Filter has been proposed as an effective pre-processing tool. More specifically, before the current input frame is sent to the encoder, MCTF uses temporal adjacent frames to remove noise for the current frame with a bilateral filter. A low-complexity hierarchical motion estimation is used in MCTF to obtain motion compensated blocks from temporal adjacent frames. These temporally adjacent motion compensation blocks and the current block are jointly treated as inputs to the temporal filter, to obtain the filtered block in the to-be-coded frame. The filter weights are determined by frame-level parameters including the relative Picture Order Count (POC) position in the GOP, the encoding parameter QP. The default MCTF with bilateral filtering can be described as follows, where I0 is the sample value of current block and Ir is the co-located sample value of the motion compensated block from temporal adjacent frames. b is the number of available neighbouring frames which are used for the filtering and i is the POC distance between current frames and neighbouring frames. In VVC, four previous and four following available neighbouring frames are used for referencing under the random-access configuration. Thus, N is set to -4 and M is set to 4. wr (b, i) is a weighting factor for the co-located sample.

[0101] I. 9 Bi-predictive temporal MVP (Bi-TMVP)

[0102] TMVP candidates based on motion trajectory crossing the current block are introduced as shown in Fig. 8.

[0103] It adds a bi-TMVP candidate (i.e., MV0 and MV1) after HMVP if a motion trajectory between a block (e.g. RefBlk0) in a reference picture (e.g. Reference picture 0) and its reference block (e.g. RefBlk1) in a reference picture (e.g. Reference picture 1) crosses the current block. For a given MV shown in a dashed line, which is a reference block MV, a pair of MV0 and MV1 is constructed and if it crosses the current block then those MVs are used as bi-TMVP.

[0104] I. 10 Local illumination compensation (LIC)

[0105] LIC is an inter prediction technique to model local illumination variation between the current block and its prediction block as a function of that between the current block template and the reference block template. The parameters of the function can be denoted by a scale α and an offset β, which forms a linear equation, that is, α*p [x] +β to compensate illumination changes, where p [x] is a reference sample pointed to by MV at a location x on reference picture. When wrap around motion compensation is enabled, the MV shall be clipped with wrap around offset taken into consideration. Since α and β can be derived based on current block template and reference block template, no signalling overhead is required for them.

[0106] The local illumination compensation proposed in JVET-O0066 is used for inter-coded CUs with the following modifications. ● Intra neighbour samples can be used in LIC parameter derivation. ● LIC is disabled for blocks with less than 32 luma samples. ● Samples of the reference block template are generated by using MC with the block MV without rounding it to integer-pel precision.

[0107] I. 11 Non-local LIC

[0108] Non-local illumination compensation (NLIC) is applied in ECM wherein the linear model is derived from the previously coded inter CUs by minimizing the difference between their reconstruction and prediction samples. When constructing the merge lists, up to 16 and 6 NLIC candidates (obtained from both spatial adjacent and non-adjacent positions) are inserted to the lists of regular merge and subblock merge respectively, and reordered with the existing merge candidates. The lengths of the output merge lists are kept unchanged. The same pattern used for non-adjacent merge mode is reused to locate the non-adjacent positions in the scheme.

[0109] I. 12 Affine Motion Compensated Prediction

[0110] In HEVC, only translational motion model is applied for motion compensation prediction (MCP) . While in the real world, there are many kinds of motion, e.g. zoom in / out, rotation, perspective motions and the other irregular motions. In VVC, a block-based affine transform motion compensation prediction is applied. As shown Figs. 9A-B, the affine motion field of the blocks 910 and 920 is described by motion information of two control point (4-parameter) in Fig. 9A or three control point motion vectors (6-parameter) in Fig. 9B.

[0111] For 4-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:

[0112] For 6-parameter affine motion model, motion vector at sample location (x, y) in a block is derived as:

[0113] Where (mv0x, mv0y) is motion vector of the top-left corner control point, (mv1x, mv1y) is motion vector of the top-right corner control point, and (mv2x, mv2y) is motion vector of the bottom-left corner control point.

[0114] In order to simplify the motion compensation prediction, block based affine transform prediction is applied. To derive motion vector of each 4×4 luma subblock, the motion vector of the centre sample of each subblock, as shown in Fig. 10, is calculated according to above equations, and rounded to 1 / 16 fraction accuracy. Then, the motion compensation interpolation filters are applied to generate the prediction of each subblock with the derived motion vector. The subblock size of chroma-components is also set to be 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luma subblocks in the collocated 8x8 luma region.

[0115] In the present invention, a scheme to improve the coding performance is introduced, where smoothing filters are applied to reference pictures to generate a refined prediction. BRIEF SUMMARY OF THE INVENTION

[0116] A method and apparatus for video coding using filtered reference pictures to improve coding performance are disclosed. According to the method, input data associated with a current block in a current frame is received, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. One or more filters to be applied to one or more reference frames associated with the current block are determined. At least one filtered reference block is generated by applying said one or more filters to at least one of said one or more reference frames. A filtered predictor for the current block is derived based on said at least one filtered reference block. The current block is encoded or decoded using the filtered predictor.

[0117] In one embodiment, the current block is coded in a bi-prediction mode associated with a first reference picture and a second reference picture, the method further comprising: generating a first filtered reference block by applying a first filter to the first reference picture; generating a second filtered reference block by applying a second filter to the second reference picture; and combining the first filtered reference block and the second filtered reference block to form the filtered predictor for the current block.

[0118] In one embodiment, one filter determined is shared by a plurality of reference frames.

[0119] In one embodiment, a plurality of reference frames in a video sequence are classified into one or more reference frame groups based on a temporal order; and wherein an individual filter is determined for each of said one or more reference frame groups. In one embodiment, the individual filter is derived for a specific reference frame group by minimizing a cost function representing differences between first samples of a current original frame and second samples of member reference frames within the specific reference frame group.

[0120] In one embodiment, at least one filter is shared by a first set of reference frames and a second set of reference frames. In one embodiment, said at least one filter is derived by minimizing a joint cost function representing (i) first differences between first samples of a current original frame and second samples of first member reference frames within the first set of reference frames and (ii) second differences between the first samples of the current original frame and third samples of second member reference frames within the second set of reference frames.

[0121] In one embodiment, at least one filter is shared by first reference frames in a first reference list and second reference frames in a second reference list. In one embodiment, said at least one filter is derived by minimizing a joint cost function representing (i) first differences between first samples of a current original frame and second samples of the first reference frames within the first reference list and (ii) second differences between the first samples of the current original frame and third samples of second reference frames in the second reference list.

[0122] In one embodiment, prior to encoding or decoding the current frame, each of said one or more reference frames associated with the current frame is processed by a corresponding prediction filter to generate a plurality of filtered reference frames; and said at least one filtered reference block is generated from the plurality of filtered reference frames.

[0123] In one embodiment, a plurality of additional reference frames is generated by applying said one or more filters to said one or more reference frames associated with the current block; and wherein the filtered predictor for the current block is derived using a combination of samples selected from the plurality of additional reference frames and said one or more reference frames associated with the current block. In one embodiment, a filtered sample in one of the plurality of additional reference frames and a corresponding un-filtered sample in one of said one or more reference frames are combined using a weighting factor to derive a target sample of the filtered predictor.

[0124] In one embodiment, applying said one or more filters to said one or more reference frames comprises selecting filter coefficients based on a local metric; and wherein the local metric comprises at least one selected from a group consisting of: an intensity of a current sample, an average intensity of the current sample and neighbouring samples, and an image gradient of a region surrounding the current sample.

[0125] In one embodiment, a first rule is used to select first filter coefficients for the current block coded in a uni-prediction mode and a second rule is used to select second filter coefficients for the current block coded in a bi-prediction mode, and wherein the second rule is different from the first rule.BRIEF DESCRIPTION OF THE DRAWINGS

[0126] Fig. 1A illustrates an exemplary adaptive Inter / Intra video coding system incorporating loop processing.

[0127] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0128] Fig. 2 illustrates the ALF filter shapes for the chroma (left) and luma (right) components.

[0129] Fig. 3A illustrates the placement of CC-ALF with respect to other loop filters.

[0130] Fig. 3B illustrates a diamond shaped filter for the chroma samples.

[0131] Fig. 4 illustrates the cross 9x9 CCALF filter shape, luma residual based taps and 4 taps in a 3x3 asymmetric cross shape as additional inputs to CCALF.

[0132] Fig. 5 illustrates the filter shape of ALF in ECM-7.0.

[0133] Fig. 6 illustrates the diamond regions in the search area for multi-pass decoder-side motion vector refinement.

[0134] Fig. 7 illustrates an example of group of pictures (GOP) structure with 16 pictures in the group.

[0135] Fig. 8 illustrates an example of bi-predictive temporal MVP derivation.

[0136] Fig. 9A illustrates an example of the affine motion field of a block described by motion information of two control point (4-parameter) .

[0137] Fig. 9B illustrates an example of the affine motion field of a block described by motion information of three control point motion vectors (6-parameter) .

[0138] Fig. 10 illustrates an example of block based affine transform prediction, where the motion vector of each 4×4 luma subblock is derived from the control-point MVs.

[0139] Fig. 11 illustrates an example of prediction filtering scheme according to one embodiment of the present invention.

[0140] Fig. 12 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0141] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0142] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0143] PROPOSED METHOD

[0144] The following methods are proposed to improve the luma / chroma inter coding prediction accuracy or coding performance. In one embodiment, whether to allow or apply the following proposed methods can depend on SPS / PPS / SH / PH syntax, or CTU or CU / PU / TU level syntax or semantic. In another embodiment, whether to allow or apply the following proposed methods can also depend on implicit conditions. For example, the proposed methods described herein can be applied based on one or more of block width, block height, block area, QP, POC, temporal index, reference index, or prediction mode. In another embodiment, the following proposed prediction filtering methods can be conditionally applied to partial positions / sub-block of the current block. In another embodiment, the term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB. Any combination of the following proposed methods in this invention can be applied.

[0145] II. Prediction Filtering

[0146] In one invention, a prediction filtering scheme is proposed to improve the prediction quality or reduce the residual signalling overhead to improve the coding performance. In the proposed scheme, for a current coding frame, a filter is derived for each reference frame using regression that minimizes differences (e.g., mean squared error (MSE) ) between samples in the current original frame and corresponding samples in the reference frame. For an example, two filters are derived for the current frame if there are two reference frames for the current frame. The predictor of an inter CU can be refined by applying the filter derived from the corresponding reference frame. In Fig. 11, an example of applying prediction filters to an inter block is illustrated. Two filters (filter 0 and filter 1) are applied on the reference blocks (Refblk0 and Refblk1) from two reference pictures (reference picture 0 and reference picture 1) respectively to improve the prediction quality for the current block (CurrBlk) in the current picture before bi-predictive blending.

[0147] In one embodiment, a filter is derived for a group of reference frames. That is, a derived filter is shared among the reference frames in the group. In one example, N1 reference frames before the current frame in the temporal order are classified into one group and N2 reference frames after the current frame in the temporal order are classified into another group. A filter is derived for each group by minimizing the differences between the samples on the current original frame and the samples on the reference frames in the group of reference frames and applied to the corresponding group. In another example, every N3 (where N3 >= 1) reference frames of current frame are classified into one group.

[0148] In one embodiment, some of the filters can be shared across coding frames. In an example, a filter can be derived by minimizing the differences between the samples on original frame of first coding frame and the samples on K1 (where K1 >= 1) reference frames of the first coding frame and the differences between the samples on original frame of the second coding frame and the samples on K2 (where K2 >= 1) reference frames of the second coding frame. The derived filter is shared for K1 (where K1 >= 1) reference frames of the first coding frame and K2 (where K2 >= 1) reference frames of the second coding frame. In another example, a filter can be applied to k1, k2, …, kj reference frames of 1st, 2nd, …, jth (where j >= 1) coding frame.

[0149] In one embodiment, some of the filters can be shared across reference lists. In an example, a filter can be derived by minimizing the differences between the samples on the current original frame and the samples on D1 (where D1 >= 1) reference frames in the first reference list of the current coding frame and the differences between the samples on the current original frame and the samples on D2 (where D2 >= 1) reference frames in the second reference list of the current coding frame. The derived filter is shared for D1 (where D1 >= 1) reference frames in the first reference list of the current coding frame and D2 (where D2 >= 1) reference frames in the second reference list of the current coding frame. In another example, a filter can be applied to d1, d2, …, di reference frames in 1st, 2nd, …, ith (where i >= 1) reference lists of a coding frame.

[0150] II. 1 Filter design

[0151] II. 1.1 Filter inputs

[0152] In one embodiment, when applying a filter to a sample of a reference frame, its neighbouring samples are directly multiplied with filter coefficients.

[0153] In one embodiment, when applying a filter to a sample of a reference frame, differences between its neighbouring samples and the sample are multiplied with filter coefficients.

[0154] In one embodiment, when applying a filter to a sample of a reference frame, differences between its neighbouring samples and an offset value are multiplied with filter coefficients.

[0155] In one embodiment, when applying a filter to a sample of a reference frame, the sample value is multiplied with a filter coefficient.

[0156] In one embodiment, when applying a filter to a sample of a reference frame, one filter coefficient represents an offset term (i.e., not multiplied with any sample values) .

[0157] In one embodiment, when applying a filter to a sample of a reference frame, the high-order sample value (e.g., the square of the sample value) is multiplied with a filter coefficient.

[0158] II. 1.2 Classification

[0159] In one embodiment, when applying a filter to a sample of a reference frame, filter coefficients are selected based on statistics associated with the sample and / or its neighbouring samples. The statistics includes at least one of the following. ● Intensity of the sample. ● Average intensity of the sample and its neighbouring samples. ● Image gradient surrounding the sample.

[0160] In one embodiment, when applying a filter to a sample of a reference frame, filter coefficients are selected based on a motion vector that points to the reference frame. Specifically, one of the following information associated with the motion vector could be used. ● MV magnitude. ● MV direction. ● MV fractional part.

[0161] In one embodiment, when applying a filter to a sample of a reference frame, filter coefficients are selected based on a position information. Specifically, one of the following could be used. ● The position of the sample in the reference frame. ● The position of a current sample in the current-coded frame that would be predicted by the sample of the reference frame.

[0162] In one embodiment, when applying a filter to a sample of a reference frame, filter coefficients are selected based on some metrics mentioned above, where the rules vary with the input bit-depths of the reference frame.

[0163] For example, when uni-prediction is performed on a current-coded frame, a low bit-depth is used to fetch the samples from a reference frame; when bi-prediction is performed on a current-coded frame, a high bit-depth is used to fetch the samples from each reference frame. In such design, a first rule is used to select filter coefficients for the uni-prediction case, and a second rule different from the first rule is used to select filter coefficients for the bi-prediction case.

[0164] In the above embodiments, the filter coefficient selection is performed at sample level. That is, for each sample, a set of filter coefficients are selected.

[0165] In the above embodiments, the filter coefficient selection is performed at block level. That is, for each block, a set of filter coefficients are selected. The block could be a CU / PU / TU, or a fixed size unit (e.g., 2x2, 4x4) .

[0166] II. 1.3 Nonlinear filtering

[0167] In one embodiment, when applying a filter to a sample of a reference frame, nonlinear filtering is used. That is, a value associated with the sample and / or its neighbouring sample is clipped before multiplying with a filter coefficient.

[0168] In one embodiment, the selection of the clipping threshold for each coefficient is explicitly signalled / parsed.

[0169] In one embodiment, the clipping thresholds are modified according to a bit-depth for fetching a reference frame.

[0170] II. 2 Filter Application

[0171] In one embodiment, before the motion estimation process of a coding frame at the encoder, all reference frames of the coding frame are filtered by corresponding prediction filters. Before decoding the coding frame, all reference frames of the coding frame are filtered by corresponding prediction filters.

[0172] In one embodiment, several additional reference frames are generated by applying corresponding prediction filters to reference frames. In one example, the additional filtered reference frames are inserted before the un-filtered reference frames. In another example, the additional filtered reference frames are inserted after the un-filtered reference frames. In another example, the additional filtered reference frames are interleaved with the un-filtered reference frames.

[0173] In one embodiment, during the motion compensation process of a coding block, the predictor before or after the interpolation filtering are filtered by the prediction filter from the corresponding reference frame.

[0174] In one embodiment, a filtered sample can be blended with an un-filtered sample by a blending weight to generate the refined sample. The blending weight can be any value pair depending on or independent of any coding information (e.g., QP, coding mode, block area, block shape, etc. ) . In one example, the weight for filtered samples is 0.25 or 0.125 and the weight for un-filtered sample is 0.75 or 0.875.

[0175] Any of the foregoing proposed methods of filtered prediction generated by applying smoothing filters to reference frames can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module

[0176] With reference to the exemplary encoder and decoder in Fig. 1A and Fig 1B, the proposed methods can be implemented in the intra / inter prediction modules. For example, in the encoder side, the required processing can be implemented as part of the Inter-Pred. unit 112 or Intra Pred. unit 110 as shown in Fig. 1A. However, the encoder may also use additional processing unit to implement the required processing. For the decoder side, the required processing can be implemented as part of the MC unit 152 or Intra Pred. 150 as shown in Fig. 1B. However, the decoder may also use additional processing unit to implement the required processing. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module. While the Inter-Pred. 112 and Intra Pred. 110 in the encoder side and MC 152 and Intra Pred. 150 in the decoder side are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0177] Fig. 12 illustrates a flowchart of an exemplary video coding system that improves the coding performance by using a refined prediction generated by applying smoothing filters to reference pictures according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g. one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data associated with a current block in a current frame is received in step 1210, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side. One or more filters to be applied to one or more reference frames associated with the current block are determined in step 1220. At least one filtered reference block is generated by applying said one or more filters to at least one of said one or more reference frames in step 1230. A filtered predictor for the current block is derived based on said at least one filtered reference block in step 1240. The current block is encoded or decoded using the filtered predictor in step 1250.

[0178] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0179] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0180] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0181] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding, the method comprising:receiving input data associated with a current block in a current frame, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;determining one or more filters to be applied to one or more reference frames associated with the current block;generating at least one filtered reference block by applying said one or more filters to at least one of said one or more reference frames;deriving a filtered predictor for the current block based on said at least one filtered reference block; andencoding or decoding the current block using the filtered predictor.2.The method of claim 1, wherein the current block is coded in a bi-prediction mode associated with a first reference picture and a second reference picture, the method further comprising:generating a first filtered reference block by applying a first filter to the first reference picture;generating a second filtered reference block by applying a second filter to the second reference picture; andcombining the first filtered reference block and the second filtered reference block to form the filtered predictor for the current block.3.The method of claim 1, wherein one filter determined is shared by a plurality of reference frames.4.The method of claim 1, wherein a plurality of reference frames in a video sequence are classified into one or more reference frame groups based on a temporal order; and wherein an individual filter is determined for each of said one or more reference frame groups.5.The method of claim 4, wherein the individual filter is derived for a specific reference frame group by minimizing a cost function representing differences between first samples of a current original frame and second samples of member reference frames within the specific reference frame group.6.The method of claim 1, wherein at least one filter is shared by a first set of reference frames and a second set of reference frames.7.The method of claim 6, wherein said at least one filter is derived by minimizing a joint cost function representing (i) first differences between first samples of a current original frame and second samples of first member reference frames within the first set of reference frames and (ii) second differences between the first samples of the current original frame and third samples of second member reference frames within the second set of reference frames.8.The method of claim 1, wherein at least one filter is shared by first reference frames in a first reference list and second reference frames in a second reference list.9.The method of claim 8, wherein said at least one filter is derived by minimizing a joint cost function representing (i) first differences between first samples of a current original frame and second samples of the first reference frames within the first reference list and (ii) second differences between the first samples of the current original frame and third samples of second reference frames in the second reference list.10.The method of claim 1, wherein, prior to encoding or decoding the current frame, each of said one or more reference frames associated with the current frame is processed by a corresponding prediction filter to generate a plurality of filtered reference frames; and said at least one filtered reference block is generated from the plurality of filtered reference frames.11.The method of claim 1, wherein a plurality of additional reference frames is generated by applying said one or more filters to said one or more reference frames associated with the current block; and wherein the filtered predictor for the current block is derived using a combination of samples selected from the plurality of additional reference frames and said one or more reference frames associated with the current block.12.The method of claim 11, wherein a filtered sample in one of the plurality of additional reference frames and a corresponding un-filtered sample in one of said one or more reference frames are combined using a weighting factor to derive a target sample of the filtered predictor.13.The method of claim 1, wherein applying said one or more filters to said one or more reference frames comprises selecting filter coefficients based on a local metric; and wherein the local metric comprises at least one selected from a group consisting of: an intensity of a current sample, an average intensity of the current sample and neighbouring samples, and an image gradient of a region surrounding the current sample.14.The method of claim 1, wherein a first rule is used to select first filter coefficients for the current block coded in a uni-prediction mode and a second rule is used to select second filter coefficients for the current block coded in a bi-prediction mode, and wherein the second rule is different from the first rule.15.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block in a current frame, wherein the input data comprises pixel data for the current block to be encoded at an encoder side or coded data for decoding the current block at a decoder side;determine one or more filters to be applied to one or more reference frames associated with the current block;generate at least one filtered reference block by applying said one or more filters to at least one of said one or more reference frames;derive a filtered predictor for the current block based on said at least one filtered reference block; andencode or decode the current block using the filtered predictor.