Methods and apparatus of filter-based intra prediction for video coding system
By classifying samples into multiple classes and applying tailored filters, the video coding system enhances prediction accuracy and reduces complexity in intra mode coding, addressing inefficiencies in existing systems.
Patent Information
- Application Number
- PCT/CN2025/071173
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
Existing video coding systems face inefficiencies in filter-based intra prediction due to the use of a single filter for all to-be-predicted samples, which fails to account for varying correlations within a block, leading to suboptimal prediction accuracy and increased complexity in intra mode coding.
Classify to-be-predicted samples into multiple classes and apply different filters based on classification settings, using filter shapes and parameters derived from neighboring regions or inherited from previous blocks, to enhance prediction accuracy and reduce complexity.
Improves prediction accuracy and reduces complexity by tailoring filters to specific sample classes, resulting in more efficient video coding performance.
Smart Images

Figure CN2025071173_17072025_PF_FP_ABST
Abstract
Description
METHODS AND APPARATUS OF FILTER-BASED INTRA PREDICTION FOR VIDEO CODING SYSTEMCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 619,338, filed on January 10, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system. In particular, the present invention relates to schemes to improve the performance of filter-based intra prediction by classifying to-be-predicted samples into multiple classes and using multiple filters accordingly. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use similar or portion of the same functional blocks as the encoder except for Transform 118 and Quantization 120 since the decoder only needs Inverse Quantization 124 and Inverse Transform 126. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. Also, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.
[0008] Partitioning of the CTUs Using a Tree Structure
[0009] In VVC, the coding tree scheme supports the ability for the luma and chroma to have a separate block tree structure (or called dual tree structure) . For P and B slices, the luma and chroma CTBs in one CTU have to share the same coding tree structure (or called single tree structure) .
[0010] Intra Mode Coding with 67 Intra Prediction Modes
[0011] To capture the arbitrary edge directions presented in natural video, the number of directional intra modes in VVC is extended from 33, as used in HEVC, to 65.
[0012] Intra Mode Coding
[0013] To keep the complexity of the most probable mode (MPM) list generation low, an intra mode coding method with 6 MPMs is used by considering two available neighbouring intra modes. The following three aspects are considered to construct the MPM list: – Default intra modes – Neighbouring intra modes – Derived intra modes.
[0014] Decoder Side Intra Mode Derivation (DIMD)
[0015] To implicitly derive the intra prediction modes of a block, a texture gradient analysis is performed at both the encoder and decoder sides. This process starts with an empty Histogram of Gradient (HoG) with 65 entries, corresponding to the 65 angular modes. Amplitudes of these entries are determined during the texture gradient analysis.
[0016] In the first step, DIMD picks a template of T=3 columns and lines from respectively left side and above side of the current block. This area is used as the reference for the gradient based intra prediction modes derivation.
[0017] In the second step, the horizontal and vertical Sobel filters are applied on all 3×3 window positions, centred on the pixels of the middle line of the template. At each window position, Sobel filters calculate the intensity of pure horizontal and vertical directions as Gx and Gy, respectively. Then, the texture angle of the window is calculated as: angle=arctan (Gx / Gy) , (1) which can be converted into one of 65 angular intra prediction modes. Once the intra prediction mode index of current window is derived as idx, the amplitude of its entry in the HoG [idx] is updated by addition of: ampl = |Gx|+|Gy| (2)
[0018] Figs. 2A-C show an example of HoG, calculated after applying the above operations on all pixel positions in the template. Fig. 2A illustrates an example of selected template 220 for a current block 210. Template 220 comprises T lines above the current block and T columns to the left of the current block. For intra prediction of the current block, the area 230 at the above and left of the current block corresponds to a reconstructed area and the area 240 below and at the right of the block corresponds to an unavailable area. Fig. 2B illustrates an example for T=3 and the HoGs are calculated for pixels 260 in the middle line and pixels 262 in the middle column. For example, for pixel 252, a 3x3 window 250 is used. Fig. 2C illustrates an example of the amplitudes (ampl) calculated based on equation (2) for the angular intra prediction modes as determined from equation (1) .
[0019] Once HoG is computed, the indices with two tallest histogram bars are selected as the two implicitly derived intra prediction modes (IPMs) for the block and are further combined with the Planar mode as the prediction of DIMD mode. The prediction fusion is applied as a weighted average of the above three predictors. To this aim, the weight of planar is fixed to 21 / 64 (~1 / 3) . The remaining weight of 43 / 64 (~2 / 3) is then shared between the two HoG IPMs, proportionally to the amplitude of their HoG bars.
[0020] Template-based Intra Mode Derivation (TIMD)
[0021] Template-based intra mode derivation (TIMD) mode implicitly derives the intra prediction mode of a CU using a neighbouring template at both the encoder and decoder, instead of signalling the intra prediction mode to the decoder.
[0022] Intra Sub-Partitions (ISP)
[0023] The intra sub-partitions (ISP) divides luma intra-predicted blocks vertically or horizontally into 2 or 4 sub-partitions depending on the block size.
[0024] CCLM (Cross Component Linear Model) Overview
[0025] The main idea behind CCLM mode (sometimes abbreviated as LM mode) is as follows: chroma components of a block can be predicted from the collocated reconstructed luma samples by linear models whose parameters are derived from already reconstructed luma and chroma samples that are adjacent to the block.
[0026] MMLM Overview
[0027] As indicated by the name, the original CCLM mode employs one linear model for predicting the chroma samples from the luma samples for the whole CU, while in MMLM (Multiple Model CCLM) , there can be two models. In MMLM, neighbouring luma samples and neighbouring chroma samples of the current block are classified into two groups, each group is used as a training set to derive a linear model (i.e., particular α and β are derived for a particular group) . Furthermore, the samples of the current luma block are also classified based on the same rule for the classification of neighbouring luma samples.
[0028] Convolutional Cross-Component Model (CCCM)
[0029] In CCCM, a convolutional model is applied to improve the chroma prediction performance. The filter coefficients are calculated by minimising MSE between predicted and reconstructed chroma samples in the reference area. The MSE minimization is performed by calculating autocorrelation matrix for the luma input and a cross-correlation vector between the luma input and chroma output. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The process follows roughly the calculation of the ALF filter coefficients in Enhanced Compression Model (ECM) , however LDL decomposition was chosen instead of Cholesky decomposition to avoid using square root operations.
[0030] An Extrapolation Filter-Based Intra Prediction (EIP) Mode (from JVET-AF0080)
[0031] In JVET-AF0080 (Luhang Xu, et al., “EE2-2.7: An extrapolation filter-based intra prediction mode” , Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 32nd Meeting, Hannover, DE, 13–20 October 2023, Document: JVET-AF0080) , the extrapolation filter-based intra prediction is disclosed.
[0032] A. Obtaining the EIP Filter
[0033] The three EIP filter shapes are shown in Fig. 3, where the three filter shapes correspond to square 310, horizontal strip 320, and vertical strip 330.
[0034] There are two ways to obtain the filter coefficients for the current CU. First, the coefficients can be derived from the neighbouring reconstructed pixels. Second, they can also be inherited from the previously decoded blocks.
[0035] B. Derivation of EIP Coefficients
[0036] The decoder decodes the relevant syntax elements to determine the selected type of reconstructed area and the filter shape for the current block. The selected filter moves in the selected reconstructed area either horizontally or vertically with a one-pixel step to construct the auto-correlation matrix and the cross-correlation vector. The calculation of coefficients from the auto-correlation matrix and the cross-correlation vector is the same as that in convolutional cross-component model (CCCM) .
[0037] The three types of the reconstructed area are defined as shown in Fig. 4, where the three reconstructed areas correspond to Left-Above area (Fig. 4A) , Above area (Fig. 4B) , and Left area (Fig. 4C) . The size of the reconstructed area depends on the min (blockWidth, blockHeight) and the selected filter shape. For example, when the current block is an 8x16 block and the selected filter shape is 4x4. The aboveSize of the reconstructed area is equal to min (8, 16) + 4 –1 = 11, and the leftSize of the reconstructed area is equal to min (8, 16) + 4 –1 = 11.
[0038] C. Inheritance of the EIP Filters
[0039] The filter shape and the filter coefficients can be inherited from the previous decoded blocks with EIP or EIP merge mode. The decoder decodes an EIP merge flag to decide whether the proposed merge mode is used when the current block uses the EIP mode. A merge index is further decoded when the EIP merge flag is true. The EIP merge list includes spatial adjacent and non-adjacent candidates, temporal candidates, and history candidates. The constructed EIP merge list can include up to 12 candidates and the list will be reduced to up to 6 candidates by the reordering process based on the SAD cost measured on an L-shape template with column width and row height equal to 1. In the SAD calculation, predictions of the template area by EIP filters are generated only from reconstructed (neighbouring and template) samples, allowing the EIP filters to be applied in parallel rather than sequentially.
[0040] D. Prediction of the Current Block
[0041] The EIP mode generates prediction values for the current block from the top-left position to the bottom-right position by a diagonal prediction order, as shown in Fig. 5.
[0042] The calculation for the prediction values is shown as follows: where pred (x, y) is the predicted value at (x, y) in the current block, ci is the ith coefficient of the selected EIP filter, the index of the coefficients is from 0 to 14, is a reconstructed or a predicted value used for the current position’s prediction. offsetXi and offsetY are the position offsets to the current position along x and y directions, respectively.
[0043] E. Mapping to the LFNST / NSPT / MTS Set
[0044] The proposed method in JVET-AF0080 uses the DIMD process to derive an intra prediction mode of the current block based on the EIP predicted samples. Specifically, a horizontal gradient and a vertical gradient are calculated for each predicted sample to build a histogram of gradient (HoG) . Then the intra prediction mode corresponding to the largest histogram count is used to determine the low-frequency non-separable transform (LFNST) , non-separable primary transform (NSPT) or multiple transform selection (MTS) transform set.
[0045] F. Proposed CU Level Syntax
[0046] The EIP related syntax is signalled at CU level.
[0047] In the present invention, methods and apparatus to improve the performance of filter-based intra prediction by classifying to-be-predicted samples into multiple classes and using multiple filters accordingly are disclosed. BRIEF SUMMARY OF THE INVENTION
[0048] A method and apparatus for video coding are disclosed. According to this method, input data associated with a current block are received, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. To-be-predicted samples of the current block are classified into multiple classes. The current block is encoded or decoded by using multiple filters, wherein the to-be-predicted samples of the current block in a target class are predicted using one target filter selected from the multiple filters according to the target class.
[0049] In one embodiment, the current block is coded using EIP (Extrapolation Filter-Based Intra Prediction) .
[0050] In one embodiment, whether to apply classification setting to the current block is determined first, and said classifying the to-be-predicted samples of the current block is applied if classification setting is applied to the current block.
[0051] In one embodiment, one or more sets of filter information and / or one or more sets of threshold information of the current block are stored for the current block and / or referenced by one or more subsequent blocks. In one embodiment, said one or more sets of filter information comprise filter shape, filter parameters, partial filter parameters, or a combination thereof.
[0052] In one embodiment, input samples for the multiple filters comprise reconstructed and / or predicted samples associated with the current block, and / or reconstructed and / or predicted samples associated with a reference region reconstructed prior to the current block.
[0053] In one embodiment, input samples for at least one of the multiple filters are within an MxN region around a current to-be-predicted sample at (x, y) , excluding locations with horizontal coordinate > x and / or vertical coordinate > y, and wherein M and N are positive integers.
[0054] In one embodiment, one or more filter parameters for at least one of the multiple filters are derived using one or more templates of the current block. In one embodiment, said one or more filter parameters for at least one of the multiple filters are derived by a regression-based process. In one embodiment, the regression-based process is unified with a process of deriving cross-component models for one or more cross-component chroma modes. In one embodiment, the regression-based process corresponds to Gaussian elimination.
[0055] In one embodiment, different filters are used for the to-be-predicted samples of the current block belonging to different classes. In one embodiment, different filter shapes are used for the to-be-predicted samples of the current block belonging to different classes.
[0056] In one embodiment, said classifying the to-be-predicted samples of the current block into the multiple classes is based on pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or a combination thereof of one or more filtering inputs.
[0057] In one embodiment, filter information for at least one of the multiple filters is inherited from one or more previously-coded blocks. In one embodiment, when filter information for the current block is inherited from one or more previously-coded blocks, whether the current block is classified into the multiple classes depends on the filter information for the current block inherited from said one or more previously-coded blocks. In another embodiment, when filter information for the current block is inherited from one or more previously-coded blocks, whether the current block is classified into the multiple classes depends on one or more enabling conditions associated with block position, block width, block height, block area, or a combination thereof.
[0058] In one embodiment, when filter information for the current block is derived, whether the current block is classified into the multiple classes depends on enabling conditions and / or one or more signalled or parsed syntax elements.BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.
[0060] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0061] Fig. 2A illustrates an example of selected template for a current block, where the template comprises T lines above the current block and T columns to the left of the current block.
[0062] Fig. 2B illustrates an example for T=3 and the HoGs (Histogram of Gradient) are calculated for pixels in the middle line and pixels in the middle column.
[0063] Fig. 2C illustrates an example of the amplitudes (ampl) for the angular intra prediction modes.
[0064] Fig. 3 illustrates three types of filter shapes with fifteen inputs and generates one output for EIP process.
[0065] Figs. 4A-C illustrate three types (Fig. 4A: Left-Above area, Fig. 4B: Above area, and Fig. 4C: Left area) of reconstructed areas used to derive filter coefficients for EIP.
[0066] Fig. 5 illustrates an example of scanning order for generating predictions for different positions in the current block by a diagonal order.
[0067] Fig. 6 illustrates an example of the filter shape using a pattern around (not including) the position (x, y) of the to-be-predicted sample, where the pattern is a MxN region.
[0068] Fig. 7 illustrates an example of a square MxN region excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y.
[0069] Fig. 8 illustrates an example of a non-square region (M<N) excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y.
[0070] Fig. 9 illustrates an example of a non-square region (M>N) excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y.
[0071] Fig. 10 illustrates an example of template (i.e., neighbouring region) of the current block used for deriving filter parameters, where the template refers to top template, left template, and top-left template.
[0072] Fig. 11 illustrates an example of to-be-predicted samples near the top boundary of the current block.
[0073] Fig. 12 illustrates an example of to-be-predicted samples at the inner portion of the current block and / or far away from the top and left boundaries of the current block.
[0074] Fig. 13 illustrates an example of classifying to-be-predicted samples into two classes, where the lighter samples are classified as case 1 and the darker samples are classified as class 2.
[0075] Fig 14 illustrates an example of threshold set as the top-left sample outside of the current block.
[0076] Fig. 15 illustrates an example of different filters on the template and the current block.
[0077] Fig. 16 illustrates an exemplary processing flow for one embodiment of the present invention.
[0078] Fig. 17 illustrates a flowchart of an exemplary video coding system that classifies to-be-predicted samples into multiple classes and uses multiple filters according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0079] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0080] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0081] Proposed Method
[0082] In VVC, for an intra block obtained by partitioning as mentioned before, first, an intra prediction mode is signalled to generate the predictor of the current intra block. Also, ISP is proposed to apply an intra direction to the short-distance reconstruction of the previous-coded sub-partition to get the predictor for the current sub-partition. However, with the block texture being complex, only one intra direction, DC, planar, or a combination is not good enough for generating an accurate predictor of the current block. Furthermore, with the intra prediction mode signalling being expensive, in the emerging video coding technology, DIMD and / or TIMD appear to improve intra prediction with a pre-defined decoder-side derivation to derive one or more intra prediction modes among existing intra prediction directions, planar, DC, or a combination to generate the predictor of the current block. Moreover, filter-based intra prediction modes apply filtering to generate the predictors of the current block. For example, the filter-based intra prediction mode may refer to EIP. In this invention, the classification setting is proposed to improve a target filter (or model) -based intra prediction mode. In one embodiment, the target filter-based intra prediction mode is EIP. In another embodiment, the target filter-based intra prediction mode is not limited to EIP and / or can be any variation of filter-based intra prediction mode.
[0083] In the following, various aspects of the present invention are disclosed. In section (1) , several variations of filter-based intra prediction modes are proposed to be the target filter-based intra prediction mode. In section (2) , the classification setting is defined to derive the predictor of the current block when the target filter-based intra prediction mode is used for the current block. In section (3) , several enabling conditions are checked to determine whether the classification setting can be used for the current block coded with the target filter-based intra prediction mode. In section (4) , several methods are proposed to define the storage rules of the target filter-based intra prediction mode with the classification setting. In section (5) , exemplary processing steps of this invention are described.
[0084] (1) Target Filter-Based Intra Prediction Mode
[0085] A filter (or model) of the target filter-based intra prediction mode includes a filter shape and / or determination of filter parameters. After deciding a filter shape among one or more candidate filter shapes and / or a determination of filter parameters among one or more candidate determinations of filter parameters, the filter of the current block is obtained. For each to-be-predicted sample of the current block, the filter is applied. In one embodiment, the filter is applied following a pre-defined order such as horizontal scanning, vertical scanning, diagonal scanning, and / or any scanning order. For example, the pre-defined order is diagonal scanning. That is, the predicted sample located at the position in the first of the diagonal scanning for the current block is outputted first. Then, the predicted sample located at the position in the second, third, …, n-th of the diagonal scanning for the current block is outputted in order, where n corresponds to the block area. When applying the filter, the input of the filter is combined (or convolved) with the filter parameters to get the output of the filter. In one embodiment, the input samples refer to reconstructed and / or predicted samples associated with the current block and / or reconstructed and / or predicted samples associated with the reference region reconstructed prior to the current block. Finally, after applying the filter to each to-be-predicted sample in the current block, the prediction of the current block is obtained.
[0086] In one embodiment, the filter shape uses a pattern around (not including) the position (x, y) of the to-be-predicted sample. The pattern is a MxN region or any subset of the MxN region around the to-be-predicted sample where M is a first pre-defined positive number larger than 1 and / or N is a second pre-defined positive number larger than 1. Fig. 6 shows an example of the MxN region.
[0087] In one sub-embodiment, M and N can be the same or different. For example, M and / or N are equal to 1, 2, 3, 4, …, or any number specified in the standard. For another example, M is larger than N if the current block has the block width larger than the block height. For another example, N is larger than M if the current block has the block height larger than the block width. For another example, M and / or N are adaptive according to the position (x, y) .
[0088] In another sub-embodiment, the subset of the MxN region can be the MxN region (a) excluding the region with the horizontal coordinate > x and / or (b) excluding the region with the vertical coordinate > y. In one example, the subset of MxN is the MxN region with (a) + (b) and M is equal to N. Therefore, the MxN region is square. The proposed method of subset of the MxN region is not limited to using in the case of M equal to N and / or can be applied to the case of M > N and / or N > M. For another example, the subset of MxN is the MxN region with (a) + (b) and M > N. For another example, the subset of MxN is the MxN region with (a) + (b) and N > M.
[0089] An example of the MxN region (square) excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y is shown in Fig. 7. An example of the MxN region (non-square, M < N) excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y is shown in Fig. 8. An example of the MxN region (non-square, M > N) excluding the region with the horizontal coordinate > x and excluding the region with the vertical coordinate > y is shown in Fig. 9.
[0090] In another embodiment, multiple filter shapes are supported and the selection from the multiple filter shapes depends on the explicit signalling at the block-level, CTU-level, slice-level, tile-level, SPS-level, PPS-level, picture-level (for example, picture header) , and / or sequence-level. If the total number of the candidate filter shapes in a list containing multiple filter shapes is K, a syntax is signalled to select a filter shape. For example, the syntax is truncated unary coding. A shortest codeword is used to indicate the first candidate filter shape in the list.
[0091] In another embodiment, multiple determinations of filter parameters are supported and the selection from the multiple determinations of filter parameters depends on the explicit signalling at the block-level, CTU-level, slice-level, tile-level, SPS-level, PPS-level, picture-level (for example, picture header) , and / or sequence-level.
[0092] In another embodiment, the determination of filter parameters can be deriving filter parameters using the template (neighbouring region) of the current block and / or inheriting filter information from the previously-coded blocks. In one sub-embodiment, the selection between deriving filter parameters using the template (neighbouring region) of the current block and inheriting filter information from the previously-coded blocks depends on the explicit signalling at the block-level, CTU-level, slice-level, tile-level, SPS-level, PPS-level, picture-level (for example, picture header) , and / or sequence-level.
[0093] In one sub-embodiment, in response to deriving filter parameters using the template (neighbouring region) of the current block, the golden (i.e., a target for comparison) is the reconstructed samples on the template of the current block and the goal is to make the filtering result match the golden on the template. Any method to minimize the difference between the golden and the filtering result can be used to get the filer parameters whose filtering result achieves a good matching with the golden. For example, a regression-based method such as Gaussian elimination and / or any regression method unified with the methods of deriving cross-component models for the cross-component chroma modes (such as the cross-component chroma modes mentioned earlier or any cross-component chroma mode in the standard) is used.
[0094] In another sub-embodiment, in response to deriving filter parameters using the template (neighbouring region) of the current block, the template refers to top template, left template, top-left template, and / or any combination of the above-mentioned templates as shown in Fig. 10. The selection of the template among multiple candidate templates can depend on the width, height, area, and / or position of the current block. For example, if the position of the current block is located at the top of picture boundary and / or the top of a CTU row boundary, the top template and / or the top-left template is not used for deriving filter parameters. For example, if the position of the current block is located at the left of picture boundary and / or the left of a CTU boundary, the left template and / or the top-left template is not used for deriving filter parameters. For another example, different templates for deriving filter parameters refer to multiple determinations of filter parameters. The selection from the multiple determinations of filter parameters depends on the explicit signalling at the block-level, CTU-level, slice-level, tile-level, SPS-level, PPS-level, picture-level (for example, picture header) , and / or sequence-level. For example, the size of the template depends on the block position. When the block is at the right picture boundary or any-pre-defined boundary, the top template does not include the template region outside of the current picture or any pre-defined range. When the block is at the bottom picture boundary or any-pre-defined boundary, the left template does not include the template region outside of the current picture or any pre-defined range.
[0095] In another sub-embodiment, in response to inheriting filter information from the previously-coded blocks, the filter information for the previously-coded block is used to decide the filter of the current block. The filter information includes the filter shape, filter parameters, partial filter parameters, or any combination of them. The previously-coded block can be located in the area, reconstructed before the current block, of the picture (or frame) the same as the current block or can be located in the area of the picture reconstructed before the current picture (for example, the collocated picture of the current block) . For example, the filter information is inherited from the previously-coded block through spatial adjacent candidates (including a left neighbouring block and / or an above neighbouring block) and / or spatial non-adjacent candidates, and / or history-based candidates, and / or temporal candidates, and / or propagation candidates. Spatial adjacent candidates can be from the left, above, above-left, above-right, bottom-left neighbouring blocks of the current block, any subset of the above-mentioned positions, or any combination of them.
[0096] Spatial non-adjacent candidates can be from any pre-defined positions in a search pattern around the current block, and / or any subset of the above-mentioned positions. History-based candidates can be from a history buffer which stores filter information of the previously-coded blocks. The history buffer is empty at a pre-defined schedule. For example, the history buffer is empty at the beginning or the end of a slice, CTU / CTB, CTU / CTB row, picture, tile, sequence, and / or any pre-defined unit. Temporal candidates can be from a buffer which stores the filter information at a referred reference position in the reference frame (or reference picture) and / or a pre-defined collocated picture, and / or stores the filter information at any pre-defined positions nearing the referred reference position. For example, the referred reference position is the collocated block in the collocated picture as inter prediction. For another example, the referred reference position is indicated using the motion information of the neighbouring blocks or any pre-defined blocks associated with the current block.
[0097] Propagation candidates can be from the filter information at one or more reference positions referred by the motion information of the neighbouring blocks or any pre-defined blocks associated with the current block.
[0098] In one case, a pre-defined order is specified to check all or any subset of the above-mentioned inherited candidates. If one inherited candidate cannot find the filter information, this inherited candidate is bypassed. The first available filter information is used for the current block. In another case, a list, including multiple inherited candidates, is built. For example, the list size is fixed at a pre-defined number which specifies in the standard. A pre-defined order is set to insert all or any subset of the above-mentioned inherited candidates into the list. If all or any subset of the above-mentioned inherited candidates cannot find enough filter information to put into the list, default filter information is inserted. For another example, the list size is adaptive according to how much filter information can be found using the all or any subset of the above-mention inherited candidates.
[0099] One way to select the inherited filter information from the list is as follows. If the total number of the inherited candidates in the list containing multiple inherited candidates is K, a syntax is signalled to select an inherited filter information from the original list or a reordered list. For example, the syntax is truncated unary coding. A shortest codeword is used to indicate the first candidate in the list. One example of reordering the list is as follows. The list is reordered (or sorted) according to the measurement for each candidate on the template. For example, the measurement for a candidate depends on the distortion between the reconstructed sample of the template and the predicted samples of the template which was generated by applying the filter information of this candidate to the template. The candidate with a smaller distortion on the template can be treated as a promising candidate during the measurement. The promising candidates are reordered to be put in the front of the list. One variation is that the syntax is not required to indicate a candidate from the reordered list. During the measurement, only the most promising candidate is kept and after checking each candidate in the list, the most promising candidate which has the smallest distortion is selected for the current block.
[0100] In another embodiment, when generating the to-be-predicted samples at the top-left sample of the current block, all input samples of the filtering use the reconstructed samples neighbouring to the current block.
[0101] When generating the to-be-predicted samples near the top and / or left boundary of the current block, partial of the input samples of the filtering use the reconstructed samples neighbouring to the current block and partial of the input samples of the filtering use the previously-predicted samples within the current block. When generating the to-be-predicted samples at the inner portion of the current block and / or far away from the top / left boundaries of the current block, all input samples of the filtering use the previously-predicted samples of the current block.
[0102] Fig. 11 illustrates an example of to-be-predicted samples near the top boundary of the current block. Fig. 12 illustrates an example of to-be-predicted samples at the inner portion of the current block and / or far away from the top and left boundaries of the current block.
[0103] In another embodiment, the target filter-based intra prediction mode is used to generate the luma predictor. Therefore, the current block refers to the luma component. For an example of using single tree structure, the current block is a luma coding block (CB) in a coding unit consisting of one or more luma CBs and one or more chroma CBs. For another example of using dual tree structure, the current block is a luma coding block (CB) in a coding unit consisting of one or more luma CBs.
[0104] In another embodiment, the target filter-based intra prediction mode is used to generate the chroma predictor. Therefore, the current block refers to one or more chroma components such as Cb and / or Cr. For an example of using single tree structure, the current block is a chroma coding block (CB) in a coding unit consisting of one or more luma CBs and one or more chroma CBs. For another example of using dual tree structure, the current block is a chroma coding block (CB) in a coding unit consisting of one or more chroma CBs.
[0105] (2) Classification Setting for a Target Filter-Based Intra Prediction Mode
[0106] In the original filter-based intra prediction mode, one decided filter for the current block is used to predict each to-be-predicted sample in the current block. That is, the predictor of the current block is obtained by applying only one (or called single) filter for all to-be-predicted samples in the current block. However, different to-be-predicted samples may have different correlations with the filtering inputs. In this invention, several embodiments are proposed to classify the to-be-predicted samples according to classification checking and different filters are used for different cases in the classification checking, respectively. For an example of classification checking for two cases, a threshold is set and used for each to-be-predicted sample to define different cases in the classification checking. An example is shown in Fig. 13, where the lighter samples are classified as case 1 and the darker samples are classified as class 2. In one embodiment, in response to deriving filter parameters using the template (neighbouring region) of the current block, different filters for different cases in the classification checking are derived using different samples on the template for deriving the filter parameters. In another embodiment, in response to inheriting filter information from the previously-coded blocks, the filter information may include one or more sets of {filter shape and / or filter parameters} , one or more thresholds for classifying, or any subset of the above-mentioned.
[0107] In the current block, a to-be-predicted sample satisfying that the value of a pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset of its filtering inputs is larger than the threshold will be in the first case, and a to-be-predicted sample satisfying that the value of the pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset of its filtering inputs is smaller than or equal to the threshold will be in the second case. The classifying can depend on pixel gradient, variance, position, and / or samples of the filter (or model) inputs. Therefore, those to-be-predicted samples belonging to the first case are predicted using the first filter and those to-be-predicted samples belonging to the second case are predicted using the second filter as shown in Fig. 13. After determining the filter for a to-be-predicted sample, applying the filter for the to-be-predicted sample may be unified as the original filter.
[0108] In one embodiment, the number of cases in the classification checking for a to-be-predicted sample can be a positive integer larger than 1. For example, the number of cases is equal to 2, 3, 4, or any positive integer specified in the standard. For another example, the number of cases depends on explicit signalling at a block, SPS, PPS, tile, slice, picture, and / or sequence-level flag. For another example, the number of cases depends on the block width, height, and / or area of the current block. For a larger block, the number of cases is larger. For each case in the classification checking, a filter is used to generate the predictor.
[0109] In another embodiment, one or more thresholds are determined for the classification checking. In the example of 2 cases in total for a to-be-predicted sample in the classification checking, only one threshold is required for classifying the to-be-predicted sample into either the first case or the second case.
[0110] In one sub-embodiment, the threshold can be derived for the current block. That is, for each to-be-predicted sample of the current block, the same threshold is set to classify. For example, the threshold is set as the pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset reconstructed neighbouring samples. In one way, the reconstructed neighbouring samples are the reconstructed samples on the template region for deriving filter parameters. The template region for deriving filter parameters may include top template, left template, top-left template, and / or any subset of above-mentioned templates. The threshold derivation may depend on the signalling for determining the template region for deriving filter parameters. In another way, the reconstructed neighbouring samples are the reconstructed samples on the template region determined according to the position, width, height, or area of the current block. For another example, the threshold is set as a sample value at a pre-defined position of the reconstructed neighbouring samples where the pre-defined position is the top-left sample or any position outside of the current block. The threshold for classifying can be derived depending on pixel gradient, variance, position, and / or samples of the model inputs. Fig 14 illustrates an example of threshold set as the top-left sample outside of the current block.
[0111] In another embodiment, in response to deriving filter parameters using the template (i.e., neighbouring region) of the current block, the different filters for different cases in the classification checking are derived using different samples on the template when deriving the filter parameters. When deriving the filter parameters using the template (i.e., neighbouring region) of the current block, the golden (i.e., a target for comparison) is the reconstructed samples on the template of the current block and the goal is to make the filtering result match the golden on the template. Originally, the best matching is to achieve the total distortion from each sample on the template being minimized. With the proposed classifying, the filter parameters of the first filter are obtained by considering the total distortion from those samples, which were classified into the first case on the template, being minimized. Similarly, the filter parameters of the second filter are obtained by considering the total distortion from those samples, which were classified into the second case on the template, being minimized.
[0112] In one sub-embodiment, the threshold used in classifying the template is the same as the threshold used in classifying the current block. In another embodiment, the filter shapes of different filters for different cases in the classification checking are the same. In another embodiment, the filter shapes of different filters for different cases in the classification checking can be different. In another sub-embodiment, in the template for deriving the filter parameters, a to-be-predicted template sample satisfying that the value of the pre-defined sample, gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset of its filtering inputs is larger than the threshold will be in the first case, and a to-be-predicted template sample satisfying that the value of the pre-defined sample, gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset of its filtering inputs is smaller than or equal to the threshold will be in the second case. The classifying can depend on pixel gradient, variance, position, and / or samples of the model inputs. Therefore, the first filter parameters are derived based on the matching considering the filtering results from those to-be-predicted samples belonging to the first case and the second filter parameters are derived based on the matching considering the filtering results from those to-be-predicted samples belonging to the second case. A similar way is applied to deriving filter parameters of the nth filter when more than two cases are in the classification checking. Fig. 15 illustrates an example of different filters on the template and the current block.
[0113] In another embodiment, in response to inheriting filter information from the previously-coded blocks, only one set of {filter shape and / or filter parameters} from a previously-coded block is inherited no matter whether the classification setting is applied to the previously-coded block or not. When using the filter, which was based on the inherited filter information, for the current block, the predictor of the current block is generated as original filter-based intra prediction modes without classifying.
[0114] In another embodiment, in response to inheriting filter information from the previously-coded blocks, one set of {filter shape and / or filter parameters} from a previously-coded block is treated as one inherited candidate and another set of {filter shape and / or filter parameters} from the same previously-coded block is treated as another inherited candidate. When using the selected filter among multiple inherited candidates, for the current block, the predictor of the current block is generated as original filter-based intra prediction modes without classifying.
[0115] In another embodiment, in response to inheriting filter information from the previously-coded blocks, multiple sets of {filter shape and / or filter parameters} from a previously-coded block are inherited if the classification setting is applied to the previously-coded block. One or more thresholds for classifying the current block is derived for the current block as some proposed embodiments. For example, for each to-be-predicted sample of the current block, the same threshold is set to classify. For example, the threshold is set as the pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset reconstructed neighbouring samples. In one way, the reconstructed neighbouring samples are the reconstructed samples on the template region determined according to the position, width, height, or area of the current block. For another example, the threshold is set as a sample value at a pre-defined position of the reconstructed neighbouring samples where the pre-defined position is the top-left sample or any position outside of the current block. The threshold for classifying can be derived depending on pixel gradient, variance, position, and / or samples of the model inputs. When using the filter, which was based on the inherited filter information and the derived thresholds, for the current block, the predictor of the current block is generated as the proposed filter-based intra prediction modes with classifying.
[0116] In another embodiment, in response to inheriting filter information from the previously-coded blocks, multiple sets of {filter shape and / or filter parameters} and one or more thresholds from a previously-coded block are inherited if the classification setting is applied to the previously-coded block. When using the filter, which was based on the inherited filter information, for the current block, the predictor of the current block is generated as the proposed filter-based intra prediction modes with classifying. In one sub-embodiment, the inherited threshold can be adjusted before applying the inherited threshold to classify the current block. For example, the inherited threshold can be adjusted using the temporary threshold deriving for the current block in some proposed embodiments. For example, for each to-be-predicted sample of the current block, the same temporary threshold is set. For example, the temporary threshold is set as the pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or any information of all or any subset reconstructed neighbouring samples. In one way, the reconstructed neighbouring samples are the reconstructed samples on the template region determined according to the position, width, height, or area of the current block. For another example, the temporary threshold is set as a sample value at a pre-defined position of the reconstructed neighbouring samples where the pre-defined position is the top-left sample or any position outside the current block. The temporary threshold can be derived depending on pixel gradient, variance, position, and / or samples of the model inputs. After adjusting, the inherited threshold is used to classify the current block.
[0117] (3) Enabling Conditions of the Classification Setting
[0118] For the current block using the target filter-based intra prediction mode, with the proposed classification setting, the one or more filters for the current block may be determined in response to deriving filter parameters using the template (i.e., neighbouring region) of the current block and / or in response to inheriting filter information from the previously-coded blocks. In one embodiment, for the inheriting path, whether to classify the current block depending on the inheriting information from the previously-coded block. In one sub-embodiment, if the inheritance result indicates to use classification setting for the current block, some enabling conditions for the inheriting path are checked and the classification setting can be used for the current block only when all enabling conditions are satisfied. In another sub-embodiment, the enabling conditions for the inheriting path are checked first. If all enabling conditions for the inheriting path are satisfied and the inheritance result indicates to use classification setting, the classification setting can be used for the current block. In another embodiment, for the deriving path, whether to classify the current block depends on the enabling conditions and / or some syntax elements. Several embodiments are proposed to define the enabling conditions of the classification setting. When the proposed classification setting is applied to the current block, multiple filters are used for generating the predictor of the current block. More details can reference the section “ (2) Classification Setting for a Target Filter-Based Intra Prediction Mode” .
[0119] In one embodiment, for the inheriting path, the enabling conditions are associated with the block position, block width, block height, and / or block area. For example, if the area of the current block is smaller than a size threshold, the enabling condition is not satisfied; otherwise (i.e., the area of the current block is larger than or equal to a size threshold) , this enabling condition is satisfied.
[0120] In another embodiment, for the deriving path, the proposed classification setting is always applied to the current block using the target filter-based intra prediction mode if all enabling conditions of the classification setting are satisfied.
[0121] In one sub-embodiment, the enabling conditions are associated with the block position, block width, block height, and / or block area. For example, if the current block is located at the top and / or left boundary of the current picture, CTU row, slice, tile, and / or any pre-defined region, the enabling condition is not satisfied; otherwise (i.e., the current block is not located at the top and / or left boundary of the current picture, CTU row, slice, tile, and / or any pre-defined region) , this enabling condition is satisfied. A concrete checking of the current block not at the top and left boundaries of the current picture is that the horizontal coordinate and the vertical coordinate are both larger than 0. For the current block located at the top and left boundaries, the template of the current block is not available, so the classification setting cannot be used to classify the template samples to derive multiple filters on the template. For another embodiment, if the width and / or height and / or area of the current block is smaller than a size threshold, the enabling condition is not satisfied; otherwise (i.e., the width and / or height and / or area of the current block is larger than or equal to a size threshold) , the enabling condition is satisfied. The above-mentioned examples can be combined as follows. If the block width and / or block height multiplied with the available template height and / or template width is smaller than a size threshold, the enabling condition is not satisfied; otherwise (i.e., the block width and / or block height multiplied with the available template height and / or template width is larger than or equal to a size threshold) , this enabling condition is satisfied.
[0122] In another embodiment, for the deriving path, whether to apply the proposed classification setting to the current block using the target filter-based intra prediction mode depends on the explicit signalling. When all enabling conditions are satisfied, the explicit signalling is signalled at the encoder and / or parsed at the decoder. For example, the explicit signalling refers to a flag at the block level. When the flag indicates to apply the proposed classification setting to the current block, multiple filters are used for generating the predictor of the current block. When the flag is not signalled at the encoder and / or the flag is not parsed at the decoder, the flag is inferred as the proposed classification setting not applied to the current block.
[0123] In one sub-embodiment, the enabling conditions are associated with the block position, block width, block height, and / or block area. For example, if the current block is located at the top and / or left boundary of the current picture, CTU row, slice, tile, and / or any pre-defined region, the enabling condition is not satisfied; otherwise (the current block is not located at the top and / or left boundary of the current picture, CTU row, slice, tile, and / or any pre-defined region) , this enabling condition is satisfied. A concrete checking of the current block not at the top and left boundaries of the current picture is that the horizontal coordinate and the vertical coordinate are both larger than 0. For the current block located at the top and left boundaries, the template of the current block is not available, so the classification setting cannot be used to classify the template samples to derive multiple filters on the template. For another embodiment, if the width, height, or area of the current block, or a combination of them is smaller than a size threshold, the enabling condition is not satisfied; otherwise (i.e., the width, height, or area of the current block, or a combination of them is larger than or equal to a size threshold) , the enabling condition is satisfied. The above-mentioned examples can be combined as follows. If the block width and / or block height multiplied with the available template height and / or template width is smaller than a size threshold, the enabling condition is not satisfied; otherwise (i.e., the block width and / or block height multiplied with the available template height and / or template width is larger than or equal to a size threshold) , this enabling condition is satisfied.
[0124] (4) Storage Rule of Target Filter-Based Intra Prediction Mode with Classification
[0125] In the original design of filter-based intra prediction mode without classifying, only one set of filter information {filter shape and / or filter parameters} is stored for the current block, and / or for subsequent blocks, only one set of filter information {filter shape and / or filter parameters} can be referenced. With the proposed classification setting, one or more sets of filter information {filter shape and / or filter parameters} and / or information associated with one or more thresholds for classifying can be stored for the current block and / or can be referenced by subsequent blocks.
[0126] In one embodiment, when the classification setting is used for the current block, more than one filters are used to generate the prediction of the current block. For example, two filters are used to generate the prediction of the current block. One or more sets of filter information {filter shape and / or filter parameters} and / or information associated with one or more thresholds for classifying are stored for the current block and / or referenced by subsequent blocks. In one sub-embodiment, when storing the information associated with one or more thresholds for classifying, an offset is added to adjust the values of the thresholds and / or the adjusted values will be the information for storing. In another sub-embodiment, when subsequent blocks reference the information associated with one or more thresholds for classifying, an offset is added to adjust the values of the stored thresholds and / or the adjusted values will be used for classifying subsequent blocks. In one sub-embodiment, when storing the filter information, an offset is added to adjust the values of the filter parameters and / or the adjusted values will be the information for storing. In another sub-embodiment, when subsequent blocks reference the filter information, an offset is added to adjust the values of the stored filter information and / or the adjusted values will be used for subsequent blocks. In another sub-embodiment, only the information of one or more sets of filter information {filter shape and / or filter parameters} is stored for the current block and / or referenced by subsequent blocks. The information associated with thresholds is not stored for the current block and / or referenced by subsequent blocks. In another sub-embodiment, multiple sets of filter information and / or information associated with one or more thresholds for classifying are stored for the current block, but only one or any subset of the stored information is referenced by subsequent blocks.
[0127] In another embodiment, when the classification setting is used for the current block, multiple filters are used to generate the prediction of the current block. For example, two filters are used to generate the prediction of the current block. Only one among the multiple sets of {filter shape and / or filter parameters} is stored for the current block and / or referenced by subsequent blocks.
[0128] In another embodiment, when the classification setting is used for the current block, multiple filters are used to generate the prediction of the current block. For example, two filters are used to generate the prediction of the current block. Also, two sets of filter information {filter shape and / or filter parameters} and / or the information associated with one or more thresholds for classifying are stored for the current block and / or referenced by subsequent blocks.
[0129] (5) Example of the Present Invention
[0130] An exemplary processing flow for one embodiment of the present invention is illustrated in Fig. 16. In step 1610, it determines to use a target filter-based intra prediction mode for the current block. In step 1620, it determines whether to apply a classification setting to the current block. If the classification setting is applied to the current block (i.e., the Yes path) , it uses multiple filters to generate prediction of the current block in step 1630 and it stores one or more sets of filter information and / or threshold information of the current block and / or references the stored information by subsequent blocks in step 1632. If the classification setting is not applied to the current block (i.e., the No path) , it uses single filter to generate prediction of the current block in step 1640 and it stores only one set of filter information of the current block and / or references the stored information by subsequent blocks in step 1642.
[0131] The proposed methods in this invention can be enabled and / or disabled according to implicit rules (e.g. block width, height, or area) or according to explicit rules (e.g. syntax on block, tile, slice, picture, SPS, or PPS level) . For example, the proposed method is applied when the block area is smaller / larger than a threshold.
[0132] The term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB.
[0133] Any combination of the proposed methods in this invention can be applied.
[0134] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter, intra, IBC, prediction, transform module, or a combination of them at an encoder, and / or an inter, intra, IBC, prediction, transform module, or a combination of them at a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter, intra, IBC, prediction, transform module, or a combination of them at the encoder, and / or the inter, intra, IBC, prediction, transform module, or a combination of them at the decoder, so as to provide the information needed by the inter, intra, IBC, prediction, or transform module, or a combination of them.
[0135] With reference to the encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods can be implemented in an inter / intra / prediction / transform module (e.g. Intra Pred. 110 in Fig. 1A) of an encoder, and / or an inter / intra / prediction / transform module (e.g. Intra Pred. 150 in Fig. 1B) of a decoder.
[0136] Fig. 17 illustrates a flowchart of an exemplary video coding system that classifies to-be-predicted samples into multiple classes and uses multiple filters according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side and / or decoder side. The steps shown in the flowchart may also be implemented based on hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block are received in step 1710, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side. To-be-predicted samples of the current block are classified into multiple classes in step 1720. The current block is encoded or decoded by using multiple filters in step 1730, wherein the to-be-predicted samples of the current block in a target class are predicted using one target filter selected from the multiple filters according to the target class.
[0137] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0138] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0139] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0140] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;classifying to-be-predicted samples of the current block into multiple classes; andencoding or decoding the current block by using multiple filters, wherein the to-be-predicted samples of the current block in a target class are predicted using one target filter selected from the multiple filters according to the target class.2.The method of Claim 1, wherein the current block is coded using EIP (Extrapolation Filter-Based Intra Prediction) .3.The method of Claim 1, wherein whether to apply classification setting to the current block is determined first, and said classifying the to-be-predicted samples of the current block is applied if classification setting is applied to the current block.4.The method of Claim 1, wherein one or more sets of filter information and / or one or more sets of threshold information of the current block are stored for the current block and / or referenced by one or more subsequent blocks.5.The method of Claim 4, wherein said one or more sets of filter information comprise filter shape, filter parameters, partial filter parameters, or a combination thereof.6.The method of Claim 1, wherein input samples for the multiple filters comprise reconstructed and / or predicted samples associated with the current block, and / or reconstructed and / or predicted samples associated with a reference region reconstructed prior to the current block.7.The method of Claim 1, wherein input samples for at least one of the multiple filters are within an MxN region around a current to-be-predicted sample at (x, y) excluding locations with horizontal coordinate > x and / or vertical coordinate > y, and wherein M and N are positive integers.8.The method of Claim 1, wherein one or more filter parameters for at least one of the multiple filters are derived using one or more templates of the current block.9.The method of Claim 8, wherein said one or more filter parameters for at least one of the multiple filters are derived by a regression-based process.10.The method of Claim 9, wherein the regression-based process is unified with a process of deriving cross-component models for one or more cross-component chroma modes.11.The method of Claim 9, wherein the regression-based process corresponds to Gaussian elimination.12.The method of Claim 1, wherein different filters are used for the to-be-predicted samples of the current block belonging to different classes.13.The method of Claim 12, wherein different filter shapes are used for the to-be-predicted samples of the current block belonging to different classes.14.The method of Claim 1, wherein said classifying the to-be-predicted samples of the current block into the multiple classes is based on pre-defined sample, pixel gradient, variance, average, mean, medium, minimum, maximum, or a combination thereof of one or more filtering inputs.15.The method of Claim 1, wherein filter information for at least one of the multiple filters is inherited from one or more previously-coded blocks.16.The method of Claim 1, wherein when filter information for the current block is inherited from one or more previously-coded blocks, whether the current block is classified into the multiple classes depends on the filter information for the current block inherited from said one or more previously-coded blocks.17.The method of Claim 1, wherein when filter information for the current block is inherited from one or more previously-coded blocks, whether the current block is classified into the multiple classes depends on one or more enabling conditions associated with block position, block width, block height, block area, or a combination thereof.18.The method of Claim 1, wherein when filter information for the current block is derived, whether the current block is classified into the multiple classes depends on enabling conditions and / or one or more signalled or parsed syntax elements.19.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprise pixel data to be encoded at an encoder side or coded data associated with the current block to be decoded at a decoder side;classify to-be-predicted samples of the current block into multiple classes; andencode or decode the current block by using multiple filters, wherein the to-be-predicted samples of the current block in a target class are predicted using one target filter selected from the multiple filters according to the target class.
Citation Information
Patent Citations
Method of Filter Control for Block-Based Adaptive Loop Filtering
US20160261863A1
Adaptive Bilateral Filter In Video Coding
WO2022268184A1
Filter coefficient generation method, filtering method, video encoding method and apparatuses, video decoding method and apparatuses, and video encoding and decoding system
WO2023123512A1
Improved cross-component prediction for video coding
WO2023225013A1