Regression-based matrix-based intra prediction

Regression-based matrix-based intra-prediction methods address the inefficiencies in existing video coding standards by using neighboring L-shaped regions to derive intra-prediction matrices, enhancing prediction accuracy and compression efficiency in video coding.

WO2025152878A1PCT designated stage expired Publication Date: 2025-07-24MEDIATEK INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/071947
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-13
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in efficiently predicting pixel blocks, particularly in high-efficiency video coding (HEVC) and versatile video coding (VVC), as they rely on hybrid block-based methods that are not optimized for various local motion and texture characteristics, leading to suboptimal compression and prediction accuracy.

Method used

The implementation of regression-based matrix-based intra-prediction (MIP) methods that utilize training samples from neighboring L-shaped regions to derive an intra-prediction matrix through regression, allowing for matrix multiplication and bias term generation to enhance prediction accuracy.

Benefits of technology

Improves prediction accuracy and compression efficiency by leveraging regression-based matrix methods, which adapt to diverse local motion and texture characteristics, resulting in better video encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071947_24072025_PF_FP_ABST
    Figure CN2025071947_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of using regression-based matrix-based intra-prediction (MIP) to encode or decode pixel blocks is provided. A video coder receives data to be encoded or decoded as a current block of pixels of a current picture of a video. The video coder identifies corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape. The corresponding input and output training samples may include samples identified from one or more predefined positions relative to the current block. The video coder derives an intra-prediction matrix by performing regression based on the identified training samples. The video coder performs matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block. The video coder encodes or decodes the current block by using generated predictor.
Need to check novelty before this filing date? Find Prior Art

Description

REGRESSION-BASED MATRIX-BASED INTRA PREDICTIONCROSS REFERENCE TO RELATED PATENT APPLICATION (S)The present disclosure is part of a non-provisional application that claims the priority benefit of U.S. Provisional Patent Application No. 63 / 622,742 filed on 19 January 2024. Content of the above-listed application is herein incorporated by reference.TECHNICAL FIELDThe present disclosure relates generally to video coding. In particular, the present disclosure relates to methods of coding pixel blocks by matrix-based intra prediction.BACKGROUNDUnless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) . HEVC is based on the hybrid block-based motion-compensated DCT-like transform coding architecture. The basic unit for compression, termed coding unit (CU) , is a 2Nx2N square block of pixels, and each CU can be recursively split into four smaller CUs until the predefined minimum size is reached. Each CU contains one or multiple prediction units (PUs) .Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from the reconstructed signal, which is derived from the coded picture regions. The prediction residual signal is processed by a block transform. The transform coefficients are quantized and entropy coded together with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform on the de-quantized transform coefficients. The reconstructed signal is further processed by in-loop filtering for removing coding artifacts. The decoded pictures are stored in the frame buffer for predicting the future pictures in the input video signal.In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors (MVs) and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.A CTU can be partitioned into one or multiple non-overlapped coding units (CUs) using the quadtree (QT) with nested multi-type-tree (MTT) structure to adapt to various local motion and texture characteristics. A CU can be further split into smaller CUs using one of the five split types: quad-tree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical center-side triple-tree partitioning, horizontal center-side triple-tree partitioning.Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signaling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU. Each CU may contain one or more transform units (TUs) for representing the prediction residual blocks. A transform unit (TU) is comprised of a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples and each TB correspond to one residual block of samples from one color component. An integer transform is applied to a transform block. The level values of quantized coefficients together with other side information are entropy coded in the bitstream. The terms coding tree block (CTB) , coding block (CB) , prediction block (PB) , and transform block (TB) are defined to specify the 2-D sample array of one-color component associated with CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship is valid for CU, PU, and TU.For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information are used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.Intra block copy (IBC) or current picture referencing (CPR) refer to coding pixel blocks by referencing pixel positions within same current picture as the current block by using block vectors.SUMMARYThe following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce concepts, highlights, benefits and advantages of the novel and non-obvious techniques described herein. Select and not all implementations are further described below in the detailed description. Thus, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.Some embodiments of the disclosure provide methods of using regression-based matrix-based intra-prediction (MIP) to encode or decode pixel blocks. A video coder receives data to be encoded or decoded as a current block of pixels of a current picture of a video. The video coder identifies corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape. The input training samples are from the neighboring L-shape of the reconstructed block and the output (target) training samples are from the reconstructed region of pixels. The video coder derives an intra-prediction matrix by performing regression based on the identified training samples. The video coder performs matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block. The video coder encodes or decodes the current block by using generated predictor.In some embodiments, the training samples are identified from a plurality of reconstructed regions and their neighboring L-shape samples. Each reconstructed region may have the same or different dimensions as the current block. In some embodiments, the output training samples are identified by down-sampling the reconstructed region.In some embodiments, the corresponding input and output training samples may include samples that are identified from one or more predefined positions, such as non-adjacent positions, relative to the current block. In some embodiments, the corresponding input and output training samples may include samples identified from positions in one or more predefined search regions of the current picture, such as positions identified from a list of candidates for intra template matching prediction (IntraTMP) . In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in a reference picture that is a located by a motion vector or is a predefined collocated picture. In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in the current picture that is located by a block vector (e.g., for IBC mode. )In some embodiments, the video coder derives the intra-prediction matrix by deriving one or more linear models performs the matrix multiplication by applying the samples neighboring the current block as input to the linear models. The video coder may generate the predictor by adding a bias term, and the regression for deriving the intra-prediction matrix is also used to derive a coefficient for the bias term.In some embodiments, when a first flag is signaled (in the bitstream) to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, a second flag may be signaled to indicate whether to perform the regression based on the training samples for performing the matrix multiplication, or whether the coefficients of the matrix are derived by regression based on the training samples. In some embodiments, when a first flag is signaled to indicate whether intra template matching prediction is used to generate a predictor for the current block, a second flag may be signaled to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, and the coefficients are derived by the regression based on the training samples. In some embodiments, whether to perform the regression based on the training samples for performing the matrix multiplication is determined based on at least one of a size of the current block, a ratio of the current block, a slice type of the current block, and a color component of the current block (e.g., whether the current block is a luma block or a chroma block. )BRIEF DESCRIPTION OF THE DRAWINGSThe accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of the present disclosure. The drawings illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.FIG. 1 illustrates 67 intra predictions modes, including 65 directional or angular intra prediction modes.FIGS. 2A-D illustrate corresponding reference samples (R (x, -1) and R (-1, y) ) for PDPC applied over various angular intra-prediction modes.FIG. 3 illustrates predefined search area for intra template matching.FIG. 4 illustrates matrix-based intra prediction (MIP) process.FIG. 5 conceptually illustrates collection of data samples for training MIP linear models.FIG. 6 illustrates an example video encoder that may implement MIP.FIG. 7 illustrates portions of the video encoder that implement regression-based MIP.FIG. 8 conceptually illustrates a process for using regression-based MIP to encode pixel blocks.FIG. 9 illustrates an example video decoder that may implement MIP.FIG. 10 illustrates portions of the video decoder that implement regression-based MIP.FIG. 11 conceptually illustrates a process for using regression-based MIP to decode pixel blocks.FIG. 12 conceptually illustrates an electronic system with which some embodiments of the present disclosure are implemented.DETAILED DESCRIPTIONIn the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure.I. Intra PredictionA. Intra Prediction Modes and MPMsIntra-prediction method exploits one or more reference lines adjacent to the current prediction unit (PU) and one of the intra-prediction modes to generate the predictors for the current PU. The Intra-prediction direction can be chosen among a mode set containing multiple prediction directions, DC mode, and Planar mode. The intra prediction mode may also refer to any intra mode which determines the predictor of the current block using the spatially reconstructed samples. The number of directional intra modes may be 33 or extended to 65 direction modes. By including DC and Planar modes, the number of intra-prediction mode is 35 (or 67) . Some intra-prediction modes (e.g., 3 or 5) are identified as a set of most probable modes (MPM) for intra-prediction in current prediction block so an index may be signaled to select one of the MPMs. FIG. 1 illustrates 67 intra predictions modes, including 65 directional or angular intra prediction modes (from 2 to 66) .B. Position Dependent Intra Prediction Combination (PDPC)In VVC, the results of intra prediction of DC, planar and several angular modes may be further modified by a position dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method which invokes a combination of the boundary reference samples and HEVC style intra prediction with filtered boundary reference samples. PDPC may be applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. If the current block is Bdpcm mode or multi-reference line (MRL) index is larger than 0, PDPC is not applied.The prediction sample pred (x’ , y’ ) is predicted using an intra prediction mode (DC, planar, angular) and a linear combination of reference samples according to the following:pred (x’ , y’ ) = Clip (0, (1<<BitDepth ) –1, (wL×R-1, y + wT×Rx, -1 + (64-wL-wT) × pred (x’ , y’ ) + 32 ) >>6)where Rx, -1, R-1, y represent the corresponding reference samples located at the top and left boundaries of current sample (x, y) , respectively by an intra-prediction angular mode. wL and wT are weighting factors. The pred (x’ , y’ ) at the right side of the equation is the value of the prediction sample before PDPC modification, and pred (x’ , y’ ) at the left side of the equation is its modified value.If PDPC is applied to DC, planar, horizontal, and vertical intra modes, additional boundary filters are not needed, as required in the case of HEVC DC mode boundary filter or horizontal / vertical mode edge filters. PDPC process for DC and Planar modes is identical. For angular intra prediction modes, if the current angular mode is HOR_IDX or VER_IDX, left or top reference samples is not used, respectively. The PDPC weights and scale factors are dependent on prediction modes and the block sizes. PDPC is applied to the block with both width and height greater than or equal to 4.FIGS. 2A-D illustrate corresponding reference samples (Rx, -1 and R-1, y) for PDPC applied over various angular intra-prediction modes. FIG. 2A shows the reference samples for diagonal top-right mode. FIG. 2B shows the reference samples for diagonal bottom-left mode. FIG. 2C shows the reference samples for adjacent diagonal top-right mode. FIG. 2D shows the reference samples for adjacent diagonal bottom-left mode.The prediction sample pred (x’ , y’ ) is located at (x’ , y’ ) within the prediction block. As an example, the coordinate x of the reference sample Rx, -1 is given by: x = x’ + y’ + 1, and the coordinate y of the reference sample R-1, y is similarly given by: y = x’ + y’ + 1 for the diagonal modes. For the other angular mode, the reference samples Rx, -1 and R-1, y can be located in fractional sample position. In this case, the sample value of the nearest integer sample location is used.C. Intra Prediction FusionThe intra prediction fusion method derives predicted samples as a weighted combination of multiple predictors generated from different reference lines. In this process, multiple intra predictors are generated and then fused by weighted averaging. The process of deriving the predictors to be used in the fusion process is described as follows:(1) For angular intra prediction modes including the single mode case of TIMD and DIMD, the intra prediction fusion method derives intra prediction by weighting intra predictions obtained from multiple reference lines represented as pfusion=w0pline+w1pline+1, where pline is the intra prediction from the default reference line and pline+1 is the prediction from the line above the default reference line. The weights are set as w0=3 / 4 and w1=1 / 4.(2) For TIMD mode with blending, pline is used for the first mode (w0=1, w1=0) and pline+1 is used for the second mode (w0=0, w1=1) .(3) For DIMD mode with blending, the number of predictors selected for a weighted average is increased from 3 to 6.Intra prediction fusion may be applied to luma blocks when angular intra mode has non-integer slope (required reference samples interpolation) and the block size is greater than 16, it is used with MRL and not applied for ISP coded blocks. PDPC may be applied to the intra prediction mode using the closest to the current block reference line.D. Intra Template Matching (IntraTMP)Intra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side.The prediction signal is generated by matching the L-shaped, Top-only or Left-only causal neighbor of the current block with another block in a predefined search area. FIG. 3 illustrates predefined search area for intra template matching. The figure shows the reconstructed samples from the top and left CTUs as well as part of the reconstructed samples within the current CTU that are located above, left, bottom-left and top-right to the current block 310. As illustrated, there are 6 predefined search areas, i.e., R1 to R6.A given search order of the 6 regions is utilized, i.e., R4, R5, R6, R1, R2, and R3. Within each region, the decoder constructs a candidate list of up to 19 template matching block vectors that are ranked in ascending order according to the template cost (sum of absolute differences or SAD is used as a cost function. ) The following modes are supported:● Single predictor: A single predictor is selected from the candidate list.● Fusion of multiple predictors: multiple predictors are blended to derive the final prediction block. The blending weights are either computed from the template matching cost of each predictor, or with Wiener-filter based weight derivation method.● Sub-pel precision: When single predictor is used, sub-pel precision can be used with 1 / 2-pel precision, 1 / 4-pel precision and 3 / 4-pel precision, each with 8 possible directions.● Linear filter model: A linear filter can be learned between the reference template and current template and be applied to reference block. This mode can be used for single predictor when sub-pel precision is not used.The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportional to the block dimension (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:SearchRange_w = min (64, a*BlkW)SearchRange_h = min (64, a*BlkH)Where ‘a’ is a constant that controls the gain / complexity trade-off. In practice, ‘a’ is equal to 5. To speed-up the template matching process, the search range of all search regions is subsampled by a factor of 3. After finding the best match, a refinement process is performed. The refinement is done via a second template matching search around the best match with a reduced range.The Intra template matching tool is enabled for CUs with size less than or equal to 64 in width and height. This maximum CU size for Intra template matching is configurable. The Intra template matching prediction mode is signaled at CU level through a dedicated flag when DIMD is not used for current CU.E. Decoder Side Intra Mode Derivation (DIMD)Decoder-Side Intra Mode Derivation (DIMD) is a technique in which one or more, for example, two, intra prediction modes such as angles or directions are derived from the reconstructed neighbor samples (template) of a block, and those two predictors are combined with the non-angular predictor such as planar mode predictor with the weights derived from the gradients. The DIMD mode is used as an alternative prediction mode and / or is always checked in high-complexity RDO mode. To implicitly derive the intra prediction modes of a block, a texture gradient analysis is performed at both encoder and decoder sides. This process starts with an empty Histogram of Gradient (HoG) having 65 entries, corresponding to the 65 angular / directional intra prediction modes. Amplitudes of these entries are determined during the texture gradient analysis.In some embodiments, when DIMD is applied, up to five intra modes are derived from the reconstructed neighbor samples, and those five predictors are combined with the planar mode predictor with the weights derived from the histogram of gradients. The division operations in weight derivation are performed utilizing a lookup table (LUT) based integerization scheme. For example, the division operation in the orientation calculation Orient = Gy  / Gx may be computed by the following LUT-based scheme:x = Floor (Log2 (Gx) )normDiff = ( (Gx<< 4 ) >> x) &15x += (3 + (normDiff ! = 0) ? 1 : 0)Orient = (Gy* (DivSigTable [normDiff] | 8) + (1<< (x-1) ) ) >> xwhereDivSigTable

[0016] = {0, 7, 6, 5 , 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} .For a block of size W × H, the weight for each of the five derived modes is modified if the one the above or left histogram magnitudes is twice larger than the other one. In this case, the weights are location dependent and computed as follows:If the above histogram is twice the left, then:If the left histogram is twice the above, then:where wDimdi is the unmodified uniform weight of the DIMD selected, and Δi is pre-defined and set to 10. Derived intra modes are included into the primary list of intra most probable modes (MPM) , so the DIMD process is performed before the MPM list is constructed. The primary derived intra mode of a DIMD block is stored with a block and is used for MPM list construction of the neighboring blocks.Depending on reconstructed samples availability, the region of decoded reference samples of current W×H luma CB may be extended towards the above-right side if available, up to W additional columns. It is extended towards the bottom-left side if available, up to H additional rows.F. Template-based Intra mode Derivation (TIMD)For mode selection, template matching method can be applied by computing the cost between reconstructed samples and predicting samples. One of the examples is template-based intra mode derivation (TIMD) . TIMD is a coding method in which the intra prediction mode of a CU is implicitly derived by using a neighboring template at both encoder and decoder, instead of the encoder signaling the exact intra prediction mode to the decoder.For each intra prediction mode in MPMs, as well as the wide-angle modes if the above-right and / or bottom-left reference samples are available, SATD between the prediction and reconstruction samples of the template is calculated as cost. First two intra prediction modes with the minimum SATD are selected as the TIMD modes. These two TIMD modes are fused with the weights after applying PDPC process, and such weighted intra prediction is used to code the current CU. Position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD modes.The costs of the two selected modes are compared with a threshold, the cost factor of 2 is applied as follows:costMode2 < 2*costMode1.If this condition is true, the fusion is applied, otherwise the only mode1 is used. Weights of the modes are computed from their SATD costs as follows:weight1 = costMode2  /  (costMode1+ costMode2)weight2 = 1 -weight1The division operations may be conducted using a lookup table (LUT) based integerization scheme.G. Regression-based GPM BlendingRegression-based GPM blending mode is designed as an additional GPM implicit mode, where the two integer blending matrices (W0 and W1) are derived from the template (1 line above, 1 column left) . The blending matrices are modelled as an affine linear function of the sample positions (x, y) in the current CU:W0 (x, y) = a. x + b. y + cW1 (x, y) = 1 -W0 (x, y)The parameters (a, b, c) are derived from the reference template using a solver that minimizes mean-square-error (MSE) between corresponding input samples and output samples in the template region neighboring the current block. The MSE minimization is performed by calculating autocorrelation matrix for the input samples and a cross-correlation vector (with parameters a, b, c) between the input and output samples. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The autocorrelation matrix is calculated using the reconstructed values of luma and / or chroma samples.A list of pair of candidates is built from the regular GPM candidates and re-ordered with the template cost. The GPM implicit mode is signaled by a CU-level flag (gpm_implicit_flag) . If gpm_implicit_flag is true, a merge-idx is coded to signal the pair of GPM candidates to be used. If gpm_implicit_flag is false, the regular GPM syntax elements are signaled.II. Matrix-based Intra Prediction (MIP)Matrix-based intra prediction (MIP) method is an intra prediction technique. For predicting the samples of a rectangular pixel block of width W and height H, the video coder performing matrix-based intra prediction (MIP) takes one line of H reconstructed neighboring boundary samples from the left side of the block and one line of W reconstructed neighboring boundary samples from the upper side of the block as input. If the reconstructed samples are unavailable, they are generated as it is done in the conventional intra prediction. The generation of the prediction signal is based on (1) averaging, (2) matrix vector multiplication, and / or (3) linear interpolation.FIG. 4 illustrates matrix-based intra prediction process for a current block 400 using its neighboring boundary samples 415. As illustrated, the first step (averaging) averages the current block’s reconstructed neighboring boundary samples bdrytop and bdryleft to generate the input vector bdryred 425 for the MIP process. The second step (matrix vector multiplication) multiplies a matrix A with the input vector of boundary samples (bdryred) and then adding the bias bk to generate an initial MIP predictor 410. The third step (linear interpolation) performs interpolation or upsampling based on the initial MIP predictor 410 and boundary samples to generate a final MIP prediction 420.More generally, for some embodiments, to generate the MIP predictor, a MIP predictor predMIP is obtained by multiplying the input vector (denoted as p) and the matrix (denoted as A) and then adding the bias value b. If the dimension of the current block is not equal to the dimension of MIP predictor, another up-sampling process is required to get the final predictor predfinal.predMIP = p *A + bpredfinal = upsample (predMIP)The above matrix equation can be rewritten as the formula below, where pn is the n-th element of the input vector p, an, i is the element in the n-th row and the i-th column of the matrix A, and predMIP (i) is the i-th element in the MIP predictor predMIP.From the above equation, the linear relation between each element of MIP predictor predMIP (i) and the input vector p is observed. If there are N elements in the MIP predictor predMIP, the matrix equation may also be represented by N linear models.In some embodiments, the matrices of MIP modes are off-line trained and pre-defined, which may not be able to generate good predictions for all blocks. To improve prediction, some embodiments of the disclosure provide a regression-based MIP method which uses samples from neighboring templates as training pairs to train the matrix A. For example, in some embodiments, the elements of the matrix A (i.e., an, i) are derived from the reference template using a solver that minimizes mean-square-error (MSE) between corresponding input samples and output samples in collected training pairs. The MSE minimization is performed by calculating autocorrelation matrix for the input samples and a cross-correlation vector between the input and output samples. Autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back-substitution. The autocorrelation matrix is calculated using the reconstructed values of luma and / or chroma samples. For another example, in some embodiments, the MIP matrix equation may be derived by solving the N linear models.In some embodiments, if the regression-based MIP mode is selected, the matrix A used to generate the MIP predictor is implicitly derived by regression-based method rather than reading / loading a pre-defined matrix from memory.In some embodiments, if the regression-based MIP mode is selected, N linear models used to generate the MIP predictor are implicitly derived by a regression-based method. For each linear models, the input samples to be applied are derived from the neighboring reconstructed samples of the current block, and the corresponding output is one sample of the MIP predictor predMIP.In some embodiments, the dimension of the input vector of MIP regression model may be equal to the dimension of the input vector of MIP mode in VVC. In some embodiments, the dimension of the input vector of MIP regression model could be larger than the dimension of the input vector of MIP mode in VVC. In one embodiment, the bias term b in the MIP regression model may have a coefficient and it may be derived by the regression process with other parameters together.In some embodiments, multiple training pairs may be found from the neighboring template region of the current block and be used in the regression process of multiple MIP linear models. For example, several blocks with the same dimension as the current block and their neighboring L-shape samples may be found in the neighboring template region. The input data of one training pair may be derived from a collected L-shape samples as MIP process, and the output target of that training pair may be derived from that collected block. If the dimension of the MIP predictor is not equal to the dimension of the current block, a down-sampling operation is applied to that collected block to have a same dimension as the MIP predictor and the output target is the down-sampled collected block.In some embodiments, multiple training pairs may be found from the pre-defined non-adjacent positions (not adjacent to the current block) and be used in the regression process of multiple MIP linear models. In some embodiments, all or some of the candidates in the IntraTMP candidate list (i.e., those identified from the search regions described by reference to FIG. 3 above) may be used as (some or all of) the multiple training pairs used in the regression process of multiple MIP linear models. In some embodiments, the multiple training pairs can be from the reference blocks in the reference pictures or any pre-defined collocated pictures of the current picture which is different from the current picture.In some embodiments, if the current block is predicted by intra block copy (IBC) or intra TMP, the multiple training pairs may be determined by using the block vector information of the current block. In some embodiments, the multiple training pairs may be determined by using the block vector information and / or motion vector information from neighboring spatial adjacent or non-adjacent blocks / positions in the current picture / slice.FIG. 5 conceptually illustrates collection of data samples for training the MIP linear models. The figure shows a current block 510. To determine the MIP linear model (s) for intra-predicting the current block, training pairs may be collected from blocks or regions 520, 530, 540, 550, 560, 570 and their corresponding neighboring L-shape samples 525, 535, 545, 555, 565, 575. For example, the input data of one training pair may be derived from the L-shape samples 555, and the output target of that training pair may be derived from its corresponding block 550. In the example, the blocks or regions 520-560 (and their corresponding neighboring L-shape samples 525-565) are in the current picture and may be reconstructed samples identified by block vectors of IBC, or candidates of IntraTMP, or spatially non-adjacent predefined positions, or other predefined positions. The block 570 and its neighboring L-shape samples 575 are in a reference picture that may be located by a motion vector of the current block, or is a predefined collocated picture.In one embodiment, the regression-based MIP method could be a sub-mode of original MIP mode. If MIP flag is true for the current block, a regression-based MIP flag may be further signaled to indicate whether the regression-based MIP method is used or not. If the regression-based MIP flag is false, the MIP index and MIP transpose flag may be signaled. Otherwise, the MIP index and MIP transpose flag may be skipped.In some embodiments, the regression-based MIP method may be a sub-mode of IntraTMP. If IntraTMP flag is true for the current block, a regression-based MIP flag may be further signaled to indicate whether the proposed regression-based MIP method is used or not. If the regression-based MIP flag is true, all or some of the candidates in the IntraTMP candidate list may be used in the regression process of multiple MIP linear models, and the final predictor may be generated by multiple MIP linear models.In some embodiments, whether the regression-based MIP approach is allowed for the current block could depend on the CU size, CU ratio, slice type or color component of the current block. In some embodiments, there may be a high-level syntax signaled in SPS, PPS, PH or SH to indicate whether the regression-based MIP method is allowed for the current sequence, picture, or slice.Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / prediction module of an encoder, and / or an inter / intra / prediction module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder, so as to provide the information needed by the inter / intra / prediction module.III. Example Video EncoderFIG. 6 illustrates an example video encoder 600 that may implement matrix-based intra-prediction (MIP) . As illustrated, the video encoder 600 receives input video signal from a video source 605 and encodes the signal into bitstream 695. The video encoder 600 has several components or modules for encoding the signal from the video source 605, at least including some components selected from a transform module 610, a quantization module 611, an inverse quantization module 614, an inverse transform module 615, an intra-picture estimation module 624, an intra-prediction module 625, a motion compensation module 630, a motion estimation module 635, an in-loop filter 645, a reconstructed picture buffer 650, a MV buffer 665, and a MV prediction module 675, and an entropy encoder 690. The motion compensation module 630 and the motion estimation module 635 are part of an inter-prediction module 640. The intra-prediction module 625 and the intra-prediction estimation module 624 are part of a current picture prediction module 620, which uses current picture reconstructed samples as reference samples for prediction of the current block.In some embodiments, the modules 610 –690 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 610 –690 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 610 –690 are illustrated as being separate modules, some of the modules can be combined into a single module.The video source 605 provides a raw video signal that presents pixel data of each video frame without compression. A subtractor 608 computes the difference between the raw video pixel data of the video source 605 and the predicted pixel data 613 from the motion compensation module 630 or intra-prediction module 625 as prediction residual 609. The transform module 610 converts the difference (or the residual pixel data or residual signal 608) into transform coefficients (e.g., by performing Discrete Cosine Transform, or DCT) . The quantization module 611 quantizes the transform coefficients into quantized data (or quantized coefficients) 612, which is encoded into the bitstream 695 by the entropy encoder 690.The inverse quantization module 614 de-quantizes the quantized data (or quantized coefficients) 612 to obtain transform coefficients 618, and the inverse transform module 615 performs inverse transform on the transform coefficients 618 to produce reconstructed residual 619. The reconstructed residual 619 is added with the predicted pixel data 613 to produce reconstructed pixel data 617. In some embodiments, the reconstructed pixel data 617 is temporarily stored in a line buffer 627 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 645 and stored in the reconstructed picture buffer 650. In some embodiments, the reconstructed picture buffer 650 is a storage external to the video encoder 600. In some embodiments, the reconstructed picture buffer 650 is a storage internal to the video encoder 600.The intra-picture estimation module 624 performs intra-prediction based on the reconstructed pixel data 617 to produce intra prediction data. The intra-prediction data is provided to the entropy encoder 690 to be encoded into bitstream 695. The intra-prediction data is also used by the intra-prediction module 625 to produce the predicted pixel data 613.The motion estimation module 635 performs inter-prediction by producing MVs to reference pixel data of previously decoded frames stored in the reconstructed picture buffer 650. These MVs are provided to the motion compensation module 630 to produce predicted pixel data.Instead of encoding the complete actual MVs in the bitstream, the video encoder 600 uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 695.The MV prediction module 675 generates the predicted MVs based on reference MVs that were generated for encoding previously video frames, i.e., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 675 retrieves reference MVs from previous video frames from the MV buffer 665. The video encoder 600 stores the MVs generated for the current video frame in the MV buffer 665 as reference MVs for generating predicted MVs.The MV prediction module 675 uses the reference MVs to create the predicted MVs. The predicted MVs can be computed by spatial MV prediction or temporal MV prediction. The difference between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame (residual motion data) are encoded into the bitstream 695 by the entropy encoder 690.The entropy encoder 690 encodes various parameters and data into the bitstream 695 by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding. The entropy encoder 690 encodes various header elements, flags, along with the quantized transform coefficients 612, and the residual motion data as syntax elements into the bitstream 695. The bitstream 695 is in turn stored in a storage device or transmitted to a decoder over a communications medium such as a network.The in-loop filter 645 performs filtering or smoothing operations on the reconstructed pixel data 617 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 645 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.FIG. 7 illustrates portions of the video encoder 600 that implement regression-based MIP. For some embodiments, the figure illustrates the components of the intra-prediction module 620 of the video encoder 600. As illustrated, the intra-prediction module 620 uses neighboring samples 705 of the current block provided by the line buffer 627 as input samples. A matrix multiplication module 730 performs matrix multiplication on the input samples to generate a MIP predictor 740 as the predicted pixel data 613.The matrix multiplication 730 uses matrix parameters 725 as elements of the matrix or as linear models to perform the matrix multiplication. The matrix parameters 725 may also be generated by a regression module 720, which performs data regression on training samples 715 to generate the matrix parameters.A training sample collector 710 provides the training samples 715 by following certain predetermined rules to identify and collect specific sets of reconstructed samples stored in the reconstructed picture buffer 650 or the line buffer 627 as the training samples. Specifically, the training sample collector 710 identifies one or more blocks and their corresponding neighboring L-shapes at certain positions as the source of the corresponding training data, with the training output target sample (s) collected from the block and the training input samples collected from the block’s neighboring L-shape. Examples of the positions from which to identify the blocks and their neighboring L-shapes includes neighboring template regions, intraTMP candidates, non-adjacent neighbors, IBC reference blocks identified by block vectors, reference blocks in collocated pictures or reference pictures identified by motion vectors, or other predefined positions.FIG. 8 conceptually illustrates a process 800 for using regression-based MIP to encode pixel blocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the encoder 600 performs the process 800 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the encoder 600 performs the process 800.The encoder receives (at block 810) data to be encoded as a current block of pixels in a current picture. The encoder identifies (at block 820) corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape. The input training samples are from the neighboring L-shape of the reconstructed region and the output (target) training samples are from the reconstructed region of pixels.In some embodiments, the training samples are identified from a plurality of reconstructed regions and their neighboring L-shape samples, as described by reference to FIG. 5. Each reconstructed region may have the same or different dimensions as the current block. In some embodiments, the output training samples are identified by down-sampling the reconstructed region.In some embodiments, the corresponding input and output training samples may include samples that are identified from one or more predefined non-adjacent positions relative to the current block. In some embodiments, the corresponding input and output training samples may include samples identified from positions in one or more predefined search regions of the current picture, such as positions identified from a list of candidates for IntraTMP as described by reference to FIG. 3. In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in a reference picture that is a located by a motion vector or is a predefined collocated picture. In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in the current picture that is located by a block vector (e.g., for IBC mode. )The encoder derives (at block 830) an intra-prediction matrix by performing regression based on the identified training samples. In some embodiments, the encoder derives the intra-prediction matrix by deriving one or more linear models performs the matrix multiplication by applying the samples neighboring the current block as input to the linear models. The encoder may generate the predictor by adding a bias term, and the regression for deriving the intra-prediction matrix is also used to derive a coefficient for the bias term.In some embodiments, when a first flag is signaled (in the bitstream) to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, a second flag may be signaled to indicate whether to perform the regression based on the training samples for performing the matrix multiplication, or whether the coefficients of the matrix are derived by regression based on the training samples. In some embodiments, when a first flag is signaled to indicate whether intra template matching prediction is used to generate a predictor for the current block, a second flag may be signaled to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, and the coefficients are derived by the regression based on the training samples. In some embodiments, whether to perform the regression based on the training samples for performing the matrix multiplication is determined based on at least one of a size of the current block, a ratio of the current block, a slice type of the current block, and a color component of the current block (e.g., whether the current block is a luma block or a chroma block. )The encoder performs (at block 840) matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block. The encoder encodes (at block 850) the current block by using generated predictor to produce prediction residuals.IV. Example Video DecoderIn some embodiments, an encoder may signal (or generate) one or more syntax element in a bitstream, such that a decoder may parse said one or more syntax element from the bitstream.FIG. 9 illustrates an example video decoder 900 that may implement matrix-based intra-prediction (MIP) . As illustrated, the video decoder 900 is an image-decoding or video-decoding circuit that receives a bitstream 995 and decodes the content of the bitstream into pixel data of video frames for display. The video decoder 900 has several components or modules for decoding the bitstream 995, including some components selected from an inverse quantization module 914, an inverse transform module 915, an intra-prediction module 925, a motion compensation module 930, an in-loop filter 945, a decoded picture buffer 950, a MV buffer 965, a MV prediction module 975, and a parser 990. The motion compensation module 930 is part of an inter-prediction module 940. The intra-prediction module 925 is part of a current picture prediction module 920, which uses current picture reconstructed samples as reference samples for prediction of the current block.In some embodiments, the modules 914 –990 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device. In some embodiments, the modules 914 –990 are modules of hardware circuits implemented by one or more ICs of an electronic apparatus. Though the modules 914 –990 are illustrated as being separate modules, some of the modules can be combined into a single module.The parser 990 (or entropy decoder) receives the bitstream 995 and performs initial parsing according to the syntax defined by a video-coding or image-coding standard. The parsed syntax element includes various header elements, flags, as well as quantized data (or quantized coefficients) 912. The parser 990 parses out the various syntax elements by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding.The inverse quantization module 914 de-quantizes the quantized data (or quantized coefficients) 912 to obtain transform coefficients 918, and the inverse transform module 915 performs inverse transform on the transform coefficients 918 to produce reconstructed residual signal 919. The reconstructed residual signal 919 is added with predicted pixel data 913 from the intra-prediction module 925 or the motion compensation module 930 to produce decoded pixel data 917. The decoded pixels data are filtered by the in-loop filter 945 and stored in the decoded picture buffer 950. In some embodiments, the decoded picture buffer 950 is a storage external to the video decoder 900. In some embodiments, the decoded picture buffer 950 is a storage internal to the video decoder 900.The intra-prediction module 925 receives intra-prediction data from bitstream 995 and according to which, produces the predicted pixel data 913 from the decoded pixel data 917 stored in a line buffer 927 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction.In some embodiments, the content of the decoded picture buffer 950 is used for display. A display device 905 either retrieves the content of the decoded picture buffer 950 for display directly, or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 950 through a pixel transport.The motion compensation module 930 produces predicted pixel data 913 from the decoded pixel data 917 stored in the decoded picture buffer 950 according to motion compensation MVs (MC MVs) . These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 995 with predicted MVs received from the MV prediction module 975.The MV prediction module 975 generates the predicted MVs based on reference MVs that were generated for decoding previous video frames, e.g., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 975 retrieves the reference MVs of previous video frames from the MV buffer 965. The video decoder 900 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 965 as reference MVs for producing predicted MVs.The in-loop filter 945 performs filtering or smoothing operations on the decoded pixel data 917 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 945 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.FIG. 10 illustrates portions of the video decoder 900 that implement regression-based MIP. For some embodiments, the figure illustrates the components of the intra-prediction module 920 of the video decoder 900. As illustrated, the intra-prediction module 920 uses neighboring samples 1005 of the current block provided by the line buffer 927 as input samples. A matrix multiplication module 1030 performs matrix multiplication on the input samples to generate a MIP predictor 1040 as the predicted pixel data 913.The matrix multiplication 1030 uses matrix parameters 1025 as elements of the matrix or as linear models to perform the matrix multiplication. The matrix parameters 1025 may also be generated by a regression module 1020, which performs data regression on training samples 1015 to generate the matrix parameters.A training sample collector 1010 provides the training samples 1015 by following certain predetermined rules to identify and collect specific sets of reconstructed samples stored in the decoded picture buffer 950 or the line buffer 927 as the training samples. Specifically, the training sample collector 1010 identifies one or more blocks and their corresponding neighboring L-shapes at certain positions as the source of the corresponding training data, with the training output target sample (s) collected from the block and the training input samples collected from the block’s neighboring L-shape. Examples of the positions from which to identify the blocks and their neighboring L-shapes includes neighboring template regions, intraTMP candidates, non-adjacent neighbors, IBC reference blocks identified by block vectors, reference blocks in collocated pictures or reference pictures identified by motion vectors, or other predefined positions.FIG. 11 conceptually illustrates a process 1100 for using regression-based MIP to decode pixel blocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the decoder 900 performs the process 1100 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the decoder 900 performs the process 1100.The decoder receives (at block 1110) data to be decoded as a current block of pixels in a current picture. The decoder identifies (at block 1120) corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape. The input training samples are from the neighboring L-shape of the reconstructed region and the output (target) training samples are from the reconstructed region of pixels.In some embodiments, the training samples are identified from a plurality of reconstructed regions and their neighboring L-shape samples, as described by reference to FIG. 5. Each reconstructed region may have the same or different dimensions as the current block. In some embodiments, the output training samples are identified by down-sampling the reconstructed region.In some embodiments, the corresponding input and output training samples may include samples that are identified from one or more predefined non-adjacent positions relative to the current block. In some embodiments, the corresponding input and output training samples may include samples identified from positions in one or more predefined search regions of the current picture, such as positions identified from a list of candidates for IntraTMP as described by reference to FIG. 3. In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in a reference picture that is a located by a motion vector or is a predefined collocated picture. In some embodiments, the corresponding input and output training samples may include samples identified from a reference block and its neighboring L-shape in the current picture that is located by a block vector (e.g., for IBC mode. )The decoder derives (at block 1130) an intra-prediction matrix by performing regression based on the identified training samples. In some embodiments, the decoder derives the intra-prediction matrix by deriving one or more linear models performs the matrix multiplication by applying the samples neighboring the current block as input to the linear models. The decoder may generate the predictor by adding a bias term, and the regression for deriving the intra-prediction matrix is also used to derive a coefficient for the bias term.In some embodiments, when a first flag is signaled (in the bitstream) to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, a second flag may be signaled to indicate whether to perform the regression based on the training samples for performing the matrix multiplication, or whether the coefficients of the matrix are derived by regression based on the training samples. In some embodiments, when a first flag is signaled to indicate whether intra template matching prediction is used to generate a predictor for the current block, a second flag may be signaled to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, and the coefficients are derived by the regression based on the training samples. In some embodiments, whether to perform the regression based on the training samples for performing the matrix multiplication is determined based on at least one of a size of the current block, a ratio of the current block, a slice type of the current block, and a color component of the current block (e.g., whether the current block is a luma block or a chroma block. )The decoder performs (at block 1140) matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block. The decoder reconstructs (at block 1150) the current block by using the generated predictor and the corresponding prediction residuals. The decoder may then provide the reconstructed current block for display as part of the reconstructed current picture.V. Example Electronic SystemMany of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random-access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the present disclosure. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.FIG. 12 conceptually illustrates an electronic system 1200 with which some embodiments of the present disclosure are implemented. The electronic system 1200 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an electronic system includes various types of computer readable media and interfaces for various other types of computer readable media. Electronic system 1200 includes a bus 1205, processing unit (s) 1210, a graphics-processing unit (GPU) 1215, a system memory 1220, a network 1225, a read-only memory 1230, a permanent storage device 1235, input devices 1240, and output devices 1245.The bus 1205 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 1200. For instance, the bus 1205 communicatively connects the processing unit (s) 1210 with the GPU 1215, the read-only memory 1230, the system memory 1220, and the permanent storage device 1235.From these various memory units, the processing unit (s) 1210 retrieves instructions to execute and data to process in order to execute the processes of the present disclosure. The processing unit (s) may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by the GPU 1215. The GPU 1215 can offload various computations or complement the image processing provided by the processing unit (s) 1210.The read-only-memory (ROM) 1230 stores static data and instructions that are used by the processing unit (s) 1210 and other modules of the electronic system. The permanent storage device 1235, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 1200 is off. Some embodiments of the present disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 1235.Other embodiments use a removable storage device (such as a floppy disk, flash memory device, etc., and its corresponding disk drive) as the permanent storage device. Like the permanent storage device 1235, the system memory 1220 is a read-and-write memory device. However, unlike storage device 1235, the system memory 1220 is a volatile read-and-write memory, such a random access memory. The system memory 1220 stores some of the instructions and data that the processor uses at runtime. In some embodiments, processes in accordance with the present disclosure are stored in the system memory 1220, the permanent storage device 1235, and / or the read-only memory 1230. For example, the various memory units include instructions for processing multimedia clips in accordance with some embodiments. From these various memory units, the processing unit (s) 1210 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.The bus 1205 also connects to the input and output devices 1240 and 1245. The input devices 1240 enable the user to communicate information and select commands to the electronic system. The input devices 1240 include alphanumeric keyboards and pointing devices (also called “cursor control devices” ) , cameras (e.g., webcams) , microphones or similar devices for receiving voice commands, etc. The output devices 1245 display images generated by the electronic system or otherwise output data. The output devices 1245 include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD) , as well as speakers or similar audio output devices. Some embodiments include devices such as a touchscreen that function as both input and output devices.Finally, as shown in FIG. 12, bus 1205 also couples electronic system 1200 to a network 1225 through a network adapter (not shown) . In this manner, the computer can be a part of a network of computers (such as a local area network ( “LAN” ) , a wide area network ( “WAN” ) , or an Intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1200 may be used in conjunction with the present disclosure.Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.While the above discussion primarily refers to microprocessor or multi-core processors that execute software, many of the above-described features and applications are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.While the present disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the present disclosure can be embodied in other specific forms without departing from the spirit of the present disclosure. In addition, a number of the figures (including FIG. 8 and FIG. 11) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process. Thus, one of ordinary skill in the art would understand that the present disclosure is not to be limited by the foregoing illustrative details, but rather is to be defined by the appended claims.Additional NotesThe herein-described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are merely examples, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected" , or "operably coupled" , to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being "operably couplable" , to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interactable components.Further, with respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.Moreover, it will be understood by those skilled in the art that, in general, terms used herein, and especially in the appended claims, e.g., bodies of the appended claims, are generally intended as “open” terms, e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an, " e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more; ” the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number, e.g., the bare recitation of "two recitations, " without other modifiers, means at least two recitations, or two or more recitations. Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “asystem having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”From the foregoing, it will be appreciated that various implementations of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various implementations disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;identifying corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape;deriving an intra-prediction matrix by performing regression based on the identified training samples;performing matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block; andencoding or decoding the current block by using generated predictor.2.The video coding method of claim 1, wherein the input training samples are from the neighboring L-shape of the reconstructed region and the output training samples are from the reconstructed region.3.The video coding method of claim 1, wherein deriving the intra-prediction matrix comprises deriving one or more linear models, wherein performing the matrix multiplication comprises applying the samples neighboring the current block as input to the linear models.4.The video coding method of claim 1, wherein generating the predictor of the current block comprises adding a bias term, wherein deriving the intra-prediction matrix by performing regression comprises deriving a coefficient for the bias term.5.The video coding method of claim 1, wherein the corresponding input and output training samples are identified from a plurality of reconstructed regions and their neighboring L-shape samples.6.The video coding method of claim 5, wherein a reconstructed region in the plurality of reconstructed regions has the same dimensions as the current block.7.The video coding method of claim 1, wherein the output training samples are identified by down-sampling the reconstructed region.8.The video coding method of claim 1, wherein the corresponding input and output training samples comprise samples identified from one or more predefined non-adjacent positions relative to the current block.9.The video coding method of claim 1, wherein the corresponding input and output training samples comprise samples identified from positions in one or more predefined search regions of the current picture.10.The video coding method of claim 1, wherein the corresponding input and output training samples comprise samples identified from a reference block and its neighboring L-shape in a reference picture.11.The video coding method of claim 1, wherein the corresponding input and output training samples comprise samples identified from a reference block and its neighboring L-shape in the current picture that is located by a block vector.12.The video coding method of claim 1, wherein when a first flag is signaled to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, a second flag is signaled to indicate whether the coefficients of the matrix are derived by the regression based on the training samples.13.The video coding method of claim 1, wherein when a first flag is signaled to indicate whether intra template matching prediction is used to generate a predictor for the current block, a second flag is signaled to indicate whether an intra-prediction predictor is to be derived by performing a matrix multiplication with samples neighboring the current block, and the coefficients of the matrix are derived by the regression based on the training samples.14.The video coding method of claim 1, wherein whether to perform the regression based on the training samples for performing the matrix multiplication is determined based on at least one of a size of the current block, a ratio of the current block, a slice type of the current block, and a color component of the current block.15.An electronic apparatus comprising:a video coder circuit configured to perform operations comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;identifying corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape;deriving an intra-prediction matrix by performing regression based on the identified training samples;performing matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block; andencoding or decoding the current block by using generated predictor.16.A video decoding method comprising:receiving data to be decoded as a current block of pixels of a current picture of a video;identifying corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape;deriving an intra-prediction matrix by performing regression based on the identified training samples;performing matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block; andreconstructing the current block by using generated predictor.17.A video encoding method comprising:receiving data to be encoded as a current block of pixels of a current picture of a video;identifying corresponding input and output training samples from at least one reconstructed region of pixels and its neighboring L-shape;deriving an intra-prediction matrix by performing regression based on the identified training samples;performing matrix multiplication based on the intra-prediction matrix and samples neighboring the current block to generate a predictor of the current block; andencoding the current block by using generated predictor.

Citation Information

Patent Citations

  • Reference sampling for matrix intra prediction mode

    US20200359050A1

  • Matrix-based intra prediction device and method

    US20220078449A1

  • Method and apparatus for coding / decoding picture data

    US20220360767A1

  • Method and device for intra prediction coding of video data

    WO2021006612A1

  • Decoding method, encoding method, decoder, encoder, and encoding and decoding system

    WO2023141970A1