Video coding and decoding method and device for constraining convolution model coefficient
By introducing a cross component prediction model into a video codec and constraining the model coefficients, the problem of inefficient prediction and encoding of video pixel blocks in the prior art is solved, and a more efficient video encoding and decoding process is realized.
Patent Information
- Application Number
- CN202380057083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-02
- Filing Date
- 2023-07-28
- Publication Date
- 2025-05-16
AI Technical Summary
Existing video codec standards have inefficiency problems in the prediction and encoding of video pixel blocks, especially when dealing with complex motion and texture features.
By introducing a cross component prediction model into a video codec, the model coefficients are optimized using constraints to generate a more efficient predictor for encoding or decoding. Specific methods include cross-component prediction based on the convolutional model, constraining model coefficients by shear thresholds or limiting coefficient ranges.
Improves the efficiency of the video encoding and decoding process, especially when dealing with complex motion and texture features, and reduces the computational complexity and storage requirements of encoding and decoding.
Smart Images

Figure CN120019654A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application is part of a non-provisional application, which claims priority to U.S. Provisional Patent Application No. 63 / 370,133 filed on August 2, 2022. The entire contents of the U.S. Provisional Patent Application are incorporated herein by reference. Technical Field
[0003] The present invention relates to a video processing method and device in a video encoding and decoding system, and in particular to a method for encoding and decoding a pixel block by cross component prediction. Background Art
[0004] High-Efficiency Video Coding (HEVC) is an international video codec standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is based on a motion-compensated DCT-like transform codec architecture based on mixed pixel blocks. The basic unit of compression is called a coding unit (CU), which is a 2Nx2N square pixel block. Each CU can be recursively divided into four smaller CUs until a predefined minimum size is reached. Each CU contains one or more prediction units (PUs).
[0005] Versatile video coding (VVC) is the latest international video coding standard developed by ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 Joint Video Expert Team (JVET). The input video signal is predicted from the reconstructed signal, which is derived from the coded and decoded image region. The prediction residual signal is processed by pixel block transform. The transform coefficients are quantized and entropy coded along with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transforming the dequantized transform coefficients. The reconstructed signal is also processed by loop filtering to remove coding artifacts. The decoded image is stored in a buffer and used to predict future images in the input video signal.
[0006] In VVC, the coded image is divided into non-overlapping square pixel block areas represented by related coding tree units (CTUs). The leaf nodes of the coding tree correspond to coding units (CUs). The coded image can be represented by a set of slices, each slice including an integer number of CTUs. Each CTU in a slice is processed in raster scan order. Bi-predictive (B) slices can be decoded to predict the sample values of each pixel block using intra-picture prediction or inter-picture prediction using up to two motion vectors and reference indices. Predictive (P) slices are decoded to predict the sample values of each pixel block using intra-picture prediction or inter-picture prediction using up to one motion vector and reference index. Intra-picture (I) slices are decoded using only intra-picture prediction.
[0007] Using a quadtree (QT) with a nested multi-type-tree (MTT) structure, the CTU can be split into one or more non-overlapping codec units (CUs) to accommodate various local motion and texture features. The CU is further divided into smaller CUs using one of five partition types: quadtree partition, vertical binary tree partition, horizontal binary tree partition, vertical center-side ternary tree partition, and horizontal center-side ternary tree partition.
[0008] Each CU contains one or more prediction units (PU). The prediction unit, together with the related CU syntax, serves as a basic unit for indicating prediction sub-information. The specified prediction process is used to predict the values of related pixel samples within the PU. Each CU may contain one or more transform units (TU) representing prediction residual pixel blocks. The transform unit (TU) includes a transform pixel block (TB) of luma samples and two corresponding transform pixel blocks of chroma samples, each TB corresponding to a residual pixel block sample of a color component. An integer transform is applied to the transform pixel block. The layer values of the quantized coefficients are entropy encoded and decoded in the bitstream together with other side information. The terms Coding Tree Block (CTB), Coding Block (CB), Prediction Block (PB), and Transform Block (TB) are defined to specify a 2D sample array of a color component associated with a CTU, CU, PU, or TU, respectively. Therefore, a CTU includes a luma CTB, two chroma CTBs, and related syntax elements. Similar relationships also apply to CU, PU, and TU.
[0009] For each inter-predicted CU, the motion parameters including motion vectors, reference picture indices, and reference picture lists are indexed, and additional information is used to generate inter-predicted samples. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded or decoded using the skip mode, the CU is associated with a PU, there are no valid residual coefficients, no coded / decoded motion vector deltas, or reference picture indices. The merge mode is specified, where the motion parameters of the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates, and additional scheduling introduced in VVC. The merge mode can be applied to any inter-predicted CU. An alternative to the merge mode is the explicit transmission of motion parameters, where the motion vectors, the reference picture indices corresponding to each reference picture list, and the reference picture lists are signaled explicitly for each CU using flags and other required information. SUMMARY OF THE INVENTION
[0010] Some embodiments of the present application provide a method for performing cross-component prediction by constraining coefficients of a component prediction model. A video codec receives pixel block data to be encoded or decoded as a current block of a current picture of a video. The video codec derives a set of coefficients based on corresponding input and output component samples. The video codec constrains the derived set of coefficients according to a set of constraints. The video codec uses the constrained set of coefficients as a component prediction model to generate a predictor for the current block. The video codec encodes or decodes the current block using the generated predictor.
[0011] In some embodiments, the component prediction model is a cross-component model that generates predicted chroma samples based on reconstructed luma samples of the current block. The corresponding input and output component samples are the corresponding luma and chroma samples of a template region adjacent to the current block. The predictor may include the generated predicted chroma samples. In some embodiments, the component prediction model is a convolutional model; the set of coefficients is derived by solving a matrix equation between the corresponding input and output component samples.
[0012] In some embodiments, the encoder constrains the derived set of coefficients by clipping coefficients at a clipping threshold, or clipping different coefficients at different clipping thresholds, or limiting the coefficients within a predefined range. In some embodiments, the derived set of coefficients and the constrained set of coefficients are represented by fixed-point in the encoder, with the fractional part including N bits, and the size of the predefined range is magnified based on 1<<N. In some embodiments, when the derived coefficients exceed the predefined range, the set of coefficients is set to be equal to an identity filter, or the derived set of coefficients is not used to encode or decode the current block (CCCM mode is disabled).
[0013] Other aspects and features of the present invention will become apparent to those of ordinary skill in the art by reading the following description of specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Various embodiments of the present disclosure are described in detail with reference to the following drawings, which are presented as examples.
[0015] in:
[0016] Figure 1 Conceptual illustration of the chrominance and luminance samples used to derive linear model parameters.
[0017] Figure 2 An example of classification of adjacent samples into two groups is marked.
[0018] Figure 3 Conceptual illustration of the spatial components of a convolution filter.
[0019] Figure 4 A reference region for deriving filter coefficients of a convolutional model for the current block is shown.
[0020] Figure 5 The data path of a video codec that derives and constrains the coefficients of the model is conceptually illustrated.
[0021] Figure 6 An exemplary video encoder that can encode a block of pixels using a component prediction model is labeled.
[0022] Figure 7 The portion of the video encoder that derives and uses the component prediction model by constraining the coefficients of the model is marked.
[0023] Figure 8 The process of constraining the coefficients of a component prediction model is conceptually illustrated.
[0024] Fig. 9 An example of a video decoder that can decode a block of pixels using a component prediction model is labeled.
[0025] Fig.10 The portion of the video decoder that derives and uses the component prediction model by constraining its coefficients is marked.
[0026] Fig.11 The process of constraining the coefficients of a component prediction model is conceptually illustrated.
[0027] Fig.12 An electronic system for implementing some embodiments of the present application is conceptually described. DETAILED DESCRIPTION
[0028] It is readily understood that the components of the present invention, as generally described herein and illustrated in the accompanying drawings, may be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of the present invention, as shown in the accompanying drawings, is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention.
[0029] I. Cross-Component Linear Model (CCLM)
[0030] The cross-component linear model (CCLM) or linear model (LM) mode is a cross-component prediction mode in which the chrominance components of a block are predicted from the collocated reconstructed luminance (Luma) samples by a linear model. The parameters of the linear model (e.g., scale and offset) are derived from the already reconstructed luminance and chrominance (Chroma) samples adjacent to the block. For example, in VVC, the CCLM mode exploits the inter-channel dependencies to predict chrominance samples from reconstructed luminance samples. This prediction is performed using a linear model of the following form:
[0031] P(i,j)=α·rec′ L (i,j)+β (1)
[0032] P(i,j) in formula (1) represents the predicted chroma sample in the CU (or the predicted chroma sample of the current CU), rec′ L (i, j) represents the downsampled reconstructed luma samples of the same CU (or the corresponding reconstructed luma samples of the current CU).
[0033] The CCLM model parameters α (scaling parameter) and β (offset parameter) are derived from up to four adjacent chroma samples and their corresponding downsampled luma samples. In LM_A mode (also denoted as LM-T mode), only the upper or top adjacent templates are used to calculate the coefficients of the linear model. In LM_L mode (also denoted as LM-L mode), only the left template is used to calculate the coefficients of the linear model. In LM-LA mode (also denoted as LM-LT mode), both the left and upper templates are used to calculate the coefficients of the linear model.
[0034] Figure 1The chrominance and luma samples used to derive linear model parameters are conceptually illustrated. The diagram shows a current block 100 having luma component samples and chroma component samples in a 4:2:0 format. The luma and chroma samples adjacent to the current block are reconstructed samples. These reconstructed samples are used to derive the linear model (parameters α and β) of the cross components. Since the current block is in a 4:2:0 format, the luma samples are downsampled before being used for linear model derivation. In this example, there are 16 pairs of reconstructed luma (downsampled) and chroma samples adjacent to the current block. These 16 pairs of luma and chroma values are used to derive linear model parameters.
[0035] Assuming the current chroma block size is W×H, then W' and H' are set to:
[0036] - When LM-LT mode is applied, W'=W, H'=H;
[0037] - When LM-T mode is applied, W'=W+H;
[0038] - When LM-L mode is applied, H'=H+W.
[0039] The upper adjacent position is represented by S[0,-1]...S[W'-1,-1], and the left adjacent position is represented by S[-1,0]...S[-1,H'-1]. Then four samples are selected as:
[0040] - When LM mode is applied (both upper and left adjacent samples are available), S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4];
[0041] - When the LM-T mode is applied (only the upper adjacent samples are available), S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1];
[0042] - When LM-L mode is applied (only left neighboring samples are available), S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8], S[-1,7*H' / 8];
[0043] The four adjacent luma samples at the selected position are downsampled and compared four times to find the two larger values: x0A and x1A, and the two smaller values: x0B and x1B. Their corresponding chroma sample values are denoted as y0A, y1A, y0B, and y1B. Then, XA, XB, YA, and YB are derived as:
[0044] X a =(x 0 A+x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1 (2)
[0045] Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1 (3)
[0046] The linear model parameters α and β are obtained according to the following formula:
[0047]
[0048] β=Y b -α·X b (5)
[0049] The operation of calculating the α and β parameters according to formulas (4) and (5) can be implemented by a lookup table. In some embodiments, in order to reduce the memory required to store the lookup table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by exponential notation. For example, diff is approximated by a 4-bit effective part and an exponent. Therefore, the table of 1 / diff is simplified to 16 elements for 16 effective values, as shown below:
[0050] DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} (6)
[0051] This reduces the complexity of the calculations as well as the memory size required to store the required tables.
[0052] In some embodiments, in order to obtain more samples for calculating CCLM model parameters α and β, for LM-T mode, the upper template is expanded to include (W+H) samples, and for LM-L mode, the left template is expanded to include (H+W) samples. For LM-LT mode, both the expanded left template and the expanded upper template are used to calculate the coefficients of the linear model.
[0053] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to luma samples to achieve a 2 to 1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by a flag at the sequence parameter set (SPS) level. The two downsampling filters are as follows, corresponding to the contents of "type 0" and "type 2" ("type-0" and "type-2"), respectively.
[0054] rec′ L (i,j)=[rec L (2i-1,2j-1)+2j-1)+2j L (2i-1,2j-1)+rec L (2i+1,2j-1)+rec L (2i-1,2j)+rec L (2i+1,2j)+4]>>3 (7)
[0055] rec′ L (i,j)=[rec L (2i,2j-1)+rec L (2i-1,2j)+4*rec L (2i,2j)+rec L (2i+1,2j)+rec L (2i,2j+1)+4]>>3 (8)
[0056] In some embodiments, the α and β parameter calculations are performed as part of the decoding process, rather than just as an encoder search operation. Therefore, no syntax is used to communicate the values of α and β to the decoder.
[0057] For chroma intra mode codec, a total of 8 intra modes are allowed. These modes include five traditional intra modes and three cross-component linear modes (LM_LA, LM_A, and LM_L). Chroma intra mode codec can directly depend on the intra prediction mode of the corresponding luma block. The chroma intra mode signals and the corresponding luma intra prediction modes are shown in the following table:
[0058] Chroma intra prediction mode corresponds to luma intra prediction mode
[0059]
[0060]
[0061] Since the block partitioning structure of independent luminance and chrominance components is enabled in I slices, a chrominance block may correspond to multiple luminance blocks. Therefore, for the chrominance derivative mode (DM) mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chrominance block is directly inherited.
[0062] A unified binarization table (mapped to bin string) is used for chroma intra prediction mode according to the following table:
[0063] Chroma intra prediction mode bin string
[0064] Chroma intra prediction mode bin string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111
[0065] In this table, the first bin indicates whether it is normal mode (0) or LM mode (1). If it is LM mode, then the next bin indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next bin indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first bin of the corresponding intra_chroma_pred_mode binarization table can be discarded before entropy encoding and decoding. Or, in other words, the first bin is inferred to be 0 and therefore will not be encoded and decoded. This single binarization table is used for the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two bins in the table are context encoded and decoded with their own context model, and the remaining bins are bypass encoded and decoded.
[0066] In addition, to reduce intra chroma latency in dual trees, when the 64x64 luma codec tree node is not split (and the ISP does not have CUs for 64x64) or is split with QT, the chroma CUs in the 32x32 / 32x16 chroma codec tree nodes are allowed to use CCLM as follows:
[0067] If a 32x32 chroma node is not split or is split with QT, all chroma CUs in the 32x32 node can use CCLM.
[0068] If a 32x32 chroma node is split using horizontal BT, and the 32x16 child node is not split or uses vertical BT, all chroma CUs in the 32x16 chroma node can use CCLM.
[0069] In all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chroma CUs.
[0070] II. Multi-Mode CCLM (MMLM)
[0071] The multi-model CCLM mode (MMLM) uses two models to predict chrominance samples from the luma samples of the entire CU. Similar to CCLM, three multi-model CCLM modes (MMLM_LA, MMLM_A, and MMLM_L) are used to indicate whether to use both the upper and left neighboring samples, only the upper neighboring samples, or only the left neighboring samples in the derivation of model parameters.
[0072] In MMLM, the adjacent luminance samples and adjacent chrominance samples of the current block are divided into two groups, each of which is used as a training set to derive a linear model (i.e., derive a specific α and β for a specific group). In addition, the samples of the current luminance block are also classified according to the same rules as the classification of the adjacent luminance samples.
[0073] Figure 2 An example of classifying neighboring samples into two groups is marked. The threshold is calculated as the average of the neighboring reconstructed luma samples. Neighboring samples at [x, y] are classified into the first group if Rec'L[x, y] <= Threshold, and neighboring samples at [x, y] are classified into the second group if Rec'L[x, y]> Threshold. Therefore, the multi-model CCLM prediction of chroma samples is:
[0074] Pred c [x,y]=α1×Recˊ L [x,y]+β1if Rec' L [x,y]≤Threshold
[0075] Pred c [x,y]=α2×Recˊ L [x,y]+β2if Rec' L [x,y]>Threshold
[0076] III. Convolutional Cross-Component Model (CCCM)
[0077] In some embodiments, a convolution cross component model (CCCM) is applied to improve cross component prediction performance. For some embodiments, the convolution model has a 7-tap filter with a 5-tap plus a signed shape spatial component, a nonlinear term, and a bias term. The input of the spatial 5-tap component of the filter includes a center (C) luminance sample, which matches the chrominance sample to be predicted, and the neighboring samples of the upper side / north (N), lower side / south (S), left side / west (W), and right side / east (E) of the center luminance sample. Figure 3Conceptual illustration of the spatial components of a convolution filter. The nonlinear term (denoted as P) is expressed as a power of two of the center luminance sample C, scaled by the sample value range of the content:
[0078] P=(C*C+midVal)>>bitDepth (9)
[0079] Therefore, for 10-bit content, the nonlinear term P is calculated as:
[0080] P=(C*C+512)>>10
[0081] The bias term (denoted as B) represents a scalar offset between the input and output (similar to the bias term in CCLM) and is set to the intermediate chroma value (512 for 10-bit content). The output of the filter is calculated as the convolution between the filter coefficients ci and the input value and is clipped to the valid chroma sample range:
[0082] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B (10)
[0083] The filter coefficients ci are calculated by minimizing the MSE between the reconstructed (or target) chroma samples and the corresponding predicted chroma samples in the reference region. Each predicted chroma sample is generated from the juxtaposed luminance sample and its surrounding luminance samples using the derived component prediction model (such as formula (10)). Formula (10) is a convolution model based on the center sample and the surrounding 4 samples (C, N, S, E, W). Formula (10) can be extended to include taps of the center sample and the surrounding 8 samples (C, N, S, E, W, NE, NW, SE, SW). Formula (10) and its extended form can be called a component prediction model (because it can be used for cross-component prediction or intra-component prediction).
[0084] Figure 4 A reference region is shown for deriving filter coefficients for the convolution model of the current block. The reference region includes (reference) rows of (chrominance) samples above and to the left of the current block 400. (In this example, the current block 400 is a PU). The reference region extends one PU width to the right and one PU height below the PU boundary. The reference region can be adjusted to include only available samples. An extended area of the reference region is used to support "side samples" of signed-shaped spatial filters (e.g., Figure 3 The N, E, W, S samples, and NW, NE, SW, SE samples) are filled in when they are unavailable.
[0085] Minimization of MSE can be achieved by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is decomposed by LDL, and the final filter coefficients are calculated using the inverse permutation method. This process is similar to the calculation of the ALF filter coefficients in ECM, however, in some embodiments, LDL decomposition is selected instead of Cholesky decomposition to avoid the use of square root operations.
[0086] In some embodiments, a high-order model may be used instead of a linear model to predict chrominance samples. The high-order model may include a k-tap space term, a nonlinear term (denoted as P), and a bias term (denoted as B). The high-order model may be specified as:
[0087]
[0088] Where rec_L^'(i,j) is the downsampled reconstructed luma sample at position (i,j), neiRec_L^'(x) is a neighboring sample around rec_L^'(i,j), and a_0, a_x, b, and c are model parameters. This high-order model formula can be used to derive model parameters between color components, or between reference samples of the current frame / picture and a reference frame / picture.
[0089] III. Constraints on Model Coefficients
[0090] In some embodiments, after the matrix equation solving process, the filter / model coefficients finally derived are used to generate the CCCM predictor. However, the derived filter coefficients may be unreasonable, for example, the model overfits the reference sample, or the filter coefficients are too large. Some embodiments of the present application provide a method for constraining the derived model coefficients before generating the CCCM predictor.
[0091] Figure 5 The data path 500 of a video codec is conceptually illustrated, which derives and constrains the coefficients of a model (e.g., for CCCM). The constrained model coefficients are used to generate a predictor for a current block. As shown, the data path 500 begins with a matrix preparation module 510, which generates an autocorrelation matrix and a cross-correlation vector 515. The autocorrelation matrix is prepared based on corresponding input component samples 505 (X samples). The cross-correlation vector is prepared based on corresponding input and output component samples 505 (X samples and Y samples). The corresponding input and output component samples can be corresponding reconstructed luminance and chrominance samples of a reference template region adjacent to the current block.
[0092] The generated autocorrelation matrix and cross-correlation vector 515 are provided to a matrix equation solver module 530 to generate an optimized coefficient set 535. The coefficient constraint module 540 constrains the optimized coefficients 535 to produce a constrained final coefficient set 545. The constrained final coefficients 545 are ultimately used as a component prediction model 550 (e.g., CCLM or CCCM) to generate prediction component samples 565 (e.g., chrominance of the current block) based on reference component samples 560 (e.g., luminance of the current block). The generated prediction component samples 565 can be used as a CCCM predictor.
[0093] Different types of constraints may be applied to the optimization coefficients 535 to generate the constrained final coefficients 545. For example, in some embodiments, a predefined clipping threshold is used to clip the optimization coefficients 535 before generating the constrained final coefficients 545. In some embodiments, multiple clipping thresholds are predefined for the optimization coefficients 535, and a syntax indicating the selected threshold may be signaled. In some embodiments, multiple clipping thresholds for the optimization coefficients 535 are predefined, and the selected threshold may be explicitly derived from neighboring reconstructed samples or side information. In some embodiments, the clipping thresholds for different coefficients may be all different or partially different.
[0094] In some embodiments, the coefficients of the component prediction model 550 are represented in fixed point format, and the bit width of the integer part is limited to a certain value. In this example, before the coefficient constraint module 540 constrains, the fixed-point format of the optimization coefficient 535 is 48 bits for the integer part and 16 bits for the fractional part. After the constraint of the coefficient constraint module 540, Figure 5 In the example shown, the fixed-point format of the constraint coefficient 545 is 36 bits for the integer part and 16 bits for the fractional part. In other examples, the integer part of the constraint coefficient 545 can have different bit widths. In some embodiments, the bit widths of the integer and / or fractional part of each coefficient can be completely different or partially different. For example, the constraint coefficient can be represented in a fixed-point format with 36 bits for the integer part and 14 bits for the fractional part.
[0095] In some embodiments, the video codec may perform a clipping operation on coefficients that are out of range. In some embodiments, if the coefficient is out of range, it may be inferred that the CCCM or CCLM mode is not enabled.
[0096] In some embodiments, before generating the final predictor 565, the optimization coefficients 535 are clipped to a predefined range. For example, in terms of floating-point precision, the range can be [-1, 1), [-2, 2], [-4, 4), or [-8, 8]. If the model coefficients are represented using fixed-point precision and the number of bits in the fractional part is N bits, then the predefined range is expanded by (1 << N). For example, if the supported range for floating-point precision is [-8, 8), and the number of bits in the fractional part is 5 bits, then the supported range changes from [-8, 8) for floating-point precision to [-8 * 32, 8 * 32) for fixed-point precision with a 5-bit fractional part.
[0097] In some embodiments, the predefined ranges for different model coefficients may be different. For example, the predefined range depends on the spatial location of the model coefficients. In some embodiments, the predefined range depends on the order of the corresponding inputs. In some embodiments, all coefficients use only one predefined range. In some embodiments, when a derived coefficient exceeds the predefined range, the final coefficient 545 is derived by clipping the coefficient to the predefined range. In some embodiments, when a derived coefficient exceeds the predefined range, the final coefficient 545 is set to zero. In some embodiments, when a derived coefficient exceeds the predefined range, the coefficients in the CCCM mode are set to be equal to the identity filter (i.e., the CCCM mode does not apply filtering).
[0098] Any of the methods proposed above can be implemented in an encoder and / or a decoder. For example, any of the methods proposed can be implemented in the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder. Alternatively, any of the methods proposed can be implemented as a circuit coupled to the inter / intra / prediction module of the encoder and / or the inter / intra / prediction module of the decoder to provide the information required by the inter / intra / prediction module.
[0099] VI. Video Encoder Example
[0100] Figure 6An exemplary video encoder 600 that can encode a pixel block using a component prediction model is marked. As shown, the video encoder 600 receives an input video signal from a video source 605 and encodes the signal into a bitstream 695. The video encoder 600 has a plurality of components or modules for encoding the signal from the video source 605, including at least a transform module 610, a quantization module 611, an inverse quantization module 614, an inverse transform module 615, an intra-frame prediction estimation module 620, an intra-frame prediction module 625, a motion compensation module 630, a motion estimation module 635, an in-loop filter 645, a reconstructed picture buffer 650, an MV buffer 665, an MV prediction module 675, and some components in the entropy encoder 690. The motion compensation module 630 and the motion estimation module 635 are part of the inter-frame prediction module 640.
[0101] In some embodiments, modules 610-690 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 610-690 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 610-690 are illustrated as separate modules, some modules may be combined into a single module.
[0102] The video source 605 provides a raw video signal representing pixel data for each video frame without compression. A subtractor 608 calculates the difference between the raw video pixel data of the video source 605 and the predicted pixel data 613 from the motion compensation module 630 or the intra prediction module 625 as a prediction residual 609. The transform module 610 converts the difference (or residual pixel data or residual signal) into transform coefficients 616 (e.g., by performing a discrete cosine transform, or DCT). The quantization module 611 quantizes the transform coefficients 616 into quantized data (or quantized coefficients) 612, which are encoded into a bitstream 695 by an entropy encoder 690.
[0103] The inverse quantization module 614 dequantizes the quantized data (or quantized coefficients) 612 to obtain transform coefficients 616, and the inverse transform module 615 inversely transforms the transform coefficients 616 to generate a reconstructed residual 619. The reconstructed residual 619 is added to the predicted pixel data 613 to generate reconstructed pixel data 617. In some embodiments, the reconstructed pixel data 617 is temporarily stored in a line buffer (not shown) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 645 and stored in a reconstructed picture buffer 650. In some embodiments, the reconstructed picture buffer 650 is an external storage of the video encoder 600. In some embodiments, the reconstructed picture buffer 650 is an internal storage of the video encoder 600.
[0104] The intra estimation module 620 performs intra prediction based on the reconstructed pixel data 617 to generate intra prediction data. The intra prediction data is provided to the entropy encoder 690 to be encoded into a bitstream 695. The intra prediction data is also used by the intra prediction module 625 to generate the predicted pixel data 613.
[0105] The motion estimation module 635 performs inter-frame prediction by generating MVs to refer to pixel data of a previously decoded frame stored in the reconstructed picture buffer 650. These MVs are provided to the motion compensation module 630 to generate predicted pixel data.
[0106] The video encoder 600 uses MV prediction to generate a predicted MV, and instead of encoding the complete actual MV in the bitstream, the difference between the MV used for motion compensation and the predicted MV is encoded as residual motion data and stored in the bitstream 695.
[0107] The MV prediction module 675 generates a predicted MV based on a reference MV generated for encoding a previous video frame, i.e., a motion compensated MV for performing motion compensation. The MV prediction module 675 retrieves the reference MV of the previous video frame from the MV buffer 665. The video encoder 600 stores the MV generated for the current video frame in the MV buffer 665 as a reference MV for generating the predicted MV.
[0108] The MV prediction module 675 uses the reference MV to create a predicted MV. The predicted MV can be calculated by spatial MV prediction or temporal MV prediction. The difference (residual motion data) between the predicted MV and the motion compensation MV (MC MV) of the current frame is encoded by the entropy encoder 690 to the bit stream 695.
[0109] The entropy encoder 690 encodes various parameters and data into a bitstream 695 by using an entropy coding technique, such as context adaptive binary arithmetic coding (CABAC) or Huffman coding. The entropy encoder 690 encodes various header elements, flags, and quantized transform coefficients 612 and residual motion data as syntax elements into the bitstream 695. The bitstream 695 is in turn stored in a storage device or transmitted to a decoder via a communication medium such as a network.
[0110] The in-loop filter 645 performs filtering or smoothing operations on the reconstructed pixel data 617 to reduce coding artifacts, particularly at block boundaries. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 645 include a deblocking filter (DBF), a sample adaptive offset (SAO), and / or an adaptive loop filter (ALF).
[0111] Figure 7Identifies portions of video encoder 600 that derive and use a component prediction model by constraining the coefficients of the model. As shown in the diagram, an initial predictor generation module 720 provides an initial predictor 715 to a component prediction model 710. The initial predictor 715 may include component samples (luminance or chrominance) of a reference block for predicting component samples of a current block, or component samples of the current block for cross-component prediction. The component prediction model 710 is applied to the initial predictor 715 to generate a refined predictor 725. The samples of the refined predictor 725 can be used as the predicted pixel data 613. The initial predictor 715 may be the reconstructed luminance samples of the current block, while the refined predictor 725 may be the predicted chrominance samples of the current block.
[0112] When the current block is coded / decoded by inter prediction, a motion estimation module 635 provides an MV, and a motion compensation module 630 uses the MV to identify a reference block in a reference picture as the initial predictor 715. When the current block is coded / decoded by intra prediction, an intra prediction estimation module 620 provides an intra mode, which is used by an intra prediction module 625 to generate an intra prediction of the current block as the initial predictor 715. Then, the component samples of the initial predictor 715 can be used as the input to the component prediction model 710.
[0113] To derive the component prediction model 710, a regression data selection module 730 retrieves the required component samples from a reconstructed picture buffer 650 as regression data. The regression data can be taken from regions within and / or around the current block in the current picture, as well as regions within and / or around the reference block in the reference picture. The retrieved regression data (i.e., the required component samples) includes corresponding input (X) and output (Y) component samples for determining the coefficients or parameters of the component prediction model 710.
[0114] A model constructor 705 uses the regression data (X and Y) to derive the coefficients of the component prediction model 710 using techniques such as elimination, iteration, or factorization. In some embodiments, the model constructor 705 applies certain constraints to the derived coefficients before providing the constrained coefficients to be used as the component prediction model 710. The model constructor 705 can constrain the coefficients by clipping at a threshold or limiting the coefficients within a predefined range. In some embodiments, the model constructor 705 can apply different clipping thresholds to different coefficients. In some embodiments, the coefficients are represented in fixed-point with an N-bit fractional part, and the predefined range is scaled by 1<<N. Constraining the coefficients of the component prediction model is described in Part IV (VI) above.
[0115] Figure 8Conceptually illustrates the process 800 of constraining the coefficients of a component prediction model. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing the encoder 600 execute the flow 800 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing the encoder 600 executes the process 800.
[0116] The encoder (at block 810) receives the data to be encoded as the current block of the current picture of the video.
[0117] The encoder (at block 820) derives a set of coefficients for the component prediction model based on corresponding input and output component samples. In some embodiments, the component prediction model is a cross-component model that generates predicted chrominance samples based on the reconstructed luminance samples of the current block, and the corresponding input and output component samples are the corresponding luminance and chrominance samples of a template region adjacent to the current block. In some embodiments, the component prediction model is a convolutional model that derives the set of coefficients through factorization and back substitution based on the autocorrelation matrix between the corresponding input and output component samples.
[0118] The encoder (at block 830) constrains the derived set of coefficients according to a set of constraints. In some embodiments, the encoder constrains the derived set of coefficients by clipping the coefficients at a clipping threshold, or clipping different coefficients at different clipping thresholds, or limiting the coefficients within a predefined range. In some embodiments, the derived set of coefficients and the constrained set of coefficients are represented in floating point in the encoder. In some embodiments, the coefficients are represented in fixed point with N bits in the fractional part, and the size of the predefined range is magnified based on 1<<N. In some embodiments, when the derived coefficients exceed the predefined range, the set of coefficients is set to be equal to the identity filter, or a clipping operation is applied to the out-of-range coefficients, or the current block is not encoded or decoded using the derived set of coefficients (disabling the CCCM mode).
[0119] The encoder (at block 840) applies the constrained set of coefficients as the component prediction model to generate a predictor for the current block. The predictor may include the generated predicted chrominance samples.
[0120] The encoder (at block 850) encodes the current block by generating a prediction residual using the generated predictor.
[0121] VI. Video Decoder Example
[0122] In some embodiments, the encoder may signal (or generate) one or more syntax elements in the bitstream so that the decoder can parse the one or more syntax elements from the bitstream.
[0123] Fig. 9An example of a video decoder 900 that can decode a pixel block using a component prediction model is marked. As shown in the figure, the video decoder 900 is an image decoding or video decoding circuit that receives a bitstream 995 and decodes the content of the bitstream into pixel data of a video frame for display. The video decoder 900 has a plurality of components or modules for decoding the bitstream 995, including some components selected from an inverse quantization module 911, an inverse transform module 910, an intra-frame prediction module 925, a motion compensation module 930, an in-loop filter 945, a decoded picture buffer 950, an MV buffer 965, an MV prediction module 975, and a parser 990. The motion compensation module 930 is a part of the inter-frame prediction module 940.
[0124] In some embodiments, modules 910-990 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 910-990 are modules of hardware circuits implemented by one or more ICs of an electronic device. Although modules 910-990 are illustrated as separate modules, some of these modules may be combined into a single module.
[0125] The parser 990 (or entropy decoder) receives the bitstream 995 and performs initial parsing according to the syntax defined by the video codec or image codec standard. The parsed syntax elements include various header elements, flags, and quantization data (or quantization coefficients) 912. The parser 990 parses out various syntax elements by using entropy coding techniques, such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding.
[0126] The inverse quantization module 911 dequantizes the quantized data (or quantized coefficients) 912 to obtain transform coefficients, and the inverse transform module 910 inversely transforms the transform coefficients 916 to generate a reconstructed residual signal 919. The reconstructed residual signal 919 is added to the predicted pixel data 913 from the intra-frame prediction module 925 or the motion compensation module 930 to generate decoded pixel data 917. The decoded pixel data is filtered by the in-loop filter 945 and stored in the decoded picture buffer 950. In some embodiments, the decoded picture buffer 950 is an external storage of the video decoder 900. In some embodiments, the decoded picture buffer 950 is an internal storage of the video decoder 900.
[0127] The intra prediction module 925 receives intra prediction data from the bitstream 995 and generates predicted pixel data 913 from decoded pixel data 917 stored in the decoded picture buffer 950. In some embodiments, the decoded pixel data 917 is also stored in a line buffer (not shown) for intra prediction and spatial MV prediction.
[0128] In some embodiments, the contents of the decoded picture buffer 950 are used for display. The display device 955 directly retrieves the contents of the decoded picture buffer 950 for display, or retrieves the contents of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 950 via pixel transfer.
[0129] The motion compensation module 930 generates predicted pixel data 913 from the decoded pixel data 917 stored in the decoded picture buffer 950 according to motion compensated MVs (MC MVs). These motion compensated MVs are decoded by adding the residual motion data received from the bitstream 995 to the predicted MVs received from the MV prediction module 975.
[0130] The MV prediction module 975 generates a predicted MV based on a reference MV generated for decoding a previous video frame, for example, a motion compensated MV for performing motion compensation. The MV prediction module 975 retrieves the reference MV of the previous video frame from the MV buffer 965. The video decoder 900 stores the motion compensated MV generated for decoding the current video frame in the MV buffer 965 as a reference MV for generating the predicted MV.
[0131] The in-loop filter 945 performs filtering or smoothing operations on the decoded pixel data 917 to reduce encoding and decoding artifacts, particularly at block boundaries. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 945 include a deblocking filter (DBF), a sample adaptive offset (SAO), and / or an adaptive loop filter (ALF).
[0132] Fig.10 Portions of the video decoder 900 that derive and use the model by constraining the coefficients of the component prediction model are marked. As shown, the initial predictor generation module 1020 provides an initial predictor 1015 to the component prediction model 1010. The initial predictor 1015 may include component samples (luminance or chrominance) of a reference block for predicting component samples of a current block, or component samples of the current block for cross-component prediction. The component prediction model 1010 is applied to the initial predictor 1015 to generate a refined predictor 1025. The samples of the refined predictor 1025 may be used as the predicted pixel data 913. The initial predictor 1015 may be a reconstructed luminance sample of the current block, while the refined predictor 1025 may be a predicted chrominance sample of the current block.
[0133] When the current block is decoded by inter prediction, the entropy decoder 990 provides the MV, and the motion compensation module 930 uses the MV to identify a reference block in the reference picture as the initial predictor 1015. When the current block is decoded by intra prediction, the entropy decoder 990 provides the intra mode, which is used by the intra prediction module 925 to generate an intra prediction of the current block as the initial predictor 1015. Then, the component samples of the initial predictor 1015 can be used as inputs to the component prediction model 1010.
[0134] To derive the component prediction model 1010, the regression data selection module 1030 retrieves the required component samples from the decoded picture buffer 950 as regression data. The regression data can be taken from regions within and / or around the current block in the current picture, as well as regions within and / or around the reference block in the reference picture. The retrieved regression data (i.e., the required component samples) includes the corresponding input (X) and output (Y) component samples for determining the coefficients or parameters of the component prediction model 1010.
[0135] The model constructor 1005 uses the regression data (X and Y) to derive the coefficients of the component prediction model 1010 using techniques such as elimination, iteration, or factorization. In some embodiments, the model constructor 1005 applies certain constraints to the derived coefficients before providing the constrained coefficients to be used as the component prediction model 1010. The model constructor 1005 can constrain the coefficients by clipping at a threshold or limiting the coefficients to a predefined range. In some embodiments, the model constructor 1005 can apply different clipping thresholds to different coefficients. In some embodiments, the coefficients are represented in fixed-point with an N-bit fractional part, and the predefined range is scaled by 1<<N. The process of constraining the coefficients of the component prediction model is described in Part IV (VI) above.
[0136] Fig.11 Conceptually illustrates the process 1100 of constraining the coefficients of the component prediction model. In some embodiments, one or more processing units (e.g., processors) of the computing device implementing the decoder 900 execute the process 1100 by executing instructions stored in a computer-readable medium. In some embodiments, the electronic device implementing the decoder 900 executes the process 1100.
[0137] The decoder (at block 1110) receives the data to be decoded as the current block of the current picture of the video.
[0138] The decoder (at block 1120) derives a set of coefficients for a component prediction model based on corresponding input and output component samples. In some embodiments, the component prediction model is a cross-component model that generates predicted chroma samples based on the reconstructed luma samples of the current block, and the corresponding input and output component samples are the corresponding luma and chroma samples of a template region adjacent to the current block. In some embodiments, the component prediction model is a convolutional model that derives the set of coefficients by factorization and back substitution based on the autocorrelation matrix between the corresponding input and output component samples.
[0139] The decoder (at block 1130) constrains the derived set of coefficients according to a set of constraints. In some embodiments, the encoder constrains the derived set of coefficients by clipping coefficients at a clipping threshold, or clipping different coefficients at different clipping thresholds, or limiting the coefficients within a predefined range. In some embodiments, the derived set of coefficients and the constrained set of coefficients are represented in floating point in the encoder. In some embodiments, the coefficients are represented in fixed point with N bits in the fractional part, and the size of the predefined range is scaled based on 1<<N. In some embodiments, when the derived coefficients exceed the predefined range, the set of coefficients is set to be equal to an identity filter, or a clipping operation is applied to the out-of-range coefficients, or the derived set of coefficients is not used to encode or decode the current block (disable the CCCM mode).
[0140] The decoder (at block 1140) applies the constrained set of coefficients as a component prediction model to generate a predictor for the current block. The predictor may include the generated predicted chroma samples.
[0141] The decoder (at block 1150) reconstructs the current block by using the generated predictor. Then, the decoder may provide the reconstructed current block for display as part of the reconstructed current picture.
[0142] VII. Examples of Electronic Systems
[0143] Many of the features and applications described above are implemented as software processes that are specified as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, cores of a processor, or other processing units), they cause the processing unit to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disks, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Computer-readable media do not include carrier waves and electronic signals transmitted wirelessly or through a wired connection.
[0144] In this specification, the term "software" refers to firmware residing in read-only memory or application programs stored in magnetic memory, which can be read into memory for processing by a processor. In addition, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while maintaining different software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of independent programs that jointly implement the software inventions described herein are within the scope of this application. In some embodiments, the software program, when installed and run on one or more electronic systems, defines one or more specific machine implementations that execute and implement the operations of the software program.
[0145] Fig.12 The electronic system 1200 for implementing some embodiments of the present application is conceptually illustrated. The electronic system 1200 can be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a phone, a PDA, or any other type of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1200 includes a bus 1205, a processing unit 1210, a graphics processing unit (GPU) 1215, a system memory 1220, a network 1225, a read-only memory 1230, a permanent storage device 1235, an input device 1240, and an output device 1245.
[0146] The bus 1205 is collectively referred to as all system, peripheral device, and chipset buses, which communicatively connect the numerous internal devices of the electronic system 1200. For example, the bus 1205 communicatively connects the processing unit 1210 with the GPU 1215, the read-only memory 1230, the system memory 1220, and the permanent storage device 1235.
[0147] From these various storage units, the processing unit 1210 retrieves instructions to be executed and data to be processed to perform the processes of the present application. In different embodiments, the processing unit can be a single processor or a multi-core processor. Some instructions are passed to and executed by the GPU 1215. The GPU 1215 can offload various calculations or supplement the image processing provided by the processing unit 1210.
[0148] Read-only memory (ROM) 1230 stores static data and instructions and is used by processing unit 1210 and other modules of the electronic system. On the other hand, permanent storage device 1235 is a read-write storage device. This device is a non-volatile storage unit that can store instructions and data even when electronic system 1200 is turned off. Some embodiments of the present application use a large-capacity storage device (such as a magnetic or optical disk and its corresponding disk drive) as permanent storage device 1235.
[0149] Other embodiments use removable storage devices (such as floppy disks, flash memory devices, etc., and their corresponding disk drives) as permanent storage devices. Like permanent storage device 1235, system memory 1220 is a read-write memory device. However, unlike storage device 1235, system memory 1220 is a volatile read-write memory, such as random access memory. System memory 1220 stores some instructions and data used by the processor at runtime. In some embodiments, processes according to the contents of the present application are stored in system memory 1220, permanent storage device 1235, and / or read-only memory 1230. For example, various memory units include instructions for processing multimedia clips according to some embodiments. From these different storage units, processing unit 1210 retrieves instructions to be executed and data to be processed in order to execute the processes of some embodiments.
[0150] Bus 1205 also connects to input and output devices 1240 and 1245. Input device 1240 enables a user to communicate information and select commands to the electronic system. Input device 1240 includes an alphanumeric keyboard and pointing device (also known as a "cursor control"), a camera (such as a web camera), a microphone or similar device for receiving voice commands, etc. Output device 1245 displays images generated by the electronic system or otherwise outputs data. Output device 1245 includes a printer and display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD), and a speaker or similar audio output device. Some embodiments include devices such as a touch screen that serve as both input and output devices.
[0151] Finally, if Fig.12 As shown, bus 1205 also couples electronic system 1200 to a network 1225 via a network adapter (not shown). In this manner, the computer can be part of a computer network, such as a local area network ("LAN"), a wide area network ("WAN"), or an intranet, or a network of networks, such as the Internet. Any or all components of electronic system 1200 may be used in conjunction with the present disclosure.
[0152] Some embodiments include electronic components, such as a microprocessor, a storage device, and a memory, which stores computer program instructions to a machine-readable medium or a computer-readable medium (optionally referred to as a computer-readable storage medium, a machine-readable medium, or a machine-readable storage medium). Some examples of computer-readable media include RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD-RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini SD card, micro SD card, etc.), magnetic and / or solid-state hard drives, read-only and recordable Computer readable media may store computer programs executed by at least one processing unit and include sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as that produced by a compiler, and files containing high-level code executed by a computer, electronic component, or microprocessor using an interpreter.
[0153] While the above discussion refers primarily to microprocessors or multi-core processors that execute software, many of the above functions and applications are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuits themselves. In addition, some embodiments execute software stored in programmable logic devices (PLDs), ROMs, or RAM devices.
[0154] The terms "computer," "server," "processor," and "memory" as used in this application and any claims hereof refer to electronic or other technological devices. These terms do not include a person or group of persons. For the purposes of this specification, the terms displayed or shown refer to displaying on an electronic device. The terms "computer-readable medium," "computer-readable media," and "machine-readable medium" as used in this application and any claims hereof are entirely limited to tangible, physical objects that store information in a computer-readable form. These terms do not include any wireless signals, wired download signals, and any other transient signals.
[0155] Although the present disclosure has been described with reference to many specific details, those skilled in the art will recognize that the present disclosure may be embodied in other specific forms without departing from the spirit of the present disclosure. Figure 8 and Fig.11 ) conceptually illustrate the processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed continuously in a series of operations, and different specific operations may be performed in different embodiments. In addition, the process may be implemented using several sub-processes or as part of a larger macro process. Therefore, it will be understood by those skilled in the art that the present disclosure should not be limited by the above illustrative details, but should be defined by the appended claims.
[0156] Additional Notes
[0157] The subject matter described herein sometimes illustrates different components contained in or connected to different other components. It should be understood that these depicted architectures are only examples, and many other architectures can actually be implemented to achieve the same functionality. Conceptually, any arrangement of components to achieve the same functionality is effectively "associated" to achieve the desired functionality. Therefore, any two components combined to achieve a specific function herein can be regarded as "associated" to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components that are so associated can also be regarded as "operably connected" or "operably coupled" to achieve the desired functionality, and any two components that can be so associated can also be regarded as "operably coupled" to achieve the desired functionality. Specific examples of operable coupling include, but are not limited to, physically matable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interacting components.
[0158] In addition, with respect to the use of almost any plural and / or singular terms herein, those skilled in the art may convert from the plural to the singular and / or from the singular to the plural as the context and / or application requires. For clarity, various singular / plural permutations may be explicitly listed herein.
[0159] Furthermore, those skilled in the art will appreciate that generally, the terms used herein, particularly in the appended claims, such as the bodies of the appended claims, are generally intended to be “open” terms, e.g., the word “including” should be interpreted as “including but not limited to,” the word “having” should be interpreted as “having at least,” the word “comprising” should be interpreted as “including but not limited to,” etc. Those skilled in the art will also appreciate that if a specific number of an introduced claim recitation is intended, that intent will be expressly stated in the claim, and in the absence of such a statement, no such intent is present. For example, to aid understanding, the following appended claims may contain the use of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of these phrases should not be construed as implying that a claim recitation introduced by the indefinite article “a” or “an” limits any claim containing such an introduction to only one such recitation, even if the same claim includes the introductory phrases “one or more” or “at least one” and an indefinite article such as “a” or “an,” e.g., “a” and / or “an” should be interpreted as “at least one” or “one or more”; the same applies to definite articles used to introduce claim recitations. Furthermore, even if a specific number of an introduced claim recitation is explicitly stated, one skilled in the art will recognize that such a statement should be interpreted as at least that number, e.g., simply stating "two recitations," without other modifiers, means at least two recitations, or two or more recitations. Furthermore, where a convention similar to "at least one of A, B, and C, etc." is used, such a construction is generally understood by one skilled in the art, e.g., "a system having at least one of A, B, and C" would include, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together. Where a convention similar to "at least one of A, B, or C, etc." is used, such a construction is generally understood by one skilled in the art, e.g., "a system having at least one of A, B, or C" would include, but is not limited to, systems having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together. One skilled in the art will also understand that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to include one term, either term, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0160] As can be seen from the above, various embodiments of the present disclosure are described herein for illustrative purposes, and various modifications may be made without departing from the scope and spirit of the present disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, and the true scope and spirit are indicated by the following claims.
Claims
1. A method for video encoding and decoding, the method comprising: Receiving data of a pixel block to be encoded or decoded as a current block of a current picture of a video; Deriving a set of coefficients based on corresponding input and output component samples; Constraining the derived set of coefficients according to a set of constraints; Applying the constrained set of coefficients as a component prediction model to generate a predictor for the current block; And Encoding or decoding the current block by using the generated predictor.
2. The method according to claim 1, characterized in that The component prediction model is a cross-component model that generates predicted chrominance samples based on reconstructed luminance samples of the current block.
3. The method according to claim 1, characterized in that: The corresponding input and output component samples are corresponding luminance and chrominance samples of a template region adjacent to the current block.
4. The method according to claim 1, characterized in that: The component prediction model is a convolutional model; the set of coefficients is derived by solving a matrix equation between the corresponding input and output component samples.
5. The method according to claim 1, characterized in that The step of constraining the derived set of coefficients includes: clipping coefficients at a clipping threshold.
6. The method according to claim 1, characterized in that The step of constraining the derived set of coefficients includes: clipping different coefficients at different clipping thresholds.
7. The method according to claim 1, characterized in that The derived set of coefficients and the constrained set of coefficients are represented in floating point in a video codec.
8. The method according to claim 1, characterized in that The step of constraining the derived set of coefficients includes: restricting the coefficients within a predefined range.
9. The method according to claim 8, characterized in that The derived set of coefficients and the constrained set of coefficients are represented in fixed point in a video codec, and its fractional part includes N bits, and the size of the predefined range is magnified based on 1<<N.
10. The method according to claim 8, characterized in that When the derived coefficients exceed the predefined range, the set of coefficients is set to be equal to an identity filter.
11. The method according to claim 8, characterized in that When the derived coefficients exceed the predefined range, a clipping operation is applied to the out-of-range coefficients.
12. The method according to claim 11, characterized in that When the derived coefficients exceed the predefined range, the derived set of coefficients is not used to encode or decode the current block.
13. An electronic device, the electronic device comprising: A video encoding and decoding circuit configured to perform including: Receiving data of a pixel block to be encoded or decoded as a current block of a current picture of a video; Deriving a set of coefficients based on corresponding input and output component samples; Constraining the derived set of coefficients according to a set of constraints; Applying the constrained set of coefficients as a component prediction model to generate a predictor for the current block; and Encoding or decoding the current block by using the generated predictor.
14. A video decoding method, the method comprising the steps of: Receiving data of a pixel block to be decoded as a current block of a current picture of a video; Deriving a set of coefficients based on corresponding input and output component samples; Constraining the derived set of coefficients according to a set of constraints; Applying the constrained set of coefficients as a component prediction model to generate a predictor for the current block; And Reconstructing the current block by using the generated predictor.
15. A video encoding method, the method comprising the following steps: Receiving data of a pixel block to be encoded as a current block of a current picture of a video; Deriving a set of coefficients based on corresponding input and output component samples; Constraining the derived set of coefficients according to a set of constraints; applying the constrained set of coefficients as a component prediction model to generate a predictor for the current block; as well as The current block is encoded by using the generated predictor.