Implicit linear model derivation method and apparatus for cross component prediction using multiple reference lines
By introducing multi-reference lines and cross-component linear model predictors in the video encoding and decoding system, the problem of low encoding and decoding efficiency in the prior art is solved, and more efficient video data processing and prediction capabilities are achieved.
Patent Information
- Application Number
- CN202380070588.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2023-09-28
- Publication Date
- 2025-05-13
AI Technical Summary
When processing video data, existing video codec systems are difficult to effectively utilize multi-reference lines and cross-component linear model predictors, resulting in low encoding and decoding efficiency.
A prediction method in a video codec system is proposed. By receiving input data associated with the current block, determining the template of the current block, selecting the first and second reference lines based on available reference lines, derives the parameter set of cross component model for generating prediction samples and calculating template costs, and ultimately decides the codec mode and reference lines used.
By mixing cross-component linear model and multi-reference line technology, the efficiency of video encoding and decoding is improved, the prediction ability of video data is enhanced, and the error in the encoding and decoding process is reduced.
Smart Images

Figure CN119999191A_ABST
Abstract
Description
[0001] Cross-references
[0002] The present invention claims priority to U.S. Provisional Patent Application No. 63 / 378,704 filed on October 7, 2022, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present invention relates to a video coding system, and more particularly to a hybrid cross-component linear model predictor and a novel video coding tool using multiple reference lines in a video coding system. Background Art
[0004] Versatile video coding (VVC) is the latest international video coding standard developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Video Experts Team (JVET). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Codec representation of immersive media - Part 3: Versatile video coding, published in February 2021. VVC is developed on the basis of its predecessor High Efficiency Video Coding (HEVC), adding more codec tools to improve codec efficiency, as well as to handle various types of video sources including three-dimensional (3D) video signals.
[0005] Figure 1AAn exemplary adaptive inter / intra video codec system including a loop process is shown. For intra prediction 110, prediction data is derived based on previously encoded and decoded video data in the current picture. For inter prediction 112, motion estimation (ME) is performed on the encoder side, and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects intra prediction 110 or inter prediction 112, and provides the selected prediction data to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by transform (T) 118 and then by quantization (Q) 120. The entropy encoder 122 then encodes the transformed and quantized residual for inclusion in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged together with auxiliary information such as motion and coding modes associated with intra prediction and inter prediction, and other information such as parameters associated with loop filters applied to underlying image areas. As Figure 1A As shown, auxiliary information associated with intra prediction 110, inter prediction 112, and loop filter 130 is provided to entropy encoder 122. When inter prediction mode is used, the reference picture or pictures must also be reconstructed at the encoder end. Therefore, the transformed and quantized residual is processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to recover the residual. The residual is then added back to the prediction data 136 at reconstruction (REC) 128 to reconstruct the video data. The reconstructed video data can be stored in a reference picture buffer 134 and used to predict other frames.
[0006] like Figure 1A As shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subjected to various impairments due to the series of processes. Therefore, a loop filter 130 is usually applied to the reconstructed video data before storing it in the reference picture buffer 134 to improve the video quality. For example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) may be used. The loop filter information may need to be merged into the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 to be merged into the bitstream. In Figure 1A In , a loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in a reference picture buffer 134 . Figure 1AThe system in is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to a High Efficiency Video Coding (HEVC) system, VP8, VP9, H.264, VVC, or any other video codec standard.
[0007] like Figure 1B As shown, the decoder can use similar or partially identical functional blocks as the encoder, except for transform 118 and quantization 120, because the decoder only needs inverse quantization 124 and inverse transform 126. Instead of using the entropy encoder 122, the decoder uses an entropy decoder 140 to decode the video bitstream into quantized transform coefficients and required coding information (such as ILPF information, intra-frame prediction information, and inter-frame prediction information). The intra-frame prediction 150 on the decoder side does not need to perform a pattern search. Instead, the decoder only needs to generate an intra-frame prediction based on the intra-frame prediction information received from the entropy decoder 140. In addition, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from the entropy decoder 140, without the need for motion estimation.
[0008] According to VVC, the input picture is divided into non-overlapping square areas, called coding tree units (CTUs), similar to HEVC. Each CTU can be divided into one or more smaller coding units (CUs). The resulting CU partitions can be square or rectangular. In addition, VVC divides CTUs into prediction units (PUs) as a unit to apply prediction processing, such as inter-frame prediction, intra-frame prediction, etc.
[0009] In the present disclosure, various new coding tools are proposed to improve the coding efficiency beyond VVC. Specifically, a new coding tool involving a hybrid cross-component model mode and using multiple reference lines is disclosed. Summary of the invention
[0010] The present invention discloses a prediction method and device in a video coding system. According to the method, input data associated with a current block including a first color component and a second color component is received, wherein the input data includes pixel data to be encoded on the encoder side or encoded data associated with the current block to be decoded on the decoder side. One or more templates of the current block are determined, wherein the one or more templates of the current block are located in a previously coded area of the current block. One or more first reference lines are selected from the available reference lines of the current block. Based on the first color sample and the second color sample in the one or more first reference lines, a cross component model parameter set of each cross component in one or more cross component models is derived. By applying the cross component model parameter set associated with each cross component model in the one or more cross component models to the first color component of the one or more templates, the predicted samples of the one or more templates associated with each cross component model in the one or more cross component models of the second color component of the one or more templates are derived. Based on the predicted samples of the one or more templates associated with each cross component model in the one or more cross component models and the reconstructed samples of the one or more templates, the template cost associated with each cross component model in the one or more cross component models is determined. One or more target modes from the one or more cross-component models for a second color component of a current block are determined based on a template cost associated with the one or more cross-component models. One or more second reference lines are selected from available reference lines of the current block, wherein the one or more second reference lines are allowed to be the same as the one or more first reference lines. If the one or more second reference lines are different from the one or more first reference lines, a cross-component model parameter set of the one or more target modes is derived based on a first color sample and a second color sample in the one or more second reference lines. The second color component of the current block is encoded or decoded using codec information including the one or more target modes and a cross-component model parameter set derived from the one or more second reference lines.
[0011] In one embodiment, the template cost associated with each of the one or more cross-component models corresponds to the distortion between the predicted samples of the one or more templates associated with each of the one or more cross-component models and the reconstructed samples of the one or more templates. In one embodiment, the one or more target modes selected from the one or more cross-component models have the smallest template cost among the template costs associated with the one or more cross-component models.
[0012] In one embodiment, the current block is encoded or decoded using a final predictor by mixing a first predictor associated with a first prediction mode and a second predictor associated with a second prediction mode, wherein at least one of the first prediction mode and the second prediction mode is implicitly determined based on a template cost associated with the one or more cross-component models. In one embodiment, the first prediction mode is explicitly determined by sending or parsing an index, and the one or more cross-component models are associated with a cross-component model set, and the second prediction mode is implicitly determined based on a template cost associated with the one or more cross-component models in the cross-component model set. In one embodiment, the final predictor is generated as a weighted sum of the first predictor and the second predictor. In one embodiment, a weight factor of the weighted sum of the first predictor and the second predictor is implicitly determined based on a first template cost associated with the first prediction mode and a second template cost associated with the second prediction mode. In one embodiment, the first template cost is determined by deriving a first cross-component parameter set based on samples on the one or more first reference lines and the sent first prediction mode, and then determined based on predicted samples of the one or more templates and reconstructed samples of the one or more templates, wherein the predicted samples of the one or more templates are generated using the derived first cross-component parameter set according to the sent first prediction mode.
[0013] In one embodiment, a flag is sent or parsed to indicate whether the current block has multiple reference line (MRL) mode enabled. In one embodiment, when the flag indicates that the current block has MRL mode enabled, a second prediction mode is determined and a weighting factor of a weighted sum of a first predictor and a second predictor is derived. In one embodiment, the first predictor is generated based on a cross-component parameter set, and the cross-component parameter set is derived based on the first prediction mode and the one or more second reference lines. In one embodiment, the one or more first reference lines from the available reference lines correspond to one or more external neighboring reference lines of the one or more templates of the current block, and the second reference line is the same as the first reference line.
[0014] In one embodiment, the one or more cross-component models are associated with a second cross-component model set, and the flag is sent only when the first prediction mode is in the second cross-component model set. In one embodiment, if the first prediction mode is not in the second cross-component model set, the flag is not sent and the MRL mode is implicitly determined to be disabled.
[0015] In one embodiment, the one or more cross-component models are associated with a first cross-component model set and a second cross-component model set, wherein the first target mode is derived from the first cross-component model set and the second target mode is derived from the second cross-component model set. In one embodiment, the final predictor is generated as a weighted sum of the first predictor and the second predictor. In one embodiment, a weighting factor of the weighted sum of the first predictor and the second predictor is implicitly determined based on a first template cost associated with the first prediction mode and a second template cost associated with the second prediction mode. In one embodiment, a flag is issued or parsed to indicate whether a multi-reference line (MRL) mode is enabled for the current block. In one embodiment, when the flag indicates that the MRL mode is enabled for the current block, the first prediction mode and the second prediction mode are determined, and the weighting factor of the weighted sum of the first predictor and the second predictor is derived.
[0016] In one embodiment, the first cross-component model set and the second cross-component model set contain only linear model modes. In one embodiment, the first prediction mode is implicitly determined from the first cross-component model set and the second cross-component model set. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A An exemplary adaptive inter / intra video coding system incorporating loop processing is shown.
[0018] Figure 1B Show Figure 1A The corresponding decoder of the encoder in .
[0019] Figure 2 Examples of direction (angle), plane, and DC mode of intra prediction according to the VVC standard are shown.
[0020] Figure 3A -B shows blocks that are wider than they are tall ( Figure 3A ) and blocks that are taller than they are wide ( Figure 3B ) is an example of wide-angle intra prediction.
[0021] Figure 4 An example of the top left sample of the current block and the positions of the samples involved in the CCLM_LT mode is shown.
[0022] Figure 5 An example of a template-based intra mode derivation (TIMD) mode is shown, where TIMD implicitly derives the intra prediction mode of a CU using neighboring templates at the encoder and decoder.
[0023] Figure 6 An example of the spatial part of the convolution filter of the CCCM is shown.
[0024] Figure 7 An example of a reference region (and its filling) used to derive CCCM filter coefficients is shown.
[0025] Figure 8 Examples of four gradient modes of the Gradient Linear Model (GLM) are shown.
[0026] Fig. 9 An example is shown in which an outer neighboring line of a template is selected as a reference line of the template and an outer neighboring line of a current block is selected as a reference line of a coding block.
[0027] Fig.10 It is shown that the outer neighboring lines of the template are selected as the reference lines of the template and the reference lines of the coding block are also examples of the outer neighboring lines of the template.
[0028] Fig.11 It is shown that both the reference line of the template and the reference line of the current block are examples of outer neighboring lines of the current block.
[0029] Fig.12 It is shown that both the reference line of the template and the reference line of the current block are examples of outer neighboring lines of the current block.
[0030] Fig.13 An example of multiple reference lines of the current block to derive cross-component model parameters is shown.
[0031] Fig.14 A flowchart of an example video encoding system for implicitly deriving a linear model predictor based on Template-based Intra Mode Derivation (TIMD) according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0032] It is readily understood that the components of the present invention as generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of the embodiments of the systems and methods of the present invention shown in the drawings is not intended to limit the scope of the present invention, but is merely representative of selected embodiments of the present invention. References in this specification to "one embodiment," "embodiment," or similar language mean that a particular feature, structure, or characteristic associated with an embodiment may be included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification do not necessarily all refer to the same embodiment.
[0033] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the relevant art will recognize that the present invention may be implemented without one or more specific details, or implemented using other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the present invention. The illustrated embodiments of the present invention will be best understood by reference to the accompanying drawings, in which like parts are always represented by like numbers. The following description is by way of example only, and only certain selected apparatus and method embodiments consistent with the invention claimed herein are described.
[0034] The VVC standard combines various new codec tools to further improve the codec efficiency of the HEVC standard. Among the various new codec tools, some codec tools related to the present invention are reviewed below.
[0035] Intra-mode codec with 67 intra-prediction modes
[0036] In order to capture any edge direction present in natural videos, the number of directional intra modes in VVC is expanded from 33 used in HEVC to 65. New directional (angle) modes not in HEVC are introduced in Figure 2 Indicated by dashed arrows in , planar and DC modes remain unchanged. These more dense directional intra prediction modes apply to all block sizes and both luma and chroma intra prediction.
[0037] In VVC, for non-square blocks, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes.
[0038] In HEVC, each intra-coded block has square shape, where the length of each side is a power of 2. Therefore, no division operation is required to generate intra predictors using DC mode. In VVC, blocks can have rectangular shapes, which in general requires the use of division operations for each block. To avoid division operations for DC prediction, only the longer sides are used to calculate the average for non-square blocks.
[0039] Wide-angle intra prediction for non-square blocks
[0040] The traditional angle intra prediction direction is defined as clockwise from 45 degrees to -135 degrees. In VVC, several traditional angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are transmitted using the original mode index and remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding and decoding method remains unchanged.
[0041] To support these prediction directions, the top reference of length 2W+1 and the left reference of length 2H+1 are respectively as follows Figure 3A and Figure 3B shown.
[0042] The number of selectable modes in the wide-angle direction mode depends on the aspect ratio of the block. The selectable intra prediction modes are shown in Table 1.
[0043] Table 1 – Intra prediction modes replaced by Wide mode
[0044]
[0045] In VVC, 4:2:2 and 4:4:4 chroma formats are supported as well as 4:2:0. The chroma derived mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, and the number of entries was expanded from 35 to 67 to keep pace with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135° and above 45°, the luma intra prediction modes from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing specific values of the mapping table entries to more accurately convert the prediction angles of the chroma blocks.
[0046] Cross Component Linear Model
[0047] The main idea behind the CCLM mode (sometimes abbreviated to LM mode) is as follows: the chrominance components of a block can be predicted from the co-located reconstructed luma samples by a linear model whose parameters come from the reconstructed luma and chroma samples neighboring the block.
[0048] In VVC, the CCLM mode exploits inter-channel dependency by predicting chrominance samples from reconstructed luma samples. This prediction is performed using a linear model of the following form:
[0049] P(i,j)=a·rec′ L (i,j)+b. (1)
[0050] Here, P(i,j) represents the predicted chrominance samples in CU, and d rec′ L (i,j) represents the reconstructed luma samples of the same CU, which are downsampled in non-4:4:4 color formats. The model parameters a and b are derived based on the adjacent luma and chroma samples reconstructed at the encoder and decoder ends and do not need to be sent explicitly.
[0051] Three CCLM modes are specified in VVC, namely CCLM_LT, CCLM_L and CCLM_T. These three modes differ in the location of the reference samples used for model parameter derivation. The CCLM_T mode only involves samples from the top boundary, while the CCLM_L mode only involves samples from the left boundary. In the CCLM_LT mode, samples from the top boundary and the left boundary are used.
[0052] In general, the forecasting process of CCLM model includes three steps:
[0053] 1) Downsample the luminance block and its adjacent reconstructed samples to match the size of the corresponding chrominance block;
[0054] 2) deriving model parameters based on the reconstructed adjacent samples;
[0055] 3) Apply model equation (1) to generate chroma intra prediction samples.
[0056] Downsampling of luma components: To match the chroma sample positions of a 4:2:0 or 4:2:2: color format video sequence, two types of downsampling filters can be applied to luma samples, both with a 2:1 downsampling ratio in both horizontal and vertical directions. These two filters correspond to "type 0" and "type 2" 4:2:0 chroma format content, respectively, and are given by the following formulas
[0057]
[0058] A two-dimensional 6-tap (i.e., f2) or 5-tap (i.e., f1) filter is applied to the luma samples in the current block and its neighboring luma samples according to the sequence parameter set (SPS) level flag information. The SPS level refers to the sequence parameter set level. An exception occurs if the top line of the current block is a CTU boundary. In this case, a one-dimensional filter [1,2,1] / 4 is applied to the above neighboring luma samples to avoid using multiple luma lines above the CTU boundary.
[0059] Model parameter derivation process: The model parameters a and b derived from equation (1) are based on adjacent luminance and chrominance samples reconstructed at the encoder and decoder to avoid any signaling overhead. In the initial CCLM mode version, a linear minimum mean square error (LMMSE) estimator was used to derive the parameters. However, in the final design, only four samples are involved to reduce the computational complexity. Figure 4 The relative sample positions of an MxN chroma block 410, a corresponding 2Mx2N luma block 420, and neighboring samples (shown as filled circles and triangles) of their "type 0" content are shown.
[0060] Figure 4 In the example, the four samples used in CCLM_LT mode are marked with triangles and are located at the M / 4 and M·3 / 4 positions of the upper boundary and the N / 4 and N·3 / 4 positions of the left boundary. In CCLM_T and CCLM_L modes, the upper boundary and the left boundary are extended to (M+N) sample sizes, and the four samples used for model parameter derivation are located at (M+N) / 8, (M+N)·3 / 8, (M+N)·5 / 8, and (M+N)·7 / 8 positions.
[0061] After selecting 4 samples, 4 comparison operations are used to determine the minimum two brightness sample values and the maximum two brightness sample values. l Represents the average of the two maximum brightness sample values, let X s Represents the average of the two minimum brightness sample values. Similarly, let Y l and Y s Represents the average value of the corresponding chromaticity sample value. Then the linear model parameters are obtained according to the following formula:
[0062]
[0063] b=Y s -a·X s (3)
[0064] In this equation, the division operation to calculate parameter a is implemented using a lookup table. To reduce the memory required to store this table, the diff value (the difference between the maximum and minimum values) and parameter a are expressed in exponential notation. Here, the value of diff is approximated using a 4-bit significant part and an exponent. Therefore, the table for 1 / diff contains only 16 elements. This reduces both the complexity of the calculation and the size of the memory required to store the table.
[0065] MMLM Overview
[0066] As the name implies, the original CCLM mode uses a linear model to predict chrominance samples from the luma samples of the entire CU, while in the Multiple Model CCLM (MMLM), there can be two models. In MMLM, the adjacent luma samples and adjacent chroma samples of the current block are divided into two groups, each group is used as a training set to derive a linear model (i.e., a specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified according to the same rules as the adjacent luma samples.
[0067] o The threshold is calculated as the average of adjacent reconstructed luminance samples. Adjacent samples with Rec′L[x,y]<=threshold are grouped into group 1, while adjacent samples with Rec′L[x,y]>threshold are grouped into group 2.
[0068] o Accordingly, the chrominance prediction is obtained using a linear model:
[0069]
[0070] Chroma Intra Mode Codec
[0071] For chroma intra mode encoding and decoding, a total of 8 intra modes are allowed for chroma intra mode encoding and decoding. These modes include five traditional intra modes and three cross-component linear model modes ({LM_LA, LM_L and LM_A} or {CCLM_LT, CCLM_L and CCLM_T}). In the present disclosure, the terms {LM_LA, LM_L, LM_A} and {CCLM_LT, CCLM_L, CCLM_T} are used interchangeably. Table 2 shows the chroma mode signal transmission and derivation process. Chroma mode encoding and decoding depends directly on the intra prediction mode of the corresponding luminance block. Since the separate block partition structure of luminance and chrominance components is enabled in the I segment, one chroma block may correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited.
[0072] Table 2 – Chroma prediction modes derived from luma mode when sps_cclm_enabled_flag is true
[0073]
[0074] Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 3.
[0075] Table 3 - Unified binarization table of chrominance prediction modes
[0076]
[0077] In Table 3, the first bin indicates whether it is normal (0) or LM mode (1). If it is LM mode, the next bin indicates whether it is LM_LA (0). If it is not LM_LA, the next 1bin indicates whether it is LM_L (0) or LM_A (1). For this case, when sps_cclm_enabled_flag is 0, the first bin of the binarization table of the corresponding intra_chroma_pred_mode can be discarded before entropy encoding and decoding. Or in other words, the first bin is inferred to be 0 and therefore not encoded and decoded. This single binarization table is used for the cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two bins in Table 3 use their own context model for context encoding and decoding, and the remaining bins are bypass encoded and decoded.
[0078] Multi-Hypothesis Prediction (MHP)
[0079] In the multi-hypothesis inter-frame prediction mode (JVET-M0425), in addition to the traditional bidirectional prediction signal, one or more additional motion compensation prediction signals are also sent. The final overall prediction signal is obtained by weighted superposition of samples. bi and the first additional inter-frame prediction signal / hypothesis h3, the final prediction signal p3 can be obtained as follows:
[0080] p3=(1-α)p bi +αh3 (4)
[0081] The weighting factor α is specified by a new syntax element add_hyp_weight_idx, according to the following mapping (Table 4):
[0082] Table 4. Mapping α to add_hyp_weight_idx
[0083] add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0084] Similar to the above, multiple additional prediction signals may be used. The final overall prediction signal is iteratively accumulated through each additional prediction signal.
[0085] p n+1 = (1-α n+1 )p n +α n+1 h n+1 (5)
[0086] The final overall prediction signal is the last p n (i.e., the p with the largest index n n ). For example, at most two additional prediction signals may be used (ie, n is limited to 2).
[0087] The motion parameters for each additional prediction hypothesis can be signaled either explicitly by specifying the reference index, motion vector predictor index and motion vector difference, or implicitly by specifying the merge index. A separate multi-hypothesis merge flag distinguishes these two signaling modes.
[0088] For inter-AMVP mode, MHP is applied only when non-equal weights in BCW are selected in bi-prediction mode. Detailed information on MHP for VVC can be found in JVET-W2025 (Muhammed Coban et al., Algorithmic Description of Enhanced Compression Model 2 (ECM 2), Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29, 23rd Meeting, by Teleconference, July 7-16, 2021, Document: JVET-W2025).
[0089] MHP and BDOF can be used in combination, but BDOF is only applicable to the bidirectional prediction signal part of the prediction signal (ie, the ordinary first two hypotheses).
[0090] Template-based Intra Mode Derivation (TIMD)
[0091] The template-based intra mode derivation (TIMD) mode uses neighboring templates at the encoder and decoder to implicitly derive the intra prediction mode of the CU, instead of sending the intra prediction mode to the decoder. Figure 5 As shown, the prediction samples of the template (512 and 514) of the current block 510 are generated using the reference samples (520 and 522) of the template of each candidate mode. The cost is calculated as the sum of the absolute transform differences (SATD) between the prediction samples and the reconstructed samples of the template. The intra prediction mode with the lowest cost is selected as the TIMD mode, and used for intra prediction of the CU. The candidate mode can be 67 intra prediction modes like in VVC, or it can be extended to 131 intra prediction modes. Typically, MPM can provide clues to indicate the directional information of the CU. Therefore, in order to reduce the intra mode search space and utilize the characteristics of the CU, the intra prediction mode can be implicitly derived from the MPM list.
[0092] For each intra prediction mode in MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The first two intra prediction modes with the smallest SATD are selected as TIMD modes. After applying the PDPC process, the two TIMD modes are fused with weights, and the weighted intra prediction is used to encode and decode the current CU. The position dependent intra prediction combination (PDPC) is included in the derivation of the TIMD mode.
[0093] The costs of the two selected patterns are compared with a threshold, and in the test, a cost factor of 2 is applied as shown below:
[0094] costMode2<2*costMode1.
[0095] If this condition is true, fusion is applied, otherwise only mode 1 is used. The weights of the modes are calculated based on their SATD costs as follows:
[0096] weight1=costMode2 / (costMode1+costMode2)
[0097] weight2=1-weight1.
[0098] Fusion of Chroma Intra Prediction Modes
[0099] In the development of emerging video codec systems, it is proposed to merge the DM mode and the four default modes with the MMLM_LT mode as shown below:
[0100] pred=(w0*pred0+w1*pred1+(1<<(shift-1)))>>shift.
[0101] In the above formula, pred0 is the predictor obtained by applying the non-LM mode, pred1 is the predictor obtained by applying the MMLM_LT mode, and pred is the final predictor of the current chroma block. The two weights w0 and w1 are determined by the intra prediction mode of the adjacent chroma block, and shift is set to 2. Specifically, when the upper adjacent block and the left adjacent block are both coded and decoded using the LM mode, {w0,w1}={1,3}; when the upper adjacent block and the left adjacent block are both coded and decoded using the non-LM mode, {w0,w1}={3,1}; otherwise, {w0,w1}={2,2}.
[0102] For syntax design, if non-LM mode is selected, a flag is emitted to indicate whether fusion is applied, and the proposed fusion is only applied to I segments.
[0103] Convolutional Cross-component Model (CCCM)
[0104] The convolutional cross-component model (CCCM) is used to predict chroma samples from reconstructed luma samples, similar to the current CCLM mode. Like CCLM, when chroma subsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid.
[0105] In addition, similar to CCLM, you can choose to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one for samples above the average brightness reference value and another for the remaining samples (following the spirit of CCLM design). For PUs with at least 128 reference samples, the multi-model CCCM mode can be selected.
[0106] In CCCM, a convolutional model is adopted to improve the chrominance prediction performance. The convolutional model has a 7-tap filter consisting of a 5-tap plus-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of a center (C) luma sample that is co-located with the chrominance sample to be predicted and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as Figure 6 shown.
[0107] The non-linear term (denoted as P) is expressed as a power of 2 of the center luminance sample C, and scaled to the sample value range of the content:
[0108] P=(C*C+midVal)>>bitDepth.
[0109] For example, for 10-bit content, the nonlinear term is calculated as follows:
[0110] P=(C*C+512)>>10
[0111] The bias term (denoted as B) represents a scalar offset between input and output (similar to the offset term in CCLM), and is set to the intermediate chrominance value (512 for 10-bit content).
[0112] The output of the filter is calculated as the convolution between the filter coefficients ci and the input value, and is clipped to the valid chroma sample range:
[0113] predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B
[0114] The filter coefficients ci are calculated by minimizing the mean squared error (MSE) between the predicted and reconstructed chrominance samples in the reference region. Figure 7 An example of a reference region is shown, consisting of 6 strips of chroma samples above and to the left of PU 710. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only available samples. The region needs to be expanded (indicated as a "light grey square") to support Figure 7 The "side samples" of the plus-sign spatial filter (i.e. Figure 5 N, S, W, and E samples in ), and fill in the unavailable areas.
[0115] MSE minimization is performed by computing the autocorrelation matrix of the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back substitution. The process roughly follows the calculation of the ALF filter coefficients in ECM, but the LDL decomposition is chosen instead of the Cholesky decomposition to avoid the use of square root operations.
[0116] Bitstream sending
[0117] The usage of the mode is signaled using a PU level flag for the CABAC codec. A new CABAC context is introduced to support this. In terms of transmission, CCCM is treated as a sub-mode of CCLM. That is, the CCCM flag is sent only when the intra prediction mode is LM_CHROMA (this mode indicates that LM mode is being used).
[0118] Gradient linear model (GLM)
[0119] Compared with CCLM, GLM does not use downsampled brightness values, but uses brightness sample gradients to infer linear models. In other words, in the CCLM process, the gradient G is replaced instead of low-pass filtered brightness samples. The other designs of CCLM (such as parameter derivation, linear transformation of prediction samples) remain unchanged.
[0120] C=α·G+β
[0121] Figure 8 It is shown that G can be calculated by one of four Sobel-based gradient modes (810, 820, 830, and 840).
[0122] For transmission, when the current CU enables CCLM mode, two flags are sent for Cb / Cr components respectively to indicate whether GLM is enabled for the component; if GLM is enabled for a component, a syntax element is further sent to select Figure 8 The gradient is calculated in one of the four gradient modes.
[0123] Proposed method
[0124] In order to improve the efficiency of video coding and decoding, many coding tools are designed for more efficient transmission and / or generation of better block predictors. In the traditional mechanism, the inter-frame mode uses temporal information to predict the current block. For intra-frame blocks, spatially adjacent reference samples are used to predict the current block. In the present invention, a coding tool that creates a new hybrid mode and has multiple reference line selections to form a better predictor is provided. The concept of the coding tool is described as follows.
[0125] In one embodiment, the proposed encoding / decoding tool may be any cross-component tool that predicts and / or reconstructs samples of a currently encoded / decoded color component based on information of one or more other color components.
[0126] - For example, the cross-component tool generates prediction and / or reconstructed samples of Cb / Cr based on the luminance information.
[0127] - For another example, the cross-component tool generates prediction and / or reconstructed samples of Cr based on the information of Cb.
[0128] - As another example, the cross-component tool refers to CCLM, MMLM and / or GLM.
[0129] In another embodiment, an additional hypothesis for intra prediction (denoted as H2) is mixed with the existing intra prediction hypothesis (denoted as H1) at a predetermined weight to form a final prediction for the current block. H1 is generated by a mode (denoted as M1) selected from a first mode set (denoted as Set1), and H2 is generated by a mode (denoted as M2) selected from a second mode set (denoted as Set2). M1 and M2 can be any intra mode. Intra mode refers to a mode used to generate intra block predictions. For example, the previous 67 intra prediction modes, CCLM mode, and MMLM mode.
[0130] In another embodiment, M1 and M2 can be any LM mode. LM mode refers to a mode in which the chrominance component of a block is predicted by a linear model using co-located reconstructed luminance samples, where the parameters come from reconstructed luminance and chrominance samples adjacent to the block. LM modes include CCLM and MMLM.
[0131] In one embodiment, Set1 includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T, or any subset of the above modes, or any extension of the LM series.
[0132] In one embodiment, Set2 includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T, or any subset of the above modes, or any extension of the LM series.
[0133] In one embodiment, M1 is one of the CCLM modes, and M2 is one of the CCLM modes. Therefore, Set1 includes CCLM_LT, CCLM_L, CCLM_T, and Set2 also includes CCLM_LT, CCLM_L, CCLM_T. The combination of {M1, M2} (or {M2, M1}) includes:
[0134] ·{CCLM_LT, CCLM_L}, {CCLM_LT, CCLM_T}, {CCLM_L,
[0135] CCLM_T}
[0136] In another embodiment, M1 is one of the CCLM modes, and M2 is one of the MMLM modes. Therefore, Set1 includes CCLM_LT, CCLM_L, CCLM_T, and Set2 includes MMLM_LT, MMLM_L, MMLM_T. The combination of {M1, M2} includes:
[0137] ·{CCLM_LT,MMLM_LT},{CCLM_LT,MMLM_L},
[0138] {CCLM_LT,MMLM_T},
[0139] ·{CCLM_L,MMLM_LT},{CCLM_L,MMLM_L},{CCLM_L,MMLM_T}
[0140] ·{CCLM_T,MMLM_LT},{CCLM_T,MMLM_L},{CCLM_T,MMLM_T}
[0141] In another embodiment, M1 is one of the MMLM modes, and M2 is one of the MMLM modes. Therefore, Set1 includes MMLM_LT, MMLM_L, MMLM_T, and Set2 also includes MMLM_LT, MMLM_L, MMLM_T. The combination of {M1, M2} (or {M2, M1}) includes:
[0142] ·{MMLM_LT, MMLM_L}, {MMLM_LT, MMLM_T}, {MMLM_L, MMLM_T}
[0143] In another embodiment, M1 is one of the CCLM modes, and M2 can be one of the CCLM or MMLM modes. Therefore, Set1 includes CCLM_LT, CCLM_L, CCLM_T, and Set2 includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T. The combination of {M1, M2} (or {M2, M1}) includes:
[0144] ·{CCLM_LT,CCLM_L},{CCLM_LT,CCLM_T},{CCLM_L,CCLM_T},
[0145] The combination of {M1,M2} also includes:
[0146] ·{CCLM_LT, MMLM_LT}, {CCLM_LT, MMLM_L}, {CCLM_LT, MMLM_T},
[0147] ·{CCLM_L, MMLM_LT}, {CCLM_L, MMLM_L}, {CCLM_L, MMLM_T}
[0148] ·{CCLM_T, MMLM_LT}, {CCLM_T, MMLM_L}, {CCLM_T, MMLM_T}
[0149] In another embodiment, M1 can be one of CCLM or MMLM modes, and M2 is one of CCLM or MMLM modes. When M1 is CCLM mode, M2 must also be CCLM mode. When M1 is MMLM mode, M2 must be one of MMLM modes. The combination of {M1, M2} (or {M2, M1}) includes:
[0150] ·{CCLM_LT,CCLM_L},{CCLM_LT,CCLM_T},{CCLM_L,CCLM_T},
[0151] ·{MMLM_LT,MMLM_L},{MMLM_LT,MMLM_T},
[0152] {MMLM_L,MMLM_T}
[0153] In another embodiment, M1 is one of CCLM or MMLM modes, and M2 can be one of MMLM modes. Therefore, Set1 includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T, and Set2 includes MMLM_LT, MMLM_L, MMLM_T. The combination of {M1, M2} includes:
[0154] ·{CCLM_LT, MMLM_LT}, {CCLM_LT, MMLM_L}, {CCLM_LT, MMLM_T},
[0155] ·{CCLM_L, MMLM_LT}, {CCLM_L, MMLM_L}, {CCLM_L, MMLM_T}
[0156] ·{CCLM_T, MMLM_LT}, {CCLM_T, MMLM_L}, {CCLM_T, MMLM_T}
[0157] The combination of {M1, M2} (or {M2, M1}) also includes:
[0158] ·{MMLM_LT, MMLM_L}, {MMLM_LT, MMLM_T}, {MMLM_L, MMLM_T}
[0159] In another embodiment, M1 is one of CCLM or MMLM modes, and M2 can be one of CCLM or MMLM modes. Therefore, Set1 includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T, and Set2 also includes CCLM_LT, CCLM_L, CCLM_T, MMLM_LT, MMLM_L, MMLM_T. The combination of {M1, M2} (or {M2, M1}) includes:
[0160] ·{CCLM_LT, CCLM_L}, {CCLM_LT, CCLM_T}, {CCLM_L, CCLM_T}
[0161] ·{CCLM_LT, MMLM_LT}, {CCLM_LT, MMLM_L}, {CCLM_LT, MMLM_T},
[0162] ·{CCLM_L, MMLM_LT}, {CCLM_L, MMLM_L}, {CCLM_L, MMLM_T}
[0163] ·{CCLM_T, MMLM_LT}, {CCLM_T, MMLM_L}, {CCLM_T, MMLM_T}
[0164] ·{MMLM_LT, MMLM_L}, {MMLM_LT, MMLM_T}, {MMLM_L, MMLM_T}
[0165] In one embodiment, the codec tool is applied to blocks of any color component. For example, it can be applied to a luminance component and / or a chrominance component. For another example, it can be applied to a luminance component and / or a Cb component. For another example, it can be applied to a luminance component and / or a Cr component.
[0166] In another embodiment, the codec tool is applied to blocks of chrominance components. For example, it can be applied to Cb and / or Cr components.
[0167] In one embodiment, M1 and M2 are derived using the TIMD method, where the TIMD process is described as follows. Figure 5 As shown, the linear model of each allowed mode of M1 and M2 is derived using the reference sample of the template. Each linear model is applied to samples of one or more other color components of the template to generate predicted samples of the current encoded / decoded color component of the template. The TIMD cost is then calculated as the distortion between the predicted and reconstructed samples of the template. The mode with the smallest TIMD cost among the allowed modes is selected. For example, if M1 is allowed as one of the CCLM modes and M2 is allowed as one of the MMLM modes. The mode with the smallest TIMD cost among CCLM_LT, CCLM_L, and CCLM_T is selected as M1, and the mode with the smallest TIMD cost among MMLM_LT, MMLM_L, and MMLM_T is selected as M2. For another example, if M1 and M2 are allowed as one of the CCLM modes at the same time, the modes with the smallest and second smallest TIMD costs among CCLM_LT, CCLM_L, and CCLM_T are selected as M1 and M2.
[0168] In another embodiment, M1 is sent explicitly, and M2 is derived using the TIMD method. The linear model of each allowed mode of M2 is derived using the reference samples of the template. Each linear model is applied to samples of one or more other color components of the template to generate predicted samples of the current encoded / decoded color component of the template. The TIMD cost is then calculated as the distortion between the predicted and reconstructed samples of the template. The mode with the minimum TIMD cost in the allowed modes is selected as M2. For example, if M2 is allowed as one of the MMLM modes. The mode with the minimum TIMD cost in MMLM_LT, MMLM_L, MMLM_T is selected as M2.
[0169] In one embodiment, the blending weights of H1 and H2 are implicitly determined. For example, the weights may be inferred to be equal weights.
[0170] In another embodiment, the weight is implicitly determined by the TIMD cost. The TIMD cost is defined as follows. For each mode, Figure 5As shown, using a linear model derived from a reference sample based on the template, a predicted sample of the currently encoded / decoded color component of the template is generated by applying the linear model to samples of one or more other color components of the template. The TIMD cost is calculated as the SATD between the predicted sample and the reconstructed sample of the template. The TIMD cost can be determined at the decoder and does not need to be sent. For example, if the mode that generates H1 has a larger TIMD cost, H1 uses a smaller weight during mixing. If the mode that generates H2 has a larger TIMD cost, H2 uses a smaller weight during mixing. For another example, the weight is inversely proportional to the TIMD cost.
[0171] w h1 =TIMD_cost h2 / (TIMD_cost h1 +TIMD_cost h2 )
[0172] w h2 =TIMD_cost h1 / (TIMD_cost h1 +TIMD_cost h2 )
[0173] In another embodiment, if the mode that generates H1 has a much larger TIMD cost than H2, then only H2 is used to form the final prediction. For example, let k be a value greater than one. When TIMD_cost h1 >k*TIMD_cost h2 When , only H2 is used to form the final prediction.
[0174] In another embodiment, if the mode that generates H2 has a much larger TIMD cost than H1, only H1 is used to form the final prediction. For example, let k be a value greater than one. When TIMD_cost h2 >k*TIMD_cost h1 When , only H1 is used to form the final forecast.
[0175] In another embodiment, the weights are selected from a predetermined weight set. For example, the weight set includes {(0,4), (1,3), (2,2), (3,1), (4,0)}. After selecting modes M1 and M2, the template prediction based on M1 and the template prediction based on M2 are mixed according to the weights in the weight set, and the distortion corresponding to each weight in the weight set (i.e., the difference between the mixed template prediction generated according to the weight and the reconstructed sample of the template) is calculated, and then the weight with the smallest distortion from the weight set is selected.
[0176] In one embodiment, when the current block has multiple reference lines, such as Fig. 9As shown, the outer adjacent line (920 / 922) of the template (910 / 912) is selected as the reference line of the template. The linear model applied to the template is derived from the reference line of the template. Fig. 9 As shown, the outer neighboring lines (930 / 932) of the block are selected as reference lines for encoding the block. The linear model applied to the current block is derived from the reference lines of the current block.
[0177] In another embodiment, when the current block has multiple reference lines, such as Fig.10 As shown, the outer adjacent lines (920 / 922) of the template are selected as reference lines of the template (910 / 912). Fig.10 As shown, the reference line (920 / 922) of the coding block is also the outer neighbor line of the template. The linear model applied to the template and the linear model applied to the current block are both derived from the same reference line.
[0178] In one sub-embodiment, the selection of reference lines (used to generate prediction samples within the template and / or current block) depends on the template size (e.g., L2 for the top template and / or L1 for the left template), and the template size is adaptive based on the width, height and / or area of the current block, and / or defined based on explicit signaling. In more detail, the template includes a top template and / or a left template. The area of the top template is defined as the width of the current block multiplied by L2, while the area of the left template is defined as L1 multiplied by the height of the current block. In the present invention, L1 and L2 are jointly set to the same value or individually set to the same or different values. The method proposed in this embodiment can be applied to adaptively set L1 and / or L2.
[0179] In one example, for a block whose area is greater than a predetermined threshold (e.g., 4, 8, 16, 32, 64, 128, 256, or any positive integer), the settings of L1 and / or L2 may be as follows:
[0180] o On the one hand, the template size becomes larger (e.g., initially 2, and expanded to 4).
[0181] o Another way, the template size becomes smaller (e.g., initially 4, reduced to 2).
[0182] o Alternatively, the number of candidate template sizes is increased and a longer indication is derived / sent to specify the template size to use.
[0183] o Alternatively, the number of candidate template sizes is reduced and a shorter indication is derived / sent to specify the template size to use.
[0184] In another example, for blocks whose size (e.g., width or height) is greater than a predetermined threshold (e.g., 4, 8, 16, 32, 64, 128, 256, the maximum transform size specified in the standard, or any positive integer), the settings of L1 and / or L2 may be as follows:
[0185] o On the one hand, the template size becomes larger (for example, initially 2 and expanded to 4).
[0186] o Another way, the template size becomes smaller (for example, initially 4, and reduced to 2).
[0187] o Alternatively, the number of candidate template sizes is increased and a longer indication is derived / sent to specify the template size to use.
[0188] o Alternatively, the number of candidate template sizes is reduced and a shorter indication is derived / sent to specify the template size to use.
[0189] In another embodiment, the reference lines of the multiple color components used are aligned or related. That is, after determining the reference line of one color component used, the reference lines of the other color components used can be derived accordingly. Take CCLM / MMLM as an example. When deriving model parameters (such as scaling factor a and / or offset b), the input includes a downsampled reference brightness line and a reference chromaticity (Cb or Cr) line, and after determining the downsampled reference brightness line, the reference chromaticity line can be derived accordingly. In one embodiment, the reference chromaticity line is aligned with the downsampled reference brightness line. In another embodiment, the reference chromaticity line is derived by the downsampled reference brightness line and the offset, wherein the offset is implicitly set to a predetermined value, adaptively determined according to the width, height, and area of the current block, and / or explicitly determined according to the transmission of the block, fragment, picture, SPS or picture parameter set (Picture Parameter Set, PPS for short) level. In another approach, the downsampled reference luma line is derived from a reference chroma line and an offset, where the offset is implicitly set to a predetermined value, adaptively determined based on the width, height, area of the current block, and / or explicitly determined based on the transmission of the block, fragment, picture, SPS or PPS level.
[0190] In another embodiment, the term “block” in the present invention may refer to TU / TB, CU / CB, PU / PB, CTU / CTB, or any predetermined area.
[0191] In another embodiment, when the current block has multiple reference lines, the reference lines of the template and the reference lines of the current block are both outer adjacent lines of the current block. The linear model (1110 / 1112) applied on the template and the linear model applied on the current block are both from the same reference line (930 / 932), such as Fig.11The template is outside the reference lines.
[0192] In another embodiment, when the current block has multiple reference lines, the reference line (930 / 932) of the template and the reference line (930 / 932) of the current block are both outer adjacent lines of the current block, such as Fig.12 As shown, the linear model applied on the template and the linear model applied on the current block are both from the same reference line (930 / 932). The template (910 / 912) is located immediately outside the current block.
[0193] In one embodiment, a flag is sent / parsed at a block level to indicate that "MRL (multi-reference line) mode" applies to the current block. For example, the flag may be sent at a CU level and / or a PU level and / or a CTU level. For another example, the flag may be sent at a CB level, a PB level, a CTB, a TU / TB, any predetermined region level, or any combination thereof.
[0194] In one embodiment, when the flag indicates that "MRL mode" is disabled, the original syntax for sending / parsing the intra prediction mode of the current block is followed. When the flag indicates that "MRL mode" is enabled, the hybrid mode is applied and M1, M2 and hybrid weights are derived.
[0195] In another embodiment, M1 is sent explicitly, and a flag is sent / parsed to indicate whether "MRL mode" is applied to the current block.
[0196] In one sub-embodiment, when the "MRL Mode" flag is enabled, blending mode is enabled and M2 and blending weights are derived. When the flag is disabled, blending is not used.
[0197] In a sub-embodiment, when the "MRL mode" flag is enabled, and M1 is one of the CCLM modes (CCLM_LT, CCLM_L, CCLM_T), blending mode is enabled, and M2 and blending weights are derived. When the "MRL mode" flag is enabled, and M1 is one of the MMLM modes (MMLM_LT, MMLM_L, MMLM_T), only M1 is used to generate the predictors, and blending is not used. The reference lines used to derive the linear model are sent. When this flag is disabled, the reference lines used for derivation are the outer neighboring lines of the block (such as Fig.13 The first reference line shown) and are not sent.
[0198] In one embodiment, the mode selected for generating H1 and H2 is implicitly determined. Figure 5 ) are unavailable, the blending mode is disabled.
[0199] In one embodiment, the tool proposed in the present invention is not used with CCCM and GLM. If the "MRL mode" flag is on, the syntax associated with CCCM and GLM is bypassed. If the "MRL mode" flag is off, the syntax associated with CCCM and GLM is sent. For another example, if CCCM or GLM mode is used, the syntax of the "MRL mode" flag is bypassed. If CCCM or GLM mode is used, the syntax of the "MRL mode" flag is sent.
[0200] In another embodiment, the tool proposed in the present invention is used with CCCM and GLM. Regardless of whether the "MRL mode" flag is on or off, the syntax associated with CCCM and GLM will be sent. If the "MRL mode" flag is on, the settings associated with CCCM and GLM are applied. For example, the CCCM or GLM model is used instead of the CCLM or MMLM model. For another example, the number of reference lines used is 6 instead of 1. If the "MRL mode" flag is off, the decoder operates in the manner indicated by the CCCM or GLM syntax.
[0201] Any of the above proposed implicit linear model derivations with multiple reference lines can be implemented in an encoder and / or decoder. For example, any of the proposed methods can be implemented in an intra-frame coding module (e.g. Figure 1A Intra-frame prediction 110 in the decoder) or the intra-frame coding module of the decoder (e.g. Figure 1B Alternatively, any of the proposed methods may be implemented as a circuit coupled to an intra-frame coding module of an encoder and / or a decoder.
[0202] Fig.14A flowchart of an exemplary video coding system for implicitly deriving a linear model predictor based on template-based intra mode derivation (TIMD) according to an embodiment of the present invention is shown. The steps shown in the flowchart can be implemented as program codes that can be executed on one or more processors (e.g., one or more CPUs) on the encoder side. The steps shown in the flowchart can also be implemented based on hardware, such as one or more electronic devices or processors arranged to execute the steps in the flowchart. According to the method, in step 1410, input data associated with a current block including a first color component and a second color component is received, wherein the input data includes pixel data to be encoded on the encoder side or encoded data associated with the current block to be decoded on the decoder side. In step 1420, one or more templates of the current block are determined, wherein the one or more templates of the current block are located in a previously encoded and decoded area of the current block. In step 1430, one or more first reference lines are selected from the available reference lines of the current block. In step 1440, a cross-component model parameter set of each cross-component model in one or more cross-component models is derived based on the first color sample and the second color sample in the one or more first reference lines. In step 1450, the prediction samples of the one or more templates associated with each of the one or more cross-component models for the second color component of the one or more templates are derived by applying the cross-component model parameter set associated with each of the one or more cross-component models to the first color component of the one or more templates. In step 1460, the template cost associated with each of the one or more cross-component models is determined based on the prediction samples of the one or more templates associated with each of the one or more cross-component models and the reconstructed samples of the one or more templates. In step 1470, one or more target modes from the one or more cross-component models are determined for the second color component of the current block based on the template cost associated with the one or more cross-component models. In step 1480, one or more second reference lines are selected from the available reference lines of the current block, wherein the one or more second reference lines are allowed to be the same as the one or more first reference lines. In step 1490, if the one or more second reference lines are different from the one or more first reference lines, the cross-component model parameter set of the one or more target modes is derived based on the first color samples and the second color samples in the one or more second reference lines. In step 1495, the second color component of the current block is encoded or decoded using the encoding and decoding information including the one or more target modes and the cross-component model parameter set derived from the one or more second reference lines.
[0203] The flow chart shown is intended to illustrate an example of video encoding according to the present invention. Those skilled in the art may modify each step, rearrange the steps, split the steps or merge the steps to practice the present invention without departing from the spirit of the present invention. In the present disclosure, specific syntax and semantics have been used to illustrate examples of implementing embodiments of the present invention. Technicians can practice the present invention by replacing syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0204] The above description is intended to enable one of ordinary skill in the art to practice the present invention in the context of a specific application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not limited to the specific embodiments shown and described, but should be given the widest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details are described in order to provide a thorough understanding of the present invention. Nevertheless, it should be understood by those skilled in the art that the present invention can be practiced.
[0205] The above-mentioned embodiments of the present invention can be implemented in various hardware, software codes or a combination of the two. For example, the embodiments of the present invention can be one or more circuits integrated into a video compression chip or a program code integrated into a video compression software to perform the processing described herein. The embodiments of the present invention can also be a program code to be executed on a digital signal processor (Digital Signal Processor, referred to as DSP) to perform the processing described herein. The present invention can also be related to multiple functions performed by a computer processor, a digital signal processor, a microprocessor or a field programmable gate array (fieldprogrammable gate array, referred to as FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software codes or firmware codes that define the specific methods embodied in the present invention. The software code or firmware code can be developed in different programming languages and different formats or styles. The software code can also be compiled for different target platforms. However, the different code formats, styles and languages of the software code and other ways of configuring the code to perform tasks according to the present invention will not deviate from the spirit and scope of the present invention.
[0206] The present invention may be implemented in other specific forms without departing from its spirit or essential characteristics. The examples described should be considered in all respects as illustrative only and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. All changes in the equivalent meaning and scope of the claims should be included within their scope.
Claims
1. A method for encoding and decoding a color image, the method comprising: Receiving input data associated with a current block including a first color component and a second color component, wherein the input data includes pixel data to be encoded at an encoder side or encoded data associated with the current block to be decoded at a decoder side; Determining one or more templates of the current block, wherein the one or more templates of the current block are located in a previously encoded area of the current block; selecting one or more first reference lines from a plurality of available reference lines of the current block; deriving a set of cross-component model parameters for each of the one or more cross-component models based on the first color samples and the second color samples in the one or more first reference lines; deriving a plurality of prediction samples of the one or more templates associated with each of the one or more cross-component models for the second color component of the one or more templates by applying the set of cross-component model parameters associated with each of the one or more cross-component models to the first color component of the one or more templates; determining a template cost associated with each of the one or more cross-component models based on the plurality of predicted samples of the one or more templates associated with each of the one or more cross-component models and a plurality of reconstructed samples of the one or more templates; determining one or more target modes for the second color component of the current block from the one or more cross-component models based on the template costs associated with the one or more cross-component models; selecting one or more second reference lines from the plurality of available reference lines of the current block, wherein the one or more second reference lines are allowed to be the same as the one or more first reference lines; deriving a plurality of cross-component model parameter sets for the one or more target patterns based on the first color samples and the second color samples in the one or more second reference lines if the one or more second reference lines are different from the one or more first reference lines; and The second color component of the current block is encoded or decoded using encoding and decoding information including the one or more target modes and the cross-component model parameter set derived from the one or more second reference lines.
2. The method for encoding and decoding a color image according to claim 1, characterized in that: The template cost associated with each of the one or more cross-component models corresponds to distortion between the multiple predicted samples of the one or more templates associated with each of the one or more cross-component models and the multiple reconstructed samples of the one or more templates.
3. The method for encoding and decoding a color image according to claim 2, wherein: The one or more target patterns selected from the one or more cross-component models have a minimum template cost among the plurality of template costs associated with the one or more cross-component models.
4. The method for encoding and decoding a color image according to claim 1, characterized in that: The current block is encoded or decoded using a final predictor by mixing a first predictor associated with a first prediction mode and a second predictor associated with a second prediction mode, and wherein at least one of the first prediction mode and the second prediction mode is implicitly determined based on the multiple template costs associated with the one or more cross-component models.
5. The method for encoding and decoding a color image according to claim 4, characterized in that: The first prediction mode is explicitly determined by sending or parsing an index, and the one or more cross-component models are associated with a cross-component model set, and the second prediction mode is implicitly determined based on the multiple template costs associated with the one or more cross-component models in the cross-component model set.
6. The method for encoding and decoding a color image according to claim 5, characterized in that: The final predictor is generated by a weighted sum of the first predictor and the second predictor.
7. The method for encoding and decoding a color image according to claim 6, characterized in that: A weighting factor of the weighted sum of the first predictor and the second predictor is implicitly determined based on a first template cost associated with the first prediction mode and a second template cost associated with the second prediction mode.
8. The method for encoding and decoding a color image according to claim 7, characterized in that: The first template cost is determined by deriving a first cross-component parameter set based on the multiple samples on the one or more first reference lines and the transmitted first prediction mode, and then determined based on multiple prediction samples of the one or more templates and multiple reconstructed samples of the one or more templates, and wherein the multiple prediction samples of the one or more templates are generated according to the transmitted first prediction mode using the derived first cross-component parameter set.
9. The method for encoding and decoding a color image according to claim 6, wherein: It is characterized in that A flag is sent or parsed to indicate whether the current block enables the multi-reference line mode, and when the flag indicates that the current block enables the multi-reference line mode, the second prediction mode is determined and a weighting factor of the weighted sum of the first predictor and the second predictor is derived.
10. The method for encoding and decoding a color image according to claim 9, characterized in that: The first predictor is generated based on a cross-component parameter set, and the cross-component parameter set is derived based on the first prediction mode and the one or more second reference lines.
11. The method for encoding and decoding a color image according to claim 9, characterized in that: The one or more first reference lines of the plurality of available reference lines correspond to one or more outer neighboring reference lines of the one or more templates of the current block, and the second reference line is the same as the first reference line.
12. The method for encoding and decoding a color image according to claim 9, characterized in that: The one or more cross-component models are associated with a second set of cross-component models, and a flag is issued only when the first prediction mode is in the second set of cross-component models.
13. The method for encoding and decoding a color image according to claim 12, characterized in that: If the first prediction mode is not in the second cross-component model set, the flag is not sent and the multi-reference line mode is implicitly determined to be disabled.
14. The method for encoding and decoding a color image according to claim 4, characterized in that: The one or more cross-component models are associated with a first set of cross-component models and a second set of cross-component models, and wherein a first target pattern is derived from the first set of cross-component models and a second target pattern is derived from the second set of cross-component models.
15. The method for encoding and decoding a color image according to claim 14, characterized in that: The final predictor is generated by a weighted sum of the first predictor and the second predictor.
16. The method for encoding and decoding a color image according to claim 15, characterized in that: A weighting factor of the weighted sum of the first predictor and the second predictor is implicitly determined based on a first template cost associated with the first prediction mode and a second template cost associated with the second prediction mode.
17. The method for encoding and decoding a color image according to claim 16, characterized in that: A flag is sent or parsed to indicate whether the current block enables the multi-reference line mode, and when the flag indicates that the current block enables the multi-reference line mode, the first prediction mode and the second prediction mode are determined, and the weighting factor of the weighted sum of the first predictor and the second predictor is derived.
18. The method for encoding and decoding a color image according to claim 14, characterized in that: The first cross-component model set and the second cross-component model set only contain multiple linear model patterns.
19. The method for encoding and decoding a color image according to claim 14, characterized in that: The first prediction mode is implicitly determined from the first cross-component model set, and the second prediction mode is implicitly determined from the second cross-component model set.
20. A video encoding and decoding apparatus, the apparatus comprising one or more electronic devices or processors, configured to: receiving input data associated with a current block including a first color component and a second color component, wherein The input data includes pixel data to be encoded at the encoder side or encoded data associated with the current block to be decoded at the decoder side; Determining one or more templates of the current block, wherein the one or more templates of the current block are located in a previously encoded area of the current block; selecting one or more first reference lines from a plurality of available reference lines of the current block; deriving a set of cross-component model parameters for each of the one or more cross-component models based on the first color samples and the second color samples in the one or more first reference lines; deriving a plurality of prediction samples of the one or more templates associated with each of the one or more cross-component models for the second color component of the one or more templates by applying the set of cross-component model parameters associated with each of the one or more cross-component models to the first color component of the one or more templates; determining a template cost associated with each of the one or more cross-component models based on the plurality of predicted samples of the one or more templates associated with each of the one or more cross-component models and a plurality of reconstructed samples of the one or more templates; determining one or more target modes for the second color component of the current block from the one or more cross-component models based on the template costs associated with the one or more cross-component models; selecting one or more second reference lines from the plurality of available reference lines of the current block, wherein the one or more second reference lines are allowed to be the same as the one or more first reference lines; deriving a plurality of cross-component model parameter sets for the one or more target patterns based on the first color samples and the second color samples in the one or more second reference lines if the one or more second reference lines are different from the one or more first reference lines; and The second color component of the current block is encoded or decoded using encoding and decoding information including the one or more target modes and the cross-component model parameter set derived from the one or more second reference lines.