Simplification of coding mode based on adjacent sample related parameter model
By simplifying the decoding mode design based on the parameter model of neighboring samples in video encoding and decoding technology, the inefficiency problems under the influence of complex design and noise in the prior art are solved, and a more efficient decoding process is achieved.
Patent Information
- Application Number
- CN202510346157.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-13
- Filing Date
- 2019-11-01
- Publication Date
- 2025-05-27
AI Technical Summary
In the existing video encoding and decoding technologies, the decoding mode design based on the parameter model of neighboring samples is complex, and the decoding efficiency is low under the influence of noise.
The decoding mode design based on the parameter model of the adjacent sample is simplified, and the parameters of the linear parameter model are derived by replacing the minimum mean square method, a unified simplification process is adopted, and alternative decoding modes or correction terms are used if necessary.
It improves decoding efficiency, reduces computational complexity, and enhances performance stability in noisy environments.
Smart Images

Figure CN120050432A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 201980071930.7, titled "Simplification of Decoding Modes Based on Adjacent Sample Correlation Parameter Models", filed on November 1, 2019, the content of the parent application is incorporated herein by reference. Technical Field
[0002] At least one embodiment of the present invention mainly relates to methods or apparatuses for video encoding or decoding, compression or decompression. Background Art
[0003] To achieve high compression efficiency, image and video coding schemes typically employ prediction, which includes motion vector prediction and transforms to exploit spatial and temporal redundancies in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image and the predicted image, which is usually represented as a prediction error or prediction residual, is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction. Summary of the Invention
[0004] At least one embodiment of the present invention mainly relates to a method or apparatus for video encoding or decoding, and more particularly, to a method or apparatus for simplifying decoding modes based on adjacent sample correlation parameter models.
[0005] According to a first aspect, a method is provided. The method includes the steps of: determining a prediction of samples in a current block, the determination being based on at least one of the adjacent samples in the current block and on a parameter model calculated from the adjacent samples in the current block and reference samples in a reference frame; and encoding the samples in the current block based on the prediction.
[0006] According to a second aspect, a method is provided. The method includes the following steps: determining a prediction of samples in a current block, the determination being based on at least one of the adjacent samples in the current block and on a parameter model calculated from the adjacent samples in the current block and reference samples in a reference frame; and decoding the samples in the current block based on the prediction.
[0007] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to encode a block of video or decode a bitstream by executing any of the above methods.
[0008] According to another main aspect of at least one embodiment, there is provided an apparatus comprising: means according to any of the decoding embodiments; and (i) an antenna configured to receive a signal comprising a video block, (ii) a band limiter configured to limit the received signal to a band comprising the video block, or (iii) a display configured to display an output representing the video block.
[0009] According to another main aspect of at least one embodiment, there is provided a non-transitory computer-readable medium comprising data content generated according to any of the described encoding embodiments or variations.
[0010] According to another main aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the described encoding embodiments or variations.
[0011] According to another main aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.
[0012] According to another main aspect of at least one embodiment, there is provided a computer program product comprising instructions which, when executed by a computer, cause the computer to perform any of the described decoding embodiments or variations.
[0013] These and other aspects, features and advantages of the main aspects will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Example locations for deriving samples of a and b are shown.
[0015] Figure 2 An example of the LM_A mode is shown.
[0016] Figure 3 An example of the LM_L mode is shown.
[0017] Figure 4 A diagram of the LIC mode in JEM is shown.
[0018] Figure 5 A standard general video compression scheme is shown
[0019] Figure 6 A standard, general video decompression scheme is shown.
[0020] Figure 7 Samples selected from the top line at the rightmost position and the left column at the bottommost position are shown.
[0021] Figure 8 An exemplary block diagram showing the use of two samples at specific locations in the neighborhood.
[0022] Figure 9 Samples selected from the top row of the rightmost position and the leftmost column of the bottommost are shown.
[0023] Figure 10 Samples selected from the top row of the rightmost and leftmost positions are shown.
[0024] Figure 11 Samples selected from the left column at the topmost and bottommost positions are shown.
[0025] Figure 12 Samples selected at the lower left, upper left, and upper right positions are shown.
[0026] Figure 13 Samples at the rightmost and leftmost positions selected from the top row are shown.
[0027] Figure 14 Samples selected from the left column at the topmost and bottommost positions are shown.
[0028] Figure 15 Samples selected (a) at three or more positions, (b) at the top two positions and two left positions, and (c) at the top three positions and three left positions are shown.
[0029] Figure 16 An exemplary block diagram for testing the reliability of linear model derivation is shown.
[0030] Figure 17 Weights used in intra - inter prediction in the hybrid frame are shown.
[0031] Figure 18 An embodiment of a method according to the described aspects is shown.
[0032] Figure 19 An exemplary processor - based subsystem for implementing the mainly described aspects is shown.
[0033] Figure 20 A block diagram of the CCLM / MDLM process is shown.
[0034] Figure 21 A modified block diagram of the CCLM / MDLM process according to the first embodiment is shown.
[0035] Figure 22 A modified block diagram of the CCLM / MDLM process according to the second embodiment is shown.
[0036] Figure 23 A modified block diagram of the CCLM / MDLM process according to a variation of the second embodiment is shown.
[0037] Figure 24 A modified block diagram of the CCLM / MDLM process according to the third embodiment is shown.
[0038] Figure 25 Another embodiment of a method according to the described aspects is shown.
[0039] Figure 26 An example apparatus according to the described aspects is shown. Detailed Description
[0040] The embodiments described herein pertain to the field of video compression and mainly relate to video compression as well as video encoding and decoding.
[0041] To achieve high compression efficiency, image and video coding schemes typically employ prediction (including motion vector prediction) and transformation to exploit the spatial and temporal redundancies in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image and the predicted image, which is typically represented as a prediction error or prediction residue, is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through inverse processes corresponding to the entropy coding, quantization, transformation, and prediction.
[0042] In the HEVC (High Efficiency Video Coding, ISO / IEC 23008-2, ITU-T H.265) video compression standard, motion compensated temporal prediction is employed to exploit the redundancy present between consecutive pictures in a video.
[0043] To this end, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented by a coding tree in the compressed domain. This is a quadtree partitioning of the CTU, where each leaf is referred to as a coding unit (CU).
[0044] Then, each CU is given some intra-frame or inter-frame prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. Intra-frame or inter-frame coding modes are assigned at the CU level.
[0045] In the JVET (Joint Video Exploration Team) proposal for a new video compression standard, called the Joint Exploration Model (JEM), a quadtree - binary tree (QTBT) block partitioning structure has been proposed due to its high compression performance. By dividing a block in a binary tree (BT) horizontally or vertically in the middle, it can be divided into two equal - sized sub - blocks. Thus, a BT block can have a rectangular shape with unequal width and height, which is different from the blocks in a QT, where the blocks always have a square shape with equal height and width. In HEVC, the angular intra - prediction directions are defined from 45° to - 135° at 180° angles, and they have been maintained in JEM, which has made the definition of the angular directions independent of the target block shape.
[0046] To encode these blocks, intra - prediction is used to provide an estimated version of the block using previously reconstructed neighboring samples. Then, the difference between the source block and the prediction is encoded. In the above - mentioned classical codecs, a single - row reference sample is used on the left and top of the current block.
[0047] In HEVC (High - Efficiency Video Coding, H.265), the encoding of frames of a video sequence is based on a quadtree (QT) block partitioning structure. A frame is divided into square coding tree units (CTUs), and all CTUs undergo quadtree - based partitioning based on the rate - distortion (RD) criterion and are divided into multiple coding units (CUs). Each CU is either intra - predicted (i.e., it is predicted spatially from causally adjacent CUs) or inter - predicted (i.e., it is predicted temporally from decoded reference frames). In an I - slice, all CUs are intra - predicted, while in P - and B - slices, CUs can be either intra - or inter - predicted. For intra - prediction, HEVC defines 35 prediction modes, which include a planar mode (indexed as mode 0), a DC mode (indexed as mode 1), and 33 angular modes (indexed as modes 2 to 34). The angular modes are associated with prediction directions ranging from 45 degrees to - 135 degrees in the clockwise direction. Since HEVC supports a quadtree (QT) block partitioning structure, all prediction units (PUs) have a square shape. Thus, from the perspective of the PU (prediction unit) shape, the definition of the prediction angles from 45 degrees to - 135 degrees is reasonable. For a target prediction unit of size N×N pixels, the top reference array and the left reference array each have a size of 2N + 1 samples, which is required to cover the above - mentioned angular range of all target pixels. Considering that the height and width of the PU have equal lengths, it is also reasonable that the lengths of the two reference arrays are equal.
[0048] The present invention belongs to the field of video compression. More specifically, the present invention focuses on multiple modes that use a parametric model to perform prediction of a given block, where the parameters of the model are derived from the neighboring samples of the block. Two examples of such modes are the "Cross-Component Linear Model" (CCLM) mode and the "Local Illumination Compensation" (LIC) mode. The object of the present invention is to simplify and improve the design of these modes.
[0049] Description of CCLM and Variants
[0050] The following section describes different variants of CCLM.
[0051] Basic CCLM Mode Description
[0052] In its initial version (cf_VET_K1002), the CCLM mode consists in predicting chrominance samples based on the reconstructed luminance samples of the same block or CU by using the following linear model:
[0053] pred C (i,j) = a.rec L ’(i,j) + b (Equation 1)
[0054] where pred C (i,j) represents the predicted chrominance sample in the CU, and rec L ’(i,j) represents the downsampled reconstructed luminance sample of the same CU. The parameters a and b are derived by minimizing the regression error between the neighboring reconstructed luminance and chrominance samples around the current block as follows:
[0055] a = (SLC – SL.SC) / (SLL – SL.SL) (Equation 2)
[0056] b = SC – a.SL (Equation 3)
[0057] where L(i,j) represents the downsampled top and left neighboring reconstructed luminance samples, C(i,j) represents the top and left neighboring reconstructed chrominance samples, N is equal to twice the minimum of the width and height of the current chrominance decoding block, and SL, SC, SLL, SLC are defined as follows (the symbol represents the sum over the top and left neighboring samples):
[0058] ● SL = ∑L(n)
[0059] ● SC = ∑C(n)
[0060] ● SLC = N·∑(L(n)·C(n))
[0061] ● SLL = N·∑(L(n)·L(n))
[0062] For a decoded block with a square shape, the above two equations are directly applied. For a non-square decoded block, the adjacent samples on the longer boundary are first subsampled to obtain the same number of samples as the shorter boundary. Figure 1 Shows the positions of the samples of the current block and the left and upper samples involved in the CCLM mode.
[0063] When decoding a CU using the CCLM mode, the least mean square (LMS) method is performed during the decoding process. As a result, no syntax is used to transmit the a and b values to the decoder.
[0064] MDLM mode
[0065] The MDLM mode is an improvement to the basic CCLM design proposed in JVET-L0338, where in addition to the (top + left) reference sample template, it is also possible to select only the left or only the top template to derive the linear model coefficients α and β. This means that 2 new CCLM modes called LM_A and LM_L are added.
[0066] In the LM_A mode (see Figure 2 ), only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to (W + H), where W is the width of the block and H is its height. In the LM_L mode (see Figure 3 ), only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H + W).
[0067] For non-square blocks, the upper template is extended to W + W, and the left template is extended to H + H.
[0068] If the above / left template is not available, the LM_A / LM_L mode will not be checked or signaled. If the number of available samples is not large enough, the rightmost (for the top template) sample or the bottommost (for the left template) sample is copied to the nearest log2 number to fill the template.
[0069] CCLM / MDLM with line buffer constraints
[0070] During the current CCLM mode coefficient derivation process, 2 luminance line buffers are used in principle for subsampling to obtain the top template of the CCLM mode (CCLM or MDLM), while only 1 luminance line buffer is used in ordinary luminance component intra prediction. To reduce the line buffers, only the LM_L mode is used for the CU and the CTU top boundary. In this case, no additional line buffer is required.
[0071] Description of local illumination compensation
[0072] In this tool, the decoder calculates some prediction parameters ( Figure 4 ) based on some reconstructed picture samples located on the left and / or top of the current block to be predicted and reference picture samples located on the left and / or top of the motion-compensated block. In the considered prior art codec (JEM), the use of LIC for a given block depends on a flag associated with the block, which is called the LIC flag.
[0073] The LIC parameters (a, b) are based on minimum mean square minimization, which minimizes the following distortion:
[0074] dist = ∑ r∈Vcur,s∈Vref (Rcur(r) - a.Rref(s) - b) 2 (Equation 4)
[0075] where Rcur(r) is the adjacent reconstructed sample and Rref(s) is the reference sample. The derivation of a and b is similar to the way a and b were derived in the previous section (Equations 2 and 3).
[0076] Once the encoder or decoder obtains the LIC parameters for the current CU, the prediction pred(i,j) of the current CU includes the following (unidirectional prediction case):
[0077] pred(i,j) = a.ref(i,j) + b (Equation 5)
[0078] where ref(i,j) is the reference block for the temporal prediction of the current block.
[0079] The main aspect described herein aims to improve and simplify the design of modes similar to CCLM or LIC based on the adjacent sample correlation parameter model. The proposed modifications relate to the way the parameters of the parameter model are derived and how to design the prediction tool based on the parameter model included in the codec in a unified and simplified manner compared to the prior art.
[0080] One method proposes to simplify the CCLM process for deriving linear parameters. A replacement for the LMS method is proposed to derive the parameters a and b as the parameters of a straight line passing through two points corresponding to the minimum and maximum luminance values among all luminance adjacent reconstructed samples.
[0081] The values of a and b are derived as follows:
[0082] a = (C B – C A ) / (L B – L A ) (Equation 6)
[0083] b = C A – a.LA (Equation 7)
[0084] where (L A , C A ) is a pair of luminance and chrominance values in adjacent reconstructed samples, and L A has the minimum value among all the luminance values, while (L B , C B ) is a pair of luminance and chrominance values in adjacent reconstructed samples, and L B has the maximum value among all the luminance values.
[0085] This method still requires performing multiple checks to identify the minimum and maximum luminance values. Moreover, it may face problems when L A and L A are close to each other.
[0086] The LMS method used in the initial CCLM and LIC schemes has other problems. An important problem is that when the input samples are corrupted by noise, LMS will cause bias, which is obviously present because the samples are obtained from decoding or prediction. This may reduce the decoding efficiency of the tool.
[0087] The main aspects described in this article propose various changes:
[0088] – Simplify the sample selection for deriving the parameters of the parametric model: Retrieve samples from predefined locations
[0089] – Use an alternative decoding mode when the derivation of the parameters of the parametric model is unreliable
[0090] – Insert correction terms that may be signaled in the bitstream during the derivation of the parameters of the parametric model
[0091] – Unify the derivation process of the parameters of the parametric model between LIC and CCLM
[0092] – Extend CCLM for inter-frame blocks and hybrid intra-inter blocks.
[0093] Consider the general problem of predicting the current block of the sample Pcur(p) at position p in a block of size W columns × H rows (based on its co-located reference sample Rref(p)). Also consider using a bit depth of B bits to represent the sample. In CCLM, the reference sample is a reconstructed luminance sample. In LIC, the reference sample is a sample from a motion-compensated block in a reference picture. Then, also consider that in the neighborhood of the block to be predicted, the reconstructed current sample (Rcur) and the reconstructed reference sample (Rref) are available. This is shown in Figure 7 . The adjacent samples are not necessarily in the nearest row / column of the block.
[0094] The goal is to derive Pcur(p) for p in the block from Rref(p) in the block and from a parametric model, where the parametric model is calculated from samples Rcur and Rref located in the neighborhood of the block (usually the upper row and the left column outside the block).
[0095] Example 1 - Using 2 samples directly selected at specific positions in the reference sample array
[0096] In one embodiment, to simplify the derivation of the parameters of the parametric model, the parameters are derived from at least 2 samples of adjacent samples, and the at least two samples are selected such that the samples are spatially distant from each other.
[0097] In one implementation, the following procedure ( Figure 8 the procedure described in
[0098] – If both the top and left samples are available (step 401), then select the available samples (Rref A , Rcur A ) in the outer top row at the rightmost position, and select the available samples (Rref B , Rcur B ) in the outer left column at the bottommost position (step 403) (see the illustration in Figure 9 ).
[0099] – Otherwise, if only the top sample is available (step 402), then select the available samples (Rref A , Rcur A ) in the outer top row at the rightmost position, and select the available samples (Rref B , Rcur B ) in the outer top row at the leftmost position (step 405) (see the illustration in Figure 10 ).
[0100] – Otherwise, if only the left sample is available (step 404), then select the available samples (Rref A , Rcur A ) in the outer left column at the bottommost position, and select the available samples (Rref B , Rcur B ) in the outer left column at the topmost position (step 407) (see the illustration in Figure 11 ).
[0101] – Otherwise, do not apply the CCLM mode (step 406).
[0102] The parameters a and b are derived as follows:
[0103] a = (Rcur B – Rcur A ) / (Rref B – Rref A ) (Equation 8)
[0104] b = Rcur A – a.Rref A (Equation 9)
[0105] And the prediction for any position p in the block is calculated as:
[0106] Pcur(p) = a.Rref(p) + b (Equation 10)
[0107] Compared with JVET-L0191, this solution avoids multiple checks required to identify the minimum and maximum values of reference samples in the neighborhood.
[0108] The same concept can be directly applied to the MDLM mode. For example, when selecting the top sample for MDLM, the sample as Figure 10 shown is used. When selecting the left sample for MDLM, the sample as Figure 11 shown is used.
[0109] Example 2 - Using 3+ samples directly selected at specific positions in the reference sample array
[0110] In this example, to simplify the derivation of the parameters of the parameter model, the parameters are derived from at least 3 samples of adjacent samples, and the at least 3 samples are selected such that the samples are spatially far apart as Figure 12 shown.
[0111] This concept also applies to the MDLM case, as Figure 13 and Figure 14 shown.
[0112] Different from the method mentioned above that compares all samples to find the minimum and maximum luminance value samples, only 3 or more samples are used to calculate the minimum and maximum luminance values. The worst case is limited to two comparisons in the case of three samples.
[0113] According to the previous method, the linear model parameters are calculated according to Equations 6 and 7.
[0114] As Figure 15 (a) shown, more than three samples can be selected at specific positions.
[0115] In another variant, up to 4 samples are used as follows. For the reference samples of the top row of size Wtop, the samples at positions x = 0, x = Wtop - 1 are used. For the reference samples of the left column of size Hleft, the samples at positions y = 0, y = Hleft - 1 are used. This is shown in Figure 15 (b).
[0116] In another variant, up to 6 samples are used as follows. For the reference samples of the top row of size Wtop, the samples at positions x = 0, x = Wtop - 1 and one sample in the middle (e.g., position x = Wtop / 2) are used. For the reference samples of the left column of size Hleft, the samples at positions y = 0, y = Hleft - 1 and one sample in the middle (e.g., position y = Hleft / 2) are used. This is shown in Figure 15 (c).
[0117] In an embodiment, in the case of calculating the linear parameter based on the minimum and maximum values of the reference samples (L A and L B ) as in the submitted JVET - L0191, only the selected reference luminance samples are used to calculate these minimum and maximum values. Since in the above embodiments, the maximum number of reference samples is reduced to 2, 3, 4, 5, or 6, this significantly limits the number of checks required to identify the minimum and maximum luminance sample values. In the submitted JVET - L0191, in the worst case, for a given block with Wtop top reference samples and Hleft left reference samples, the number of such checks is equal to (Wtop + Hleft) x 2. Using the present invention, this number is reduced to 2 x 2, 3 x 2, 4 x 2, 5 x 2, or 6 x 2.
[0118] Embodiment 3 - Using an alternative mode when the linear model is not well - defined
[0119] The calculation of the linear parameter involves division. In the case of LMS, it is:
[0120] a = (SLC – SL.SC) / (SLL – SL.SL) (Equation 11)
[0121] In the previous method, it is
[0122] a = (Rcur B – Rcur A ) / (Rref B – Rref A ) (Equation 12)
[0123] In both cases, a is obtained as a = Num / Den, where Num is the numerator and Den is the denominator of the division. When Den has a small magnitude, this can be problematic, which may lead to an unstable estimation of the linear parameter.
[0124] It can also be considered that for blocks of too small a size, the number of samples used to derive the linear parameter is insufficient to obtain a reliable estimation.
[0125] In one embodiment, the prediction based on the linear model is used only when the linear parameter derivation is considered to be well-defined. Otherwise, an alternative mode ( Figure 16 example block diagram in
[0126] Different ways of checking the reliability of the linear parameter derivation can be used. For example, if one of the following conditions is true, the linear parameter derivation is applied:
[0127] - If Den > T1,
[0128] ○ where T1 is a predefined threshold, which may depend on the block size, and B is the sample bit depth. For example,
[0129] ■ T1 = T2 * W * H * 2^B
[0130] ■ where T2 is a predefined threshold
[0131] - If (WxH > Nmin), the linear parameter derivation is applied
[0132] ○ where W and H are the width and height of the block
[0133] Otherwise, a simplified model is used.
[0134] The thresholds T1 or T2 can also be signaled at various levels (e.g., per SPS, PPS, slice, tile group, tile, CTU, or CU). A specific threshold can be signaled for each block size.
[0135] The alternative mode can be based on using a simplified model:
[0136] ● Additive model: a is forced to be 1, and only b is derived.
[0137] Pcur(p) = Rref(p) + b
[0138] ● Scaling model: b is forced to be 0, and only a is derived.
[0139] Pcur(p) = a.Rref(p)
[0140] Embodiment 4 - Using a correction parameter in the derivation of linear parameters
[0141] In one embodiment, a correction parameter CP is introduced into the formula for deriving the linear parameter.
[0142] Compared with the prior art, due to the flexibility introduced by multiple possible correction parameters CP, the advantage of this correction parameter is that it improves the decoding efficiency.
[0143] When deriving the scaling parameter a of the linear model in additive or multiplicative mode, CP can be used to correct the numerator or denominator. For example, the following correction modes can be applied:
[0144] - Num’ = CP * Num and a = Num’ / Den
[0145] - Num’ = (Num + CP * sign(Num)) and a = Num’ / Den
[0146] - Den’ = CP * Den and a = Num / Den’
[0147] - Den’ = (Den + CP * sign(Den)) and a = Num / Den’
[0148] The correction parameter CP can be signaled at various levels (e.g., according to SPS, PPS, slice, tile group, tile, CTU, or CU).
[0149] The parameter can be retrieved from a finite set of K possible predefined values {CP 0 , CP 1 , …, CP K-1}. Only the indices corresponding to the indices of the values in this set need to be decoded.
[0150] CP can depend on Num or Den. In particular, when CP is additive, CP can increase as the value considered increases:
[0151] - Num’ = (Num + (abs(Num) >> K2) * sign(Num)) and a = Num’ / Den
[0152] - Den’ = (Den + (abs(Den) >> K2) * sign(Den)) and a = Num / Den’
[0153] Alternatively, negative correction can be used:
[0154] - Num’ = (Num - (abs(Num) >> K2) * sign(Num)) and a = Num’ / Den
[0155] -Den’ = (Den - (abs(Den) >> K2) * sign(Den)) and a = Num / Den’
[0156] Alternatively, negative correction can be used
[0157] -Num’ = (Num - (abs(Num) >> K2) * sign(Num)) and a = Num’ / Den
[0158] -Den’ = (Den - (abs(Den) >> K2) * sign(Den)) and a = Num / Den’
[0159] where K2 is a given predetermined value. For example, K2 = 6, which corresponds to CP = k / 64. abs(x) is a function that returns the modulus of x.
[0160] Example 4a - Using a modified lookup table to generate the division
[0161] To simplify implementation, the division involved in the derivation of linear parameters (which may increase implementation complexity) can be done through a lookup table.
[0162] In fact, the division
[0163] a = Num / Den
[0164] can be implemented without any division as:
[0165] a = (Num * Int((1 << K0) / Den) + offset0) >> K0
[0166] where K0 is a given value corresponding to the division precision, offset0 is a given offset value, usually equal to (1 << (K0 - 1)), and int() is the integer or base operator (rounded to the nearest lower integer value).
[0167] More generally, it can be implemented as follows:
[0168] a = (Num * (1 << Int(Den / (1 << K1))) * Int((1 << K0) / (Den % K1)) + offset0) >> K0
[0169] where K1 is a given parameter that fixes the maximum size of the LUT (equal to (1 << K1)), and "%" is the modulo operator.
[0170] The value Int((1 << K0) / k) can be stored in the lookup table divLUT[k].
[0171] In one embodiment, a correction parameter CP is used to modify the look-up table divLUT[k] to introduce a bias in the estimation. For example, the following correction patterns can be applied:
[0172] divLUT[k] = Int(2^K0 / (k + CP)) (Equation 13)
[0173] divLUT[k] = Int(2^K0 / (k * CP)) (Equation 14)
[0174] divLUT[k] = Int((2^K0 + CP) / k) (Equation 15)
[0175] divLUT[k] = Int((2^K0 * CP) / k) (Equation 16)
[0176] CP can depend on k. In particular, when CP is added (in the case of Equation 13 or 15), the CP module can increase with k.
[0177] In one example,
[0178] CP = k >> K2,
[0179] or
[0180] CP = -k >> K2,
[0181] where K2 is a given predetermined value. For example, K2 = 6, which is equivalent to CP = k / 64 or CP = -k / 64.
[0182] The LUT can be stored in the decoder. Alternatively, it can be calculated on-the-fly, and the correction parameter CP or K2 can be signaled in the stream at various levels (e.g., for each SPS, PPS, slice, tile group, tile, CTU, or CU).
[0183] Embodiment 5 - Unified LIC and CCLM
[0184] In the current design of LIC, the LMS process is applied to derive the linear parameters. While in the current CCLM, the linear parameters are derived from two sets of samples corresponding to the minimum and maximum values of the reference luminance samples.
[0185] In one embodiment, the derivation of LIC parameters and CCLM parameters is unified and uses the same simplified process. For example, the same derivation process based on identifying two sets of samples is used in both tools.
[0186] In one embodiment, the derivation of both LIC and CCLM linear parameters consists of identifying the two sets of sample sets (Rref A, Rcur A ) and (Rref B , Rcur B ), where Rref A and Rref B correspond to the minimum and maximum values of adjacent reference samples.
[0187] In another embodiment, both LIC and CCLM linear parameter derivations consist of identifying two sets of samples (Rref A , Rcur A ) and (Rref B , Rcur B ) obtained at extreme positions among available adjacent sample positions.
[0188] In both cases, the linear parameters are derived as:
[0189] a = (Rcur B – Rcur A ) / (Rref B – Rref A ) (Equation 17)
[0190] b = Rcur A – a.Rref A (Equation 18)
[0191] And the prediction for any position p in the block is calculated as:
[0192] Pcur(p) = a.Rref(p) + b (Equation 19)
[0193] The variations discussed in Embodiments 2 and 3 can also be applied to both cases.
[0194] Embodiment 6 - Extending CCLM to Inter - Frame Blocks
[0195] In the current design, CCLM is only applied to intra - frame CUs or blocks.
[0196] In one embodiment, CCLM is enabled to predict the chrominance components of inter - frame CUs. Thus, a new mode, namely, hybrid inter - frame CCLM, is introduced here. The mode can be signaled on a per - CU basis using a CU - level flag.
[0197] - Decode the luminance component using an inter - frame mode.
[0198] - Perform the entire process of prediction and reconstruction of the luminance component samples.
[0199] ○ Perform the complete reconstruction process until the complete reconstruction of the luminance block samples.
[0200] - Predict the chrominance component samples of the block by using the reconstructed luma samples of the block and by using the CCLM mode (i.e., using linear parameters calculated from the neighboring reconstructed luma and chrominance samples of the block).
[0201] ○ This means that no temporal prediction is used to construct the chrominance component samples of the block.
[0202] In terms of the operation pipeline, this new mode has the same problem as the traditional CCLM mode. Since reconstructed samples from the neighborhood and reconstructed luma samples from the current block are required, the processing of the blocks decoded in the hybrid inter-CCLM mode is preferably delayed once all intra and inter luma blocks have been processed.
[0203] Example 7 - Extending CCLM to Intra-Inter Blocks
[0204] In the VTM (Versatile Video Model), a new mode, i.e., hybrid intra-inter, is introduced. This mode combines an intra prediction and a merge index temporal prediction. In a merge CU, a flag is signaled for the merge mode to select an intra mode from the intra candidate list when the flag is true. For the luma component, the intra candidate list is derived from four intra prediction modes including DC, planar, horizontal, and vertical modes, and the size of the intra candidate list can be 3 or 4, depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. A weighted average is used to combine an intra prediction mode selected by the intra mode index and a merge index prediction selected by the merge index. For the chrominance component, DM is always applied without additional signaling.
[0205] The weights used to combine the predictions are described as follows (also shown in Figure 17 . Equal weights are applied when the DC or planar mode is selected or when the width or height of the block is less than 4. For those blocks with width and height greater than or equal to 4, when the horizontal / vertical mode is selected, a block is first divided vertically / horizontally into four equal-area regions. Each weighted set (denoted as (w_intra i , w_inter i ), where i is from 1 to 4, and (w_intra 1 , w_inter 1 ) = (6, 2), (w_intra 2 , w_inter 2 ) = (5, 3), (w_intra 3 , w_inter 3)=(3, 5), and (w_intra 4 , w_inter 4 )=(2, 6)) are for the region closest to the reference sample, while (w_intra 4 , w_inter 4 ) are for the region farthest from the reference sample. Then, the combined prediction can be calculated by adding the two weighted predictions and shifting right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the prediction value can be retained for subsequent adjacent CU reference.
[0206] In the proposed embodiment, CCLM is enabled to predict the chrominance component of the hybrid intra - inter CU. Thus, a new mode, namely, hybrid inter - CCLM, is introduced. The mode can be signaled for each CU using a CU - level flag. This flag indicates whether to use the DM or CCLM mode.
[0207] Alternatively, instead of applying the DM mode to chrominance as done in other methods, the CCLM mode is applied instead of DM.
[0208] - Decode the luma component using the hybrid intra - inter mode.
[0209] - Perform the entire process of prediction and reconstruction of the luma component samples.
[0210] - Predict the chrominance component samples of the block by using the reconstructed luma samples of the block and by using the CCLM mode (i.e., using the linear parameters calculated from the adjacent reconstructed luma and chrominance samples of the block).
[0211] In the first version, there is no mixing of intra and inter prediction of the chrominance component, and the CCLM mode is used to fully predict the chrominance block.
[0212] In a variant, a weighted mixing of intra and inter prediction is still applied to the chrominance component as done when DM corresponds to the horizontal or vertical mode, which means the final prediction of chrominance is a mixture of inter prediction and CCLM. The mixing method described in the prior art can be applied.
[0213] Alternatively, as done in the prior art when DM corresponds to the DC and planar modes, the same weight can be used for the entire chrominance block.
[0214] As in the previous embodiment, in terms of the pipelining of operations, once all intra, inter, and hybrid intra - inter luma blocks have been processed, the processing of the blocks decoded in the hybrid inter - CCLM mode can be delayed.
[0215] Reduction of memory size for CCLM
[0216] In its actual implementation, the CCLM process in submission JVET-L0191 is implemented as follows (where B represents the bit depth of the luminance and chrominance signals).
[0217] Once the minimum and maximum luminance values L A , L B and their associated chrominance values C A , C B are identified, the linear parameters are derived as follows. The pseudocode for each step is used in the block diagram of Figure 20 to illustrate the process.
[0218] The variables a, b, and shift_pred are derived as follows:
[0219] - The parameters shift, add, diff, and k are derived as follows:
[0220] ■ If (B > 8), shift is set to equal (B - 9), otherwise shift is set to equal 0 (step 501)
[0221] ■ If (shift > 0), add is set to equal (1 << (shift - 1)), otherwise it is set to equal (step 502)
[0222] ■ diff = (L B - L A + add) >> shift (step 503)
[0223] ■ shift_pred = 16
[0224] - If diff is greater than 0 (step 504), then the following steps are applied:
[0225] ■ div = ((C B - C A ) × LUT_low[diff - 1] + 2 15 ) >> 16 (step 505)
[0226] ■ a = ((C B - C A ) × LUT_high[diff - 1] + div + add) >> shift (step 506)
[0227] - Otherwise (step 504), then the following steps are applied:
[0228] ■ a = 0 (step 507)
[0229] - b is derived as follows (step 508)
[0230] ■b = C A -((a × L A ) >> shift_pred)
[0231] LUT_high and LUT_low are two lookup tables with 512 elements, and each element is derived as follows.
[0232] LUT_high[x] = Floor(2 16 / diff)
[0233] LUT_low[x] = Floor(2 32 / diff) - Floor(2 16 / diff) × 2 16
[0234] Floor(x) is the largest integer less than or equal to x
[0235] For any p in the chrominance block, the predicted sample Pcur(p) is derived as follows (step 509):
[0236] ■Pcur(p) = ((pRef(p) × a) >> shift_pred) + b
[0237] Clipping is also applied to keep the signal within the allowed range defined by the signal bit depth.
[0238] The following problems are observed:
[0239] - Two lookup tables with 512 integers are required, namely, LUT_high and LUT_low
[0240] - For signals greater than 8 bits, a right shift (B - 9) is applied to derive the parameter a, which may result in a loss of precision
[0241] - When generating the predicted sample Pcur(p), a right shift of the parameter k is applied to the first term of the formula,
[0242] which may result in a loss of accuracy
[0243] The following embodiments are intended to solve these problems. They can be combined together.
[0244] Embodiment 8 - Removal of one of the lookup tables
[0245] In one embodiment, the process is simplified by removing the lookup table LUT_low. The parameter a is derived as follows.
[0246] a = ((C B - C A)×LUT_high[diff - 1]+add)>>shift
[0247] In one variant, LUT_high[x] is derived as follows:
[0248] LUT_high[x] = Floor((2 16 +(diff / 2)) / diff)
[0249] This enables a 2 - fold reduction in memory requirements.
[0250] Figure 21 The modified process is shown, where the changed boxes are indicated in bold. The new box is step 606, which replaces the previous step 506. The previous step 505 is removed.
[0251] Example 9 - Modifying Access to the Look - up Table
[0252] In one embodiment, access to the look - up table is modified as follows.
[0253] shift=(L B - L A ) / 2 K
[0254] Or equivalently
[0255] shift=(L B - L A )>>K
[0256] where K is an integer value less than B.
[0257] This results in:
[0258] - Reducing the size of the look - up table to 2 K elements. When K = 8, this limits the table to 256 elements instead of 512 elements in the reference implementation of JVET - L0191.
[0259] - When (L B - L A ) is less than 2 K even when 2 B is higher than the look - up table size, a higher precision of the calculation of a can be obtained. This is not the case in the reference implementation of JVET - L0191, where once 2 B is higher than the actual look - up table size (512), then (L B - L A ) is divided by (B - 9).
[0260] Figure 22A modified process is shown, where the changed boxes are indicated in bold. The new box is step 701, which replaces the previous step 501.
[0261] In one embodiment, an additional step 701a is introduced after step 701 and before step 502 to modify the shift value as follows.
[0262] - If shift > 0, then shift = 1 + Floor(Log2(shift))
[0263] where Log2(x) is the base-2 logarithm of x.
[0264] This change is shown in Figure 23 .
[0265] For example, for the case of K = 8 (a table of size 2 K = 256 elements) and an input signal bit depth of B = 10, the following results are obtained:
[0266] - If ((L B - L A ) is from 0 to 255, then shift is set equal to 0
[0267] - Otherwise, if ((L B - L A ) is from 256 to 511, then shift is set equal to 1
[0268] - Otherwise, if ((L B - L A ) is from 512 to 1023, then shift is set equal to 2
[0269] This process ensures that the value of (diff - 1) remains within the maximum table index value.
[0270] Example 10 - Adaptation of Linear Prediction
[0271] In one embodiment, to obtain higher precision in the calculation of the predicted signal, the parameter b is calculated as follows:
[0272] b = (C A << shift_pred) - (a × L A ) + (1 << (shift_pred - 1))
[0273] And the linear prediction is performed as follows.
[0274] Pcur(p) = (pRef(p) × a + b) >> shift_pred
[0275] Figure 24A modified process is shown, where the changed boxes are indicated in bold. The new boxes are step 808, which replaces previous step 508, and step 809, which replaces previous step 509.
[0276] Figure 18 An embodiment of method 1800 according to the main aspects described herein is shown. The method begins at start block 1801, and control proceeds to block 1810 for predicting samples in a current block based on at least one of the neighboring samples in the current block and based on a parametric model calculated from the neighboring samples in the current block and reference samples in a reference frame. Control proceeds from block 710 to block 720 to encode the block using the predicted samples.
[0277] Figure 25 Another embodiment of method 2500 according to the main aspects described herein is shown. The method begins at block 2501, and control proceeds to block 2510 for predicting samples in a current block based on at least one of the neighboring samples in the current block and based on a parametric model calculated from the neighboring samples in the current block and reference samples in a reference frame. Control proceeds from block 2510 to block 2520 to decode the block using the predicted samples.
[0278] Figure 26 An embodiment of apparatus 2600 for encoding, decoding, compressing, or decompressing video data using a simplification of a neighboring sample-dependent parametric model-based decoding mode is shown. The apparatus includes a processor 2610 and may be interconnected with a memory 2620 via at least one port. Both the processor 2610 and the memory 2620 may also have one or more additional interconnections to external connections.
[0279] The processor 2610 is further configured to insert or receive information in a bitstream and to compress, encode, or decode using any of the described aspects.
[0280] This application describes multiple aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described as having particularities and are generally described in a way that may sound restrictive, at least to show individual characteristics. However, this is for the purpose of clear description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide additional aspects. Additionally, these aspects can also be combined and interchanged with aspects described in earlier documents.
[0281] The aspects described and contemplated in this application can be implemented in many different forms. The following Figure 5 、 6 and 19 provide some embodiments, but other embodiments can be envisioned, and for Figure 5, 6 The discussion of 19 does not limit the breadth of implementation. At least one of the aspects mainly relates to video encoding and decoding, and at least one other aspect mainly relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing a bitstream generated according to any of the described methods.
[0282] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture" and "frame" may be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0283] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions may be modified or combined.
[0284] The various methods and other aspects described in this application can be used to modify modules, for example, Figure 5 and Figure 6 the intra prediction, entropy coding, and / or decoding modules (160, 260, 145, 230) shown. In addition, the present invention is not limited to VVC or HEVC, and can be applied to, for example, other standards and proposals (whether pre-existing or future-developed) and extensions of any such standards and proposals (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application can be used alone or in combination.
[0285] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0286] Figure 5 An encoder 100 is shown. Variants of the encoder 100 can be envisioned, but for clarity, the encoder 100 is described below without describing all the expected variants.
[0287] Before being encoded, a video sequence can undergo pre-encoding processing (101), for example, applying a color transformation to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the preprocessing and appended to the bitstream.
[0288] In encoder 100, as described below, pictures are encoded by encoder elements. The picture to be encoded is partitioned (102) and processed in units such as CUs. Each unit is encoded using, for example, an intra or inter mode. When a unit is encoded in the intra mode, intra prediction (160) is performed. In the inter mode, motion estimation (175) and compensation (170) are performed. The encoder determines (105) which one of the intra mode or the inter mode to use for encoding the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.
[0289] Then, the prediction residual is transformed (125) and quantized (130). Entropy coding (145) is performed on the quantized transform coefficients, motion vectors, and other syntax elements to output a bitstream. The encoder may skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass both the transformation and quantization, that is, directly decode the residual without applying the transformation or quantization process.
[0290] The encoder decodes the decoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse-transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).
[0291] Figure 6 A block diagram of video decoder 200 is shown. In decoder 200, as described below, the bitstream is decoded by decoder elements. Video decoder 200 generally performs a decoding process that is inverse to the encoding process as Figure 5 described. Encoder 100 generally also performs video decoding as part of encoding video data.
[0292] In particular, the input to the decoder includes a video bitstream, which may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other decoded information. The picture partitioning information indicates how the picture is partitioned. The decoder can thus partition (235) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined (255) with the prediction block to reconstruct the image block. The prediction block can be obtained (270) from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0293] The decoded picture may further undergo post-decoding processing (285), e.g., an inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the remapping process performed in the precoding process (101). The post-decoding processing may use metadata derived in the precoding process and signaled in the bitstream.
[0294] Figure 19 A block diagram of an example of a system in which various aspects and embodiments are implemented is shown. System 1000 may be implemented as a device including the various components described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described herein.
[0295] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing various aspects described herein, for example. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., volatile memory devices and / or non-volatile memory devices). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0296] The system 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents the module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding module and the decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000 or may be incorporated within the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0297] The program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to execute the various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during the execution of the processes described herein. These stored items may include but are not limited to input video, decoded video or portions of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.
[0298] In some embodiments, the memory within the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Team of experts JVET).
[0299] As shown in block 1130, input to the elements of the system 1000 can be provided through various input devices. Such input devices include, but are not limited to: (i) an RF section that receives RF signals transmitted over the air, for example, by a broadcaster, (ii) component (COMP) input terminals (or a set of component input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high definition multimedia interface (HDMI) input terminals. Figure 19 Other examples not shown include composite video.
[0300] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection to a desired band by filtering, down-converting, and filtering again. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0301] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 1000 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Solomon error correction) may be implemented, as needed, within, for example, a separate input processing IC or processor 1010. Similarly, various aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 1010. The streams of demodulation, error correction, and demultiplexing are provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0302] The various elements of system 1000 may be disposed within an integrated housing. Within this integrated housing, the various elements may be interconnected using suitable connection arrangements (e.g., internal buses known in the art, including inter-integrated circuit (I2C) buses, wiring, and printed circuit boards) and data may be transmitted therebetween.
[0303] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0304] In various embodiments, a wireless network (e.g., a Wi-Fi network, such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to the system 1000. The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other cloud-based communications. Other embodiments use a set-top box that transfers data via an HDMI connection of the input box 1130 to provide streaming data to the system 1000. Still other embodiments use an RF connection of the input box 1130 to provide streaming data to the system 1000. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0305] The system 1000 may provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes one or more of the following: for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 may be used for a television, a tablet, a laptop computer, a cellular phone (mobile phone), or other devices. The display 1100 may also be integrated with other components (e.g., as in a smart phone), or separate (e.g., an external monitor for a laptop computer). In various examples of the embodiments, the other peripheral devices 1120 include one or more of the following: a standalone digital video disc (or digital versatile disc) (DVR, for both), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of the system 1000. For example, a disc player performs the function of playing the output of the system 1000.
[0306] In various embodiments, control signals are transmitted between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link (AV Link), Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speaker 1110 may be integrated in a single unit in an electronic device (e.g., a television) together with other components of system 1000. In various embodiments, display interface 1070 includes a display driver, such as a timing controller ((T Con) chip.
[0307] For example, if the RF portion of input 1130 is part of a separate set-top box, the display 1100 and speaker 1110 may alternatively be separated from one or more of the other components. In various embodiments where the display 1100 and speaker 1110 are external components, the output signal may be provided via a dedicated output connection, such as including an HDMI port, a USB port, or a COMP output.
[0308] These embodiments may be implemented by processor 1010 or by computer software implemented in hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. Memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, as non-limiting examples, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 1010 may be of any type suitable for the technical environment and, as non-limiting examples, may include one or more of the following: a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0309] Various implementations involve decoding. As used in this application, "decoding" may include, for example, all or part of the processing performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include processes performed by the decoders of the various implementations described in this application.
[0310] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0311] Various implementations involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can include, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as, partitioning, differential encoding, transformation, quantization, and entropy coding. In various embodiments, such processes also or alternatively include processes performed by the encoders of the various implementations described in this application.
[0312] As a further example, in one embodiment, "encoding" refers only to entropy coding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy coding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0313] Note that the grammatical elements used herein are descriptive terms. Thus, they do not exclude the use of other grammatical element names.
[0314] When the drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding method / process.
[0315] Various embodiments relate to parametric models. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually with a constraint on computational complexity. It can be measured by a rate-distortion optimization (RDO) metric, or by least mean square (LMS), mean absolute error (MAE), or other such measurements. The rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, which includes all considered modes or decoder parameter values, and a complete evaluation of their decoding costs and the associated distortion of the reconstructed signal after decoding and decoding. Faster methods can also be used to save coding complexity, especially by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, for example, by using approximate distortion only for some possible coding options and full distortion for other coding options. Other methods only evaluate a subset of the possible coding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the decoding cost and the associated distortion.
[0316] The implementations and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (e.g., only discussed as a method), the implementation of the features discussed can also be implemented in other forms (e.g., an apparatus or a program). For example, an apparatus can be implemented in appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, which includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end users.
[0317] References to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation” and any other variations thereof that occur in various places in this application do not necessarily all refer to the same embodiment.
[0318] In addition, this application may relate to “determining” various information. Determining the information can include, for example, one or more of the following: estimating the information, calculating the information, predicting the information, or retrieving the information from a memory.
[0319] In addition, the present application may relate to "accessing" various information. Accessing the information may include, for example, one or more of the following: receiving the information, retrieving the information (e.g., retrieving the information from a memory), storing the information, moving the information, copying the information, computing the information, determining the information, predicting the information, or estimating the information.
[0320] Additionally, the present application may refer to "receiving" various information. Like "accessing", receiving is intended to be a broad term. Receiving the information may include, for example, one or more of the following: accessing the information or retrieving the information (e.g., from a memory). Further, during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.
[0321] It should be understood that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of the first and second-listed options (A and B), or the selection of the first and third-listed options (A and C), or the selection of the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.
[0322] In addition, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a particular one of a plurality of decoding modes or flags. Thus, in one embodiment, the same parameters are used on the encoder side and the decoder side. Accordingly, for example, an encoder may send (explicitly signal) a particular parameter to a decoder such that the decoder may use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as other parameters, signaling may be used without transmission (implicitly signal) simply to allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling may be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the term "signal", the term "signal" may also be used as a noun herein.
[0323] As will be apparent to those of ordinary skill in the art, implementations may generate various signals that are formatted to carry information such as may be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over various different wired or wireless links. The signal may be stored on a processor-readable medium.
[0324] We have described a number of embodiments. The features of these embodiments may be provided individually or in any combination. Additionally, across various claim categories and types, embodiments may individually or in any combination include one or more of the following features, devices, or aspects:
[0325] ● Modify the decoding mode processing applied in the decoder and / or encoder.
[0326] ● Enable a number of advanced decoding mode prediction methods in the decoder and / or encoder.
[0327] ● Insert syntax elements in the signaling such that the decoder can identify the decoding mode prediction method to use.
[0328] ● Based on these syntax elements, select the decoding mode prediction method to apply to the decoder.
[0329] ● Apply the decoding mode prediction method at the decoder for deriving the mode.
[0330] ● Derive a parameter by using the above prediction process and by using the removal of a lookup table.
[0331] ● Derive a parameter by using the above prediction process and by using the modification of a lookup table.
[0332] ● Use linear prediction to derive a prediction parameter.
[0333] ● Adapt the residual at an encoder according to any of the embodiments discussed.
[0334] ● A bitstream or signal including one or more of the described syntax elements or a variant thereof.
[0335] ● A bitstream or signal including a syntax that conveys information generated according to any of the described embodiments.
[0336] ● Create and / or transmit and / or receive and / or decode according to any of the described embodiments.
[0337] ● A method, process, device, medium storing instructions, medium storing data, or signal according to any of the described embodiments.
[0338] ● Insert into the signaling a syntax element that enables the decoder to determine the decoding mode in a manner corresponding to the way used by the encoder.
[0339] ● Create and / or transmit and / or receive and / or decode a bitstream or signal including one or more of the described syntax elements or a variant thereof.
[0340] ● A TV, set - top box, cellular phone, tablet computer, or other electronic device that performs decoding mode determination according to any of the described embodiments.
[0341] ● A TV, set - top box, cellular phone, tablet computer, or other electronic device that performs decoding mode determination according to any of the described embodiments and displays (e.g., using a monitor, screen, or other type of display) the resulting image.
[0342] ● A TV, set - top box, cellular phone, tablet computer, or other electronic device that selects, band - limits, tunes (e.g., using a tuner) a channel to receive a signal including an encoded image and performs decoding mode determination according to any of the described embodiments.
[0343] ● A TV, set - top box, cellular phone, tablet computer, or other electronic device that receives (e.g., using an antenna) a signal including an encoded image over - the - air and performs decoding mode determination.
Claims
1. A method, comprising: determining a prediction of a sample in the current block based on at least one of adjacent samples in the current block and based on a parametric model, the parametric model being calculated from the adjacent samples in the current block and reference samples in a reference frame; encoding the sample in the current block based on the prediction.
2. An apparatus, comprising: a processor configured to: determine a prediction of a sample in the current block based on at least one of adjacent samples in the current block and based on a parametric model, the parametric model being calculated from the adjacent samples in the current block and reference samples in a reference frame; encode the sample in the current block based on the prediction.
3. A method, comprising: determining a prediction of a sample in the current block based on at least one of adjacent samples in the current block and based on a parametric model, the parametric model being calculated from the adjacent samples in the current block and reference samples in a reference frame; decoding the sample in the current block based on the prediction.
4. An apparatus, comprising: a processor configured to: determine a prediction of a sample in the current block based on at least one of adjacent samples in the current block and based on a parametric model, the parametric model being calculated from the adjacent samples in the current block and reference samples in a reference frame; decode the sample in the current block based on the prediction.
5. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the parametric model is derived from a linear model.
6. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the parameters of the parametric model are derived using a look-up table.
7. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the parameters of the parametric model are derived from at least two samples of adjacent samples with a spatial distance constraint.
8. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the parameters of the parametric model are derived from at least three adjacent samples, wherein the three samples are respectively at the rightmost of the top row of adjacent samples above the block, at the bottom of the column of the left adjacent sample, and at the intersection of the top reference row and the left reference column.
9. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein, if the linear parameter derivation is well-defined, a prediction based on a linear model is used, otherwise an alternative mode is used.
10. The method according to claim 1 or 3, or the apparatus according to claim 2 or 4, wherein the derivation of the parameters of the parametric model includes calibration parameters.
11. The method according to claim 1 or claim 3, or the apparatus according to claim 2 or claim 4, wherein, a cross-component linear model is enabled for predicting the chrominance components of an inter-frame decoded block.
12. A device, comprising: the apparatus according to any one of claims 4 to 11; and At least one of the following: (i) an antenna configured to receive a signal that includes a video block, (ii) a band limiter configured to limit the received signal to a band that includes the video block, and (iii) a display configured to display an output representative of the video block.
13. A non-transitory computer-readable medium comprising data content generated by the method according to any one of claims 1 and 5 to 11 or by the apparatus according to any one of claims 2 and 5 to 11 for playback using a processor.
14. A signal comprising video data generated by the method according to any one of claims 1 and 5 to 11 or by the apparatus according to any one of claims 2 and 5 to 11 for playback using a processor.
15. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1, 3, and 5 to 11.