Constrained Upsampling Process in Matrix-Based Intra Prediction
By employing matrix-based intra prediction methods, specifically ALWIP modes, the challenges of high bandwidth demand in video coding are addressed, resulting in improved coding efficiency and reduced complexity for both existing and future video standards.
Patent Information
- Application Number
- JP2023190715
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2023-11-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-05-28
AI Technical Summary
Existing video coding technologies face challenges in efficiently processing and compressing digital video, particularly due to increasing demands for higher resolution videos, which leads to higher bandwidth requirements.
The implementation of matrix-based intra prediction methods for video coding, including affine linear weighted intra prediction (ALWIP) modes, which perform conversions between video blocks and bitstream representations using boundary downsampling, matrix-vector multiplication, and upsampling operations.
This approach enhances video coding efficiency by improving runtime performance and reducing computational complexity, while maintaining or improving video quality across various video standards.
Smart Images

Figure 0007688092000057 
Figure 0007688092000058 
Figure 0007688092000059
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application is based on International Patent Application No. PCT / CN2020 / 092906 filed on May 28, 2020 and Japanese Patent Application No. 2021 - 570148 filed on November 25, 2021, and claims the priority and benefit of International Patent Application No. PCT / CN2019 / 089590 filed on May 31, 2019. For all purposes under the law, the entire disclosure of the above application is incorporated by reference as part of the disclosure of this application.
[0002] [Technical Field] This patent document relates to video coding technology, devices, and systems.
Background Art
[0003] Despite the progress of video compression, digital video still occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use is expected to continue to increase.
Summary of the Invention
[0004] Digital video coding, and specifically, apparatuses, systems, and methods related to matrix - based intra - prediction methods for video coding are described. The described methods can be applied to both existing video coding standards (e.g., HEVC (High Efficiency Video Coding)) and future video coding standards (e.g., VVC (Versatile Video Coding)) or codecs.
[0005] In a representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix based intra prediction (MIP) mode, in which the predicted block of the current video block performs a boundary downsampling operation on samples previously coded in the video, followed by a matrix vector multiplication operation, and then optionally an upsampling operation to determine, and the conversion includes performing a boundary downsampling operation in a single step in which the reduced boundary samples of the current video block are generated according to a rule based at least on the reference boundary samples of the current video block, and the conversion includes performing a matrix vector multiplication operation using the reduced boundary samples of the current video block.
[0006] In another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix based intra prediction (MIP) mode, in which the final predicted block of the current video block performs a boundary downsampling operation on samples previously coded in the video, followed by a matrix vector multiplication operation, and then an upsampling operation to determine, and the conversion includes performing an upsampling operation determined by the final predicted block using the reduced predicted block of the current video block and using the reconstructed adjacent samples of the current video block according to a rule, and the reduced predicted block is obtained by performing a matrix vector multiplication operation on the reduced boundary samples of the current video block.
[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, constructing at least a portion of an MPM list for the ALWIP mode based on at least a portion of an MPM list for non-ALWIP intra modes based on the determination, and performing a conversion between the current video block and a bitstream representation of the current video block based on the MPM list for the ALWIP mode.
[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes determining that a luminance component of a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, estimating a chroma intra mode based on the determination, and performing a conversion between the current video block and a bitstream representation of the current video block based on the chroma intra mode.
[0009] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, and performing a conversion between the current video block and a bitstream representation of the current video block based on the determination.
[0010] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes determining that a current video block is coded using a coding mode different from the affine linear weighted intra prediction (ALWIP) mode, and based on this determination, performing a conversion between the current video block and the bitstream representation of the current video block.
[0011] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes generating a first prediction for a current video block using the affine linear weighted intra prediction (ALWIP) mode, generating a second prediction for the current video block using a position dependent intra prediction combination (PDPC) based on the first prediction, and based on the second prediction, performing a conversion between the current video block and the bitstream representation of the current video block.
[0012] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. This exemplary method includes determining that a current video block is coded using the affine linear weighted intra prediction (ALWIP) mode, predicting a plurality of sub-blocks of the current video block based on the ALWIP mode, and based on this prediction, performing a conversion between the current video block and the bitstream representation of the current video block.
[0013] In yet another representative aspect, a method of video processing is disclosed. The method includes determining a context of a flag indicating use of an affine linear weighted intra prediction (ALWIP) mode based on rules for a current video block during conversion between the current video block and a bitstream representation of the current video block; predicting a plurality of sub-blocks of the current video block based on the ALWIP mode; and performing a conversion between the current video block and the bitstream representation of the current video block based on the prediction.
[0014] In yet another representative aspect, a method of video processing is disclosed. The method includes determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode; and performing at least two filtering stages on samples of the current video block in an upsampling process associated with the ALWIP mode during conversion between the current video block and a bitstream representation of the current video block, wherein a first accuracy of samples in a first filtering stage of the at least two filtering stages is different from a second accuracy of samples in a second filtering stage of the at least two filtering stages.
[0015] In yet another aspect, a method of video processing is disclosed. The method includes determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode; and performing at least two filtering stages on samples of the current video block in an upsampling process associated with the ALWIP mode during conversion between the current video block and a bitstream representation of the current video block, wherein the upsampling process is performed in a fixed order when both vertical upsampling and horizontal upsampling are performed.
[0016] In yet another further aspect, a method of video processing is disclosed. The method includes determining that a current video block is to be coded using an Affine Linear Weighted Intra Prediction (ALWIP) mode, and performing at least two filtering stages on samples of the current video block in an upsampling process associated with the ALWIP mode during conversion between the current video block and a bitstream representation of the current video block, the conversion including performing a transpose operation prior to the upsampling process.
[0017] In yet another representative further aspect, the above-described method is embodied in the form of processor-executable code and stored in a computer-readable program medium.
[0018] In yet another representative further aspect, an apparatus configured or operable to perform the above-described method is disclosed. The apparatus may include a processor programmed to implement this method.
[0019] In yet another representative further aspect, a video decoder device may implement the method described herein.
[0020] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
[0022] Due to the increasing demands for higher resolution videos, video coding methods and technologies have become ubiquitous in modern technology. A video codec typically includes electronic circuitry or software for compressing or decompressing digital video and is constantly being improved to provide higher coding efficiency. A video codec converts uncompressed video into a compressed format or vice versa. There is a complex relationship among video quality, the amount of data used to represent the video (determined by the bitrate), the complexity of the encoding and decoding algorithms, the sensitivity to data loss and errors, the ease of editing, random access, and end-to-end latency. Compression formats usually conform to standard video compression specifications such as, for example, the HEVC (High Efficiency Video Coding) standard (also known as H.265 or MPEG-H Part 2), the upcoming VVC (Versatile Video Coding) standard, or other current and / or future video coding standards.
[0023] Embodiments of the disclosed technology can be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve runtime performance. In this document, section headings are used to improve readability of the description, but those section headings do not limit the description or embodiments (and / or implementations) to only their respective sections.
[0024] 1 A brief review of HEVC 1.1 Intra prediction in HEVC / H.265 Intra prediction involves generating samples for a given TB (transform block) using previously reconstructed samples in the color channel under consideration. Intra prediction modes are signaled separately for the luma channel and the chroma channels, and the chroma channel intra prediction mode optionally depends on the luma channel intra prediction mode via the ‘DM_CHROMA’ mode. Intra prediction modes are signaled at the PB (prediction block) level, but the intra prediction process is applied at the TB level according to the remaining quadtree hierarchy of the CU, thereby enabling the coding of one TB to affect the coding of the next TB within the CU and thus shortening the distance to the samples used as reference values.
[0025] HEVC includes 35 intra prediction modes, namely the DC mode, the planar mode, and 33 directional or ‘angular’ intra prediction modes. The 33 angular intra prediction modes are shown in Figure 1.
[0026] For PBs related to the chroma color channels, the intra prediction mode is defined as either the planar mode, the DC mode, the horizontal mode, the vertical mode, the ‘DM_CHROMA’ mode, or sometimes the diagonal mode ‘34’.
[0027] Note that in the chroma formats 4:2:2 and 4:2:0, chroma PB can (respectively) overlap with two or four luma PB. In this case, the luma direction in DM_CHROMA is taken from the upper left of these luma PB.
[0028] The DM_CHROMA mode indicates that the intra prediction mode of the luma color channel PB is applied to the chroma color channel PB. Since this is relatively common, the most accurate mode coding scheme for intra_chroma_pred_mode is biased to favor the selection of this mode.
[0029] 2 Examples of Intra Prediction in VVC 2.1 Intra Mode Coding with 67 Intra Prediction Modes To capture any edge direction shown in natural video, the number of directional intra modes is extended from 33 to 65 as used in HEVC. The additional direction modes are shown as red dotted arrows in Figure 2, and the planar mode and DC mode remain the same. These more dense directional intra prediction modes are applied to all block sizes and to both luma intra prediction and chroma intra prediction.
[0030] 2.2 Example of Cross-Component Linear Model (CCLM) In some embodiments, also, to reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode (also referred to as LM) is used in JEM, where chroma samples are:
Equation
[0031] Here, pred C(i, j) represents the predicted chroma sample within the CU, and recL’(i, j) represents the downsampled reconstructed luma sample of the same CU. The linear model parameters α and β are derived from the relationship between the luma values and chroma values from two samples, which are the luma sample with the minimum sample value and the luma sample with the maximum sample value within the set of downsampled adjacent luma samples and their corresponding chroma samples. Figure 3 shows an example of the sample arrangement of the current block involved in the CCLM mode and the positions of the left sample and the top sample.
[0032] This parameter calculation is performed as part of the decoding process, not simply as an encoder search process. As a result, there is no syntax used to convey the α and β values to the decoder.
[0033] In chroma intra mode coding, a total of eight intra modes are made available for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). Chroma mode coding directly depends on the intra prediction mode of the corresponding luma block. In an I slice, since a separate block partitioning structure is possible for the luma component and the chroma component, one chroma block may correspond to multiple luma blocks. Therefore, in the chroma DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0034] 2.3 Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 4 shows an example of four reference lines, where samples in segments A and F are padded with the closest samples from segments B and E respectively, rather than fetched from the reconstructed adjacent samples. HEVC intra picture prediction uses the closest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used. The index (mrl_idx) of the selected reference line is signaled and used to generate the intra predictor. For a reference line idx greater than 0, only the additional reference line modes within the MPM list are included and only the mpm index is signaled without the remaining modes.
[0035] 2.4 Intra sub-partition (ISP) The Intra Sub-Partitions (ISP) tool divides the luma intra prediction block into two or four sub-partitions either vertically or horizontally depending on the block size. For example, the minimum block size in ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided by four sub-partitions. Figure 5 shows two examples of possibilities. All sub-partitions satisfy the condition of having at least 16 samples.
[0036] For each sub - partition, a reconstructed sample is obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization, and inverse transformation, for example. Therefore, the reconstructed sample values of each sub - partition are available for generating the prediction of the next sub - partition, and each sub - partition is processed iteratively. Further, the first sub - partition to be processed is the one containing the top - left sample of the CU, and then it is followed by downward (horizontal split) or right - ward (vertical split). As a result, the reference samples used for generating the sub - partition prediction signal are located only on the left and above the line. All sub - partitions share the same intra - mode.
[0037] 2.5 Affine Linear Weighted Intra Prediction (ALWIP or Matrix - based Intra Prediction) Affine Linear Weighted Intra Prediction (ALWIP, also known as Matrix - based Intra Prediction (MIP)) was proposed in JVET - N0217.
[0038] Two tests are conducted in JVET - N0217. In Test 1, ALWIP is designed with an 8 - KB memory constraint and a maximum of 4 multiplications per sample. Test 2 is similar to Test 1, but further simplifies the design in terms of memory requirements and model architecture: · A single set of matrices and offset vectors for all block shapes. · Reducing the number of modes to 19 for all block shapes. · Reducing the memory requirement to 5760 10 - bit values, i.e., up to 7.20 kilobytes. · Linear interpolation of prediction samples is performed in a single step per direction, replacing the iterative interpolation as in the first test.
[0039] 2.5.1 Test 1 of JVET - N0217 To predict a sample of a rectangular block of width W and height H, Affine Linear Weighted Intra Prediction (ALWIP) takes as input H reconstructed adjacent boundary samples of one line to the left of the block and W reconstructed adjacent boundary samples of one line above the block. If these reconstructed samples are not available, they are generated as done in conventional intra prediction. The generation of the prediction signal is based on the following three steps.
[0040] From the boundary samples, four samples in the case of W = H = 4 and eight samples in all other cases are extracted by averaging.
[0041] Using these averaged samples as input, matrix-vector multiplication and subsequent addition of an offset are performed. The result is a reduced prediction signal for a subsampled set of samples within the original block.
[0042] From the prediction signal for the subsampled set, the prediction signals for the remaining positions are generated by linear interpolation, which is linear interpolation with a single step in each direction.
[0043] The matrices and offset vectors required to generate the prediction signal are taken from three sets of matrices S 0 , S 1 , S 2 . Set S 0 consists of 18 matrices A 0 i , i ∈ {0, …, 17}, each having 16 rows and 4 columns, and 18 offset vectors b 0 i , i ∈ {0, …, 17}, each having size 16. This set of matrices and offset vectors is used for 4×4 blocks. Set S 1 consists of 10 matrices A 1 i , i ∈ {0, …, 9}, each having 16 rows and 8 columns, and 10 offset vectors b 1 i, which consists of \(i\in\{0,\ldots,9\}\). The matrices and offset vectors of this set are used for blocks of sizes \(4\times8\), \(8\times4\), and \(8\times8\). Finally, the set \(S\) 2 consists of six matrices \(A\) 2 i , \(i\in\{0,\ldots,5\}\) each having 64 rows and 8 columns and six offset vectors \(b\) 2 i , \(i\in\{0,\ldots,5\}\) each having size 64. The matrices and offset vectors of this set, or a part of these matrices and offset vectors, are used for all other block shapes.
[0044] The total number of multiplications required for calculating the matrix - vector product is always less than or equal to \(4\times W\times H\). In other words, the multiplications required in the ALWIP mode are at most 4 per sample.
[0045] 2.5.2 Boundary Averaging In the first step, the input boundaries \(bdry\) top and \(bdry\) left are reduced to smaller boundaries \(bdry\) red top and \(bdry\) red left . Here, \(bdry\) red top and \(bdry\) red left both consist of 2 samples in the case of \(4\times4\) blocks and both consist of 4 samples in all other cases.
[0046] In the case of \(4\times4\) blocks, for \(0\leq i\lt2\), [Equation] is defined, and \(bdry\) red left is defined similarly.
[0047] In other cases, when the block width \(W\) is given as \(W = 4\cdot2^k\), for \(0\leq i\lt4\), [Equation] is defined as, and bdry red left is defined similarly.
[0048] These two reduced boundaries bdry red top and bdry red left are concatenated into the reduced boundary vector bdry red which is thus of size 4 for a 4×4 block in shape and size 8 for all other shaped blocks. When the mode points to the ALWIP mode, this concatenation is defined as follows:
Number
[0049] Finally, for the interpolation of the subsampled prediction signal, a second version of the averaged boundary is required for large blocks. That is, if min(W,H)>8 and W≧H, then W is described as W = 8×2l, where 0≦i<8,
Number
[0050] If min(W,H)>8 and H>W, then bdry redII left is defined similarly.
[0051] 2.5.3 Generation of Reduced Prediction Signal by Matrix-Vector Multiplication From the reduced input vector bdry red the reduced prediction signal pred red is generated. The latter signal is for a downsampled block of width W red and height H red . Here, W red and H red are:
Number
[0052] The reduction prediction signal pred red is calculated by computing a matrix-vector product and adding a correction value: pred red = A·bdry red + b
[0053] where A is a matrix having the rows of W red ·H red and having 4 columns when W = H = 4, and 8 columns in all other cases, and b is a vector of size W red ·H red .
[0054] The matrix A and the vector b are taken from one of the following sets S 0 , S 1 , S 2 . Define the index idx = idx(W, H) as follows:
Number
[0055] Also, let m be the following:
Number
[0056] And if idx ≤ 1 or idx = 2 and min(W, H)>4, then A = A idx m and b = b idx m . When idx = 2 and min(W, H) = 4, in the case of W = 4, A corresponding to the odd x - coordinates in the downsampled block idx m or in the case of H = 4, A corresponding to the odd y - coordinates in the downsampled block idx m Let A be the matrix resulting from excluding all rows of.
[0057] Finally, the reduced prediction signal is replaced by its transpose matrix in the following cases: · W = H = 4 and mode ≥ 18 · max(W, H) = 8 and mode ≥ 10 · max(W, H) > 8 and mode ≥ 6.
[0058] pred red The number of multiplications required for the calculation of is 4 in the case of W = H = 4. This is because, in this case, A has 4 columns and 16 rows. In all other cases, A has 8 columns and W red · H red rows, and it can be immediately verified that in these cases 8 · W red · H red ≤ 4W · H multiplications are required, that is, also in these cases, the number of multiplications required to calculate pred red is at most 4 per sample.
[0059] 2.5.4 Illustration of the entire ALWIP process The entire process of averaging, matrix-vector multiplication, and linear interpolation is illustrated for different shapes in FIGS. 6 - 9. The remaining shapes are treated as in one of the illustrated cases.
[0060] 1. Given a 4×4 block, ALWIP takes two averages along each axis of the boundary. The resulting 4 input samples enter a matrix-vector multiplication. A matrix is taken from set S 0 After adding an offset, this produces 16 final prediction samples. No linear interpolation is required to generate this prediction signal. Thus, a total of (4 · 16) / (4 · 4) = 4 multiplications are performed per sample.
[0061] 2. Given an 8×8 block, ALWIP takes four averages along each axis of the boundary. The resulting 8 input samples enter a matrix-vector multiplication. A matrix is taken from set S 1A matrix is taken from . This produces 16 samples at the odd positions of the prediction block. Thus, a total of (8·16) / (8·8) = 2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically by using the reduced upper boundary. Horizontal interpolation continues by using the original left boundary.
[0062] 3. Given an 8×4 block, ALWIP takes four means along the horizontal axis of the boundary and four original boundary values on the left boundary. The resulting eight input samples enter a matrix-vector multiplication. Set S 1 A matrix is taken from . This produces 16 samples at the odd horizontal positions and each vertical position of the prediction block. Thus, a total of (8·16) / (8·4) = 4 multiplications are performed per sample. After adding the offset, these samples are interpolated horizontally by using the original left boundary.
[0063] 4. Given a 16×16 block, ALWIP takes four means along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. To generate the four means, a two-stage downsampling operation is utilized. First, every two consecutive samples are used to derive one downsampling value, thus obtaining eight means of the upper / left neighbors. Next, the eight means per side are further downsampled to generate four means per side. The four means are used to derive the reduced prediction signal. Additionally, the eight means per side are further utilized for generating the final prediction samples via subsampling of the reduced prediction signal. Set S 2A matrix is taken from it. This produces 64 samples at odd positions of the prediction block. Thus, a total of (8·64) / (16·16) = 2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically by using the 8 averages of the upper boundary. Horizontal interpolation continues by using the original left boundary. The interpolation process does not add any multiplications in this case. Thus, in total, 2 multiplications per sample are required to compute the ALWIP prediction.
[0064] For larger shapes, the procedure is basically the same, and it is easy to confirm that the number of multiplications per sample is less than 4.
[0065] For a W×8 block with W>8, only horizontal interpolation is required since samples are given at odd horizontal positions and each vertical position.
[0066] Finally, for a W×4 block with W>8, let matrix A_k be the matrix resulting from excluding all rows corresponding to odd entries along the horizontal axis within the downsampled block. Thus, the output size is 32, and in this case as well, only horizontal interpolation remains to be performed.
[0067] The transposed cases are also handled according to these.
[0068] 2.5.5 Single-step linear interpolation For a W×H block with max(W,H)≥8, the prediction signal is generated by linear interpolation from the reduced prediction signal pred for W red ×H red . Linear interpolation is performed vertically, horizontally, or in both directions depending on the block shape. When linear interpolation is applied in both directions, it is applied first horizontally when W<H, and first vertically otherwise. red
[0069] Without loss of generality, consider a W×H block where max(W,H)≧8 and W≧H. Then, one-dimensional linear interpolation is performed as follows. Without loss of generality, it is sufficient to describe the linear interpolation in the vertical direction. First, the prediction signal of reduction is extended upward by the boundary signal. The vertical upsampling factor U ver =H / H red is defined, and U ver =2 uver >1 is described. Then, the extended prediction signal of reduction is defined by
Number
[0070] And from this extended prediction signal of reduction, the prediction signal linearly interpolated in the vertical direction is generated for 0≦x<W red , 0≦y<H red , and 0≦k<U ver by
Number
[0071] 2.5.6 Signaling of the proposed intra prediction mode For each coding unit (CU) in the intra mode, a flag indicating whether the ALWIP mode is applied to the corresponding prediction unit (PU) is sent in the bitstream. The signaling of the latter index is coordinated with MRL in the same way as in JVET-M0043. When the ALWIP mode is applied, the index predmode of the ALWIP mode is signaled using an MPM list with three MPMS.
[0072] Here, the derivation of MPM is performed using the intra modes of the upper and left PUs as follows. Three fixed tables map_angular_to_alwip Angular that assign the ALWIP mode to each of the conventional intra prediction modes predmode idx , idx∈{0,1,2} exist: predmode ALWIP = map_angular_to_alwip idx [predmode Angular
[0073] Index indicating from which of the three sets the ALWIP parameters are taken for each PU of width W and height H, as in Section 2.5.3: idx(PU) = idx(W,H) ∈ {0,1,2} is defined.
[0074] Upper prediction unit PU above is available, belongs to the same CTU as the current PU, and is in the intra mode, and when idx(PU) = idx(PU above ) and ALWIP is applied to PU ALWIP above in predmode above then
Number
[0075] When the upper PU is available, belongs to the same CTU as the current PU, and is in the intra mode, and the conventional intra prediction mode predmode Angular above is applied to the upper PU
Number
[0076] In all other cases
Number
[0077] Finally, three fixed default lists, each containing three different ALWIP modes, list idx , idx ∈ {0, 1, 2} is provided. The default list list idx(PU) as well as the mode mode ALWIP above and mode ALWIP left are used to construct three different MPMs by replacing -1 with the default value and eliminating duplicates.
[0078] The left and upper adjacent blocks used for ALWIP MPM list construction are A1 and B1 shown in FIG. 10.
[0079] 2.5.7 Adaptive MPM List Derivation for Conventional Luma and Chroma Intra Prediction Modes The proposed ALWIP mode is reconciled with the MPM - based coding of the conventional intra prediction mode as follows. The luma and chroma MPM list derivation processes for the conventional intra prediction mode map the ALWIP mode predmode ALWIP for a given PU to one of the conventional intra prediction modes using a fixed table map_alwip_to_angular idx , idx ∈ {0, 1, 2}: predmode Angular = map_alwip_to_angular idx(PU) [predmode ALWIP
[0080] In luma MPM list derivation, whenever an adjacent luma block using the ALWIP mode predmode ALWIP is encountered, this block is treated as if it were using the conventional intra prediction mode predmode Angular . In chroma MPM list derivation, whenever the current luma block uses the LWIP mode, the ALWIP mode is converted to the conventional intra prediction mode using the same mapping.
[0081] 2.5.8 Corresponding Revised Working Draft In some embodiments, as described in this section, based on the embodiments of the disclosed technology, portions related to intra_lwip_flag, intra_lwip_mpm_flag, intra_lwip_mpm_idx, and intra_lwip_mpm_remainder have been added to the working draft.
[0082] In some embodiments, as described in this section, for representing additions and changes to the working draft based on the embodiments of the disclosed technology, <begin>Tag and <end>Tags are used. Syntax Table Coding unit syntax [Table 1] TIFF0007688092000015.tif18166 Semantic <begin>intra_lwip_flag[x0][y0] equal to 1 specifies that the intra prediction type for the luma sample is the affine linear weighted intra prediction. intra_lwip_flag[x0][y0] equal to 0 specifies that the intra prediction type for the luma sample is not the affine linear weighted intra prediction. If intra_lwip_flag[x0][y0] does not exist, it is assumed to be equal to 0. The syntax elements intra_lwip_mpm_flag[x0][y0], intra_lwip_mpm_idx[x0][y0], and intra_lwip_mpm_remainder[x0][y0] specify the affine linear weighted intra prediction mode for the luma sample. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. If intra_lwip_mpm_flag[x0][y0] is equal to 1, the affine linear weighted intra prediction mode is inferred from the adjacent intra predicted coding units according to Section 8.4.X. If intra_lwip_mpm_flag[x0][y0] does not exist, it is assumed to be equal to 1. <end> The intra_subpartitions_split_flag[x0][y0] defines whether the intra subpartition split type is horizontal or vertical. If intra_subpartitions_split_flag[x0][y0] does not exist, it is estimated as follows: - If intra_lwip_flag[x0][y0] is equal to 1, intra_subpartitions_split_flag[x0][y0] is estimated to be equal to 0. - Otherwise, the following applies: - If cbHeight is greater than MaxTbSizeY, intra_subpartitions_split_flag[x0][y0] is estimated to be equal to 0; - Otherwise (if cbWidth is greater than MaxTbSizeY), intra_subpartitions_split_flag[x0][y0] is assumed to be equal to 1. Decoding Process 8.4.1 General decoding process for coding units coded in intra prediction mode The inputs to this process are as follows: - The luma position (xCb, yCb) that defines the top-left sample of the current coding block with respect to the top-left luma sample of the current picture, - The variable cbWidth that defines the width of the current coding block in luma samples, - The variable cbHeight that defines the height of the current coding block in luma samples, - The variable treeType that defines whether a single tree is used or a dual tree is used, and if a dual tree is used, whether the current tree corresponds to the luma component or the chroma component. The output of this process is the modified reconstructed picture before in-loop filtering. The derivation process of quantization parameters defined in Section 8.7.1 is called with the luma position (xCb, yCb), the width cbWidth of the current coding block in the luma samples, the height cbHeight of the current coding block in the luma samples, and the variable treeType as inputs. When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_LUMA, the decoding process for luma samples is defined as follows: - When pcm_flag[xCb][yCb] is equal to 1, the reconstructed picture is changed as follows: SL[xCb+i][yCb+j]= pcm_sample_luma[(cbHeight*j)+i]<<(BitDepthY-PcmBitDepthY), (8-6) with i = 0..cbWidth-1, j = 0..cbHeight-1 - Otherwise, the following applies: 1. The luma intra prediction mode is derived as follows: - When intra_lwip_flag[xCb][yCb] is equal to 1, the derivation process of the affine linear weighted intra prediction mode defined in Section 8.4.X is called with the luma position (xCb, yCb), the width cbWidth of the current coding block in the luma samples, and the height cbHeight of the current coding block in the luma samples as inputs; - Otherwise, the derivation process of the luma intra prediction mode defined in Section 8.4.2 is called with the luma position (xCb, yCb), the width cbWidth of the current coding block in the luma samples, and the height cbHeight of the current coding block in the luma samples as inputs; The general decoding process for intra-blocks defined in section 2.8.4.4.1 is called with variables nTbW set equal to the luma position (xCb, yCb), tree type treeType, cbWidth, variable nTbH set equal to cbHeight, variable predModeIntra set equal to IntraPredModeY[xCb][yCb], and variable cIdx set equal to 0 as inputs, and its output is the modified reconstructed picture before in-loop filtering. … <begin> 8.4.X Derivation Process of Affine Linear Weighted Intra Prediction Mode The input to this process is as follows: - The luma position (xCb, yCb) that defines the top-left sample of the current coding block with respect to the top-left luma sample of the current picture, - The variable cbWidth that defines the width of the current coding block in the luma samples, - The variable cbHeight that defines the height of the current coding block in the luma samples. In this process, the affine linear weighted intra prediction mode IntraPredModeY[xCb][yCb] is derived. IntraPredModeY[xCb][yCb] is derived by the following ordered steps: 1. The adjacent positions (xNbA, yNbA) and (xNbB, yNbB) are set equal to (xCb - 1, yCb) and (xCb, yCb - 1), respectively; 2. For X that is replaced by either A or B, the variable candLwipModeX is derived as follows: - The block availability derivation process [Ed.(BB): Neighboring block availability check process tbd] defined in section 6.4.X is called with the position (xCurr, yCurr) set equal to (xCb, yCb) and the adjacent position (xNbY, yNbY) set equal to (xNbX, yNbX), and its output is assigned to availableX; - The candidate affine linear weighted intra prediction mode candLwipModeX is derived as follows: - If one or more of the following conditions are true, candLwipModeX is set equal to -1; - The variable availableX is equal to FALSE; - CuPredMode[xNbX][yNbX] is not equal to MODE_INTRA and mh_intra_flag[xNbX][yNbX] is not equal to 1; - pcm_flag[xNbX][yNbX] is equal to 1; - X is equal to B, and yCb - 1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY); - Otherwise, the following applies: - The block size type derivation process defined in section 8.4.X.1 is called with the width cbWidth of the current coding block in luma samples and the height cbHeight of the current coding block in luma samples as inputs, and its output is assigned to the variable sizeId; - If intra_lwip_flag[xNbX][yNbX] is equal to 1, the block size type derivation process defined in section 8.4.X.1 is called with the width nbWidthX of the adjacent coding block in luma samples and the height nbHeightX of the adjacent coding block in luma samples as inputs, and its output is assigned to the variable sizeIdX; - If sizeId is equal to sizeIdX, candLwipModeX is set equal to IntraPredModeY[xNbX][yNbX]; - Otherwise, candLwipModeX is set equal to -1; - Otherwise, candLwipModeX is derived using IntraPredModeY[xNbX][yNbX] and sizeId as defined in Table 8 - X1; 3. candLwipModeList[x] for x = 0..2 is derived as follows using lwipMpmCand[sizeId] as defined in Table 8 - X2: - If both candLwipModeA and candLwipModeB are equal to -1, the following applies: candLwipModeList[0]=lwipMpmCand[sizeId][0] (8 - X1) candLwipModeList[1]=lwipMpmCand[sizeId][1] (8 - X2) candLwipModeList[2] = lwipMpmCand[sizeId][2] (8 - X3) - Otherwise, the following applies: - If candLwipModeA is equal to candLwipModeB, or if either candLwipModeA or candLwipModeB is equal to -1, the following applies: candLwipModeList[0] = (candLwipModeA!=-1)? candLwipModeA : candLwipModeB (8 - X4) - If candLwipModeList[0] is equal to lwipMpmCand[sizeId][0], the following applies: candLwipModeList[1] = lwipMpmCand[sizeId][1] (8 - X5) candLwipModeList[2] = lwipMpmCand[sizeId][2] (8 - X6) - Otherwise, the following applies: candLwipModeList[1] = lwipMpmCand[sizeId][0] (8 - X7) candLwipModeList[2] = (candLwipModeList[0]!=lwipMpmCand[sizeId][1])? lwipMpmCand[sizeId][1] : lwipMpmCand[sizeId][2] (8 - X8) - Otherwise, the following applies: candLwipModeList[0] = candLwipModeA (8 - X9) candLwipModeList[1] = candLwipModeB (8 - X10) - If neither candLwipModeA nor candLwipModeB is equal to lwipMpmCand[sizeId][0], the following applies: candLwipModeList[2] = lwipMpmCand[sizeId][0] (8-X11) - Otherwise, the following applies: - If neither candLwipModeA nor candLwipModeB is equal to lwipMpmCand[sizeId][1], the following applies: candLwipModeList[2] = lwipMpmCand[sizeId][1] (8-X12) - Otherwise, the following applies: candLwipModeList[2] = lwipMpmCand[sizeId][2] (8-X13) 4. IntraPredModeY[xCb][yCb] is derived by applying the following procedure: - If intra_lwip_mpm_flag[xCb][yCb] is equal to 1, IntraPredModeY[xCb][yCb] is set equal to candLwipModeList[intra_lwip_mpm_idx[xCb][yCb]; - Otherwise, IntraPredModeY[xCb][yCb] is derived by applying the following ordered steps: 1. For i = 0..1 and for each i, j=(i + 1)..2, if candLwipModeList[i] is greater than candLwipModeList[j], the two values are exchanged as follows: (candLwipModeList[i], candLwipModeList[j]) = Swap(candLwipModeList[i], candLwipModeList[j]) (8-X14) 2. IntraPredModeY[xCb][yCb] is derived by the following ordered steps: i. IntraPredModeY[xCb][yCb] is set equal to intra_lwip_mpm_remainder[xCb][yCb]; ii. When i, including both ends, is equal to 0 to 2 and IntraPredModeY[xCb][yCb] is greater than or equal to candLwipModeList[i], the value of IntraPredModeY[xCb][yCb] is incremented by only 1. The variable IntraPredModeY[x][y] for x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1 is set to be equal to IntraPredModeY[xCb][yCb]. 8.4.X.1 Derivation Process of Prediction Block Size Type The input to this process is as follows: - A variable cbWidth that defines the width of the current coding block in luma samples, - A variable cbHeight that defines the height of the current coding block in luma samples. The output of this process is the variable sizeId. The variable sizeId is derived as follows: - If both cbWidth and cbHeight are equal to 4, sizeId is set equal to 0; - Otherwise, if both cbWidth and cbHeight are less than or equal to 8, sizeId is set equal to 1; - Otherwise, sizeId is set equal to 2. [Table 2] [Table 3] <end> 8.4.2 Derivation Process of Luma Intra Prediction Mode The inputs to this process are as follows: - The luma position (xCb, yCb) that defines the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - The variable cbWidth that defines the width of the current coding block in the luma sample, - The variable cbHeight that defines the height of the current coding block in the luma sample. In this process, the luma intra prediction mode IntraPredModeY[xCb][yCb] is derived. Table 8-1 defines the values and related names for the intra prediction mode IntraPredModeY[xCb][yCb]. [Table 4] IntraPredModeY[xCb][yCb] is derived by the following ordered steps: 1. The adjacent positions (xNbA, yNbA) and (xNbB, yNbB) are set equal to (xCb - 1, yCb + cbHeight - 1) and (xCb + cbWidth - 1, yCb - 1), respectively; 2. For X replaced by either A or B, the variable candIntraPredModeX is derived as follows: - <begin>6.4.X Section Block Availability Derivation Process [Ed.(BB): Neighboring Block Availability Check Process tbd] <end>is called with the position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring position (xNbY, yNbY) set equal to (xNbX, yNbX), and its output is assigned to availableX; - The candidate intra prediction mode candIntraPredModeX is derived as follows: - If one or more of the following conditions are true, candIntraPredModeX is set equal to INTRA_PLANAR; - The variable availableX is equal to FALSE; - CuPredMode[xNbX][yNbX] is not equal to MODE_INTRA and clip_flag[xNbX][yNbX] is not equal to 1; - pcm_flag[xNbX][yNbX] is equal to 1; - X is equal to B and yCb - 1 is less than ((yCb >> CtbLog2SizeY) << CtbLog2SizeY); - Otherwise, candIntraPredModeX is derived as follows: - If intra_lwip_flag[xCb][yCb] is equal to 1, candIntraPredModeX is derived by the following ordered steps: i. The block size type derivation process specified in section 8.4.X.1 is called with the width cbWidth of the current coding block in luma samples and the height cbHeight of the current coding block in luma samples as inputs, and its output is assigned to the variable sizeId; ii. candIntraPredModeX is derived using IntraPredModeY[xNbX][yNbX] and sizeId as specified in Table 8 - X3; - Otherwise, candIntraPredModeX is set equal to IntraPredModeY[xNbX][yNbX]; 3. The variables ispDefaultMode1 and ispDefaultMode2 are defined as follows: - When IntraSubPartitionsSplitType is equal to ISP_HOR_SPLIT, ispDefaultMode1 is set equal to INTRA_ANGULAR18 and ispDefaultMode2 is set equal to INTRA_ANGULAR5; - Otherwise, ispDefaultMode1 is set equal to INTRA_ANGULAR50 and ispDefaultMode2 is set equal to INTRA_ANGULAR63. …
Table 5
Table 6
Number
Table 7
Number
Number
Table 8
Table 9
Table 10
[0083] Overview of ALWIP To predict a sample of a rectangular block of width W and height H, Affine Linear Weighted Intra Prediction (ALWIP) takes as input H reconstructed adjacent boundary samples of one line to the left of the block and W reconstructed adjacent boundary samples of one line above the block. If these reconstructed samples are not available, they are generated as is done in conventional intra prediction. ALWIP is only applied to luma intra blocks. For chroma intra blocks, conventional intra coding modes are applied.
[0084] The generation of the prediction signal is based on the following three steps.
[0085] 1. From the boundary samples, 4 samples in the case of W = H = 4 and 8 samples in all other cases are extracted by averaging.
[0086] 2. Using these averaged samples as input, matrix-vector multiplication and subsequent addition of an offset are performed. The result is a reduced prediction signal for a subsampled set of samples within the original block.
[0087] 3. From the prediction signal for the subsampled set, the prediction signals for the remaining positions are generated by linear interpolation, which is linear interpolation with a single step in each direction.
[0088] When the ALWIP mode is applied, the index predmode of the ALWIP mode is signaled using an MPM list with three MPMS. Here, the derivation of the MPM is done using the intra modes of the upper and left PUs as follows. Three fixed tables map_angular_to_alwip Angular that assign the ALWIP mode to each of the conventional intra prediction modes predmode idx , idx ∈ {0, 1, 2} exist: predmode ALWIP = map_angular_to_alwip idx [predmode Angular
[0089] Index indicating from which of the three sets the ALWIP parameters are taken for each PU of width W and height H: idx(PU) = idx(W,H) ∈ {0,1,2} is defined.
[0090] Upper prediction unit PU above is available, belongs to the same CTU as the current PU, and is in the intra mode, and idx(PU) = idx(PU above ) and ALWIP is in ALWIP mode predmode ALWIP above is applied to PU above then,
Number
[0091] When the upper PU is available, belongs to the same CTU as the current PU, and is in the intra mode, and the conventional intra prediction mode predmode Angular above is applied to the upper PU,
Number
[0092] In all other cases,
Number
[0093] Finally, three fixed default lists, each containing three different ALWIP modes, are provided. The default lists idx and idx ∈ {0, 1, 2} are provided. The default lists idx(PU) along with the mode ALWIP above and mode ALWIP left are used to construct three different MPMs by replacing the default value of -1 and eliminating duplicates.
[0094] In the luma MPM list derivation, whenever an adjacent luma block using the ALWIP mode predmode ALWIP is encountered, this block is treated as if it were using the conventional intra prediction mode predmode Angular .
Number
[0095] 3. Transform in VVC 3.1 MTS (Multiple Transform Selection) In addition to the DCT-II adopted in HEVC, the multiple transform selection (MTS) scheme is also used to residual code both inter and intra coding blocks. This uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII.
[0096] 3.2 RST (Reduced Secondary Transform) proposed in JVET-N0193 For 4×4 and 8×8 blocks, RST applies 16×16 and 16×64 non-separable transforms respectively. The first forward and inverse transforms are still performed in the same way as the two 1D horizontal / vertical transform paths. The second forward and inverse transforms are separate process steps from those of the first transform. In the encoder, first the first forward transform is performed, followed by the second forward transform, quantization, and CABAC bit coding. In the decoder, first the CABAC bit decoding and inverse quantization are performed, followed by the second inverse transform, and then the first inverse transform. RST is only applied to intra-coding TUs in both intra-slice and inter-slice.
[0097] 3.3 Unified MPM List for Intra-mode Coding in JVET-N0185 A unified (unified) 6-MPM list is proposed for intra-blocks regardless of whether the Multiple Reference Line (MRL) and Intra-Sub-Partition (ISP) coding tools are applied. This MPM list is constructed based on the intra-modes of the left and upper adjacent blocks, as in VTM4.0. Denoting the left mode as Left and the mode of the upper block as Above, the unified MPM list is as follows: · When the adjacent blocks are not available, its intra-mode is set to Planar by default. · When both modes Left and Above are non-angular modes: MPM list → {Planar, DC, V, H, V-4, V+4} · When one of the modes Left and Above is an angular mode and the other is a non-angular mode: a. Set mode Max to the larger of the modes Left and Above. b. MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2} · When both Left and Above are angular modes and are different: a. Set mode Max to the larger of the modes Left and Above; b. When the difference between the Left and Above modes is within the range of 2 to 62, inclusive of both ends: i. MPM list → {Planar, Left, Above, DC, Max - 1, Max + 1} c. Otherwise: i. MPM list → {Planar, Left, Above, DC, Max - 2, Max + 2} · When both Left and Above are in the angular mode and they are the same: a. MPM list → {Planar, Left, Left - 1, Left + 1, DC, Left - 2}
[0098] In addition, the first bin of the MPM index codeword is CABAC context - coded. A total of 3 contexts are used depending on whether the current intra - block is MRL - enabled, ISP - enabled, or a normal intra - block.
[0099] The left - adjacent block and the upper - adjacent block used for constructing the unified MPM list are A2 and B2 shown in Figure 10.
[0100] First, one MPM flag is coded. If the block is coded in one of the modes in the MPM list, an additional MPM index is coded. Otherwise, the index to the remaining modes (excluding MPM) is coded.
[0101] 4 Examples of drawbacks in existing implementations The design of ALWIP in JVET - N0217 has the following problems: 1) At the JVET meeting in March 2019, unified 6 - MPM list generation was adopted for the MRL mode, ISP mode, and normal intra - mode. However, the affine linear weighted prediction mode uses a different 3 - MPM list construction, which complicates the MPM list construction. Complicated MPM list construction may reduce the decoder throughput, especially for small blocks such as 4×4 samples; 2) ALWIP is only applicable to the luma components of blocks. For the chroma components of the ALWIP coding block, the chroma mode index is coded and sent to the decoder, which may result in unnecessary signaling; 3) The interaction between ALWIP and other coding tools should be considered; 4) In the following formula:
Number
[0102] 5 Exemplary method for matrix-based intra-coding Embodiments of the technology disclosed herein address the drawbacks of existing implementations, thereby providing video coding with higher coding efficiency and lower computational complexity. The matrix-based intra prediction method for video coding can enhance both existing and future video coding standards, as described in this document, and will become apparent in the following examples described for various implementations. The examples of the disclosed technology provided below are for illustrative purposes and are not meant to be construed as limiting. In one example, unless explicitly stated otherwise, the various features described in these examples can be combined.
[0103] In the following description, the intra prediction mode refers to the angular intra prediction mode (including DC, planar, CCLM, and other possible intra prediction modes), and the intra mode refers to the normal intra mode, or MRL, or ISP, or ALWIP.
[0104] In the following description, "other intra modes" may refer to one or more intra modes other than ALWIP, such as, for example, the normal intra mode, or MRL, or ISP.
[0105] In the following description, SatShift(x,n) is
Equation
[0106] In one example, offset0 and / or offset1 are set to (1<<n)>>1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.
[0107] In another example, offset0=offset1=((1<<n)>>1)-1 or ((1<<(n-1)))-1.
[0108] Clip3(min, max, x) is defined as
Number
[0109] MPM List Construction for ALWIP 1. It is proposed that all or part of the MPM list for ALWIP can be constructed according to all or part of the procedure for constructing the MPM list for non - ALWIP intramodes (such as normal intramode, MRL, or ISP, etc.); a. In one example, the size of the MPM list for ALWIP can be the same as the size of the MPM list for non - ALWIP intramode; i. For example, in both ALWIP and non - ALWIP intramodes, the size of the MPM list is 6; b. In one example, the MPM list for ALWIP can be derived from the MPM list for non - ALWIP intramode; i. In one example, first, the MPM list for non - ALWIP intramode can be constructed. Then, some or all of them can be converted into MPMs that can be further added to the MPM list for ALWIP coding blocks; 1) Alternatively, pruning may be applied when adding the converted MPM to the MPM list for ALWIP coding blocks; 2) The default mode may be added to the MPM list for ALWIP coding blocks; a. In one example, the default mode can be added before the converted ones from the MPM list of non - ALWIP intramode; b. Instead, the default mode can be added after the converted ones from the MPM list of non - ALWIP intramode; c. Instead, the default mode can be added to interleave with the converted ones from the MPM list of non - ALWIP intramode; d. In one example, the default mode can be fixed to be the same for all types of blocks; e. Alternatively, the default mode may be determined according to the coded information such as, for example, the availability of adjacent blocks, the mode information of adjacent blocks, the block size, etc.; ii. In one example, one intra prediction mode in the MPM list for the non - ALWIP intra mode can be converted to the corresponding ALWIP intra prediction mode when it is put into the MPM list for ALWIP; 1) Or, all intra prediction modes in the MPM list for the non - ALWIP intra mode may be converted to the corresponding ALWIP intra prediction modes before being used to construct the MPM list for ALWIP; 2) Or, when the MPM list for the non - ALWIP intra mode can be further used to derive the MPM list for ALWIP, all candidate intra prediction modes (which may include intra prediction modes from adjacent blocks and default intra prediction modes such as planar and DC) may be converted to the corresponding ALWIP intra prediction modes before being used to construct the MPM list for the non - ALWIP intra mode; 3) In one example, two converted ALWIP intra prediction modes can be compared; a. In one example, if they are the same, only one of them can be put into the MPM list for ALWIP; b. In one example, if they are the same, only one of them can be put into the MPM list for the non - ALWIP intra mode; iii. In one example, K out of S intra prediction modes in the MPM list for the non - ALWIP intra mode can be selected as the MPM list for the ALWIP mode. For example, K is equal to 3 and S is equal to 6; 1) In one example, the first K intra prediction modes in the MPM list for the non - ALWIP intra mode may be selected as the MPM list for the ALWIP mode; It is proposed that one or more adjacent blocks used to derive the MPM list for ALWIP can also be used to derive the MPM list for non-ALWIP intra modes (such as normal intra mode, MRL, or ISP, etc.); a. In one example, the adjacent block to the left of the current block used to derive the MPM list for ALWIP should be the same as that used to derive the MPM list for non-ALWIP intra mode; i. Assuming that the upper left corner of the current block is (xCb, yCb) and the width and height of the current block are W and H, in one example, the left adjacent block used to derive the MPM list for both ALWIP and non-ALWIP intra modes can cover the position (xCb - 1, yCb). In an alternative example, the left adjacent block used to derive the MPM list for both ALWIP and non-ALWIP intra modes can cover the position (xCb - 1, yCb + H - 1); ii. For example, the left adjacent block and the upper adjacent block used in constructing the unified MPM list are A2 and B2 shown in FIG. 10; b. In one example, the adjacent block above the current block used to derive the MPM list for ALWIP should be the same as that used to derive the MPM list for non-ALWIP intra mode; i. Assuming that the upper left corner of the current block is (xCb, yCb) and the width and height of the current block are W and H, in one example, the upper adjacent block used to derive the MPM list for both ALWIP and non-ALWIP intra modes can cover the position (xCb, yCb - 1). In an alternative example, the upper adjacent block used to derive the MPM list for both ALWIP and non-ALWIP intra modes can cover the position (xCb + W - 1, yCb - 1); ii. For example, the left adjacent block and the upper adjacent block used in constructing the unified MPM list are A1 and B1 shown in FIG. 10; 3. It is proposed that the MPM list for ALWIP can be constructed in different ways according to the current width and / or height of the block; a. In one example, different adjacent blocks can be accessed for different block dimensions; 4. It is proposed that the MPM list for ALWIP and the MPM list for non - ALWIP intra - mode can be constructed using the same procedure but different parameters; a. In one example, K out of S intra - prediction modes in the MPM list construction procedure for non - ALWIP intra - mode can be derived for the MPM list used in the ALWIP mode. For example, K is equal to 3 and S is equal to 6; i. In one example, the first K intra - prediction modes in the MPM list construction procedure can be derived for the MPM list used in the ALWIP mode; b. In one example, the first mode in the MPM list can be different; i. For example, the first mode in the MPM list for non - ALWIP intra - mode can be Planar, but it can be Mode X0 in the MPM list for ALWIP; 1) In one example, X0 can be an ALWIP intra - prediction mode converted from Planar; c. In one example, the stuffing modes in the MPM list can be different; i. For example, the first three stuffing modes in the MPM list for non - ALWIP intra - mode can be DC, Vertical, and Horizontal, but they can be Mode X1, X2, X3 in the MPM list for ALWIP; 1) In one example, X1, X2, X3 can be different for different sizeIds; ii. In one example, the number of stuffing modes can be different; d. In one example, the adjacent modes in the MPM list can be different; i. For example, the normal intra prediction modes of adjacent blocks are used to construct the MPM list for non-ALWIP intra modes. And they are converted to ALWIP intra prediction modes for constructing the MPM list for the ALWIP mode; e. In one example, the shifted modes in the MPM list may be different; i. For example, assuming X is a normal intra prediction mode and K0 is an integer, X + K0 can be put into the MPM list for non-ALWIP intra modes. And assuming Y is an ALWIP intra prediction mode and K1 is an integer, Y + K1 can be put into the MPM list for ALWIP, and K0 can be different from K1; 1) In one example, K1 may depend on the width and height; 5. When constructing the MPM list for the current block in non-ALWIP intra mode, it is proposed that adjacent blocks be treated as unavailable if they are coded in ALWIP; a. Alternatively, when constructing the MPM list for the current block in non-ALWIP intra mode, adjacent blocks are treated as if they are coded in a predetermined intra prediction mode (such as Planar etc.) if they are coded in ALWIP; 6. When constructing the MPM list for the current block in ALWIP intra mode, it is proposed that adjacent blocks be treated as unavailable if they are coded in non-ALWIP intra mode; a. Alternatively, when constructing the MPM list for the current block in ALWIP intra mode, adjacent blocks are treated as if they are coded in a predetermined ALWIP intra mode X if they are coded in non-ALWIP intra mode; i. In one example, X may depend on block dimensions such as width and / or height; 7. It is proposed to remove the storage of the ALWIP flag from the line buffer; a. In one example, when the second block to be accessed is located in a different LCU / CTU row / region from the current block, the condition check as to whether the second block is coded in ALWIP is skipped; b. In one example, when the second block to be accessed is located in a different LCU / CTU row / region from the current block, the second block is treated in the same way as the non-ALWIP mode, for example, being treated as a normal intra-coding block; 8. When encoding the ALWIP flag, K (K >= 0) or fewer contexts may be used; a. In one example, K = 1; 9. Instead of directly storing the mode index related to the ALWIP mode, it is proposed to store the transformed intra prediction mode of the ALWIP coding block; a. In one example, the decoded mode index related to one ALWIP coding block is mapped to the normal intra mode, for example, according to map_alwip_to_angular as described in Section 2.5.7; b. Alternatively, further, the storage of the ALWIP flag is completely removed; c. Alternatively, further, the storage of the ALWIP mode is completely removed; d. Alternatively, further, the condition check as to whether one adjacent / current block is coded with the ALWIP flag may be skipped; e. Alternatively, further, the conversion of the normal intra prediction related to the mode assigned to the ALWIP coding block and one block to be accessed may be skipped; ALWIP for Different Color Components 10. It is proposed that when the corresponding luma block is coded in the ALWIP mode, the estimated chroma intra mode (e.g., DM mode) can always be applied; a. In one example, when the corresponding luma block is coded in ALWIP mode, without signal transmission, the chroma intra mode is presumed to be the DM mode; b. In one example, the corresponding luma block can cover the corresponding samples of the chroma samples at a given position (e.g., the upper left of the current chroma block, the center of the current chroma block); c. In one example, the DM mode can be derived according to the intra prediction mode of the corresponding luma block, such as by mapping the (ALWIP) mode to one of the normal intra modes; 11. When the corresponding luma block of the chroma block is coded in ALWIP mode, several DM modes can be derived; 12. It is proposed that a special mode be assigned to the chroma block when one corresponding luma block is coded in ALWIP mode; a. In one example, the special mode is defined to be a given normal intra prediction mode regardless of the intra prediction mode associated with the ALWIP coding block; b. In one example, multiple different ways of intra prediction can be assigned to this special mode; 13. It is proposed that ALWIP can also be applied to the chroma component; a. In one example, the matrix and / or bias vector can be different for different color components; b. In one example, the matrix and / or bias vector can be jointly defined for Cb and Cr; i. In one example, the Cb and Cr components can be concatenated; ii. In one example, the Cb and Cr components can be interleaved; c. In one example, the chroma component can share the same ALWIP intra prediction mode as the corresponding luma block; i. In one example, when the corresponding luma block applies the ALWIP mode and the chroma block is coded in the DM mode, the same ALWIP intra prediction mode is applied to the chroma component; ii. In one example, the same ALWIP intra prediction mode is applied to the chroma component and subsequent linear interpolation can be skipped; iii. In one example, the same ALWIP intra prediction mode is applied to the chroma component using a subsampled matrix and / or a bias vector; d. In one example, the number of ALWIP intra prediction modes may be different for different components; i. For example, the number of ALWIP intra prediction modes for the chroma component can be less than that for the luma component with the same block width and height; Applicability of ALWIP 14. It is proposed that it can be signaled whether ALWIP can be applied; a. For example, it can be signaled at the sequence level (e.g., within the SPS), at the picture level (e.g., within the PPS or picture header), at the slice level (e.g., within the slice header), at the tile group level (e.g., within the tile group header), at the tile level, at the CTU row level, or at the CTU level; b. For example, if ALWIP cannot be applied, the intra_lwip_flag can be presumed to be 0 without being signaled; 15. It is proposed that whether ALWIP can be applied may depend on the block width (W) and / or height (H); c. For example, when W>=T1 (or W>T1) and H>=T2 (or H>T2) (e.g., T1=T2=32), it can be considered that ALWIP cannot be applied; i. For example, when W<=T1 (or W<T1) and H<=T2 (or H<T2) (e.g., T1=T2=32), it can be considered that ALWIP cannot be applied; d. For example, when W >= T1 (or W > T1) or H >= T2 (or H > T2) (for example, T1 = T2 = 32), it may be considered that ALWIP cannot be applied; i. For example, when W <= T1 (or W < T1) or H <= T2 (or H < T2) (for example, T1 = T2 = 32), it may be considered that ALWIP cannot be applied; e. For example, when W + H >= T (or W + H > T) (for example, T = 256), it may be considered that ALWIP cannot be applied; i. For example, when W + H <= T (or W + H < T) (for example, T = 256), it may be considered that ALWIP cannot be applied; f. For example, when W * H >= T (or W * H > T) (for example, T = 256), it may be considered that ALWIP cannot be applied; i. For example, when W * H <= T (or W * H < T) (for example, T = 256), it may be considered that ALWIP cannot be applied; g. For example, when ALWIP cannot be applied, intra_lwip_flag may be presumed to be 0 without being signaled; Calculation Problems in ALWIP 16. It is proposed that the shift operation involved in ALWIP can only shift the number left or right by S only, provided that S must be 0 or more; a. In one example, the right shift operation may be different when S is 0 or greater than 0; i. In one example, upsBdryX[x] is when uDwn > 1,
Number
Number
Number
Table 11
Equation
Equation
Number
Number
Number
Equation
Equation
Table 16
Equation
Table 17
Table 18
Equation
Table 19
[0110] The above examples can be incorporated in the context of the methods described below, such as methods 1100 - 1400 and 2300 - 2400, which can be implemented, for example, in a video encoder and / or decoder.
[0111] FIG. 11 shows a flowchart of an exemplary method for video processing. Method 1100 includes, at step 1110, determining that a current video block is to be coded using an affine linear weighted intra prediction (ALWIP) mode.
[0112] Method 1100 includes, at step 1120, constructing at least a part of an MPM list for the ALWIP mode based on at least a part of an MPM list for a non-ALWIP intra mode based on the determination.
[0113] Method 1100 includes, at step 1130, performing a conversion between the current video block and a bitstream representation of the current video block based on the MPM list for the ALWIP mode.
[0114] In some embodiments, the size of the MPM list for the ALWIP mode is the same as the size of the MPM list for a non-ALWIP intra mode. In one example, the size of the MPM list for the ALWIP mode is 6.
[0115] In some embodiments, method 1100 further has a step of inserting a default mode into the MPM list for the ALWIP mode. In one example, the default mode is inserted before a part of the MPM list for the ALWIP mode based on the MPM list for a non-ALWIP intra mode. In another example, the default mode is inserted following a part of the MPM list for the ALWIP mode based on the MPM list for a non-ALWIP intra mode. In yet another example, the default mode is inserted interleaved with a part of the MPM list for the ALWIP mode based on the MPM list for a non-ALWIP intra mode.
[0116] In some embodiments, constructing the MPM list for the ALWIP mode and the MPM list for the non-ALWIP intra mode is based on one or more adjacent blocks.
[0117] In some embodiments, constructing the MPM list for the ALWIP mode and the MPM list for the non-ALWIP intra mode is based on the current height or width of the video block.
[0118] In some embodiments, constructing the MPM list for the ALWIP mode is based on a first set of parameters that is different from a second set of parameters used to construct the MPM list for the non-ALWIP intra mode.
[0119] In some embodiments, method 1100 further includes determining that an adjacent block of the current video block is encoded in the ALWIP mode and specifying the adjacent block as not available when constructing the MPM list for the non-ALWIP intra mode.
[0120] In some embodiments, method 1100 further includes determining that an adjacent block of the current video block is encoded in a non-ALWIP mode and specifying the adjacent block as not available when constructing the MPM list for the ALWIP mode.
[0121] In some embodiments, the non-ALWIP intra mode is based on a normal intra mode, a multiple reference line (MRL) intra prediction mode, or an intra sub-partition (ISP) tool.
[0122] FIG. 12 shows a flowchart of an exemplary method for video processing. Method 1200 includes, at step 1210, determining that the luma component of the current video block is to be coded using an affine linear weighted intra prediction (ALWIP) mode.
[0123] Method 1200 includes, at step 1220, estimating the chroma intra mode based on the above determination.
[0124] Method 1200 includes, at step 1230, performing a conversion between a current video block and a bitstream representation of the current video block based on a chroma intra mode.
[0125] In some embodiments, a luma component covers a predetermined chroma sample of a chroma component. In one example, the predetermined chroma sample is the top-left sample or the center sample of the chroma component.
[0126] In some embodiments, the estimated chroma intra mode is the DM mode.
[0127] In some embodiments, the estimated chroma intra mode is the ALWIP mode.
[0128] In some embodiments, the ALWIP mode is applied to one or more chroma components of the current video block.
[0129] In some embodiments, different matrices or bias vectors of the ALWIP mode are applied to different color components of the current video block. In one example, the different matrices or bias vectors are jointly defined for the Cb and Cr components. In another example, the Cb component and the Cr component are concatenated. In yet another example, the Cb component and the Cr component are interleaved.
[0130] FIG. 13 shows a flowchart of an exemplary method for video processing. Method 1300 includes, at step 1310, determining that a current video block is to be coded using an affine linear weighted intra prediction (ALWIP) mode.
[0131] Method 1300 includes, at step 1320, performing a conversion between the current video block and a bitstream representation of the current video block based on the determination.
[0132] In some embodiments, the determination is based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a tile header, a coding tree unit (CTU) row, or a CTU region.
[0133] In some embodiments, the determination is based on the height (H) or width (W) of the current video block. In one example, W > T1 or H > T2. In another example, W ≥ T1 or H ≥ T2. In yet another example, W < T1 or H < T2. In yet another example, W ≤ T1 or H ≤ T2. In yet another example, T1 = 32 and T2 = 32.
[0134] In some embodiments, the determination is based on the height (H) or width (W) of the current video block. In one example, W + H ≤ T. In another example, W + H ≥ T. In yet another example, W × H ≤ T. In yet another example, W × H ≥ T. In yet another example, T = 256.
[0135] FIG. 14 shows a flowchart of an exemplary method for video processing. Method 1400 includes, at step 1410, determining that the current video block is to be coded using a coding mode different from the affine linear weighted intra prediction (ALWIP) mode.
[0136] Method 1400 includes, at step 1420, performing a conversion between the current video block and a bitstream representation of the current video block based on the determination.
[0137] In some embodiments, the coding mode is a combined intra-inter prediction (CIIP) mode, and method 1400 further includes performing a selection between the ALWIP mode and the normal intra prediction mode. In one example, performing the selection is based on explicit signaling within the bitstream representation of the current video block. In another example, performing the selection is based on a predetermined rule. In yet another example, the predetermined rule always selects the ALWIP mode when the current video block is coded using the CIIP mode. In yet another example, the predetermined rule always selects the normal intra prediction mode when the current video block is coded using the CIIP mode.
[0138] In some embodiments, the coding mode is a cross-component linear model (CCLM) prediction mode. In one example, the downsampling procedure for the ALWIP mode is based on the downsampling procedure for the CCLM prediction mode. In another example, the downsampling procedure for the ALWIP mode is based on a first parameter set, and the downsampling procedure for the CCLM prediction mode is based on a second parameter set different from the first parameter set. In yet another example, the downsampling procedure for the ALWIP mode or the CCLM prediction mode has at least one of a selection of a downsampling position, a selection of a downsampling filter, a rounding operation, or a clipping operation.
[0139] In some embodiments, method 1400 further includes applying one or more of RST (Reduced Secondary Transform), a secondary transform, a rotation transform, or a non-separable secondary transform (NSST).
[0140] In some embodiments, method 1400 further includes applying block-based DPCM (differential pulse coded modulation) or residual DPCM.
[0141] In some embodiments, a video processing method includes determining, during conversion between a current video block and a bitstream representation of the current video block, a context of a flag indicating use of an affine linear weighted intra prediction (ALWIP) mode based on a rule for the current video block, predicting a plurality of sub-blocks of the current video block based on the ALWIP mode, and performing a conversion between the current video block and the bitstream representation of the current video block based on the prediction. The rule may be implicitly defined using prior art or may be signaled within a coding bitstream. Other examples and aspects of this method are further described in items 37 and 38 of Section 4.
[0142] In some embodiments, a method for video processing includes determining that a current video block is to be coded using an affine linear weighted intra prediction (ALWIP) mode, and performing at least two filtering stages on samples of the current video block in an upsampling process associated with the ALWIP mode during conversion between the current video block and a bitstream representation of the current video block, wherein a first accuracy of samples in a first filtering stage of the at least two filtering stages is different from a second accuracy of samples in a second filtering stage of the at least two filtering stages.
[0143] In one example, samples of the current video block are prediction samples, intermediate samples before an upsampling process, or intermediate samples after an upsampling process. In another example, samples are upsampled in a first dimension horizontally in a first filtering stage, and samples are upsampled in a second dimension vertically in a second filtering stage. In yet another example, samples are upsampled in a first dimension vertically in a first filtering stage, and samples are upsampled in a second dimension horizontally in a second filtering stage.
[0144] In one example, the output of the first filtering stage is right-shifted or divided to produce a processed output, which is the input to the second filtering stage. In another example, the output of the first filtering stage is left-shifted or multiplied to produce a processed output, which is the input to the second filtering stage. Further examples and aspects of this method are described in more detail in item 40 of section 4.
[0145] As further described in items 41 to 43 of section 4, the video processing method includes determining that the current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, and performing at least two filtering stages on the samples of the current video block in an upsampling process related to the ALWIP mode during the conversion between the current video block and the bitstream representation of the current video block, where the upsampling process is performed in a fixed order when both vertical upsampling and horizontal upsampling are performed. As further described in items 41 to 43 of section 4, another method includes determining that the current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, and performing at least two filtering stages on the samples of the current video block in an upsampling process related to the ALWIP mode during the conversion between the current video block and the bitstream representation of the current video block, where the conversion includes performing a transpose operation before the upsampling process.
[0146] Further features of the above method are described in items 41 to 43 of section 4.
[0147] 6 Implementation Examples of the Disclosed Technology FIG. 15 is a block diagram of a video processing apparatus 1500. The apparatus 1500 may be used to implement one or more of the methods described herein. The apparatus 1500 may be embodied in a smartphone, tablet, computer, or Internet of Things (IoT) receiver. The apparatus 1500 may include one or more processors 1502, one or more memories 1504, and video processing hardware 1506. The (one or more) processors 1502 may be configured to execute one or more of the methods described in this document (including, but not limited to, methods 11100-1400 and 2100-2300). The (one or more) memories 1504 may be used to store data and code used to execute the methods and techniques described herein. The video processing hardware 1506 may be used to implement some of the techniques described in this document in hardware circuitry.
[0148] In some embodiments, the video coding method may be implemented using an apparatus implemented on a hardware platform as described with respect to FIG. 15.
[0149] Some embodiments of the disclosed techniques include making a determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in processing blocks of video, but will not necessarily change the resulting bitstream based on the use of that tool or mode. That is, the conversion from a block of video to a bitstream representation of the video will use the video processing tool or mode when it is enabled based on the determination. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been changed based on that video processing tool or mode. That is, the conversion from a bitstream representation of the video to a block of video will be performed using the video processing tool or mode that was enabled based on the determination.
[0150] Some embodiments of the disclosed technology include making a determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use that tool or mode in converting a block of video to a video bitstream. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was disabled based on the determination.
[0151] FIG. 21 is a block diagram illustrating an example of a video coding system 100 that may utilize the technology of this disclosure. As shown in FIG. 21, video coding system 100 may include a source device 110 and a destination device 120. Source device 110 may generate encoded video data and may be referred to as a video encoding device. Destination device 120 may be able to decode the encoded video data generated by source device 110 and may be referred to as a video decoding device. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0152] The video source 112 may include sources such as, for example, a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may have one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a series of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 120 through the network 130a via the I / O interface 116. The encoded video data may also be stored on the storage medium / server 130b for access by the destination device 120.
[0153] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.
[0154] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.
[0155] Video encoder 114 and video decoder 124 may operate according to video compression standards such as, for example, the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0156] FIG. 22 is a block diagram illustrating an example of a video encoder 200 that may be the video encoder 114 within the system 100 shown in FIG. 21.
[0157] Video encoder 200 may be configured to execute any or all of the techniques of this disclosure. In the example of FIG. 22, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in this disclosure.
[0158] The functional components of video encoder 200 may include a partitioning unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0159] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode where at least one reference picture is the picture where the current video block is located.
[0160] Also, some components, such as the motion estimation unit 204 and the motion compensation unit 205, are shown separately for illustrative purposes in the example of FIG. 22, but can be highly integrated.
[0161] The splitting unit 201 can split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0162] The mode selection unit 203 selects, for example, based on an error result, one of a plurality of coding modes that are intra or inter, and provides the obtained intra or inter coding block to a residual generation unit 207 that generates residual block data and a reconstruction unit 212 that reconstructs the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter predication (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select the resolution of the motion vector for a block (e.g., sub-pixel or integer pixel accuracy) in the case of inter prediction.
[0163] To perform inter prediction on the current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information of the pictures from the buffer 213 and the decoded samples other than the picture associated with the current video block.
[0164] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0165] In some examples, the motion estimation unit 204 can perform one - direction prediction on the current video block, and the motion estimation unit 204 can search for reference pictures in list 0 or list 1 for the reference video block corresponding to the current video block. Then, the motion estimation unit 204 can generate a reference index indicating the reference picture in list 0 or list 1 including the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. Based on the reference video block indicated by the motion information of the current video block, the motion compensation unit 205 can generate a predicted video block of the current block.
[0166] In other examples, the motion estimation unit 204 can perform two - direction prediction on the current video block, and the motion estimation unit 204 can search for reference pictures in list 0 for the reference video block corresponding to the current video block and can also search for reference pictures in list 1 for another reference video block corresponding to the current video block. Then, the motion estimation unit 204 can generate a reference index indicating the reference pictures in list 0 and list 1 including the reference video blocks, and a motion vector indicating the spatial displacement between those reference video blocks and the current video block. The motion estimation unit 204 can output those reference indexes and the motion vector of the current video block as the motion information of the current video block. Based on the reference video block indicated by the motion information of the current video block, the motion compensation unit 205 can generate a predicted video block of the current block.
[0167] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding process.
[0168] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0169] In one example, motion estimation unit 204 may indicate a value that indicates to video decoder 300 that the current video block has the same motion information as another video block within the syntax structure associated with the current video block.
[0170] In another example, motion estimation unit 204 may identify, within the syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may determine the motion vector of the current video block using the motion vector of the indicated video block and the motion vector difference.
[0171] As described above, video encoder 200 may signal motion vectors predictively. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0172] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks within the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0173] The residual generation unit 207 may generate residual data for a current video block by subtracting (e.g., indicated by a negative sign) one or more predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples within the current video block.
[0174] In other examples, for example, in skip mode, there may be no residual data for the current video block for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0175] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0176] After the transform processing unit 208 generates a transform coefficient video block for the current video block, the quantization unit 209 may quantize the transform coefficient video block for the current video block based on one or more quantization parameter (QP) values for the current video block.
[0177] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block for the current block stored in the buffer 213.
[0178] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts within the video block.
[0179] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0180] FIG. 19 is a block diagram illustrating an example of a video decoder 300 that may be a video decoder 124 within the system 100 shown in FIG. 17.
[0181] The video decoder 300 may be configured to execute any or all of the techniques of this disclosure. In the example of FIG. 19, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to execute any or all of the techniques described in this disclosure.
[0182] In the example of FIG. 19, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. The video decoder 300 may, in some examples, perform a decoding path that is generally inverse to the encoding path described with respect to the video encoder 200 (FIG. 18).
[0183] The entropy decoding unit 301 can extract the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., an encoded block of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-decoded video data, the motion compensation unit 302 can determine motion information including a motion vector, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and merge mode.
[0184] The motion compensation unit 302 may perform interpolation based on an interpolation filter as appropriate to generate a motion-compensated block. An identifier regarding the interpolation filter used with sub-pixel precision may be included in the syntax element.
[0185] The motion compensation unit 302 can calculate an interpolation value for sub-integer pixels of a reference block using the interpolation filter used by the video encoder 20 during encoding of the video block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information, and generate a prediction block using the interpolation filter.
[0186] The motion compensation unit 302 can use a part of the syntax information to determine the size of the block used to encode a frame and / or slice of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is divided, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.
[0187] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using an intra prediction mode, for example received within a bitstream. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided within the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0188] The reconstruction unit 306 can form a decoded block by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If desired, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. Then, the decoded video block is stored in the buffer 307, which provides a reference block for subsequent motion compensation / intra prediction and also generates a decoded video for presentation on a display device.
[0189] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation or vice versa. For example, the bitstream representation or coding representation of a current video block can correspond to bits that are co-located or bits that span at different positions within the bitstream as defined by the syntax. For example, a video block can be encoded with respect to the transformed and coded error residual values and also using bits within a header or other fields within the bitstream. Further, during the transform, the decoder can analyze the bitstream with the recognition that some fields may or may not be present based on decisions as described in the above solutions. Similarly, the encoder can also determine whether a particular syntax field is included or not and generate a coding representation by including or excluding the syntax field from the coding representation accordingly.
[0190] FIG. 20 is a block diagram showing an example of a video processing system 2000 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of system 2000. System 2000 can include an input 2002 that receives video content. The video content may be received in a raw or uncompressed format, such as, for example, 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 2002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi® or a cellular interface.
[0191] System 2000 can include a coding component 2004 that can implement various coding or encoding methods described in this document. Coding component 2004 can reduce the average bit rate of the video from input 2002 to the output of coding component 2004 to generate a coded representation of the video. Coding techniques are therefore sometimes referred to as video compression techniques or video transcoding techniques. The output of coding component 2004 can be stored or connected as represented by component 2006 and transmitted via communication. The stored or communicated bitstream (or coded) representation of the video received at input 2002 can be used by component 2008 to generate pixel values or a displayable video that is sent to display interface 2010. The process of generating a video that a user can view from the bitstream representation is sometimes referred to as video decompression. Also, although certain video processing operations may be referred to as "coding" operations or tools, it is understood that coding tools or operations are used in an encoder and corresponding decoding tools or operations that reverse the coding result are performed in a decoder.
[0192] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI (registered trademark)), a DisplayPort, etc. Examples of a storage interface include SATA (serial advanced technology attachment), PCI, an IDE interface, etc. The technology described in this document may be embodied in various electronic devices such as, for example, a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0193] In some embodiments, the ALWIP mode or the MIP mode is used to calculate a prediction block of a current video block by performing a boundary downsampling operation (or an averaging operation) on previously coded samples of the video, followed by a matrix-vector multiplication operation, and then optionally (or optionally) performing an upsampling operation (or a linear interpolation operation). In some embodiments, the ALWIP mode or the MIP mode is used to calculate a prediction block of a current video block by performing a boundary downsampling operation (or an averaging operation) on previously coded samples of the video, followed by a matrix-vector multiplication operation. In some embodiments, the ALWIP mode or the MIP mode may also perform an upsampling operation (or a linear interpolation operation) after performing the matrix-vector multiplication operation.
[0194] FIG. 23 shows a flowchart of an exemplary method 2300 for video processing. Method 2300 includes step 302 of performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra prediction (MIP) mode, in which the predicted block of the current video block performs a boundary downsampling operation on previously coded samples of the video, followed by a matrix vector multiplication operation, followed by selectively performing an upsampling operation to be determined, and the conversion includes performing a boundary downsampling operation in a single stage in which the reduced boundary samples of the current video block are generated according to a rule based at least on the reference boundary samples of the current video block, and the conversion includes performing a matrix vector multiplication operation using the reduced boundary samples of the current video block.
[0195] In some embodiments of method 2300, the reference boundary samples are reconstructed adjacent samples of the current video block. In some embodiments of method 2300, the rule defines that the reduced boundary samples are generated from the reconstructed adjacent samples of the current video block. In some embodiments of method 2300, the reconstructed adjacent samples are decoded adjacent samples without a reference filtering process that are not filtered before the reconstructed adjacent samples are used for the current video block. In some embodiments of method 2300, the reconstructed adjacent samples are decoded adjacent samples with a reference filtering process that are filtered before the reconstructed adjacent samples are used for the current video block.
[0196] In some embodiments related to method 2300, the angular intra prediction samples are generated from reconstructed neighboring samples. In some embodiments related to method 2300, the non-angular intra prediction samples are generated from reconstructed neighboring samples. In some embodiments related to method 2300, the rule stipulates that the reduced boundary samples are generated from reconstructed neighboring samples located in the row above the current video block and / or reconstructed neighboring samples located in the column to the left of the current video block. In some embodiments related to method 2300, the first number (N) of reduced boundary samples is generated from the second number (M) of reconstructed neighboring samples, and the reduced boundary samples are generated using the third number (K) of consecutive reconstructed neighboring samples. In some embodiments related to method 2300, K = M / N. In some embodiments related to method 2300, K = (M + N / 2) / N. In some embodiments related to method 2300, the reduced boundary samples are generated based on the average of consecutive reconstructed neighboring samples in the third number (K) of consecutive reconstructed neighboring samples.
[0197] In some embodiments related to method 2300, the reduced boundary samples are generated based on the weighted average of consecutive reconstructed neighboring samples in the third number (K) of consecutive reconstructed neighboring samples. In some embodiments related to method 2300, the rule stipulates that the reduced boundary samples located to the left of the current video block are generated from reconstructed neighboring samples located in the adjacent column to the left of the current video block, and the reduced boundary samples located above the current video block are generated from reconstructed neighboring samples located in the adjacent row above the current video block. In some embodiments related to method 2300, the current video block is a 16×16 video block, and the four reduced boundary samples located to the left of the 16×16 video block are generated from reconstructed neighboring samples located in the adjacent column to the left of the 16×16 video block, and the four reduced boundary samples located above the 16×16 video block are generated from reconstructed neighboring samples located in the adjacent row above the 16×16 video block.
[0198] In some embodiments related to method 2300, the rule defines that the technique by which the reduced boundary samples are generated by the boundary downsampling operation is based on the dimensions of the current video block. In some embodiments related to method 2300, the rule defines that the technique by which the reduced boundary samples are generated by the boundary downsampling operation is based on the coding information of the current video block. In some embodiments related to method 2300, the coding information includes the width of the current video block, the height of the current video block, and an indication of the intra prediction mode or the transform mode associated with the current video block. In some embodiments related to method 2300, the rule defines that the reduced boundary samples are generated regardless of the size of the current video block. In some embodiments related to method 2300, the rule defines that the process by which the reduced boundary samples located to the left of the current video block are generated is different from the process by which the reduced boundary samples located above the current video block are generated. In some embodiments related to method 2300, the current video block has a width (M) and a height (N), the first predetermined number of the reduced boundary samples located above the current video block is M, N, or the minimum of M and N, the second predetermined number of the reduced boundary samples located to the left of the current video block is M, N, or the minimum of M and N, the first number of reduced boundary samples is generated based on adjacent samples or by copying adjacent samples located in the row above the current video block, and the second number of reduced boundary samples is generated by copying adjacent samples or based on adjacent samples located in the column to the left of the current video block.
[0199] FIG. 24 shows a flowchart of an exemplary method 2400 for video processing. Method 2400 includes step 2402 of performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra prediction (MIP) mode, in which the final prediction block of the current video block performs a boundary downsampling operation on previously coded samples of the video, followed by a matrix vector multiplication operation, followed by an upsampling operation to determine, and the conversion includes performing an upsampling operation determined by the final prediction block using a reduced prediction block of the current video block and using reconstructed neighboring samples of the current video block according to rules, and the reduced prediction block is obtained by performing a matrix vector multiplication operation on the reduced boundary samples of the current video block.
[0200] In some embodiments of method 2400, the rule specifies that the final prediction block is determined by using all of the reconstructed neighboring samples. In some embodiments of method 2400, the rule specifies that the final prediction block is determined by using a part of the reconstructed neighboring samples. In some embodiments of method 2400, the reconstructed neighboring samples are adjacent to the current video block. In some embodiments of method 2400, the reconstructed neighboring samples are not adjacent to the current video block. In some embodiments of method 2400, the reconstructed neighboring samples are the adjacent row above the current video block, and / or the reconstructed neighboring samples are the adjacent column to the left of the current video block.
[0201] In some embodiments related to method 2400, the rule stipulates that when determining the final prediction block, the reduced boundary samples are excluded from the upsampling operation. In some embodiments related to method 2400, the rule stipulates that the final prediction block is determined by using a set of reconstructed adjacent samples selected from the reconstructed adjacent samples. In some embodiments related to method 2400, the set of reconstructed adjacent samples selected from the reconstructed adjacent samples includes all selections of the reconstructed adjacent samples located to the left of the current video block. In some embodiments related to method 2400, the set of reconstructed adjacent samples selected from the reconstructed adjacent samples includes all selections of the reconstructed adjacent samples located above the current video block.
[0202] In some embodiments related to method 2400, the set of reconstructed adjacent samples selected from the reconstructed adjacent samples includes the selection of K out of M consecutive reconstructed adjacent samples located to the left of the current video block. In some embodiments related to method 2400, K is equal to 1 and M is equal to 2, 4, or 8. In some embodiments related to method 2400, K out of each M consecutive reconstructed adjacent samples includes the last K reconstructed adjacent samples out of each M consecutive ones. In some embodiments related to method 2400, K out of each M consecutive reconstructed adjacent samples includes the first K reconstructed adjacent samples out of each M consecutive ones. In some embodiments related to method 2400, the set of reconstructed adjacent samples selected from the reconstructed adjacent samples includes the selection of K out of M consecutive reconstructed adjacent samples located above the current video block. In some embodiments related to method 2400, K is equal to 1 and M is equal to 2, 4, or 8. In some embodiments related to method 2400, K out of each M consecutive reconstructed adjacent samples includes the last K reconstructed adjacent samples out of each M consecutive ones.
[0203] In some embodiments related to method 2400, K out of each M consecutive reconstructed adjacent samples include the first K reconstructed adjacent samples out of each M consecutive ones. In some embodiments related to method 2400, a set of reconstructed adjacent samples is selected from the reconstructed adjacent samples based on the width and / or height of the current video block. In some embodiments related to method 2400, in response to the width of the current video block being greater than or equal to the height of the current video block, the set of reconstructed adjacent samples is selected to include all of the reconstructed adjacent samples located to the left of the current video block, and / or the set of reconstructed adjacent samples is selected to include a certain number of reconstructed adjacent samples located above the current video block, and the certain number of reconstructed adjacent samples depends on the width of the current video block.
[0204] In some embodiments related to method 2400, the k-th selected reconstructed adjacent sample located above the current video block is located at a position described by (blkX+(k+1)*blkW / M-1,blkY-1), where (blkX,blkY) represents the upper left position of the current video block, M is the number of reconstructed adjacent samples, and k is greater than or equal to 0 and less than or equal to (M-1). In some embodiments related to method 2400, in response to the width being 8 or less, the number of reconstructed adjacent samples is equal to 4. In some embodiments related to method 2400, in response to the width being greater than 8, the number of reconstructed adjacent samples is equal to 8. In some embodiments related to method 2400, the set of reconstructed adjacent samples is selected to include all of the reconstructed adjacent samples located to the left of the current video block, and / or the set of reconstructed adjacent samples is selected to include a certain number of reconstructed adjacent samples located above the current video block, and the certain number of reconstructed adjacent samples depends on the width of the current video block.
[0205] In some embodiments related to method 2400, depending on the width being 8 or less, the number of reconstructed adjacent samples is equal to 4. In some embodiments related to method 2400, depending on the width being greater than 8, the number of reconstructed adjacent samples is equal to 8. In some embodiments related to method 2400, depending on the width of the current video block being less than the height of the current video block, the set of reconstructed adjacent samples is selected to include all of the reconstructed adjacent samples located above the current video block, and / or the set of reconstructed adjacent samples is selected to include a certain number of reconstructed adjacent samples located to the left of the current video block, and the certain number of reconstructed adjacent samples depends on the height of the current video block.
[0206] In some embodiments related to method 2400, the k-th selected reconstructed adjacent sample located to the left of the current video block is located at a position described by (blkX - 1, blkY + (k + 1)*blkH / M - 1), where (blkX, blkY) represents the upper left position of the current video block, M is the number of reconstructed adjacent samples, and k is from 0 to (M - 1). In some embodiments related to method 2400, depending on the height being 8 or less, the number of reconstructed adjacent samples is equal to 4. In some embodiments related to method 2400, depending on the height being greater than 8, the number of reconstructed adjacent samples is equal to 8. In some embodiments related to method 2400, the set of reconstructed adjacent samples is selected to include all of the reconstructed adjacent samples located above the current video block, and / or the set of reconstructed adjacent samples is selected to include a certain number of reconstructed adjacent samples located to the left of the current video block, and the certain number of reconstructed adjacent samples depends on the height of the current video block.
[0207] In some embodiments related to method 2400, depending on the height being 8 or less, the number of reconstructed adjacent samples is equal to 4. In some embodiments related to method 2400, depending on the height being greater than 8, the number of reconstructed adjacent samples is equal to 8. In some embodiments related to method 2400, the rule stipulates that the rule is determined by using a set of modified reconstructed adjacent samples obtained by the final prediction block modifying the reconstructed adjacent samples. In some embodiments related to method 2400, the rule stipulates that the set of modified reconstructed adjacent samples is obtained by performing a filtering operation on the reconstructed adjacent samples. In some embodiments related to method 2400, the filtering operation uses an N-tap filter. In some embodiments related to method 2400, N is equal to 2 or 3.
[0208] In some embodiments related to method 2400, the rule stipulates that the filtering operation is adaptively applied according to the MIP mode in which the final prediction block of the current video block is determined. In some embodiments related to method 2400, the technique in which the final prediction block is determined by an upsampling operation is based on the dimensions of the current video block. In some embodiments related to method 2400, the technique in which the final prediction block is determined by an upsampling operation is based on the coding information related to the current video block. In some embodiments related to method 2400, the coding information includes an indication of the intra prediction direction or transform mode related to the current video block.
[0209] In some embodiments of the method of this patent document, performing the transformation includes generating a bitstream representation from the current video block. In some embodiments of the method of this patent document, performing the transformation includes generating the current video block from the bitstream representation.
[0210] From the above, it is to be understood that while specific embodiments of the technology disclosed herein have been described herein for purposes of illustration, various changes can be made without departing from the scope of the invention. Accordingly, the technology disclosed herein is not limited except as by the appended claims.
[0211] The implementation of the subject matter and the functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware, or combinations of one or more of these, including the structures disclosed in this specification and those that are structurally equivalent thereto. The implementation of the subject matter described in this specification can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that generates a machine-readable propagated signal, or combinations of one or more of these. The terms "data processing unit" or "data processing apparatus" include, by way of example, any device, device, and machine that processes data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or combinations of one or more of these.
[0212] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program need not necessarily correspond to a file in a file system. The program may be stored in a part of a file that holds other programs or data (for example, one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in a plurality of coordinated files (for example, files that hold one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or may be deployed to be executed on multiple computers, which may be located in one place or distributed across multiple places and interconnected by a communication network.
[0213] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. These processes and logical flows can also be performed by, for example, dedicated logic circuits such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits), and the apparatus can also be implemented as such dedicated logic circuits.
[0214] Processors suitable for the execution of a computer program include, by way of example, any one or more processors of both general and special purpose microprocessors, and any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as, by way of example, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from or transfer data to such mass storage devices. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media and memory devices, including semiconductor memory devices such as, for example, EPROM, EEPROM, and flash memory devices. The processor and the memory may be supplemented or incorporated by dedicated logic circuitry.
[0215] It is intended that the specification, together with the drawings, be considered as exemplary only, exemplary meaning an example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.
[0216] This patent document contains many details, but they should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of mechanisms that may be specific to particular embodiments of a particular invention. The specific plurality of mechanisms described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various mechanisms described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although a plurality of mechanisms may be described as acting in a particular combination and may even initially be claimed as such, in some cases, one or more mechanisms can be excluded from the claimed combination, and the claimed combination can also be led to a sub-combination or a variation of a sub-combination.
[0217] Similarly, although the drawings show the processing in a particular order, this should not be understood as requiring that the operations be performed in or in the order shown, or that all of the processing shown be performed, in order to achieve the desired result. Also, the separation of the various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0218] Only a few implementations and examples are described, and other implementations, extensions, and variations can be made based on what is described and illustrated in this patent document.< / end> < / end> < / begin> < / end> < / begin> < / end> < / begin> < / end> < / begin> < / end> < / begin>
Claims
Claim 1 A method for processing video data, comprising: determining that a first intra mode is applied to a video block of the video for conversion between the video block of the video and a bitstream of the video, wherein a process in the first intra mode includes selectively performing an upsampling operation following a matrix-vector multiplication operation to generate prediction samples for the video block of the video; performing the conversion based on the prediction samples; wherein an input to the upsampling operation includes adjacent reference samples of the video block of the video; the upsampling operation includes at least one of a horizontal upsampling operation and a vertical upsampling operation, and an order of the horizontal upsampling operation and the vertical upsampling operation in the upsampling operation for the video block having a height greater than a width is the same as that for the video block having a width greater than a height. Claim 2 The process in the first intra mode further includes a downsampling operation before the matrix-vector multiplication operation based on a size of the video block; the downsampling operation is performed on the adjacent reference samples of the video block to generate reduced samples; the reduced samples are input to the matrix-vector multiplication operation and excluded from the upsampling operation, according to the method of Claim 1. Claim 3 N reduced samples are derived from M adjacent reference samples without deriving intermediate samples, and each of the N reduced samples is determined based on K consecutive samples of the M adjacent reference samples; M and N are integers, and K is determined based on M / N, according to the method of Claim 2. Claim 4 Each of the N reduced boundary samples is determined based on an average of the K consecutive samples of the M adjacent reference samples, according to the method of Claim 3. Claim 5 A one-dimensional vector array is further derived based on concatenating the reduced samples, and the one-dimensional vector array is used as an input to the matrix-vector multiplication operation to generate a two-dimensional array, according to the method of Claim 2. Claim 6 The method according to claim 2, wherein a two-dimensional array having a width of a first value and a height of a second value is transposed into an array having a width of the second value and a height of the first value before performing the upsampling operation according to the syntax element.
7. The method according to any one of claims 1 to 6, wherein the conversion includes encoding the video block into the bitstream.
8. The method according to any one of claims 1 to 6, wherein the conversion includes decoding the video block from the bitstream.
9. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to determine that a first intra mode is applied to the video block of the video for conversion between the video block of the video and the bitstream of the video, wherein the process in the first intra mode includes selectively performing an upsampling operation following a matrix-vector multiplication operation to generate a prediction sample for the video block of the video, the step; executing the conversion based on the prediction sample; and cause to be executed, wherein the input to the upsampling operation includes adjacent reference samples of the video block of the video, the upsampling operation includes at least one of a horizontal upsampling operation and a vertical upsampling operation, and the order of the horizontal upsampling operation and the vertical upsampling operation in the upsampling operation for the video block having a height greater than the width is the same as that for the video block having a width greater than the height. Apparatus.
10. Cause the processor to determine that a first intra mode is applied to the video block of the video for conversion between the video block of the video and the bitstream of the video, wherein the process in the first intra mode includes selectively performing an upsampling operation following a matrix-vector multiplication operation to generate a prediction sample for the video block of the video, the step; executing the conversion based on the prediction sample; and store instructions for causing to be executed. The input to the upsampling operation includes adjacent reference samples of the video block of the video, The upsampling operation includes at least one of a horizontal upsampling operation and a vertical upsampling operation, and the order of the horizontal upsampling operation and the vertical upsampling operation in the upsampling operation for the video block having a height greater than the width is the same as that for the video block having a width greater than the height. A non-transitory computer-readable storage medium.
11. A method for storing a bitstream of a video, Determining that a first intra mode is applied to a video block of the video, wherein the process in the first intra mode includes selectively performing an upsampling operation following a matrix-vector multiplication operation to generate prediction samples for the video block of the video. A step, Generating the bitstream based on the determination; Storing the bitstream in a non-transitory computer-readable recording medium; Including, The input to the upsampling operation includes adjacent reference samples of the video block of the video, The upsampling operation includes at least one of a horizontal upsampling operation and a vertical upsampling operation, and the order of the horizontal upsampling operation and the vertical upsampling operation in the upsampling operation for the video block having a height greater than the width is the same as that for the video block having a width greater than the height. Method.
Citation Information
Patent Citations
Block-based prediction
JP2022531902A