Video processing method, device, storage medium and storage method
Matrix-based intra-prediction methods address the bandwidth challenges in video coding by converting video blocks and generating mode lists, resulting in improved efficiency and performance across different video coding standards.
Patent Information
- Application Number
- JP2021559287
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-12
- Filing Date
- 2020-04-13
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2040-04-13
AI Technical Summary
Current video coding technologies face challenges in efficiently managing bandwidth demand due to the increasing number of connected devices and the complexity of video compression algorithms.
The implementation of matrix-based intra-prediction methods for video coding, which involves converting video blocks between current and coded representations, generating most probable mode lists, and applying boundary downsampling and selective upsampling processes to predict video blocks.
This approach enhances video coding efficiency by optimizing bandwidth usage and improving runtime performance across various video coding standards, including HEVC and VVC.
Smart Images

Figure 0007676314000040 
Figure 0007676314000041 
Figure 0007676314000042
Abstract
Description
[Background technology]
[0001] Related Applications This application is a national phase derivative of International Patent Application No. PCT / CN2020 / 084488, filed on April 13, 2020, which is incorporated herein by reference in its entirety. Priority and benefit of International Patent Application No. PCT / CN2019 / 082424 filed on April 12, 2019 Mainly Zhang death For all legal purposes, the entire disclosure of the aforementioned application is hereby incorporated by reference as part of the disclosure of this application.
[0002] Technical Field This patent document relates to video coding techniques, devices and systems.
[0003] background Despite advances in video compression, digital video still accounts for the largest amount of bandwidth on the Internet and other digital communications networks. It is expected that the bandwidth demands for digital video usage will continue to increase as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention
[0004] Devices, systems and methods related to digital video coding, and in particular matrix-based intra prediction methods for video coding, are described. The described methods are applicable to both existing video coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding standards (e.g., Versatile Video Coding (VVC)) or codecs.
[0005] A first embodiment of a video processing method includes generating a first most probable mode (MPM) list for transforming between a current video block of a video and a coded representation of the current video block using a first procedure based on a rule; and performing the transform between the current video block and the coded representation of the current video block using the first MPM list, where the transform of the current video block uses a matrix-based intra prediction (MIP) mode, where in the MIP mode, a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; the rule specifies that the first procedure used to generate the first MPM list is the same as a second procedure used to generate a second MPM list for transforming other video blocks of the video that are coded using a non-MIP intra mode different from the MIP mode; and at least a portion of the first MPM list is generated based on at least a portion of the second MPM list.
[0006] A video processing method of a second embodiment includes the steps of: generating a Most Probable Mode (MPM) list based on a rule for transforming between a current video block of a video and a coded representation of the current video block, the rule being based on whether a neighboring video block of the current video block is coded in a matrix-based intra prediction (MIP) mode, where in the MIP mode, a prediction block of the neighboring video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation, followed by a selective upsampling operation; and performing a transform between the current video block and the coded representation of the current video block using the MPM list, where the transform applies a non-MIP mode to the current video block, the non-MIP mode being different from the MIP mode.
[0007] A video processing method of a third embodiment includes the steps of: decoding a current video block of a coded video in a coded representation of the current video block using a matrix-based intra-prediction (MIP) mode, in which a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation, followed by a selective upsampling operation; and updating line buffers associated with the decoding without storing information in the line buffers indicating whether the current video block is coded using the MIP mode.
[0008] A video processing method of a fourth embodiment includes performing a conversion between a current video block and a bitstream representation of the current video block, where the current video block is coded using a matrix-based intra-prediction (MIP) mode, where in the MIP mode, a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation followed by a selective upsampling operation, where for at most K contexts in the arithmetic encoding or decoding process, a flag is coded into the bitstream representation, the flag indicating whether the current video block is coded using the MIP mode, where K is greater than or equal to zero.
[0009] A video processing method of a fifth embodiment relates to a conversion between a current video block of a video and a bitstream representation of the current video block, the method including the steps of: generating an intra-prediction mode for a current video block coded in a matrix-based intra-prediction (MIP) mode, where in the MIP mode, a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; determining a rule for storing information indicating the intra-prediction mode based on whether the current video block is coded in the MIP mode; and performing the conversion according to the rule, where the rule defines that a syntax element of the intra-prediction mode is stored in the bitstream representation of the current video block and the rule defines that a mode index of the MIP mode of the current video block is not stored in the bitstream representation.
[0010] A sixth embodiment video processing method includes the steps of: making a first determination that a luma video block of a video is coded using a matrix-based intra prediction (MIP) mode, in which a prediction block of the luma video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; making a second determination regarding a chroma intra mode to be used for a chroma video block associated with the luma video block based on the first determination; and performing a conversion between the chroma video block and a bitstream representation of the chroma video block based on the second determination.
[0011] A video processing method of a seventh embodiment includes performing a transformation between a current video block of a video and a coded representation of the current video block, the transformation being based on a determination of whether to code the current video block using a matrix-based intra-prediction (MIP) mode, in which a predicted block for the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then selectively performing an upsampling operation after a matrix-vector multiplication operation.
[0012] A video processing method of an eighth embodiment includes the steps of: determining, according to a rule, whether a coding mode different from a matrix-based intra-prediction (MIP) mode and a MIP mode is to be used to encode a current video block of a video into a bitstream representation of the current video block, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video and then selectively performing an upsampling operation after a matrix-vector multiplication operation; and adding the encoded representation of the current video block to the bitstream representation based on the determination.
[0013] A video processing method of a ninth embodiment includes the steps of determining that a current block of video is to be coded into a bitstream representation using a matrix-based intra-prediction (MIP) mode and a coding mode different from MIP, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then selectively performing an upsampling operation; and generating a decoded representation of the current video block by analyzing and decoding the bitstream representation.
[0014] A video processing method of a tenth embodiment includes the steps of: making a determination regarding the applicability of a loop filter to a reconstructed block of a current video block of video, in a conversion between a coded representation of the video and a current video block, the current video block being coded using a matrix-based intra-prediction (MIP) mode, in which a predicted block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation, followed by a selective upsampling operation; and processing the current video block in accordance with the determination.
[0015] A video processing method of an eleventh embodiment includes the steps of: determining, according to a rule, a type of neighboring samples of a current video block to be used to encode a current video block of a video into a bitstream representation of the current video block; and adding the encoded representation of the current video block to the bitstream representation based on the determination, wherein the current video block is encoded using a matrix-based intra-prediction (MIP) mode, and in the MIP mode, a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then selectively performing an upsampling operation after performing a matrix-vector multiplication operation.
[0016] A video processing method of a twelfth embodiment includes the steps of: determining according to a rule that a current video block of a video is coded into a bitstream representation using a matrix-based intra-prediction (MIP) mode and using a type of neighboring samples of the current video block, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video and then selectively performing an upsampling operation after a matrix-vector multiplication operation; and generating a decoded representation of the current video block by analyzing and decoding the bitstream representation.
[0017] A video processing method of a thirteenth embodiment includes performing a conversion between a current video block of a video and a bitstream representation of the current video block, the conversion including generating a prediction block for the current video block by selecting and applying a matrix multiplication using a matrix of samples using a matrix-based intra prediction (MIP) mode and / or by selecting and adding an offset using an offset vector for the current video block, the samples being obtained from row and column wise averages of previously coded samples of the video, the selection being based on reshaping information associated with applying a luma mapping with chroma scaling (LMCS) technique with respect to a reference picture of the current video block.
[0018] A video processing method of a fourteenth embodiment includes the steps of: determining that a current block should be coded using a matrix-based intra-prediction (MIP) mode, in which a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; and performing a conversion between the current video block and a bitstream representation of the current video block based on the determination, wherein the performing of the conversion is based on rules for co-application of the MIP mode and other coding techniques.
[0019] A video processing method of a fifteenth embodiment includes a step of performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra-prediction (MIP) mode, where performing the conversion using the MIP mode includes generating a prediction block by applying matrix multiplication using a matrix of samples obtained from row-wise and column-wise averages of previously coded samples of the video, the matrix being dependent on a bit depth of the samples.
[0020] A video processing method of a sixteenth embodiment includes the steps of: generating an intermediate prediction signal for a current video block of a video using a matrix-based intra-prediction (MIP) mode, in which a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; generating a final prediction signal based on the intermediate prediction signal; and performing a conversion between the current video block and a bitstream representation of the current video block based on the final prediction signal.
[0021] A video processing method of a seventeenth embodiment includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra-prediction (MIP) mode; performing the conversion includes using an interpolation filter in an upsampling process for the MIP mode; in the MIP mode, a matrix multiplication is applied to a first set of samples obtained from row and column wise averaging of previously coded samples of the video, and the interpolation filter is applied to a second set of samples obtained from the matrix multiplication; and the interpolation filter excludes a bilinear interpolation filter.
[0022] A video processing method of an eighteenth embodiment includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to rules, the rules specifying an application relationship of a matrix-based intra-prediction (MIP) mode or a conversion mode during the conversion; the MIP mode including determining a predictive block for the current video block by selectively performing an upsampling process after performing a boundary downsampling process on previously coded samples of the video and a matrix-vector multiplication process; and the conversion mode specifying use of a conversion process to determine a predictive block for the current video block.
[0023] A video processing method of a 19th embodiment includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra-prediction (MIP) mode, in which a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and then selectively performing an upsampling operation after a matrix-vector multiplication operation; performing the conversion includes deriving boundary samples according to a rule by applying a left bit-shift operation or a right bit-shift operation to a sum of at least one reference boundary sample; and the rule determines whether to apply the left bit-shift operation or the right bit-shift operation.
[0024] A video processing method of a twentieth embodiment includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra-prediction (MIP) mode; in the MIP mode, a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and a matrix-vector multiplication operation followed by a selective upsampling operation; the prediction samples predSamples[xHor+dX][yHor] are calculated in the upsampling operation according to the following formula: predSamples[ xHor + dX ][ yHor ] = ( ( upHor - dX ) * predSamples[ xHor ][ yHor ] + dX * predSamples[ xHor + upHor ][ yHor ] + offsetHor) / upHor, and predSamples[ xVer ][ yVer + dY ] = ( ( upVer - dY ) * predSamples[ xVer ][ yVer ] + dY * predSamples[ xVer ][ yVer + upVer ]+ offsetVer ) / upVer where offsetHor and offsetVer are integers, upHor is a function of a predetermined value based on the size of the current video block and the width of the current video block, upVer is a function of a predetermined value based on the size of the current video block and the height of the current video block, dX is 1···upHor-1, dY is 1···upVer-1, and xHor is a position based on upHor and yHor is a position based on upVer.
[0025] A video processing method of a twenty-first embodiment relates to a conversion between a current video block of a video and a bitstream representation of the current video block, the method comprising the steps of: generating an intra-prediction mode for a current video block coded using a matrix-based intra-prediction (MIP) mode, in which a predictive block for the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation; determining a rule for storing information indicative of the intra-prediction mode based on whether the current video block is coded in the MIP mode; and performing the conversion according to the rule, the rule specifying that the bitstream representation excludes storage of information indicative of the MIP mode associated with the current video block.
[0026] In yet another exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, constructing at least a portion of a Most Probable Mode (MPM) list for the ALWIP mode based on at least a portion of an MPM list for a non-ALWIP intra mode based on the determination, and performing a conversion between the current video block and a bitstream representation of the current video block based on the MPM list for the ALWIP mode.
[0027] In yet another exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining that a luma component of a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode, estimating a chroma intra mode based on the determination, and performing a conversion between the current video block and a bitstream representation of the current video block based on the chroma intra mode.
[0028] In yet another exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining that a current video block is to be coded using an affine linear weighted intra prediction (ALWIP) mode and, based on the determination, performing a conversion between the current video block and a bitstream representation of the current video block.
[0029] In yet another exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining that a current video block is to be coded using a coding mode other than an affine linear weighted intra prediction (ALWIP) mode, and, based on the determination, performing a conversion between the current video block and a bitstream representation of the current video block.
[0030] In yet another representative aspect, the above-described methods are embodied in the form of processor executable code and stored on a computer readable program medium.
[0031] In yet another representative aspect, a device configured or operable to perform the above-mentioned method is disclosed. The device may include a processor programmed to implement the method.
[0032] In yet another exemplary embodiment, a video decoder device may implement the methods described herein.
[0033] These and other aspects and features of the disclosed technology are explained in more detail in the drawings, specification and claims. [Brief description of the drawings]
[0034] [Figure 1] Here are some examples of the 33 intra prediction directions:
[0035] [Diagram 2] 67 examples of intra prediction modes are shown.
[0036] [Diagram 3] 1 shows an example of sample locations used to derive the weights of a linear model.
[0037] [Figure 4] An example of four reference lines adjacent to a predicted block is shown.
[0038] [Figure 5A] An example of sub-partitions depending on block size is shown below. [Figure 5B] An example of sub-partitions depending on block size is shown below.
[0039] [Figure 6] An example of ALWIP for a 4x4 block is shown below.
[0040] [Figure 7] Here is an example of ALWIP for an 8x8 block.
[0041] [Figure 8] An example of ALWIP for an 8x4 block is shown below.
[0042] [Figure 9] Here is an example of ALWIP for a 16x16 block.
[0043] [Figure 10] An example of adjacent blocks used in an MPM list configuration is shown below.
[0044] [Figure 11] 1 illustrates a flowchart of an exemplary method for matrix-based intra prediction in accordance with the disclosed techniques.
[0045] [Figure 12] 13 shows a flowchart of another exemplary method for matrix-based intra prediction in accordance with the disclosed techniques.
[0046] [Figure 13] 13 shows a flowchart of yet another exemplary method for matrix-based intra prediction in accordance with the disclosed techniques.
[0047] [Figure 14]13 shows a flowchart of yet another exemplary method for matrix-based intra prediction in accordance with the disclosed techniques.
[0048] [Figure 15] 1 is a block diagram of an example hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.
[0049] [Figure 16] FIG. 1 is a block diagram illustrating an example video processing system capable of implementing various techniques disclosed herein.
[0050] [Figure 17] FIG. 1 is a block diagram illustrating an example video coding system in which techniques of the present disclosure can be utilized.
[0051] [Figure 18] FIG. 1 is a block diagram illustrating an example of a video encoder.
[0052] [Figure 19] FIG. 2 is a block diagram illustrating an example of a video decoder.
[0053] [Figure 20] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 21] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 22] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 23] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 24]13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 25] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 26] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 27A] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 27B] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 28] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 29A] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 29B] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 30] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 31] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 32] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 33] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 34] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 35] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Diagram 36] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 37] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. [Figure 38] 13 shows an illustrative flowchart of an additional example method for matrix-based intra prediction according to the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0054] Due to the ever-increasing demand for higher resolution video, video coding methods and techniques are ubiquitous in modern technology. Video codecs typically include electronic circuits or software that compress or decompress digital video and are constantly being improved to provide higher coding efficiency. Video codecs convert uncompressed video to compressed form and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, the sensitivity to data loss and errors, the ease of editing, random access, and end-to-end delay (latency). Compressed formats usually comply with a standard video compression standard, such as the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2), the finalized Versatile Video Coding (VVC) standard, or other current and / or future video coding standards.
[0055] Embodiments of the disclosed technology may be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve runtime performance. Section headings are used in this patent document to improve the readability of the description and do not in any way limit the description or embodiments (and / or implementations) to any particular section.
[0056] 1. A brief review of HEVC 1.1 Intra Prediction in HEVC / H.265 Intra prediction involves creating samples for a given TB (transform block) using previously reconstructed samples in the assumed color channels. Intra prediction modes are signaled separately for luma and chroma channels, and the chroma channel intra prediction mode optionally depends on the luma channel intra prediction mode via the 'DM_CHROMA' mode. Although the intra prediction modes are signaled at the PB (prediction block) level, the intra prediction process is applied at the TB level according to the residual quadtree hierarchy of the CU, which allows the coding of one TB to influence the coding of the next TB in the CU, thus reducing the distance to the samples used as references.
[0057] HEVC includes 35 intra-prediction modes (DC mode, planar mode, and three directional or "angular" intra-prediction modes). The 33 angular intra-prediction modes are shown in FIG.
[0058] For PBs associated with chroma color channels, the intra prediction mode is specified as either planar, DC, horizontal, vertical, 'DM_CHROMA' mode, or sometimes the diagonal mode '34'.
[0059] For chroma formats 4:2:2 and 4:2:0, a chroma PB may overlap with two or four luma PBs (respectively); in this case the luma direction in DM_CHROMA is taken from the top-left of these luma PBs.
[0060] The DM_CHROMA mode indicates that the intra prediction mode of the luma color channel PB is applied to the chroma color channel PB. Since this is relatively common, the most probable mode coding scheme of intra_chroma_pred_mode is biased towards this mode being selected.
[0061] 2. Example of intra prediction in VVC 2.1 Intra-mode coding with 67 intra-prediction modes To capture any edge orientation presented in natural video, the number of directional intra modes is expanded from 33 as used in HEVC to 65. The additional directional modes are depicted as red dotted arrows in Figure 2, while the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and for both luma and chroma intra prediction.
[0062] 2.2 Example of a Cross-Component Linear Model (CCLM) In some implementations, to reduce cross-component redundancy, a Cross-Component Linear Model (CCLM) prediction mode (also called LM) is used in JEM, where chroma samples are predicted based on the reconstructed luma samples of the same CU by using a linear model as follows:
[0063]
number
[0064] Here, pred C(i,j) represents the predicted chroma sample in CU, and rec L '(i,j) represents the downsampled reconstructed luma sample of the same CU. The linear model parameters α and β are derived from the relationship between luma and chroma values from two samples, the minimum and maximum luma samples in the set of downsampled adjacent luma samples, and their corresponding chroma samples. Figure 3 shows an example of the location of the samples of the current block and the left and top samples involved in the CCLM mode.
[0065] This parameter calculation is performed as part of the decoding process and is not exactly the same as the encoder's search operation, and as a result no syntax is used to communicate the values of α and β to the decoder.
[0066] For chroma intra mode coding, a total of eight chroma intra modes are allowed for chroma intra mode coding. These modes include five traditional intra modes and three cross-component linear models (CCLM, LM_A, and LM_L). The chroma mode coding directly depends on the intra prediction mode of the corresponding luma block. Since separate block partitioning structures are enabled for luma and chroma components in I slices, one chroma block may correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra prediction mode of the corresponding luma block that covers the center position of the current chroma block is directly inherited.
[0067] 2.3 Multiple Reference Line (MRL) Intra Prediction Multiple Reference Line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 4, an example of four reference lines is depicted, where samples of segments A and F are padded with the nearest samples from segments B and E, respectively, instead of being fetched from reconstructed neighboring samples. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used. The index (mrl_idx) of the selected reference line is signaled and used to generate the intra predictor. For reference line idx greater than 0, we just include the additional reference line modes in the MPM list and signal the mpm index without the remaining modes.
[0068] 2.4 Intra Sub-Partition (ISP) The Intra Sub-Partition (ISP) tool divides a luma intra prediction block vertically or horizontally into two or four sub-partitions depending on the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided by four sub-partitions. Figure 5 shows two possible examples. All sub-partitions satisfy the condition that they have at least 16 samples.
[0069] For each sub-partition, reconstructed samples are obtained by adding the residual signal to the prediction signal, where the residual signal is generated by processes such as entropy decoding, inverse quantization, and inverse transform. Thus, the reconstructed sample values of each sub-partition are available to generate the prediction of the next sub-partition, and each sub-partition is processed iteratively. Furthermore, the first sub-partition processed is the one that contains the top-left samples of the CU, and then the one that follows downwards (horizontal division) or to the right (vertical division). As a result, the reference samples used to generate the sub-partition prediction signal are only located on the left and upper sides of the line. All sub-partitions share the same intra mode.
[0070] 2.5 Affine Linear Weighted Intra Prediction (ALWIP or Matrix-Based Intra Prediction) Affine Linear Weighted Intra Prediction (ALWIP, also known as Matrix-Based Intra Prediction (MIP)) is proposed in JVET-N0217.
[0071] In JVET-N0217, two tests are performed. In test 1, ALWIP is designed with at most 4 multiplications per sample, with a memory limit of 8K bytes. Test 2 is similar to test 1, but further simplifies the design in terms of memory requirements and model architecture.
[0072] ○ Single set of matrices and offset vectors for all block shapes
[0073] ○ Reduce the number of modes to 19 for all block shapes.
[0074] ○ Reduces memory requirements to 5760 10-bit values, or 7.20 kilobytes.
[0075] o A linear interpolation of the predicted samples is performed in a single step per direction, replacing the iterative interpolation as in the first test.
[0076] 2.5.1 Test 1 of JVET-N0217 To predict samples of a rectangular block of width W and height H, Affine Linear Weighted Intra Prediction (ALWIP) takes as input one line with H reconstructed adjacent boundary samples to the left of the block and one line with W reconstructed adjacent boundary samples above the block. If reconstructed samples are not available, they are generated similarly to regular intra prediction. The generation of the prediction signal is based on three steps:
[0077] Of the boundary samples, 4 samples are extracted by averaging if W=H=4, and 8 samples in all other cases.
[0078] Taking the averaged samples as input, a matrix-vector multiplication followed by the addition of an offset is performed. The result is a reduced prediction signal for the set of subsampled samples in the original block.
[0079] The prediction signals at the remaining positions are generated from the prediction signals for the subsampled set by linear interpolation, which is a single-step linear interpolation in each direction.
[0080] The matrices and offset vectors required to generate the prediction signal are taken from three sets of matrices S0, S1, and S2. Set S0 contains 18 matrices A0, each with 16 rows and 4 columns. i ,i∈{0,...,17} and 18 offset vectors b0, each of size 16 i ,i∈{0,...,17}. The matrices and offset vectors belonging to the set are used for blocks of size 4 × 4. The set S1 consists of 10 matrices A1, each with 16 rows and 8 columns. i ,i∈{0,...,9} and 10 offset vectors b1, b2, b3, b4, b5, b6, b7, b8, b9, b10, b11, b12, b13, b14, b15, b16, b17, b18, b19, b20, b210, b22, b23, b24, b25, b36, b37, b46, b47, b i,i∈{0,...,9}. The matrices and offset vectors belonging to the set are used for blocks of size 4×8, 8×4, 8×8. Finally, the set S2 consists of 6 matrices A2 each with 64 rows and 8 columns. i ,i∈{0,...,5} and 6 offset vectors b2, each of size 64 i ,i∈{0,...,5}. Matrices and offset vectors belonging to that set, or portions of these matrices and offset vectors, are used for all other block shapes.
[0081] The total number of multiplications required to compute the matrix-vector product is always less than or equal to 4×W×H. In other words, for the ALWIP mode, a maximum of 4 multiplications are required per sample.
[0082] 2.5.2 Boundary Equalization In the first step, the input boundary bdry top and bdry left is the smaller boundary bdry red top and bdry red left where bdry red top and bdry red left For 4x4 blocks, both consist of 2 samples, and for all other cases, both consist of 4 samples. In the case of 4x4 blocks, for 0≦i<2, it is defined as follows:
[0083]
number
[0084] Similarly, bdry red left is determined. In all other cases, the block width W is W=4·2 k , 0≦i<4, it is defined as follows:
[0085]
number
[0086] Similarly, bdry red left is determined.
[0087] 2 reduced borders bdry red top and bdry red left is the reduced boundary vector bdry red which is of size 4 for blocks of shape 4x4 and of size 8 for blocks of all other shapes. If the mode refers to the ALWIP mode, this concatenation is defined as follows:
[0088]
number
[0089] Finally, for the interpolation of the subsampled prediction signal, for large blocks a second version of the averaged boundary is needed: min(W,H)>8 and W≧H, where W=8*2 l Then for 0≦i<8, it can be defined as:
[0090]
number
[0091] If min(W,H)>8 and H>W, then bdry redII left is defined similarly. 2.5.3 Generating reduced prediction signals by matrix-vector multiplication The reduced input vector bdry red From the reduced prediction signal pred red The latter signal has a width W red and height Hred where W red and height H red is defined as follows:
[0092]
number
[0093]
number
[0094] The reduced prediction signal pred red is calculated by taking the matrix-vector product and adding an offset:
[0095]
number
[0096] Here, A is W red H red B is a matrix with W rows and 4 columns if W=H=4, and 8 columns in all other cases. red H red It is a vector of size
[0097] The matrix A and vector b are taken from one of the sets S0, S1, S2 as follows: Let index idx=idx(W,H) be defined as:
[0098]
number
[0099] Additionally, set m to:
[0100]
number
[0101] Then, if idx≦1 or idx=2 and min(W,H)>4, then A=A idx m and b=b idx m If idx=2 and min(W,H)=4, then A is idx m Let W be the matrix resulting from removing all rows of H = 4, corresponding to odd x-coordinates in the downsampled block, or H = 4, corresponding to odd y-coordinates in the downsampled block.
[0102] Finally, the downsized prediction signal is replaced by its transpose if:
[0103] ○ W=H=4 and mode≧18
[0104] ○ max(W,H)=8 and mode≧10
[0105] ○ max(W,H)>8 and mode≧6
[0106] pred red The number of multiplications required to compute A is 4 if W=H=4, because then A has 4 columns and 16 rows. In all other cases, A has 8 columns and W red H red rows, in which 8·Wred·Hred≦4·W·H multiplications are required, i.e., in this case, pred red It can be quickly verified that at most four multiplications are required per sample to compute
[0107] 2.5.4 Overall ALWIP process description The overall process of averaging, matrix vector multiplication, and linear interpolation is illustrated for various shapes in Figures 6-9. Note that the remaining shapes are treated similarly to one of the cases shown.
[0108] Given a 1.4x4 block, ALWIP takes two averages along each axis of the boundary. The resulting four input samples go into a matrix-vector multiplication. The matrix is taken from set S0. After adding an offset, this results in 16 final predicted samples. No linear interpolation is required to generate the predicted signal. Thus, a total of (4·16) / (4·4)=4 multiplications are performed per sample.
[0109] Given a 2.8x8 block, ALWIP takes 4 averages along each axis of the boundary. The matrix is taken from set S1. This results in 16 samples in odd positions of the prediction block. Thus, a total of (8·16) / (8·8)=2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically by using the reduced top boundary. Horizontal interpolation follows by using the original left boundary.
[0110] Given a 3.8x4 block, ALWIP takes the four averages along the horizontal axis of the boundary and the four original boundary values at the left boundary. The resulting eight input samples go into a matrix vector multiplication. The matrix is taken from set S1. This results in 16 samples at odd horizontal positions and at each vertical position of the predicted block. Thus, a total of (8·16) / (8·4)=4 multiplications are performed per sample. After adding the offset, these samples are horizontally interpolated by using the original left boundary.
[0111] Assuming a 4.16×16 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is taken from set S2. This results in 64 samples at the odd positions of the prediction block. Thus, a total of (8·64) / (16·16) = 2 multiplications per sample are performed. After adding the offset, these samples are interpolated vertically by using the eight averages of the upper boundary. Horizontal interpolation follows by using the original left boundary. In this case, the interpolation process does not add any multiplications. Thus, a total of two multiplications per sample are required to compute the ALWIP prediction.
[0112] For larger shapes, the procedure is essentially the same, and it is easy to verify that the number of multiplications per sample is less than 4.
[0113] For a W×8 (W>8) block, only horizontal interpolation is required because the samples are given at each vertical position at odd horizontal positions.
[0114] Finally, for a W×4 (W>8) block, let A_kbe be the matrix resulting from excluding all rows corresponding to odd entries along the horizontal axis of the downsampled block. Thus, the output size is 32, and again, only horizontal interpolation is performed.
[0115] The transposed case is handled accordingly.
[0116] 2.5.5. Single-Step Linear Interpolation For a W×H block (max(W,H)≧8), the prediction signal results from the downsampled prediction signal by linear interpolation with respect to W red ×H red Linear interpolation is performed vertically, horizontally, or bidirectionally depending on the shape of the block. If linear interpolation is to be applied bidirectionally, it is applied horizontally first if W<H, and vertically first otherwise.
[0117] Without loss of generality, consider a W×H block, where max(W,H)≧8 and W≧H. Then, a one-dimensional linear interpolation is performed as follows: Without loss of generality, it is sufficient to describe the linear interpolation in the vertical direction. First, the downscaled prediction signal is expanded to the top by the border signal. The vertical upsampling factor U ver =H / H red Define U ver =2 uver >1. Then we define the extended downscaled prediction signal as:
[0118]
number
[0119] A vertically linearly interpolated prediction is then generated from this expanded downscaled prediction as follows:
[0120]
number
[0121] 2.5.6 Proposed Signaling of Intra Prediction Modes
[0122] For each coding unit (CU) in intra mode, a flag is transmitted in the bitstream indicating whether the ALWIP mode should be applied to the corresponding prediction unit (PU). The signaling of the latter index is aligned with the MRL as in JVET-M0043. If the ALWIP mode should be applied, the ALWIP mode index predmode is signaled using an MPM list with 3MPMS.
[0123] Here, the derivation of the MPM is performed using the intra mode of the top and left PUs as follows: Three fixed tables map_angular_to_alwip idx, idx∈{0,1,2}, which are the conventional intra prediction modes predmode Angular Assign ALWIP mode to.
[0124]
number
[0125] For each PU of width W and height H, define the indices:
[0126] idx(PU) = idx(W,H)∈{0,1,2}
[0127] This index specifies that the ALWIP parameters are taken from one of three sets as in section 2.5.3.
[0128] Prediction Unit PU above is available, belongs to the same CTU as the current PU, and is in intra mode, then idx(PU) = idx(PU above ) and the ALWIP mode is ALWIP mode predmode ALWIP above PU above If it applies to you, set it as follows:
[0129]
number
[0130] Conventional intra prediction mode predmode if the prediction unit above is available, belongs to the same CTU as the current PU, and is in intra mode Angular above If applies to the above PU, set it as follows:
[0131]
number
[0132] In all other cases, set it as follows:
[0133]
number
[0134] This means that this mode is not available. Similarly, the mode mode can be used without the restriction that the left PU must belong to the same CTU as the current PU. ALWIP left Derive:
[0135] Finally, there are three fixed default lists: idx , idx∈{0,1,2}, each containing three different ALWIP modes. idx(PU) and mode ALWIP above and mode ALWIP left We configure three different MPMs by replacing the default values with -1 and omitting repetitions.
[0136] The left and top neighboring blocks used in the ALWIP MPM list construction are A1 and B1 as shown in FIG.
[0137] 2.5.7 Deriving Adaptive MPM Lists for Traditional Luma and Chroma Intra Prediction Modes The proposed ALWIP mode is consistent with MPM-based coding of conventional intra prediction modes as follows: The derivation process of luma and chroma MPM lists for conventional intra prediction modes is based on a fixed table map_alwip_to_angular idx , idx∈{0,1,2} to map the ALWIP mode for a given PU to one of the conventional intra prediction modes:
[0138]
number
[0139] For Luma MPM list derivation, ALWIP mode predmode ALWIP Whenever a neighboring luma block using the predmode Angular For chroma MPM list derivation, whenever the current luma block uses ALWIP mode, the same mapping is used to convert ALWIP modes to conventional intra prediction modes.
[0140] 2.5.8 Working Drafts, amended accordingly In some embodiments, portions relating to intra_lwip_flag, intra_lwip_mpm_flag, intra_lwip_mpm_idx, and intra_lwip_mpm_remainder have been added to the working draft based on embodiments of the disclosed technology, as described in this section.
[0141] In some embodiments, as described in this section, <begin>Tags and <end>Tags are used to indicate additions and modifications to working drafts that are based on embodiments of the disclosed technology. Syntax Table Coding Unit Syntax
[0142]
number
[0000] = lwipMpmCand[ sizeId ]
[0000] (8-X1) candLwipModeList
[0001] = lwipMpmCand[ sizeId ]
[0001] (8-X2) candLwipModeList
[0002] = lwipMpmCand[ sizeId ]
[0002] (8-X3) - Otherwise the following applies: - If candLwipModeA is equal to candLwipModeB, or if candLwipModeA or candLwipModeB is equal to -1, the following applies: candLwipModeList
[0000] =( candLwipModeA != -1 ) ? candLwipModeA : candLwipModeB (8-X4) - If candLwipModeList
[0000] is equal to lwipMpmCand[ sizeId ]
[0000] , the following applies: candLwipModeList
[0001] = lwipMpmCand[ sizeId ]
[0001] (8-X5) candLwipModeList
[0002] = lwipMpmCand[ sizeId ]
[0002] (8-X6) - Otherwise the following applies: candLwipModeList
[0001] = lwipMpmCand[ sizeId ]
[0000] (8-X7) candLwipModeList
[0002] = ( candLwipModeList
[0000] != lwipMpmCand[ sizeId ]
[0001] ) ? lwipMpmCand[ sizeId ]
[0001] : lwipMpmCand[ sizeId ]
[0002] (8-X8) - Otherwise the following applies: candLwipModeList
[0000] = candLwipModeA (8-X9) candLwipModeList
[0001] = candLwipModeB (8-X10) - If both candLwipModeA and candLwipModeB are not equal to lwipMpmCand[ sizeId ]
[0000] , the following applies: candLwipModeList
[0002] = lwipMpmCand[ sizeId ]
[0000] (8-X11) - Otherwise the following applies: - If both candLwipModeA and candLwipModeB are not equal to lwipMpmCand[ sizeId ]
[0001] , the following applies: candLwipModeList
[0002] = lwipMpmCand[ sizeId ]
[0001] (8-X12) - Otherwise the following applies: candLwipModeList
[0002] = lwipMpmCand[ sizeId ]
[0002] (8-X13) 4. IntraPredModeY[xCb][yCb] is derived by applying the following steps: - if intra_lwip_mpm_flag[xCb][yCb] is equal to 1, IntraPredModeY[xCb][yCb] is set equal to candLwipModeList[intra_lwip_mpm_idx[xCb][yCb]]. - Otherwise, IntraPredModeY[xCb][yCb] is derived by applying the following step sequence: 1. If candLwipModeList[ i ] is greater than candLwipModeList[ j ] for i=0..1 and for i,j=(i+1)..2, then both values are swapped as follows: ( candLwipModeList[ i ], candLwipModeList[ j ] ) = Swap( candLwipModeList[ i ], candLwipModeList[ j ] ) (8-X14) 2. IntraPredModeY[xCb][yCb] is derived by applying the following sequence of steps: i. IntraPredModeY[xCb][yCb] is set equal to intra_lwip_mpm_remainder[xCb][yCb]. ii. For i equal to 0 through 2, inclusive, if IntraPredModeY[xCb][yCb] is greater than or equal to candLwipModeList[i], then the value of IntraPredModeY[xCb][yCb] is incremented by one. The variable IntraPredModeY[x][y] is set equal to IntraPredModeY[xCb][yCb], where x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1. 8.4.X.1 Prediction block size type derivation process The inputs to this process are: - Variable cbWidth: Specifies the width of the current coding block in luma samples. - Variable cbHeight: Specifies the height of the current coding block in luma samples. The output of this process is the variable sizeId. The variable sizeId is derived as follows: - If both cbWidth and cbHeight are equal to 4, then sizeId is set equal to 0. - Otherwise, if both cbWidth and cbHeight are less than or equal to 8, sizeId is set equal to 1. - Otherwise, sizeId is set equal to 2. Table 8-X1 - Specification of mapping between intra prediction and affine linear weighted intra prediction modes
[0143] [Table 1] Table 8 - X2-Affine Linear Weighted Intra Prediction Candidate Mode Specifications
[0144] [Table 2] <end> 8.4.2 Luma intra prediction mode derivation process The inputs to this process are: - Luma Position(xCb, yCb): Specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture. - Variable cbWidth: Specifies the width of the current coding block in luma samples. - Variable cbHeight: Specifies the height of the current coding block in luma samples. In this process, the luma intra prediction mode IntraPredModeY[xCb][yCb] is derived. Table 8-1 specifies values for the intra prediction modes IntraPredModeY[xCb][yCb] and associated names. Table 8-1 Specification of intra prediction modes and associated names
[0145] [Table 3] NOTE - : The intra prediction modes INTRA_LT_CCLM, INTRA_L_CCLM and INTRA_T_CCLM are applicable to chroma components only. IntraPredModeY[xCb][yCb] is derived in the following step sequence: 1. The adjacent positions (xNbA, yNbA) and (xNbB, yNbB) are set equal to (xCb - 1, yCb + cbHeight - 1) and (xCb + cbWidth - 1, yCb - 1), respectively. 2. If X is replaced by A or B, the variable candIntraPredModeX is derived as follows: - The availability derivation process for the block as specified in Section 6.4.X. <begin>[Ed.(BB):Adjacent block availability check process tbd] <end>is called with input a position (xCurr, yCurr) set equal to (xCb, yCb) and an adjacent position (xNbY, yNbY) set equal to (xNbX, yNbX), and the output is assigned to availableX. - The candidate intra prediction modes candIntraPredModeX are derived as follows: - candIntraPredModeX is set equal to INTRA_PLANAR if one or more of the following conditions are true: - The variable availableX is equal to FALSE. - CuPredMode[xNbX][yNbX] is not equal to MODE_INTRA and ciip_flag[xNbX][yNbX] is not equal to 1. - pcm_flag[xNbX][yNbX] is equal to 1. - X is equal to B and yCb-1 is less than ( ( yCb >> CtbLog2SizeY ) << CtbLog2SizeY ). - Otherwise, candIntraPredModeX is derived as follows: - If intra_lwip_flag[xCb][yCb] is equal to 1, candIntraPredModeX is derived in the following step order: i. The size type derivation process for the block as specified in Clause 8.4.X.1 is called with input the width of the current coding block in luma samples, cbWidth, and the height of the current coding block in luma samples, cbHeight, and the output is assigned to the variable sizeId. ii. candIntraPredModeX is derived using IntraPredModeY[xNbX][yNbX] and sizeId as specified in Table 8-X3. - Otherwise, candIntraPredModeX is set equal to IntraPredModeY[xNbX][yNbX]. 3. The variables ispDefaultMode1 and ispDefaultMode2 are defined as follows: - If IntraSubPartitionsSplitType is equal to ISP_HOR_SPLIT, then ispDefaultMode1 is set equal to INTRA_ANGULAR18 and ispDefaultMode2 is set equal to INTRA_ANGULAR5. - Otherwise, ispDefaultMode1 is set equal to INTRA_ANGULAR50 and ispDefaultMode2 is set equal to INTRA_ANGULAR63. ... Table 8-X3 - Specification of the mapping between affine linear weighted intra prediction modes and intra prediction modes
[0146] [Table 4] 8.4.3 Chroma Intra Prediction Mode Derivation Process The inputs to this process are: - Luma Position(xCb, yCb): Specifies the top-left sample of the current chroma coding block relative to the top-left luma sample of the current picture. - Variable cbWidth: Specifies the width of the current coding block in luma samples. - Variable cbHeight: Specifies the height of the current coding block in luma samples. In this process, the chroma intra prediction mode IntraPredModeC[xCb][yCb] is derived. - If intra_lwip_flag[xCb][yCb] is equal to 1, lumaIntraPredMode is derived by the following step sequence: i. The size type derivation process for the block as specified in Clause 8.4.X.1 is called with input the width of the current coding block in luma samples, cbWidth, and the height of the current coding block in luma samples, cbHeight, and the output is assigned to the variable sizeId. ii. The luma intra prediction mode is derived using IntraPredModeY[ xCb + cbWidth / 2 ][ yCb + cbHeight / 2 ] and sizeId as specified in Table 8-X3. - Otherwise, lumaIntraPredMode is set equal to IntraPredModeY[ xCb + cbWidth / 2 ][ yCb + cbHeight / 2 ]. The chroma intra prediction mode IntraPredModeC[xCb][yCb] is derived using intra_chroma_pred_mode[xCb][yCb] and lumaIntraPredMode as specified in Tables 8-2 and 8-3. ... xxx. Intra-sample prediction <begin> The inputs to this process are: - Sample position (xTbCmp, yTbCmp): Specifies the top-left sample of the current transform block relative to the top-left sample of the current picture. - Variable predModeIntra: Specifies the intra prediction mode. - Variable nTbW: Specifies the width of the transformation block. - Variable nTbH: Specifies the height of the transformation block. - Variable nCbW: Specifies the width of the coding block. - Variable nCbH: Specifies the height of the coding block. - Variable cIdx: Specifies the color components of the current block. The output of this process is the predicted samples predSamples[x][y] for x=0..nTbW-1, y=0..nTbH-1. The predicted samples predSamples[ x ][ y ] are derived as follows: - If intra_lwip_flag[xTbCmp][yTbCmp] is equal to 1 and cIdx is equal to 0 then the affine linear weighted intra sample prediction process as specified in clause 8.4.4.2.X1 is called with input position (xTbCmp, yTbCmp), intra prediction mode predModeIntra, transform block width nTbW and height nTbH, and the output is predSamples. - Otherwise, the general intra sample prediction process as specified in 8.4.4.2.X1 is called with inputs the position (xTbCmp, yTbCmp), the intra prediction mode predModeIntra, the transform block width nTbW and height nTbH, the coding block width nCbW and height nCbH, and the variable cIdx, and the output is predSamples. 8.4.4.2.X1 Affine Linear Weighted Intra-Sample Prediction The inputs to this process are: - Sample position (xTbCmp, yTbCmp): Specifies the top-left sample of the current transform block relative to the top-left sample of the current picture. - Variable predModeIntra: Specifies the intra prediction mode. - Variable nTbW: Specifies the width of the transformation block. - Variable nTbH: Specifies the height of the transformation block. The output of this process is the predicted samples predSamples[x][y] for x=0..nTbW-1, y=0..nTbH-1. The size-type derivation process for the block as specified in clause 8.4.X.1 is called with the transformed block width, nTbW, and the transformed block height, nTbH, as input, and the output is assigned to the variable sizeId. The variables numModes, boundarySize, predW, predH, and predC are derived using sizeId as specified in Table 8-X4. Table 8-X4 - Specifications for number of modes, boundary sample sizes, and prediction sizes depending on sizedId
[0147] [Table 5] The transposed flag is derived as follows: isTransposed = ( predModeIntra > ( numModes / 2 ) ) ? 1 : 0 (8-X15) The flags needUpsBdryHor and needUpsBdryVer are derived as follows: needUpsBdryHor = ( nTbW > predW ) ? TRUE : FALSE (8-X16) needUpsBdryVer = ( nTbH > predH ) ? TRUE : FALSE (8-X17) The variables upsBdryW and upsBdryH are derived as follows: upsBdryW = ( nTbH > nTbW ) ? nTbW : predW (8-X18) upsBdryH = ( nTbH > nTbW ) ? predH : nTbH (8-X19) The variables lwipW and lwipH are derived as follows: lwipW = ( isTransposed = = 1) ? predH : predW (8-X20) lwipH = ( isTransposed = = 1) ? predW : predH (8-X21) For generation of reference samples refT[x] for x = 0..nTbW - 1 and reference samples refL[y] for y = 0..nTbH - 1, the reference sample derivation process as specified in clause 8.4.4.2.X2 is called with the sample positions (xTbCmp, yTbCmp), the transform block width nTbW and the transform block height nTbH as inputs and the top and left reference samples refT[x] for x = 0..nTbW - 1 and refL[y] for y = 0..nTbH - 1, respectively, as outputs. For generation of boundary samples p[x] for x = 0..2 * boundarySize - 1, the following applies: - the boundary reduction process as specified in clause 8.4.4.2.X3 is called with the block size nTbW, the reference sample refT, the boundary size boundarySize, the upsampling boundary flag needUpsBdryVer, the upsampling boundary size upsBdryW for the reference sample above, as inputs, the reduced boundary samples redT[ x ] for x = 0..boundarySize - 1, and the upsampling boundary samples upsBdryT[ x ] for x = 0..upsBdryW - 1. - the boundary reduction process as specified in clause 8.4.4.2.X3 is called with the block size nTbH, the reference sample refL, the boundary size boundarySize, the upsampling boundary flag needUpsBdryHor, the upsampling boundary size upsBdryH for the left reference sample, and with the reduced boundary samples redL[ x ] for x = 0..boundarySize - 1, and the upsampling boundary samples upsBdryL[ x ] for x = 0..upsBdryH - 1 as outputs. - The reduced top and left boundary samples redT and redL are assigned to the boundary sample array p as follows: - if isTransposed is equal to 1, then p[x] is set equal to redL[x] for x = 0..boundarySize - 1, and p[x + boundarySize] is set equal to redT[x] for x = 0..boundarySize - 1. - Otherwise, p[ x ] is set equal to redT[ x ] for x = 0..boundarySize - 1, and p[ x + boundarySize ] is set equal to redL[ x ] for x = 0..boundarySize - 1. For an intra-sample prediction process according to predModeIntra, the following step sequence applies: 1. The affine linear weighted samples for x = 0..lwipW - 1, y = 0..lwipH - 1 are derived as follows: - The variable modeId is derived as follows: modeId = predModeIntra - ( isTransposed = = 1) ? ( numModes / 2 ) : 0 (8-X22) - The weight matrix mWeight[ x ][ y ] for x = 0..2 * boundarySize - 1, y = 0..predC * predC - 1 is derived using sizeId and modeId as specified in Table 8-XX [TBD: Add Weight Matrix]. - The bias vector vBias[ y ] for y = 0..predC * predC - 1 is derived using sizeId and modeId as specified in Table 8-XX [TBD:Add Bias Vector]. The variable sW is derived using the sizeId and modeId as specified in Table 8-X5. The affine linear weighted sample predLwip[ x ][ y ] for x = 0..lwipW - 1, y = 0..lwipH - 1 is derived as follows:
[0148]
number
[0149] [Table 6] 8.4.4.2.X2 Reference sample derivation process The inputs to this process are: - sample position ( xTbY, yTbY ): Specifies the top-left luma sample of the current transform block relative to the top-left luma sample of the current picture. - Variable nTbW: Specifies the conversion block width. - Variable nTbH: Specifies the transformation block height. The output of this process are the top and left reference samples, refT[x] for x = 0..nTbW - 1 and refL[y] for y = 0..nTbH - 1, respectively. The adjacent samples, refT[ x ] for x = 0..nTbW - 1 and refL[ y ] for y = 0..nTbH - 1, are the constructed samples preceding the in-loop filter process and are derived as follows: - the upper and left adjacent luma positions (xNbT, yNbT) and (xNbL, yNbL) are ( xNbT, yNbT ) = ( xTbY + x, yTbY - 1 ) (8-X28) ( xNbL, yNbL ) = ( xTbY - 1, yTbY + y ) (8-X29) It is specified by - The availability derivation process for the block as specified in Clause 6.4.X [Ed. (BB): Adjacent Block Availability Checking Process tbd] is called with input the current luma position (xCurrr, yCurrr) set equal to (xTbY, yTbY) and the neighboring luma position above (xNbT, yNbT), and the output is assigned to availTop[ x ] for x = 0..nTbW-1. - The availability derivation process for the block as specified in Article 6.4.X [Ed. (BB): Adjacent Block Availability Checking Process tbd] is called with input the current luma position (xCurrr, yCurrr) set equal to (xTbY, yTbY) and the left adjacent luma position (xNbL, yNbL), and the output is assigned to availLeft[ y ], for y = 0..nTbH-1. The above reference sample refT[x] for x = 0..nTbW - 1 is derived as follows: - If all availTop[x] for x = 0..nTbW - 1 is equal to TRUE, then the sample at position (xNbT, yNbT) is assigned to refT[x] for x = 0..nTbW - 1. - Otherwise, if availTop
[0000] is equal to FALSE, then for all refT[ x ] for x = 0..nTbW - 1, 1 << ( BitDepth Y - 1). - Otherwise, the reference sample refT[x] for x = 0..nTbW - 1 is derived by the following sequence of steps: 1. The variable lastT is set equal to FALSE, the position x of the first element of the sequence availTop[x], for x = 1..nTbW - 1. 2. For all x = 0..lastT - 1, the sample at position (xNbT, yNbT) is assigned to refT[x]. 3. For all x = lastT..nTbW - 1, set refT[ x ] equal to refT[ lastT - 1 ]. - The left reference sample refL[y] for x = 0..nTbH - 1 is derived as follows: - if all availLeft[y] for y = 0..nTbH - 1 is equal to TRUE, the sample at position (xNbL, yNbL) is assigned to refL[y] for y = 0..nTbH - 1. - Otherwise, if availLeft
[0000] is equal to FALSE, then for all refL[ y ] for y = 0..nTbH - 1, 1 << ( BitDepth Y - 1). Otherwise, the reference samples refL[y] for y = 0..nTbH - 1 are derived by the following sequence of steps: 1. The variable lastL is set equal to FALSE, the position y of the first element of the sequence availLeft[y], for y = 1..nTbH - 1. 2. For all y = 0..lastL - 1, the sample at position (xNbL, yNbL) is assigned to refL[y]. 3. For all y = lastL..nTbH - 1, set refL[ y ] equal to refL[ lastL - 1 ]. Specification of the boundary reduction process The inputs to this process are: - Variable nTbX: Specifies the conversion block size. - Reference sample refX[ x ] : x = 0..nTbX - 1. - Variable boundarySize: Specifies the downsampled boundary size. - Flag needUpsBdryX: Specifies whether intermediate boundary samples are needed for upsampling. - Variable upsBdrySize: Specifies the boundary size for upsampling. The output of this process is the reduced boundary samples redX[x] for x = 0..boundarySize - 1, and the upsampled boundary samples upsBdryX[x] for x = 0..upsBdrySize - 1. The upsampling boundary samples upsBdryX[ x ] for x = 0..upsBdrySize - 1 are derived as follows: - If needUpsBdryX is equal to TRUE and upsBdrySize is less than nTbX, then apply:
[0150]
number
[0151]
number
[0152] [Table 7] Table 9-15 - Mapping of ctxInc to syntax elements using context coding bins
[0153] [Table 8] Table 9-16 - ctxInc Specification Using Left and Top Syntax Elements
[0154] [Table 9] <end>
[0155]
[0142] Overview of ALWP To predict samples of a rectangular block of width W and height H, Affine linear weighted intra prediction (ALWIP) takes as input one line of H reconstructed adjacent boundary samples on the left side of the block and one line of W reconstructed adjacent boundary samples on the top side of the block. If reconstructed samples are not available, they are generated similarly to regular intra prediction. ALWIP only applies to luma intra blocks. For chroma intra blocks, conventional intra coding modes are applied.
[0143] The generation of the prediction signal is based on the following three steps:
[0144] 1. Of the boundary samples, 4 samples are extracted by averaging if W=H=4, and 8 samples otherwise.
[0145] 2. A matrix-vector multiplication followed by the addition of an offset is performed using the averaged samples as input. The result is a reduced prediction signal for the set of subsampled samples in the original block.
[0146] 3. Predictions at the remaining positions are generated from the predictions for the subsampled set by linear interpolation, which is a single-step linear interpolation in each direction.
[0147] If the ALWIP mode should be applied, the ALWIP mode index predmode is signaled using MPM-list together with 3MPMS, where the derivation of MPM is performed using the intra modes of the top and left PUs as follows: Conventional intra prediction mode predmode Angular Three fixed tables map_angular_to_alwip that assign ALWIP modes to each of idx , idx∈{0,1,2}.
[0148] predmode ALWIP = map_angular_to_alwip idx [predmode Angular ]
[0149] -
[0151] For each PU of width W and height H, the index idx(PU) = idx(W,H) ∈{0,1,2} which specifies the set of three from which the ALWIP parameters are drawn.
[0152] The prediction unit PU above is available, belongs to the same CTU as the current PU, and is in intra mode, and idx(PU) = idx(PU above ) and ALWIP mode predmode ALWIP above Together with PU above If ALWIP applies to the following, set it as follows:
[0153] modeALWIPabove=predmode ALWIP above
[0154] If the upper PU is available, belongs to the same CTU as the current PU, and is in intra mode, the conventional intra prediction mode predmode Angular above If applies to the PU above, set it as follows:
[0155] mode ALWIP above = map_angular_to_alwip idx(PUabove) [predmode Angular above ]
[0156] Otherwise, set it as follows:
[0157] mode ALWIP above = -1
[0158] This means that the mode is unavailable. Similarly, the mode mode can be used without the constraint that the left PU must belong to the same CTU as the current PU. ALWIP left Derive.
[0159] Finally, there are three fixed default lists: idx , idx∈{0,1,2}, each of which contains three different ALWIP modes. idx(PU) and mode ALWIP above , mode ALWIP left From, three different MPMs can be constructed by replacing the default values with -1 and omitting the repetitions.
[0160] For luma MPM list derivation, adjacent luma blocks are in ALWIP mode predmode ALWIP When a block is encountered that uses the traditional intra prediction mode predemode Angular will be treated as if they were using
[0161] predemode Angular = map_alwip_to_angular idx(PU) [predmode ALWIP ]
[0162] 3 Conversion in VVC 3.1 Multiple Transform Selection (MTS) In addition to the DCT-II adopted in HEVC, the Multiple Transform Selection (MTS) method is used for residual coding of both inter- and intra-coded blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII.
[0163] 3.2 Reduced Quadratic Transformation (RST) proposed in JVET-N0193 The reduced secondary transform (RST) applies 16x16 and 16x64 non-separable transforms to 4x4 and 8x8 blocks, respectively. The primary forward and inverse transforms are still performed in the same manner as the two 1-D horizontal / vertical transform passes. The secondary forward and inverse transforms are separate process steps from the primary transform. For the encoder, the primary forward transform is performed first, followed by the secondary forward transform and quantization, and then the CABAC bit encoding. For the decoder, CABAC bit decoding and inverse quantization, and then the secondary inverse transform are performed first, followed by the primary inverse transform. The RST is only applied to intra-coded TUs, both intra-slice and inter-slice.
[0164] 3.3 Unified MPM List for Intra-mode Coding in JVET-N0185 A unified 6-MPM list is proposed for intra blocks, whether Multiple Reference Line (MRL) and Intra Sub-Partition (ISP) coding tools are applied or not. The MPM list is constructed based on the top-left intra modes of the left and top adjacent blocks, as in VTM4.0. Assuming the left mode is Left and the above block's mode is Above, the unified MPM list is constructed as follows: If no adjacent blocks are available, the Intra mode defaults to Planar. If both modes Left and Above are non-angle modes: a. MPM list → {Planar, DC, V, H, V-4, V+4} If one of the modes Left and Above is a non-angle mode and the other is a non-angle mode: a. Set mode Max as the greater of Left and Above. b. MPM list → {Planar, Max, DC, Max -1, Max +1, Max -2} If both Left and Above are in angle mode and they are different: a. Set mode Max as the greater of Left and Above. b. If the difference between modes Left and Above is between 2 and 62 inclusive: i. MPM list → {Planar, Left, Above, DC, Max-1, Max+1} c. In other cases i. MPM list → {Planar, Left, Above, DC, Max-2, Max+2} If both Left and Above are in angle mode and they are the same: a. MPM list → {Planar, Left, Left-1, Left+1, DC, Left-2}
[0165] In addition, the first bin of the MPM index codeword is CABAC context coded. A total of three contexts are used depending on whether the current intra block is MRL-enabled, ISP-enabled, or a normal intra block.
[0166] The left and top neighboring blocks used in the unified MPM list construction are A2 and B2 as shown in FIG.
[0167] One MPM flag is coded first. If the block is coded with one of the modes in the MPM list, the MPM index is coded further. Otherwise, the indices of the remaining modes (excluding MPM) are coded.
[0168] 4 Examples of shortcomings in existing implementations The design of ALWIP in JVET-N0217 has the following problems:
[0169] 1) In the JVET meeting in March 2019, a unified 6-MPM list generation was adopted for MRL, ISP, and normal intra modes. However, the Affine Linear Weighted Prediction mode uses a different 3-MPM list configuration, which makes the MPM list configuration more complicated. The complex MPM list configuration may impair the decoder throughput, especially for small blocks such as 4x4 samples.
[0170] 2) ALWIP only applies to the luma component of a block. For the chroma components of an ALWIP coded block, a chroma mode index is coded and transmitted to the decoder, which may result in unnecessary signaling.
[0171] 3) The interaction of ALWIP with other coding tools should be considered. 4) When calculating upsBdryX using the following formula:
[0172]
number
[0173] 5) When upsampling the prediction samples, no rounding is applied.
[0174] 6) In the deblocking process, ALWIP coded blocks are treated as normal intra blocks.
[0175] 5. Example of a method for matrix-based intra-coding Embodiments of the techniques disclosed herein overcome shortcomings of existing implementations, thereby resulting in higher coding efficiency but less computational complexity for video coding. The matrix-based intra prediction method for video coding, as described herein, may improve both existing and future video coding standards, as elucidated in the following examples in which various implementations are described. The examples of the disclosed techniques provided below illustrate general concepts and are not intended to be construed as limiting. In the examples, various features described in these examples may be combined, unless expressly indicated to the contrary.
[0176] In the following description, intra prediction mode refers to angular intra prediction mode (including DC, planar, CCLM, and other possible intra prediction modes); intra mode refers to normal intra mode, MRL, ISP, or ALWIP.
[0177] In the following description, "other intra modes" may refer to one or more intra modes other than ALWIP, such as normal intra mode, MRL, or ISP.
[0178] In the following discussion, SatShift(x,n) is defined as:
[0179]
number
[0180] Shift(x, n) is defined as Shift(x,n) = (x + offset0)>>n.
[0181] In one example, offset0 and / or offset1 are (1<<n)> In another example, offset0 and / or offset1 are set to 0.
[0182] Another example is offset0=offset1= ((1<<n)> >1)-1 or ((1<<(n-1)))-1. In another example, offset0=offset1= ((1<<n)> >1)-1 or ((1<<(n-1)))-1. Clip3(min, max, x) is defined as follows:
[0183]
number
[0184] Building an MPM list for ALWIP 1. It is proposed that all or part of the MPM list for ALWIP is constructed according to all or part of the procedures for constructing MPM lists for non-ALWIP intra modes (e.g. normal intra mode, MRL, ISP). In one example, the size of the MPM list for ALWIP may be the same as the size of the MPM list for non-ALWIP intra mode. i. For example, the size of the MPM list is 6 for both ALWIP and non-ALWIP intra modes. b. In one example, the MPM list for ALWIP may be derived from the MPM list for non-ALWIP intra-mode. In one example, an MPM list for non-ALWIP intra modes may be built first, after which some or all of them may be converted to MPM and then added to the MPM list for ALWIP coded blocks. 1) Furthermore, alternatively, pruning may be applied when adding transformed MPMs to the MPM list of an ALWIP coded block. 2) The default mode may be added to the MPM list of ALWIP-coded blocks. In one example, a default mode may be added before the one converted from the MPM list of non-ALWIP intra modes. b. Alternatively, the default mode may be added after the one converted from the MPM list of non-ALWIP intra modes. c. Alternatively, the default modes may be added in an interleaved manner with those converted from the MPM list of non-ALWIP intra modes. d. In one example, the default mode may be fixed to be the same for all block types. e. Alternatively, the default mode may be determined according to coded information such as availability of neighboring blocks, mode information of neighboring blocks, block dimensions. ii. In one example, an intra-prediction mode in an MPM list for non-ALWIP intra modes may be converted to a corresponding ALWIP intra-prediction mode when it is populated into an ALWIP MPM list. 1) Alternatively, all intra-prediction modes in the MPM list for non-ALWIP intra modes may be converted to corresponding ALWIP intra-prediction modes before being used to construct the ALWIP MPM list. 2) Alternatively, in the case where the MPM list of non-ALWIP intra modes may be further used to derive the MPM list of ALWIP, all candidate intra prediction modes (which may include intra prediction modes from neighboring blocks and default intra prediction modes such as Planar and DC) may be converted to corresponding ALWIP intra prediction modes before being used to construct the MPM list of non-ALWIP intra modes. 3) In one example, two transformed ALWIP intra prediction modes may be compared. In one example, if they are the same, only one of them can be populated into the MPM list of ALWIP. b. In one example, if they are the same, only one of them can be populated in the non-ALWIP intra-mode MPM list. In one example, K intra prediction modes out of S in the MPM list of the non-ALWIP intra modes may be selected as the MPM list of the ALWIP modes, e.g., K is equal to 3 and S is equal to 6. 1) In one example, the first K intra prediction modes of the MPM list of the non-ALWIP intra modes may be selected as the MPM list of the ALWIP modes. 2. It is proposed that one or more neighboring blocks used to derive an MPM list for ALWIP may also be used to derive an MPM list for a non-ALWIP intra mode (e.g. normal intra mode, MRL, or ISP). In one example, the adjacent block to the left of the current block used to derive the ALWIP MPM list should be the same as the one used to derive the non-ALWIP intra-mode MPM list. i. Assuming the top-left corner of the current block is (xCb, yCb) and the width and height of the current block are W and H, in one example, the left neighboring block used to derive the MPM list for both ALWIP and non-ALWIP intra modes may cover position (xCb-1, yCb). In another example, the left neighboring block used to derive the MPM list for both ALWIP and non-ALWIP intra modes may cover position (xCb-1, yCb+H-1). ii. For example, the left neighboring block and the top neighboring block used in the unified MPM list construction are A2 and B2 as shown in FIG. b. In one example, the adjacent blocks above the current block used to derive the ALWIP MPM list should be the same as those used to derive the non-ALWIP intra-mode MPM list. i. Assuming the top-left corner of the current block is (xCb, yCb) and the width and height of the current block are W and H, in one example, the upper neighboring block used to derive the MPM list for both ALWIP and non-ALWIP intra modes may cover position (xCb, yCb-1). In another example, the upper neighboring block used to derive the MPM list for both ALWIP and non-ALWIP intra modes may cover position (xCb+W-1, yCb-1). ii. For example, the left neighboring block and the top neighboring block used in the unified MPM list construction are A1 and B1 as shown in FIG. 3. It is proposed that the MPM list for ALWIP may be constructed in different ways depending on the width and / or height of the current block. In one example, for different block sizes, different adjacent blocks may be accessed. 4. It is proposed that the MPM list for ALWIP and the MPM list for non-ALWIP intra-mode may be configured with different parameters, although the procedure is the same. In one example, in an MPM list construction procedure for a non-ALWIP intra mode, K out of S intra prediction modes may be derived for an MPM list used in an ALWIP mode, where K is equal to 3 and S is equal to 6. i. In one example, the first K intra-prediction modes in the MPM list construction procedure may be derived for the MPM list used in the ALWIP mode. b. In one example, the first modes of the MPM list may be different. i. For example, the first mode in the MPM list for non-ALWIP intra modes may be Planar, but it may also be Mode X0 in the MPM list for ALWIP. 1) In one example, X0 may be an ALWIP intra prediction mode converted from Planar. c. In one example, the stuffing modes in the MPM list may be different. i. For example, the first three stuffing modes in the MPM list for non-ALWIP intra modes may be DC, Vertical and Horizontal, but for ALWIP the MPM list may be Modes X1, X2, X3. 1) In one example, X1, X2, X3 may be different for different sizeId. ii. In one example, the number of stuffing modes may be different. d. In one example, adjacent modes in an MPM list may be different. For example, the normal intra prediction modes of neighboring blocks are used to build an MPM list for non-ALWIP intra prediction modes, and then converted to ALWIP intra prediction modes to build an MPM list for ALWIP modes. e. In one example, the shifted modes in the MPM list may be different. For example, X+K0 (where X is a normal intra prediction mode and K0 is an integer) may be put into the MPM list for non-ALWIP intra prediction modes, and Y+K1 (where Y is an ALWIP intra prediction mode and K1 is an integer) may be put into the MPM list for ALWIP, where K0 may be different from K1. 1) In one example, K1 may depend on the width and height. 5. When building the MPM list for the current block in non-ALWIP intra mode, it is proposed that if a neighboring block is coded with ALWIP, the neighboring block is treated as unavailable. Alternatively, when constructing the MPM list for the current block in non-ALWIP intra modes, the neighboring blocks are treated as being coded in a predefined intra prediction mode (e.g., Planar) if they are coded in ALWIP. 6. When constructing the MPM list for the current block in ALWIP mode, it is proposed that if a neighboring block is coded in non-ALWIP intra mode, the neighboring block is treated as unavailable. Alternatively, when constructing the MPM list for the current block in ALWIP mode, the neighboring blocks are treated as being coded in a predefined ALWIP intra-prediction mode X if they are coded in a non-ALWIP intra-mode. i. In one example, X may depend on a block dimension, such as width and / or height. 7. It is proposed to remove the storage of the ALWIP flag from the line buffer. In one example, if the second block to be accessed is located in a different LCU / CTU row / region compared to the current block, the conditioned check of whether the second block is coded with ALWIP is skipped. b. In one example, if the second block to be accessed is located in a different LCU / CTU row / region compared to the current block, the second block is treated similarly to the non-ALWIP mode, such that it is treated as a normal intra-coded block. 8. When encoding the ALWIP flag, at most K contexts (K>=0) may be used. In one example, K=1. 9. Instead of directly storing the mode index associated with the ALWIP mode, it is proposed to store the transformed intra-prediction mode of the ALWIP coded block. a. In one example, the decoded mode index associated with one ALWIP coded block is mapped to a normal intra mode according to map_alwip_to_angular as described in Section 2.5.7. b. Additionally or alternatively, the storage of the ALWIP flag may be eliminated entirely. c. Additionally or alternatively, ALWIP mode storage may be eliminated entirely. d. Further, alternatively, the condition check whether one adjacent / current block is coded with the ALWIP flag may be skipped. e. Additionally, alternatively, mode conversions assigned to ALWIP coded blocks and normal intra predictions associated with one accessed block may be skipped. ALWIP for different color components 10. It is proposed that an estimated chroma intra mode (e.g., DM mode) may always be applied if the corresponding luma block is coded in ALWIP mode. In one example, a chroma intra mode is inferred to be a DM mode without signaling if the corresponding luma block is coded in ALWIP mode. b. In one example, the corresponding luma block may be one that covers corresponding samples of chroma samples located at a given location (e.g., top-left of the current chroma block, center of the current chroma block). c. In one example, the DM mode may be derived according to the intra prediction mode of the corresponding luma block, for example, by mapping the (ALWIP) mode to one of the normal intra modes. 11. If the corresponding luma block of a chroma block is coded in ALWIP mode, several DM modes may be derived. 12. It is proposed that a special mode be assigned to a chroma block if one corresponding luma block is coded in ALWIP mode. In one example, the special mode is defined to be a given normal intra-prediction mode regardless of the intra-prediction mode associated with the ALWIP coded block. b. In one example, various methods of intra prediction may be assigned to this special mode. 13. It is proposed that ALWIP may be applied to chroma components. In one example, the matrices and / or bias vectors may be different for different color components. b. In one example, matrices and / or bias vectors may be predefined jointly for Cb and Cr. i. In one example, the Cb and Cr components may be concatenated. ii. In one example, the Cb and Cr components may be interleaved. c. In one example, a chroma component may share the same ALWIP intra prediction mode as the corresponding luma block. i. In one example, if the corresponding luma block applies the ALWIP mode and the chroma block is coded in DM mode, the same ALWIP intra prediction mode is applied to the chroma components. ii. In one example, the same ALWIP intra prediction mode is applied to the chroma components, and subsequent linear interpolation can be skipped. iii. In one example, the same ALWIP intra prediction mode is applied to the chroma components along with a subsampled matrix and / or bias vector. d. In one example, the number of ALWIP intra prediction modes for different components may be different. i. For example, the number of ALWIP intra prediction modes for a chroma component may be less than the number for the luma component of the same block width and height. Applying ALWIP 14. It is proposed that it may be possible to signal whether ALWIP is applicable. a. For example, it may be signaled at the sequence level (e.g., in the SPS), at the picture level (e.g., in the PPS or picture header), at the slice level (e.g., in the slice header), at the tile group level (e.g., in the tile group header), at the tile level, at the CTU row level, or at the CTU level. b. For example, if ALWIP cannot be applied, intra_lwip_flag may not be signaled and may be presumed to be 0. 15. It is proposed that whether ALWIP can be applied may depend on the block width (W) and / or height (H). c. For example, ALWIP may not be applied when W >= T1 (or W > T1) and H >= T2 (or H > T2). For example, T1 = T2 = 32. i. For example, ALWIP may not be applied when W <= T1 (or W < T1) and H <= T2 (or H < T2). For example, T1 = T2 = 32. d. For example, ALWIP may not be applied when W >= T1 (or W > T1) or H >= T2 (or H > T2). For example, T1 = T2 = 32. i. For example, ALWIP may not be applied when W <= T1 (or W < T1) or H <= T2 (or H < T2). For example, T1 = T2 = 32, or T1 = T2 = 8. e. For example, ALWIP may not be applied when W + H >= T (or W * H > T). For example, T = 256. i. For example, ALWIP may not be applied when W + H <= T (or W + H < T). For example, T = 256. f. For example, ALWIP may not be applied when W * H >= T (or W * H > T). For example, T = 256. i. For example, ALWIP may not be applied when W * H <= T (or W * H < T). For example, T = 256. g. For example, if ALWIP cannot be applied, intra_lwip_flag may not be signaled and may be presumed to be 0. Computational issues in ALWP 16. It is proposed that any shift operation involving ALWIP can only be a left shift or a right shift by an amount equal to S, where S must be greater than or equal to 0. In one example, if S is greater than or equal to 0, the right shift operation may be different. i. In one example, upsBdryX[x] should be calculated as follows:
[0185]
number
[0186]
number
[0185] The above-described embodiments may be incorporated in the context of the methods described below, for example methods 1100 to 1400 and 2000 to 3800, which may be implemented in a video encoder and / or decoder.
[0186] Figure 11 shows a flowchart of an example method for video processing. The method 1100 includes, at step 1102, determining that a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode.
[0187] Based on the determination, the method 1100 includes, at step 1104, constructing a Most Probable Mode (MPM) list for the ALWIP mode based at least in part on the MPM list for the non-ALWIP intra mode.
[0188] The method 1100 includes, at step 1106, performing a conversion between the current video block and a bitstream representation of the current video block based on the ALWIP mode MPM list.
[0189] In some embodiments, the size of the MPM list in ALWIP mode is the same as the size of the MPM list in non-ALWIP intra mode. In one example, the size of the MPM list in ALWIP mode is 6.
[0190] In some embodiments, the method 1100 further includes inserting the default mode into the MPM list for the ALWIP mode. In one example, the default mode is inserted before a portion of the MPM list for the ALWIP mode that is based on the MPM list for the non-ALWIP intra mode. In another example, the default mode is inserted after a portion of the MPM list for the ALWIP mode that is based on the MPM list for the non-ALWIP intra mode. In yet another example, the default mode is inserted in an interleaved manner with a portion of the MPM list for the ALWIP mode that is based on the MPM list for the non-ALWIP intra mode.
[0191] In some embodiments, construction of the MPM list for ALWIP mode and the MPM list for non-ALWIP intra mode is based on one or more adjacent blocks.
[0192] In some embodiments, construction of the ALWIP mode MPM list and the non-ALWIP intra mode MPM list is based on the height or width of the current video block.
[0193] In some embodiments, construction of the MPM list for the ALWIP mode is based on a first set of parameters that differ from a second set of parameters used to construct the MPM list for the non-ALWIP intra mode.
[0194] In some embodiments, the method 1100 further includes determining that a neighboring block of the current video block is coded in ALWIP mode and designating the neighboring block as unavailable when constructing the MPM list for non-ALWIP intra modes.
[0195] In some embodiments, the method 1100 further includes determining that a neighboring block of the current video block is coded in a non-ALWIP intra mode and designating the neighboring block as unavailable when constructing the MPM list for the ALWIP mode.
[0196] In some implementations, the non-ALWIP intra mode is based on a normal intra mode, a multiple reference line (MRL) intra prediction mode, or an intra sub-partition (ISP) tool.
[0197] 12 illustrates a flowchart of an example method for video processing. The method 1200 includes, at step 1210, determining that a luma component of a current video block is coded using an affine linear weighted intra prediction (ALWIP) mode.
[0198] The method 1200 includes, at step 1220, estimating a chroma intra mode based on the determination.
[0199] The method 1200 includes, at step 1230, performing a conversion between the current video block and a bitstream representation of the current video block based on the chroma intra mode.
[0200] In some implementations, the luma component covers a given chroma sample of a chroma component. In one example, the given chroma sample is the top left sample or the center sample of the chroma component.
[0201] In some implementations, the estimated chroma intra mode is a DM mode.
[0202] In some embodiments, the estimated chroma intra mode is an ALWIP mode.
[0203] In some embodiments, the ALWIP mode is applied to one or more chroma components of the current video block.
[0204] In some embodiments, different matrices or bias vectors of the ALWIP mode are applied to different color components of the current video block. In one example, different matrices or bias vectors are defined together in advance for the Cb and Cr components. In another example, the Cb and Cr components are concatenated. In yet another example, the Cb and Cr components are interleaved.
[0205] FIG. 13 shows a flowchart of an exemplary method for video processing. Method 1300 includes, at step 1302, determining that the current video block is to be coded using an Affine Linear Weighted Intra Prediction (ALWIP) mode.
[0206] Method 1300 includes, based on the determination, at step 1304, performing a conversion between the current video block and a bitstream representation of the current video block.
[0207] In some embodiments, the determination is based on signaling in a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), a slice header, a tile group header, a tile header, a Coding Tree Unit (CTU) row, or a CTU region.
[0208] In some embodiments, the determination is based on the height (H) or width (W) of the current video block. In one example, W > T1 or H > T2. In another example, W ≥ T1 or H ≥ T2. In yet another example, W < T1 or H < T2. In yet another example, W ≤ T1 or H ≤ T2. In yet another example, T1 = 32 and T2 = 32.
[0209] In some embodiments, the determination is based on the height (H) or width (W) of the current video block. In one example, W+H≦T. In another example, W+H≧T. In yet another example, W×H≦T. In yet another example, W×H≧T. In yet another example, T=256.
[0210] 14 illustrates a flowchart of an example method for video processing. The method 1400 includes, at step 1402, determining that a current video block is to be coded using a coding mode other than an affine linear weighted intra prediction (ALWIP) mode.
[0211] The method 1400 includes, based on the determination, performing a conversion at step 1404 between the current video block and a bitstream representation of the current video block.
[0212] In some embodiments, the coding mode is a combined intra- and inter-prediction (CIIP) mode, and method 1400 further includes performing a selection between the ALWIP mode and a normal intra prediction mode. In one example, performing the selection is based on explicit signaling in a bitstream representation of the current video block. In another example, performing the selection is based on a predetermined rule. In yet another example, the predetermined rule always selects the ALWIP mode if the current video block is coded using a CIIP mode. In yet another example, the predetermined rule always selects the normal intra prediction mode if the current video block is coded using a CIIP mode.
[0213] In some embodiments, the coding mode is a cross-component linear model (CCLM) prediction mode. In one example, the downsampling procedure of the ALWIP mode is based on the downsampling procedure of the CCLM prediction mode. In another example, the downsampling procedure of the ALWIP mode is based on a first set of parameters, and the downsampling procedure of the CCLM prediction mode is based on a second set of parameters different from the first set of parameters. In yet another example, the downsampling procedure for the ALWIP mode or the CCLM prediction mode includes at least one of a selection of downsampled positions, a selection of a downsampling filter, a rounding operation, or a clipping operation.
[0214] In some embodiments, the method 1400 further includes applying one or more of a shrinking quadratic transform (RST), a quadratic transform, a rotation transform, or a non-separable quadratic transform (NSST).
[0215] In some embodiments, the method 1400 further includes applying block-based differential pulse coded modulation (DPCM) or residual DPCM.
[0216] 6. Examples of implementation of the disclosed technology FIG. 15 is a block diagram of a video processing device 1500. The device 1500 may be used to implement one or more methods described herein. The device 1500 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 1500 may include one or more processors 1502, one or more memories 1504, and video processing hardware 1506. The processor 1502 may be configured to implement one or more methods described herein, including but not limited to methods 1100 to 1400 and 2000 to 3800. The memories 1504 may be used to store data and codes used to implement the methods and techniques described herein. The video processing hardware 1506 may be used to implement some techniques described herein in a hardware circuit.
[0217] In some embodiments, the video coding method may be implemented using an apparatus implemented on a hardware platform such as that described in connection with FIG.
[0218] FIG. 16 is a block diagram illustrating an example video processing system 1600 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1600. System 1600 may include an input 1602 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or in a compressed or encoded format. Input 1602 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[0219] The system 1600 may include a coding component 1604 capable of implementing various coding or encoding methods described herein. The coding component 1604 may reduce the average bit rate of the video from the input 1602 to the output of the coding component 1604 to generate a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the coding component 1604 may be stored or transmitted via a communication connection as represented by component 1606. The stored or communicated bitstream (or coded) representation of the video received at the input 1602 may be used by component 1608 to generate pixel values or displayable video that are sent to the display interface 1610. The process of generating a user viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, although certain video processing operations are referred to as "coding" operations or tools, it will be understood that the coding tools or operations are used at the encoder and the corresponding decoding tools or operations that process the results of the coding in the reverse direction are performed at the decoder.
[0220] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI), Displayport, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The technology described in this document may be embodied in a variety of electronic devices, such as a mobile phone, a laptop, a smart phone, or other device capable of performing digital data processing and / or video display.
[0221] Some embodiments of the disclosed techniques include making a judgment or decision to enable a video processing tool or mode. In one example, if a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in processing blocks of video, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, conversion of blocks of video to a bitstream representation of video will use the video processing tool or mode if enabled based on the judgment or decision. In another example, if a video processing tool or mode is enabled, a decoder will process the bitstream with information that the bitstream has been modified based on the video processing tool or mode. That is, conversion of a bitstream representation of video to blocks of video will be performed using the video processing tool or mode that is enabled based on the judgment or decision.
[0222] Some embodiments of the disclosed techniques include making a judgment or decision to disable a video processing tool or mode. In one example, if a video processing tool or mode is disabled, an encoder will not use the tool or mode in converting blocks of video to a bitstream representation of the video. In another example, if a video processing tool or mode is disabled, a decoder will process the bitstream with information that the bitstream has not been modified using the video processing tool or mode that was disabled based on the judgment or decision.
[0223] FIG. 17 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As shown in FIG. 17, the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 can generate encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0224] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include a coded picture and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the I / O interface 116 over the network 130a. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0225] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0226] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, in which case the destination device is configured to interface with the external display device.
[0227] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.
[0228] FIG. 18 is a block diagram of an example video encoder 200, which may be the video encoder 114 in the system 100 shown in FIG.
[0229] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 18, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0230] Functional components of the video encoder 200 may include a partition unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0231] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0232] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are depicted separately in the example of FIG. 18 for illustrative purposes.
[0233] The partition unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support a variety of video block sizes.
[0234] The mode selection unit 203 selects one of the coding modes, inter or intra, based on, for example, an error result, and provides the resulting intra-coded or inter-coded block to the residual generation unit 207 for residual block data generation and to the reconstruction unit 212 for reconstruction of the coded block to be used as a reference picture. In some examples, the mode selection unit 203 may select a combined intra & inter prediction (CIIP) mode, in which prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy) in the case of inter prediction.
[0235] To perform inter prediction with respect to the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0236] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on, for example, whether the current video block is an I slice, a P slice, or a B slice.
[0237] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for reference picture blocks for the current video block. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0238] In another example, motion estimation unit 204 may perform bi-directional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 for a reference video block for the current video block and may search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indexes indicating reference pictures in lists 0 and 1 that contain the reference video blocks and motion vectors indicating spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference index and the motion vector of the current video block as motion information of the current video block. Motion compensation unit 205 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0239] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.
[0240] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block with reference to motion information of other video blocks. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0241] In one example, motion estimation unit 204 may specify a value in a syntax structure associated with a current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0242] In another example, motion estimation unit 204 may identify another video block and a motion vector differential (MVD) in a syntax structure associated with the current video block. The motion vector differential indicates a difference between the motion vector of the current video block and the motion vector of the designated video block. Video decoder 300 may determine the motion vector of the current video block using the motion vector of the designated video block and the motion vector differential.
[0243] As mentioned above, video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0244] Intra prediction unit 206 may perform intra prediction on the current video block. If intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the video block to be predicted and various syntax elements.
[0245] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., as indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.
[0246] In other examples, such as in skip mode, for the current video block, residual data for the current video block may not exist and residual generation unit 207 may not perform the subtraction operation.
[0247] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to a residual video block related to the current video block.
[0248] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0249] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block related to the current block, for storage in buffer 213.
[0250] After reconstruction unit 212 reconstructs the video blocks, a loop filtering operation may be performed to reduce video blocking artifacts in the video blocks.
[0251] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. Once the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0252] FIG. 19 is a block diagram illustrating an example of a video decoder 300, which may be the video decoder 114 in the system 100 shown in FIG.
[0253] The video decoder 300 may be configured to perform any or all of the techniques described in this disclosure. In the example of FIG. 19, the video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0254] 19, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. The video decoder 300 may, in some examples, perform a decoding path that is generally the inverse of the encoding path described with respect to the video encoder 200 (FIG. 18).
[0255] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy encoded video data, and from the entropy decoded video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 may determine such information, for example, by implementing AMVP and merge mode.
[0256] The motion compensation unit 302 may generate a motion compensated block by performing interpolation, possibly based on an interpolation filter. An identifier for an interpolation filter used with sub-pixel precision may be included in the syntax element.
[0257] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of the reference block using an interpolation filter as used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 according to received syntax information and generate the predictive block using the interpolation filter.
[0258] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to code the frames and / or slices of the coded video sequence, partition information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is coded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.
[0259] The intra prediction unit 303 may form a prediction block from spatially adjacent blocks, for example, using an intra prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0260] Reconstruction unit 306 may sum the residual blocks with corresponding prediction blocks generated by motion compensation unit 202 or intra prediction unit 303 to form decoded blocks. If desired, a deblocking filter may be applied to filter the decoded blocks to remove blockiness artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also generates decoded video for presentation on a display device.
[0261] In some embodiments, the ALWIP mode or the MIP mode is used to calculate a prediction block of a current video block by performing a boundary downsampling (or averaging operation) on previously coded samples of the video, performing a matrix-vector multiplication operation, and then selectively (or optionally) performing an upsampling operation (or linear interpolation operation). In some embodiments, the ALWIP mode or the MIP mode is used to calculate a prediction block of a current video block by performing a boundary downsampling (or averaging operation) on previously coded samples of the video, and then performing a matrix-vector multiplication operation. In some embodiments, the ALWIP mode or the MIP mode can perform an upsampling operation (or linear interpolation operation) after performing a matrix-vector multiplication operation.
[0262] 20 illustrates an example flow chart of an example method 2000 for matrix-based intra prediction. Operation 2002 includes generating a first most probable mode (MPM) list based on a rule utilizing a first procedure for transforming between a current video block of a video and a coded representation of the current video block. Operation 2004 includes performing a transform between the current video block and the coded representation of the current video block using the first MPM list, the transform of the current video block using a matrix-based intra prediction (MIP) mode, in which a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation, the rule specifying that the first procedure used to generate the first MPM list is the same as a second procedure used to generate a second MPM list for transforming other video blocks of the video coded using a non-MIP intra mode different from the MIP mode, and at least a portion of the first MPM list is generated based on at least a portion of the second MPM list.
[0263] In some embodiments of method 2000, the size of the first MPM list for MIP mode is the same as the size of the second MPM list for non-MIP intra mode. In some embodiments of method 2000, the size of the first MPM list for MIP mode and the size of the second MPM list for non-MIP intra mode are six. In some embodiments of method 2000, the second MPM list for non-MIP intra mode is constructed before the first MPM list for MIP mode is constructed. In some embodiments of method 2000, all or a portion of the second MPM list for non-MIP intra mode is converted into an MPM that is added to a portion of the first MPM list for MIP mode.
[0264] In some embodiments of the method 2000, a subset of the MPMs are added to a portion of the first MPM list for the MIP mode by pruning the MPMs. In some embodiments, the method 2000 further includes adding a default intra-prediction mode to the first MPM list for the MIP mode. In some embodiments of the method 2000, the default intra-prediction mode is added to the first MPM list for the MIP mode before a portion of the first MPM list for the MIP mode that is based on a portion of the second MPM list for the non-MIP intra mode. In some embodiments of the method 2000, the default intra-prediction mode is added to the first MPM list for the MIP mode after a portion of the first MPM list for the MIP mode that is based on a portion of the second MPM list for the non-MIP intra mode. In some embodiments of method 2000, the default intra-prediction mode is added to the first MPM list for the MIP mode in an alternating manner with a portion of the first MPM list for the MIP mode based on a portion of the second MPM list for the non-MIP intra mode. In some embodiments of method 2000, the default intra-prediction mode is the same for multiple types of video blocks. In some embodiments of method 2000, the default intra-prediction mode is determined based on coded information of the current video block. In some embodiments of method 2000, the coded information includes availability of neighboring video blocks, mode information of neighboring video blocks, or block dimensions of the current video block.
[0265] In some embodiments of the method 2000, one intra prediction mode in the second MPM list for the non-MIP intra mode is converted to its corresponding MIP mode to obtain a transformed MIP mode that is added to the first MPM list for the MIP mode. In some embodiments of the method 2000, all intra prediction modes in the second MPM list for the non-MIP intra mode are converted to their corresponding MIP modes to obtain a plurality of transformed MIP modes that are added to the first MPM list for the MIP mode. In some embodiments of the method 2000, all intra prediction modes are converted to their corresponding MIP modes to obtain a plurality of transformed MIP modes that are used to construct the second MPM list for the non-MIP intra mode. In some embodiments of the method 2000, all intra prediction modes include intra prediction modes from neighboring video blocks of the current video block and a default intra prediction mode. In some embodiments of the method 2000, the default intra prediction mode includes a planar mode and a direct current (DC) mode.
[0266] In some embodiments of the method 2000, two intra prediction modes in the second MPM list for the non-MIP intra mode are transformed to their corresponding MIP modes to obtain two transformed MIP modes, and one of the two transformed MIP modes is added to the first MPM list for the MIP mode in response to the two transformed MIP modes being the same. In some embodiments of the method 2000, two intra prediction modes in the second MPM list for the non-MIP intra mode are transformed to their corresponding MIP modes to obtain two transformed MIP modes, and one of the two transformed MIP modes is added to the second MPM list for the non-MIP intra mode in response to the two transformed MIP modes being the same. In some embodiments of the method 2000, the second MPM list for the non-MIP intra mode includes S intra prediction modes, and K of the S intra prediction modes are selected to be included in the first MPM list for the MIP mode. In some embodiments of method 2000, K is 3 and S is 6. In some embodiments of method 2000, the first K intra prediction modes in the second MPM list for non-MIP intra mode are selected for inclusion in the first MPM list for MIP mode. In some embodiments of method 2000, the first MPM list for MIP mode and the second MPM list for non-MIP intra mode are derived or constructed based on one or more neighboring video blocks of the current video block. In some embodiments of method 2000, the first MPM list for MIP mode and the second MPM list for non-MIP intra mode are derived or constructed based on the same neighboring video blocks located to the left of the current video block.
[0267] In some embodiments of the method 2000, the same neighboring video block located to the left is in front of and aligned with the current video block. In some embodiments of the method 2000, the same neighboring video block is to the left and below the current video block. In some embodiments of the method 2000, the first MPM list for MIP mode and the second MPM list for non-MIP intra mode are derived or constructed based on the same neighboring video block, and the same neighboring video block is in a bottom-most position to the left of the current video block, or the same neighboring video block is in a right-most position above the current video block. In some embodiments of the method 2000, the first MPM list for MIP mode and the second MPM list for non-MIP intra mode are derived or constructed based on the same neighboring video block located above the current video block.
[0268] In some embodiments of the method 2000, the above same neighboring video block is directly above and aligned with the current video block. In some embodiments of the method 2000, the same neighboring video block is to the left and above the current video block. In some embodiments of the method 2000, the first MPM list for the MIP mode and the second MPM list for the non-MIP intra mode are derived or constructed based on the same neighboring video block, and the same neighboring video block is in a left-most position above the current video block, or the same neighboring video block is in a top-most position to the left of the current video block. In some embodiments of the method 2000, the non-MIP intra mode is based on an intra prediction mode, a multiple reference line (MRL) intra prediction mode, or an intra sub-partition (ISP) tool. In some embodiments of the method 2000, the first MPM list for the MIP mode is further based on a height or a width of the current video block.
[0269] In some embodiments of the method 2000, the first MPM list for the MIP mode is further based on a height or width of a neighboring video block of the current video block. In some embodiments of the method 2000, the first MPM list for the MIP mode is constructed based on a first set of parameters, the first set of parameters being different from a second set of parameters used to construct the second MPM list for the non-MIP intra mode. In some embodiments of the method 2000, the second MPM list for the non-MIP intra mode includes S intra prediction modes, where K of the S intra prediction modes are derived for inclusion in the first MPM list for the MIP mode. In some embodiments of the method 2000, K is 3 and S is 6.
[0270] In some embodiments of the method 2000, the first K intra prediction modes in the second MPM list for the non-MIP intra modes are derived for inclusion in the first MPM list for the MIP modes. In some embodiments of the method 2000, the first mode listed in the first MPM list for the MIP modes is different from the first mode listed in the second MPM list for the non-MIP intra modes. In some embodiments of the method 2000, the first mode listed in the first MPM list for the MIP modes is a first mode, and the first mode listed in the second MPM list for the non-MIP intra modes is a planar mode. In some embodiments of the method 2000, the first mode in the first MPM list for the MIP modes is converted from the planar mode in the second MPM list for the non-MIP intra modes.
[0271] In some embodiments of method 2000, the first MPM list for the MIP mode includes a first set of stuffing modes, and the first set of stuffing modes is different from a second set of stuffing modes included in the second MPM list for the non-MIP intra mode. In some embodiments of method 2000, the first set of stuffing modes includes a first mode, a second mode, and a third mode, and the second set of stuffing modes includes a direct current (DC) mode, a vertical mode, and a horizontal mode. In some embodiments of method 2000, the first mode, the second mode, and the third mode are included in the first set of stuffing modes based on a size of a current video block. In some embodiments of method 2000, the first MPM list for a MIP mode includes a first set of intra prediction modes of neighboring video blocks of the current video block, and the second MPM list for a non-MIP intra mode includes a second set of intra prediction modes of neighboring video blocks of the current video block, and the first set of intra prediction modes are different from the second set of intra prediction modes.
[0272] In some embodiments of the method 2000, the second set of intra prediction modes includes intra prediction modes that are converted to MIP modes included in the first set of intra prediction modes. In some embodiments of the method 2000, the first MPM list for the MIP modes includes a first set of shifted intra prediction modes and the second MPM list for the non-MIP intra modes includes a second set of shifted intra prediction modes, the first set of shifted intra prediction modes being different from the second set of shifted intra prediction modes. In some embodiments of the method 2000, the first set of shifted intra prediction modes includes MIP modes (Y) shifted by K1 according to a first equation Y+K1 and the second set of shifted intra prediction modes includes non-MIP intra modes (X) shifted by K0 according to a second equation X+K0, where K1 is different from K0. In some embodiments of the method 2000, K1 depends on a width and a height of the current video block.
[0273] In some embodiments, the method 2000 further includes making a first determination that a neighboring video block of the current video block is coded in a non-MIP intra mode; and, in response to the first determination, making a second determination that the neighboring video block is not available for constructing a first MPM list for the MIP mode.
[0274] In some embodiments, method 2000 further includes making a first determination that a neighboring video block of the current video block is coded in a non-MIP intra-prediction mode; and in response to the first determination, making a second determination that the neighboring video block is coded in a pre-determined MIP intra-prediction mode, where the first MPM list for the MIP modes is constructed using the pre-determined MIP intra-prediction mode. In some embodiments of method 2000, the pre-determined MIP intra-prediction mode depends on a width and / or a height of the current video block.
[0275] 21 illustrates an example flow chart of an example method 2100 for matrix-based intra prediction. Operation 2102 includes generating a most probable mode (MPM) list based on rules for transforming between a current video block of a video and a coded representation of the current video block, the rule being based on whether a neighboring video block of the current video block is coded in a matrix-based intra prediction (MIP) mode, in which a predictive block of the neighboring video block is determined by performing a boundary downsampling operation on previously coded samples of the video and a matrix-vector multiplication operation followed by a selective upsampling operation. Operation 2104 includes performing a transform between the current video block and the coded representation of the current video block using the MPM list, the transform applying a non-MIP mode to the current video block, the non-MIP mode being different from the MIP mode.
[0276] In some embodiments of method 2100, the rule specifies that neighboring video blocks coded in MIP mode are treated as unavailable for generating an MPM list for a current video block coded in a non-MIP mode. In some embodiments of method 2100, the rule specifies that neighboring video blocks coded in MIP mode are determined to be coded in a predefined intra-prediction mode. In some embodiments of method 2100, the rule specifies that the predefined intra-prediction modes include planar mode.
[0277] 22 illustrates an example flow chart of an example method 2200 for matrix-based intra prediction. Operation 2202 includes decoding a current video block of a coded video in a coded representation of the current video block using a matrix-based intra prediction (MIP) mode, where a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video and a matrix-vector multiplication operation followed by a selective upsampling operation. Operation 2204 includes updating line buffers associated with the decoding without storing information in the line buffers indicating whether the current video block is coded using the MIP mode.
[0278] In some embodiments, method 2200 further includes accessing a second video block of the video, the second video block being decoded prior to decoding the current video block, the second video block being located in a different largest coding unit (LCU) or coding tree unit (CTU) row or CTU region when compared to that of the current video block, and the current video block being decoded without determining whether the second video block is coded using a MIP mode. In some embodiments, method 2200 further includes accessing the second video block by determining that the second video block is coded using a non-MIP intra mode without determining whether the second video block is coded using a MIP mode, the second video block being located in a different largest coding unit (LCU) or coding tree unit (CTU) row or CTU region than that of the current video block; and decoding the current video block based on accessing the second video block.
[0279] 23 illustrates an example flow chart of an example method 2300 for matrix-based intra prediction. Operation 2302 includes performing a conversion between a current video block and a bitstream representation of the current video block, where the current video block is coded using a matrix-based intra prediction (MIP) mode, where a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation, where for at most K contexts in the arithmetic encoding or decoding process, a flag is coded into the bitstream representation, where the flag indicates whether the current video block is coded using the MIP mode, where K is equal to or greater than zero. In some embodiments of method 2300, K is 1. In some embodiments of method 2300, K is 4.
[0280] FIG. 24 illustrates an example flow chart of an example method 2400 for matrix-based intra prediction. Operation 2402 includes generating an intra prediction mode for a current video block coded using a matrix-based intra prediction (MIP) mode for converting between a current video block of a video and a bitstream representation of the current video block, where in the MIP mode, a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix vector multiplication operation and then a selective upsampling operation. Operation 2404 includes determining a rule for storing information indicating an intra prediction mode based on whether the current video block is coded in the MIP mode. Operation 2406 includes performing the conversion according to the rule, where the rule defines that a syntax element for the intra prediction mode is stored in the bitstream representation for the current video block, and where the rule defines that a mode index of the MIP mode for the current video block is not stored in the bitstream representation.
[0281] In some embodiments of method 2400, a mode index of the MIP mode is associated with an intra-prediction mode. In some embodiments of method 2400, the rule defines that the bitstream representation excludes a flag indicating that the current video block is coded in MIP mode. In some embodiments of method 2400, the rule defines that the bitstream representation excludes storage of information indicating a MIP mode associated with the current video block. In some embodiments of method 2400, method 2400 further includes, after the conversion, performing a second conversion between a second video block of the video and the bitstream representation of the second video block, the second video block being a neighboring video block of the current video block, and the second conversion being performed without determining whether the second video block is coded using MIP mode.
[0282] FIG. 25 illustrates an example flow chart of an example method 2500 for matrix-based intra prediction. Operation 2502 includes making a first determination that a luma video block of a video is coded using a matrix-based intra prediction (MIP) mode, where a prediction block of the luma video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation. Operation 2504 includes making a second determination regarding a chroma intra mode to be used for a chroma video block associated with the luma video block based on the first determination. Operation 2506 includes performing a conversion between the chroma video block and a bitstream representation of the chroma video block based on the second determination.
[0283] In some embodiments of the method 2500, making the second determination includes determining, without relying on signaling, that the chroma intra mode is a derived mode (DM). In some embodiments of the method 2500, the luma video block covers a predetermined corresponding chroma sample of the chroma video block. In some embodiments of the method 2500, the predetermined corresponding chroma sample is a top left sample or a center sample of the chroma video block. In some embodiments of the method 2500, the chroma intra mode includes a derived mode (DM). In some embodiments of the method 2500, the DM is used to derive a first intra prediction mode of the chroma video block from a second intra prediction mode of the luma video block. In some embodiments of the method 2500, a MIP mode associated with the luma video block is mapped to a predefined intra mode.
[0284] In some embodiments of the method 2500, a plurality of derived modes (DMs) are derived in response to a first determination that the luma video block is coded in a MIP mode. In some embodiments of the method 2500, the chroma intra mode comprises a specified intra prediction mode. In some embodiments of the method 2500, the predefined intra mode or the specified intra mode is a planar mode. In some embodiments of the method 2500, the chroma video block is coded using a MIP mode. In some embodiments of the method 2500, the chroma video block is coded using a first matrix or a first bias vector that is different from a second matrix or a second bias vector of a second chroma video block. In some embodiments of the method 2500, the chroma video block is a blue component, and the first matrix or the first bias vector is predefined for the blue component, and the second chroma video block is a red component, and the second matrix or the second bias vector is predefined for the red component.
[0285] In some embodiments of the method 2500, the blue and red components are concatenated. In some embodiments of the method 2500, the blue and red components are interleaved. In some embodiments of the method 2500, the chroma video blocks are coded using the same MIP mode as the MIP mode used for the luma video blocks. In some embodiments of the method 2500, the chroma video blocks are coded using a derived mode (DM). In some embodiments of the method 2500, after the chroma video blocks are coded using the MIP mode, the upsampling process (or linear interpolation technique) is skipped. In some embodiments of the method 2500, the chroma video blocks are coded with a MIP mode using a subsampled matrix and / or a bias vector. In some embodiments of the method 2500, the number of MIP modes for the luma video blocks and the chroma video blocks are different. In some embodiments of method 2500, the chroma video blocks are associated with a first number of MIP modes and the luma video blocks are associated with a second number of MIP modes, the first number of MIP modes being less than the second number of MIP modes. In some embodiments of method 2500, the chroma video blocks and the luma video blocks have the same size.
[0286] 26 illustrates an example flow chart of an example method 2600 for matrix-based intra prediction. Operation 2602 includes performing a conversion between a current video block of a video and a coded representation of the current video block, the conversion being based on (or being based on) a determination of whether to code the current video block using a matrix-based intra prediction (MIP) mode, where in the MIP mode, a predictive block of the current video block is determined by selectively performing an upsampling operation after a boundary downsampling operation and a matrix-vector multiplication operation on previously coded samples of the video.
[0287] In some embodiments of method 2600, in response to determining that the current video block is to be coded in MIP mode, a syntax element is included in the coded representation, where the syntax element indicates use of MIP mode for coding the current video block. In some embodiments of method 2600, the coded representation includes a syntax element indicating that MIP mode is enabled or disabled for the current video block, where the syntax element indicates whether MIP mode is allowed or not allowed for coding the current video block. In some embodiments of method 2600, in response to determining that the current video block is not to be coded in MIP mode, a syntax element is not included in the coded representation, where the syntax element indicates whether MIP mode is used for the current video block.
[0288] In some embodiments of the method 2600, the second syntax element indicating enabling MIP mode is further included in a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a tile header, a coding tree unit (CTU) row, or a CTU region. In some embodiments of the method 2600, the determination of whether to use MIP mode for the transform is based on a height (H) and / or a width (W) of the current video block. In some embodiments of the method 2600, the current video block is determined not to be coded using MIP mode in response to W≧T1 and H≧T2. In some embodiments of the method 2600, the current video block is determined not to be coded using MIP mode in response to W≦T1 and H≦T2. In some embodiments of the method 2600, the current video block is determined not to be coded using MIP mode in response to W≧T1 or H≧T2. In some embodiments of method 2600, the current video block is determined not to be coded using MIP mode in response to W≦T1 or H≦T2. In some embodiments of method 2600, T1=32 and T2=32. In some embodiments of method 2600, the current video block is determined not to be coded using MIP mode in response to W+H≧T. In some embodiments of method 2600, the current video block is determined not to be coded using MIP mode in response to W+H≦T. In some embodiments of method 2600, the current video block is determined not to be coded using MIP mode in response to W×H≧T. In some embodiments of method 2600, the current video block is determined not to be coded using MIP mode in response to W×H≦T. In some embodiments of method 2600, T=256.
[0289] 27A illustrates an example flowchart of an example video encoding method 2700A for matrix-based intra prediction. Operation 2702A includes determining, according to a rule, whether to use a coding mode different from a matrix-based intra prediction (MIP) mode and a MIP mode to encode a current video block of a video into a bitstream representation of the current video block, the MIP mode including determining a prediction block of the current video block by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation followed by a selective upsampling operation. Operation 2704 includes adding the encoded representation of the current video block to the bitstream representation based on the determination.
[0290] 27B illustrates an example flowchart of an example video decoding method 2700B for matrix-based intra prediction. Operation 2702B includes determining that a current block of video is encoded into the bitstream representation using a matrix-based intra prediction (MIP) mode and a coding mode different from MIP, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video and then a matrix vector multiplication operation followed by a selective upsampling operation. Operation 2704B includes generating a decoded representation of the current video block by analyzing and decoding the bitstream representation.
[0291] In some embodiments of methods 2700A and / or 2700B, the coding mode is a combined intra-and-inter prediction (CIIP) mode, and the method further includes generating an intra prediction signal for the current video block by selecting between the MIP mode and the intra prediction mode. In some embodiments of methods 2700A and / or 2700B, the selecting is based on signaling in a bitstream representation of the current video block. In some embodiments of methods 2700A and / or 2700B, the selecting is based on a predetermined rule. In some embodiments of methods 2700A and / or 2700B, the predetermined rule selects the MIP mode in response to the current video block being coded using the CIIP mode. In some embodiments of methods 2700A and / or 2700B, the predetermined rule selects the intra prediction mode in response to the current video block being coded using the CIIP mode. In some embodiments of methods 2700A and / or 2700B, the predetermined rule selects the planar mode in response to the current video block being coded using a CIIP mode. In some embodiments of methods 2700A and / or 2700B, performing the selection is based on information related to neighboring video blocks of the current video block.
[0292] In some embodiments of methods 2700A and / or 2700B, the coding mode is a cross-component linear model (CCLM) prediction mode. In some embodiments of methods 2700A and / or 2700B, the first downsampling procedure for downsampling neighboring samples of the current video block in the MIP mode uses at least a portion of a second downsampling procedure for the CCLM prediction mode in which neighboring luma samples of the current video block are downsampled. In some embodiments of methods 2700A and / or 2700B, the first downsampling procedure for downsampling neighboring samples of the current video block in the CCLM prediction mode uses at least a portion of a second downsampling procedure for the MIP prediction mode in which neighboring luma samples of the current video block are downsampled. In some embodiments of methods 2700A and / or 2700B, the first downsampling procedure for the MIP mode is based on a first set of parameters and the first downsampling procedure for the CCLM prediction mode is based on a second set of parameters that are different from the first set of parameters.
[0293] In some embodiments of the methods 2700A and / or 2700B, the first downsampling procedure for the MIP mode and the second downsampling procedure for the CCLM prediction mode include selecting a neighboring luma position or selecting a downsampling filter. In some embodiments of the methods 2700A and / or 2700B, the neighboring luma samples are downsampled using at least one of selecting a downsampled position, selecting a downsampling filter, rounding or clipping. In some embodiments of the methods 2700A and / or 2700B, the rule defines that block-based differential pulse coding modulation (BDPCM) or residual DPCM is not applied to the current video block coded in the MIP mode. In some embodiments of the methods 2700A and / or 2700B, the rule defines that in response to applying block-based differential pulse coding modulation (BDPCM) or residual DPCM, the MIP mode is not allowed to be applied to the current video block.
[0294] 28 illustrates an example flow chart of an example method 2800 for matrix-based intra prediction. Operation 2802 includes making a determination regarding applicability of a loop filter to a reconstructed block of a current video block of video in a conversion between a coded representation of the video and the current video block, the current video block being coded using a matrix-based intra prediction (MIP) mode, in which a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation. Operation 2804 includes processing the current video block according to the determination.
[0295] In some embodiments of the method 2800, the loop filter includes a deblocking filter. In some embodiments of the method 2800, the loop filter includes a sample adaptive offset (SAO). In some embodiments of the method 2800, the loop filter includes an adaptive loop filter (ALF).
[0296] 29A illustrates an example flowchart of an example video encoding method 2900A for matrix-based intra prediction. Operation 2902A includes determining, according to a rule, a type of neighboring samples of a current video block to be used to encode the current video block of the video into a bitstream representation of the current video block. Operation 2904A includes adding, based on the determination, an encoded representation of the current video block to the bitstream representation, where the current video block is encoded using a matrix-based intra prediction (MIP) mode, in which a predictive block of the current video block is determined by selectively performing an upsampling operation after performing a boundary downsampling operation on previously coded samples of the video and a matrix-vector multiplication operation.
[0297] 29B illustrates an example flowchart of an example video decoding method 2900B for matrix-based intra prediction. Operation 2902B includes determining according to a rule that a current video block of a video is coded into the bitstream representation using a matrix-based intra prediction (MIP) mode and using a type of neighboring samples of the current video block, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video and then a matrix-vector multiplication operation followed by a selective upsampling operation. Operation 2904B includes generating a decoded representation of the current video block by analyzing and decoding the bitstream representation.
[0298] In some embodiments of methods 2900A and / or 2900B, the rules define that the type of adjacent samples is an unfiltered adjacent sample. In some embodiments of methods 2900A and / or 2900B, the rules define that the type of adjacent samples including the unfiltered adjacent sample is used for an upsampling process, the rules define that the type of adjacent samples including the unfiltered adjacent sample is used for a downsampling process, and the rules define that the type of adjacent samples including the unfiltered adjacent sample is used for a matrix vector multiplication process. In some embodiments of methods 2900A and / or 2900B, the rules define that the type of adjacent samples is a filtered adjacent sample. In some embodiments of methods 2900A and / or 2900B, the rules define that the type of adjacent samples including the unfiltered adjacent sample is used for an upsampling technique, and the rules define that the type of adjacent samples including the filtered adjacent sample is used for a downsampling technique.
[0299] In some embodiments of methods 2900A and / or 2900B, the rules define that a type of adjacent sample including an unfiltered adjacent sample is used for the downsampling technique and the rules define that a type of adjacent sample including a filtered adjacent sample is used for the upsampling technique. In some embodiments of methods 2900A and / or 2900B, the rules define that a type of adjacent sample including an unfiltered upper adjacent sample is used for the upsampling technique and the rules define that a type of adjacent sample including a filtered left adjacent sample is used for the upsampling technique. In some embodiments of methods 2900A and / or 2900B, the rules define that a type of adjacent sample including an unfiltered left adjacent sample is used for the upsampling technique and the rules define that a type of adjacent sample including a filtered upper adjacent sample is used for the upsampling technique.
[0300] In some embodiments of methods 2900A and / or 2900B, the type of neighboring sample includes an unfiltered neighboring sample or a filtered neighboring sample, and the rule defines whether the unfiltered neighboring sample or the filtered neighboring sample is used based on the MIP mode of the current video block.
[0301] In some embodiments, the method 2900A and / or 2900B further comprises converting the MIP mode to an intra-prediction mode, the type of the neighboring sample comprises an unfiltered neighboring sample or a filtered neighboring sample, and the rule defines whether the unfiltered neighboring sample or the filtered neighboring sample is used based on the intra-prediction mode. In some embodiments of the method 2900A and / or 2900B, the type of the neighboring sample comprises an unfiltered neighboring sample or a filtered neighboring sample, and the rule defines whether the unfiltered neighboring sample or the filtered neighboring sample is used based on a syntax element or a signaling. In some embodiments of the method 2900A and / or 2900B, the filtered neighboring sample, the filtered left neighboring sample, or the filtered top neighboring sample are generated using the intra-prediction mode.
[0302] FIG. 30 illustrates an example flow chart of an example method 3000 for matrix-based intra prediction. The operation 3002 includes performing a conversion between a current video block of a video and a bitstream representation of the current video block, the conversion including generating a prediction block for the current video block using a matrix-based intra prediction (MIP) mode by selecting and applying a matrix multiplication using a matrix of samples and / or selecting and adding an offset using an offset vector for the current video block, the samples being obtained from row- and column-wise averages of previously coded samples of the video, the selection being based on reshaping information associated with applying a luma mapping with chroma scaling (LMCS) technique with respect to a reference picture of the current video block. In some embodiments of the method 3000, the reshaping information is associated with a matrix and / or offset vector if the reshaping technique is enabled for the reference picture, and the matrix and / or offset vector are different from those associated with the reshaping information if the reshaping technique is disabled for the reference picture. In some embodiments of the method 3000, different matrices and / or offset vectors are used for different reshaping parameters of the reshaping information. In some embodiments of the method 3000, performing the transform includes performing intra prediction on the current video block using a MIP mode in the original domain. In some embodiments of the method 3000, neighboring samples of the current video block are mapped to the original domain in response to applying the reshaping information, where the neighboring samples are mapped to the original domain before being used in the MIP mode.
[0303] 31 illustrates an example flow chart of an example method 3100 for matrix-based intra prediction. Operation 3102 includes determining that a current block is to be coded using a matrix-based intra prediction (MIP) mode, where a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation. Operation 3104 includes performing a conversion between the current video block and a bitstream representation of the current video block based on the determination, where performing the conversion is based on rules for joint application of the MIP mode and other coding techniques.
[0304] In some embodiments of the method 3100, the rules specify that the MIP mode and the other coding technique are used mutually exclusively. In some embodiments of the method 3100, the rules specify that the other coding technique is a high dynamic range (HDR) technique or a reshaping technique.
[0305] 32 shows an example flowchart of an example method 3200 for matrix-based intra prediction. Operation 3202 includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra prediction (MIP) mode, where performing the conversion using the MIP mode includes generating a prediction block by applying a matrix multiplication using a matrix of samples obtained from row- and column-wise averages of previously coded samples of the video, where the matrix is dependent on a bit depth of the samples.
[0306] FIG. 33 illustrates an example flow chart of an example method 3300 for matrix-based intra prediction. Operation 3302 includes generating an intermediate prediction signal for a current video block of a video using a matrix-based intra prediction (MIP) mode, in which a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation. Operation 3304 includes generating a final prediction signal based on the intermediate prediction signal. Operation 3306 includes performing a conversion between the current video block and a bitstream representation of the current video block based on the final prediction signal.
[0307] In some embodiments of method 3300, generating the final prediction signal is performed by applying position-dependent intra-prediction combination (PDPC) to the intermediate prediction signal. In some embodiments of method 3300, generating the final prediction signal is performed by generating previously coded samples of the video using a MIP mode and filtering the previously coded samples with neighboring samples of the current video block.
[0308] 34 illustrates an example flowchart of an example method 3400 for matrix-based intra prediction. Operation 3402 includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra prediction (MIP) mode, where performing the conversion includes using an interpolation filter in an upsampling process for the MIP mode, where in the MIP mode a matrix multiplication is applied to a first set of samples obtained from row-wise and column-wise averaging of previously coded samples of the video, and an interpolation filter is applied to a second set of samples obtained from the matrix multiplication, where the interpolation filter excludes a bilinear interpolation filter.
[0309] In some embodiments of the method 3400, the interpolation filter comprises a 4-tap interpolation filter. In some embodiments of the method 3400, motion compensation of the chroma components of the current video block is performed using a 4-tap interpolation filter. In some embodiments of the method 3400, angular intra prediction of the current video block is performed using a 4-tap interpolation filter. In some embodiments of the method 3400, the interpolation filter comprises an 8-tap interpolation filter and motion compensation of the luma components of the current video block is performed using the 8-tap interpolation filter.
[0310] 35 illustrates an example flow chart of an example method 3500 for matrix-based intra prediction. Operation 3502 includes performing a conversion between a current video block of a video and a bitstream representation of the current video block according to a rule, the rule specifying an application relationship of a matrix-based intra prediction (MIP) mode or a conversion mode during the conversion, the MIP mode including determining a predictive block of the current video block by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation followed by a selective upsampling operation, and the conversion mode specifying use of a conversion operation to determine a predictive block for the current video block.
[0311] In some embodiments of the method 3500, the transform mode includes a contraction quadratic transform (RST), a quadratic transform, a rotation transform, or a non-separable quadratic transform (NSST). In some embodiments of the method 3500, the rule specifies that in response to a MIP mode being applied to the current video block, a transform operation using the transform mode is not applied. In some embodiments of the method 3500, the rule specifies whether to apply a transform operation using the transform mode based on a height (H) or a width (W) of the current video block. In some embodiments of the method 3500, the rule specifies that in response to W≧T1 and H≧T2, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W≦T1 and H≦T2, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W≧T1 or H≧T2, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W≦T1 or H≦T2, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, T=32 and T2=32. In some embodiments of the method 3500, the rule specifies that in response to W+H≧T, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W+H≦T, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W×H≧T, a transform operation using the transform mode is not applied to the current video block. In some embodiments of the method 3500, the rule specifies that in response to W×H≦T, a transform operation using the transform mode is not applied to the current video block.In some embodiments of the method 3500, T=256.
[0312] In some embodiments of the method 3500, the rule specifies that a transformation operation using the transformation mode is applied in response to the current video block being coded in a MIP mode. In some embodiments of the method 3500, the selection of a transformation matrix or kernel for the transformation operation is based on the current video block being coded in a MIP mode. In some embodiments, the method 3500 further includes converting the MIP mode to an intra-prediction mode and selecting a transformation matrix or kernel based on the transformed intra-prediction mode. In some embodiments, the method 3500 further includes converting the MIP mode to an intra-prediction mode and selecting a transformation matrix or kernel based on a classification of the transformed intra-prediction mode. In some embodiments of the method 3500, the transformed intra-prediction mode includes a planar mode.
[0313] In some embodiments of the method 3500, the rules specify that a transform operation using a transform mode is applied to the current video block if the MIP mode is not allowed to be applied to the current video block. In some embodiments of the method 3500, the rules specify that a Discrete Cosine Transform type II (DCT-II) transform coding technique is applied to the current video block coded using the MIP mode. In some embodiments of the method 3500, the bitstream representation of the current video block excludes signaling of a transform matrix index for the DCT-II transform coding technique. In some embodiments of the method 3500, performing the transform includes deriving a transform matrix used by the DCT-II transform coding technique. In some embodiments of the method 3500, the bitstream representation includes information related to the MIP mode that is signaled after the indication of the transform matrix.
[0314] In some embodiments of the method 3500, the bitstream representation includes an indication of the MIP mode for the transformation matrix. In some embodiments of the method 3500, the rule specifies that the bitstream representation excludes an indication of the MIP mode for the predefined transformation matrix. In some embodiments of the method 3500, the rule specifies that in response to the current video block being coded in MIP mode, a transform operation using a transform skip technique is applied to the video block. In some embodiments of the method 3500, if the current video block is coded in MIP mode, the bitstream representation of the current video block excludes signaling of the transform skip technique. In some embodiments of the method 3500, the rule specifies that for a current video block coded in MIP mode, the MIP mode is converted to a predefined intra prediction mode when the transform operation is performed by selecting a mode-dependent transform matrix or kernel. In some embodiments of the method 3500, the rule specifies that transform operations using a transform skip technique are not permitted for a current video block coded in MIP mode. In some embodiments of the method 3500, the bitstream representation excludes signaling indicating the use of a transform skip technique. In some embodiments of the method 3500, the rule specifies that transform processing using the transform skip technique is permitted for a current video block that is not coded in MIP mode. In some embodiments of the method 3500, the bitstream representation excludes signaling indicating the use of a MIP mode.
[0315] 36 illustrates an example flowchart of an example method 3600 for matrix-based intra prediction. Operation 3602 includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix-based intra prediction (MIP) mode, where a prediction block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation, where performing the conversion includes deriving boundary samples according to a rule by applying a left bit-shift operation or a right bit-shift operation to a sum of at least one reference boundary sample, where the rule determines whether to apply the left bit-shift operation or the right bit-shift operation.
[0316] In some embodiments of method 3600, the rules define that in response to the number of bits to be shifted being greater than zero, the right bit shift operation is applied using a first technique, and in response to the number of bits to be shifted being equal to zero, the rules define that the right bit shift operation is applied using a second technique, the first technique being different from the second technique. In some embodiments of the method 3600, the boundary sample upsBdryX[x] is calculated using one of the following formulas:
[0317]
number
[0318]
number
[0319]
number
[0320] FIG. 38 illustrates an example flow chart of an example method 3800 for matrix-based intra prediction. Operation 3802 involves converting between a current video block of a video and a bitstream representation of the current video block, comprising generating an intra-prediction mode for a current video block coded using a matrix-based intra-prediction (MIP) mode, where in the MIP mode, a predictive block of the current video block is determined by performing a boundary downsampling operation on previously coded samples of the video followed by a matrix-vector multiplication operation and then a selective upsampling operation. Operation 3804 determines a rule for storing information indicative of an intra-prediction mode based on whether the current video block is coded in the MIP mode. Operation 3806 includes performing the conversion according to the rule, where the rule specifies that the bitstream representation excludes storage of information indicative of a MIP mode associated with the current video block.
[0321] This patent document refers to a bitstream representation to mean a coded representation, and vice versa. From the foregoing, it will be understood that, although specific embodiments of the techniques of this disclosure have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the invention. Accordingly, the techniques disclosed herein are not to be limited except as by the appended claims.
[0322] Implementations of the subject matter and functional operations described in this patent document can be realized in various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Implementations of the subject matter described herein can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by or for controlling the operation of a data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing unit" or "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. An apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, such as code constituting a processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.
[0323] A computer program (also known as a program, software, software application, script, code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in several coordinated files (e.g., a file that stores one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer or on several computers, which can be located at one site or distributed across several sites and interconnected by a communication network.
[0324] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data to generate output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as, for example, an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0325] Processors suitable for the execution of a computer program include, by way of example, both general purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, e.g., magnetic, magnetic-optical, or optical disks, for storing data, or is operatively coupled to receive data from them, transfer data to them, or both. However, a computer need not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0326] It is intended that the specification, together with the drawings, be considered exemplary only, and exemplary is meant to be an example. As used in this application, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.
[0327] Although this patent document contains many details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in this patent document in the context of separate embodiments may also be implemented in a single embodiment in combination. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in a particular combination, or even initially claimed as such, one or more features from a claimed combination may, in some cases, be carved out of the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0328] Similarly, although acts are depicted in a particular order in the figures, this should not be understood as requiring that such acts be performed in the particular order or sequence shown, or that all of the acts illustrated be performed, to achieve desired results.Furthermore, the division of various system components in the embodiments described in this patent document should not be understood as requiring such division in all embodiments.
[0329] Only a few implementations and examples have been described; other implementations, extensions and modifications can be made based on what is described and illustrated in this patent document.< / end> < / end> < / begin> < / end> < / begin> < / end> < / begin> < / end> < / begin> < / end> < / begin>
Claims
1. 13. A method of video processing comprising: determining whether to apply a first prediction mode to a chroma block of a video for conversion between the video bitstream and the chroma block; generating a predicted sample for the chroma block based on the determining; performing the conversion between the chroma blocks and the bitstream; in response to determining to apply the first prediction mode to the chroma block, the prediction samples are generated by performing a boundary downsampling operation based on a size of the chroma block and a matrix vector multiplication operation followed by a selective upsampling operation; 4. The method of claim 3, wherein a decision to apply the first prediction mode to the chroma block is based on a first corresponding luma block being coded in the first prediction mode.
2. 2) deriving a second chroma intra prediction mode based on the derived luma intra prediction mode; and 3) deriving the prediction sample by using the second chroma intra prediction mode.
3. The method of claim 1 , wherein the first corresponding luma block is a luma block covering a luma sample that corresponds to a top-left sample of the chroma block.
4. The method of claim 2, wherein the luma intra prediction mode is derived based on a second luma prediction mode of a second corresponding luma block.
5. The method of claim 4 , wherein the second corresponding luma block is a luma block that covers a luma sample that corresponds to a center sample of the chroma block.
6. 6. The method of claim 4 or 5, wherein in response to the second luma prediction mode being the same as the first prediction mode, the luma intra prediction mode is derived to be a normal intra mode, including a DC mode, a planar mode, or an angular intra mode.
7. The method of claim 1 , wherein the transforming comprises encoding the chroma blocks into the bitstream.
8. The method of claim 1 , wherein the transforming comprises decoding the chroma blocks from the bitstream.
9. 1. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions that, when executed by the processor, cause the processor to: determining whether to apply a first prediction mode to a chroma block of a video for conversion between the video bitstream and the chroma block; generating a predicted sample for the chroma block based on the determination; and performing the conversion between the chroma blocks and the bitstream; and in response to a determination of applying the first prediction mode to the chroma block, the prediction samples are generated by performing a boundary downsampling operation based on a size of the chroma block and a matrix vector multiplication operation followed by a selective upsampling operation; 4. The apparatus of claim 3, wherein the determination to apply the first prediction mode to the chroma block is based on a first corresponding luma block being coded in the first prediction mode.
10. A non-transitory computer-readable storage medium storing instructions that cause a processor to: determining whether to apply a first prediction mode to a chroma block of a video for conversion between the video bitstream and the chroma block; generating a predicted sample for the chroma block based on the determination; and performing the conversion between the chroma blocks and the bitstream; and in response to a determination of applying the first prediction mode to the chroma block, the prediction samples are generated by performing a boundary downsampling operation based on a size of the chroma block and a matrix vector multiplication operation followed by a selective upsampling operation; 4. A storage medium, comprising: a first predictive mode for applying a first prediction mode to the chroma block;
11. 1. A method for storing a video bitstream, comprising: determining whether to apply a first prediction mode to a chroma block of the video; generating a predicted sample for the chroma block based on the determining; generating the bitstream based on the predicted samples; storing the bitstream on a non-transitory computer readable recording medium; in response to determining to apply the first prediction mode to the chroma block, the prediction samples are generated by performing a boundary downsampling operation based on a size of the chroma block and a matrix vector multiplication operation followed by a selective upsampling operation; 4. The method of claim 3, wherein a decision to apply the first prediction mode to the chroma block is based on a first corresponding luma block being coded in the first prediction mode.
Citation Information
Cited By
Matrix-based intra prediction using upsampling
US12519945B2
Calculation in matrix-based intra prediction
US12526424B2
Matrix derivation in intra coding mode
US12610037B2