Mip for all channels in the case of 4:4:4-chroma format and of single tree
By applying the same MIP mode to co-located intra-predicted blocks of multiple color components, the inefficiencies in picture and video coding are addressed, enhancing correlation and reducing costs.
Patent Information
- Application Number
- JP2025087866
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-04-02
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2041-04-01
AI Technical Summary
Matrix-based intra-prediction mode (MIP) is not currently applicable to all color components of a picture, particularly in 4:4:4 chroma format, leading to inefficiencies in picture and video coding and increased bitstream signaling costs.
Apply the same MIP mode for co-located intra-predicted blocks of multiple color components, ensuring they are evenly sampled and divided into blocks, and use a block-based decoder/encoder to select intra-prediction modes based on co-located blocks, reducing complexity and signaling costs.
Enhances correlation across color components, reducing bitstream and signaling costs while improving coding efficiency.
Smart Images

Figure 2025119034000001_ABST
Abstract
Description
[Technical Field]
[0001] DETAILED DESCRIPTION OF THE INVENTION Embodiments in accordance with the present invention relate to an apparatus and method for encoding or decoding pictures or videos using matrix-based intra prediction (MIP) for all channels in the case of 4:4:4 chroma format and a single tree. [Background technology]
[0002] In the current VTM, MIP is only used for the luma component [3]. If the intra mode of a chroma intra block is direct mode (DM) and the intra mode of the co-located luma block is MIP mode, the chroma block must use planar mode to generate the intra prediction signal. The main reason for treating DM mode in this way in the MIP case is that in the 4:2:0 case or dual-tree case, the co-located luma block may have a different shape than the color blocks. Therefore, since MIP mode is not applicable to all block shapes, the MIP mode of the co-located luma block may not be applicable to the chroma block in that case.
[0003] Therefore, it is desirable to provide concepts that render picture coding and / or video coding more efficient in order to support matrix-based intra prediction. Additionally or alternatively, it is desirable to reduce the bitstream and thus the signaling cost.
[0004] This is achieved by the subject matter of the independent claims of the present application.
[0005] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application. Summary of the Invention
[0006] According to a first aspect of the present invention, the inventors of the present application have discovered that one problem encountered when attempting to use a matrix-based intra-prediction mode (MIP mode) to predict samples of a given block of a picture stems from the fact that MIP cannot currently be used for all color components of a picture. According to the first aspect of the present application, this difficulty is overcome by using the same MIP mode for intra-predicted blocks of two color components that share the same location in a picture, i.e., co-located intra-predicted blocks, for example, when two color components are evenly sampled and evenly divided into blocks. The inventors have discovered that using the same MIP mode for co-located intra-predicted blocks of two or more color components of a picture is advantageous. This is based on the idea that strong correlation of prediction signals across all color components of a picture is beneficial, and that such correlation can be enhanced when the same intra-prediction mode is used for two or more color components of intra-predicted blocks of a picture. Using the same MIP mode for co-located blocks of different color components of the same picture can potentially reduce the bitstream and, therefore, signaling costs. Furthermore, coding complexity can be reduced.
[0007] Thus, according to a first aspect of the present application, a block-based decoder / encoder is configured to divide a picture of multiple color components and color sampling formats into blocks using a division scheme in which the picture is divided equally for each color component, e.g., the first color component of a picture has the same sampling as all other color components of the picture, e.g., all color components of the picture have the same spatial resolution. A picture may be composed of a luma component and two chroma components, e.g., in a YUV, YPbPr, and / or YCbCr color space, where the luma component and two chroma components represent multiple color components. For example, in an RGB color space, a picture may also be composed of a red component, a green component, and a blue component, which represent multiple color components. A picture may be composed of, e.g., a first color component, a second color component, and optionally a third color component. It is clear that the described block-based decoder / encoder can also be used for pictures containing other color components or pictures with a different number of color components. For example, a picture is divided equally among color components by dividing the first color component into first color component blocks, the second color component into second color component blocks, and optionally the third color component into third color component blocks. The decision of whether a color component of a picture is inter-predicted (i.e., inter-coded) or intra-predicted (i.e., intra-coded) can be made at a granularity or in units of color component blocks. The block-based decoder / encoder is configured to select one of a first set of intra-prediction modes for each intra-predicted first color component block of a picture, i.e., for each first color component block associated with intra-prediction, to encode / decode the first color component of the picture to / from a data stream on a block-by-block basis, i.e., on a first-color component block-by-first-color component block basis. The first color component blocks associated with inter-prediction, i.e., inter-predicted first color component blocks, are treated differently.The first set of intra-prediction modes includes matrix-based intra-prediction modes, each mode in which the interior of a block is predicted by deriving a sample value vector from reference samples, neighboring within the block, calculating a matrix-vector product between the sample value vector and a prediction matrix associated with the respective matrix-based intra-prediction mode to obtain a prediction vector, and predicting samples within the block based on the prediction vector. Further, the block-based decoder / encoder is configured to encode / decode a second color component of the picture on a block-by-block basis by intra-predicting a given second color component block of the picture using the matrix-based intra-prediction mode selected for the co-located intra-predicted first color component block. The first color component of the picture may be a luma component of the picture, and the second color component of the picture may be a chroma component of the picture.
[0008] According to one embodiment, the number of components of the prediction vector is less than the number of samples within the block, and the block-based decoder / encoder is configured to predict samples within the block based on the prediction vector by interpolating samples based on components of the prediction vector assigned to support sample positions within the block. This is based on the idea that such a prediction vector can be obtained by reducing the total number of multiplications required to compute a matrix-vector product, which may reduce the complexity and signaling cost of the decoder / encoder.
[0009] According to one embodiment, a block-based decoder / encoder is configured to select one of a first option and a second option for each intra-predicted second-color component block of a picture, i.e., for each second-color component block associated with intra-prediction. In the first option, the intra-prediction mode for each intra-predicted second-color component block is derived based on the intra-prediction mode selected for the co-located intra-predicted first-color component block, such that the intra-prediction mode for each intra-predicted second-color component block is equal to the intra-prediction mode selected for the co-located intra-predicted first-color component block if the intra-prediction mode selected for the co-located intra-predicted first-color component block is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second-color component block is selected based on an intra-mode index present / signaled in the data stream for the respective intra-predicted second-color component block. For example, the decoder / encoder may be configured to select the first option if direct mode or residual coding color transform mode is indicated for the intra-predicted second color component block, otherwise, e.g., by selecting the second option, the mode signaled in the data stream is used for predicting the intra-predicted second color component block.
[0010] According to one embodiment, a block-based decoder / encoder is configured to select a matrix-based intra-prediction mode from the first set of intra-prediction modes to be a subset of matrix-based intra-prediction modes from a set of disjoint subsets of matrix-based intra-prediction modes according to a block dimension of each intra-predicted first color component block. For example, each subset of matrix-based intra-prediction modes may be associated with a specific block dimension. For example, only the matrix-based intra-prediction modes of the selected subset may be from the first set of intra-prediction modes. The first set of intra-prediction modes from which the intra-prediction mode of each intra-predicted first color component block is selected may differ for different block dimensions. Therefore, the decoder / encoder is configured to pre-select an intra-prediction mode suitable for each intra-predicted first color component block based on the block dimension of the intra-predicted first color component block. This may reduce the complexity and signaling cost of the decoder / encoder.
[0011] According to one embodiment, prediction matrices associated with a set of disjoint subsets of matrix-based intra-prediction modes are machine-learned, where prediction matrices of one subset of matrix-based intra-prediction modes are of equal size, and prediction matrices of two subsets of matrix-based intra-prediction modes selected for different block sizes are of different sizes. Each subset may group multiple matrix-based intra-prediction modes associated with prediction matrices of the same dimension.
[0012] According to one embodiment, the block-based decoder / encoder is configured such that the prediction matrices associated with the matrix-based intra-prediction modes comprising the first set of intra-prediction modes are of equal size to each other and are machine-learned.
[0013] According to one embodiment, the block-based decoder / encoder is configured such that the intra-prediction modes of the first set of intra-prediction modes other than the matrix-based intra-prediction mode of the first set of intra-prediction modes include a DC mode, a planar mode, and a directional mode. For example, the first set of intra-prediction modes includes a DC mode and / or a planar mode and / or a directional mode in addition to the matrix-based intra-prediction mode.
[0014] According to one embodiment, the block-based decoder / encoder is configured to select a partitioning scheme from a set of partitioning schemes, where the set of partitioning schemes includes a partitioning scheme in which a picture is partitioned for a first color component using first partitioning information present / signaled in a data stream and a further partitioning scheme in which the picture is partitioned for a second color component using second partitioning information present / signaled in a data stream separate from the first partitioning information. Thus, the set of partitioning schemes includes, for example, a partitioning scheme in which a picture is partitioned equally for each color component and a further partitioning scheme in which a picture is partitioned differently for the first color component than for the second color component.
[0015] According to one embodiment, a block-based decoder / encoder is configured to select one of the first and second options, as described above, for each intra-predicted second-color component block of a picture. In the first option, the intra-prediction mode for each intra-predicted second-color component block is derived based on the intra-prediction mode selected for the co-located intra-predicted first-color component block, such that the intra-prediction mode for each intra-predicted second-color component block is equal to the intra-prediction mode selected for the co-located intra-predicted first-color component block if the intra-prediction mode selected for the co-located intra-predicted first-color component block is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second-color component block is selected based on an intra-mode index present / signaled in the data stream for each intra-predicted second-color component block. Furthermore, the block-based decoder / encoder may be configured to divide the further picture into multiple color component and color sampling formats, where each color component is evenly sampled into further blocks using a further division scheme, i.e., the further picture is divided for the first color component differently from the second color component, e.g., as described above. The first color component of the further picture may be divided into further blocks of the first color component, and the second color component of the further picture may be divided into further blocks of the second color component. The decoder / encoder may be configured to select, for example, one from a first set of intra-prediction modes for each further block of the intra-predicted first color component of the further picture, to encode / decode the first color component of the further picture in units of further blocks. The decoder / encoder may be configured to select one from the first option and the second option for each further block of the intra-predicted second color component of the further picture.In a first option, the intra-prediction mode for each intra-predicted second-color component additional block is derived based on the intra-prediction mode selected for the co-located first-color component additional block, such that if the intra-prediction mode selected for the co-located first-color component additional block is one of the matrix-based intra-prediction modes, the intra-prediction mode for each intra-predicted second-color component additional block is equal to the planar intra-prediction mode. In a second option, the intra-prediction mode for each intra-predicted second-color component additional block is selected based on an intra-mode index present / signaled in the data stream for each intra-predicted second-color component additional block. As already outlined above, the first option indicates, for example, that the prediction mode of the intra-predicted second-color component block is directly derived from the prediction mode of the co-located intra-predicted first-color component block. This derivation may depend on the partitioning scheme selected for the picture. Therefore, the implementation of the first option may differ for a picture compared to the additional picture. This is because the picture is divided using a division scheme in which the picture is divided equally for each color component, a further picture is divided using a further division scheme in which the picture is divided for a first color component using first division information present / signaled in the data stream, and the picture is divided for a second color component using second division information present / signaled in a data stream separate from the first division information.
[0016] According to one embodiment, a block-based decoder / encoder is configured to select one of the first and second options, as described above, for each intra-predicted second-color component block of a picture. In the first option, the intra-prediction mode for each intra-predicted second-color component block is derived based on the intra-prediction mode selected for the co-located intra-predicted first-color component block, such that the intra-prediction mode for each intra-predicted second-color component block is equal to the intra-prediction mode selected for the co-located intra-predicted first-color component block if the intra-prediction mode selected for the co-located intra-predicted first-color component block is one of the matrix-based intra-prediction modes. In the second option, the intra-prediction mode for each intra-predicted second-color component block is selected based on an intra-mode index present / signaled in the data stream for each intra-predicted second-color component block. Further, the block-based decoder / encoder may be configured to partition the still further picture of multiple color components and different color sampling formats, according to which the multiple color components are differently sampled, and select one from a first set of intra-prediction modes for each of the still further blocks of the intra-predicted first color component of the still further picture to decode / encode the first other color component of the still further picture on a block-by-block basis. The block-based decoder / encoder may be configured to select one from the first option and the second option for each of the still further blocks of the intra-predicted second color component of the still further picture.In a first option, the intra-prediction mode for each further intra-predicted block of the second color component is derived based on the intra-prediction mode selected for the further co-located block of the first color component, such that if the intra-prediction mode selected for the further co-located block of the first color component is one of the matrix-based intra-prediction modes, the intra-prediction mode for each further co-located block of the second color component is equivalent to a planar intra-prediction mode. In a second option, the intra-prediction mode for each further co-located block of the second color component is selected based on an intra-mode index present / signaled in the data stream for each further co-located block of the second color component. As already outlined above, the first option indicates, for example, that the prediction mode for the intra-predicted second color component block is derived directly from the prediction mode of the co-located intra-predicted first color component block. This derivation may depend on the color sampling format of the picture. Therefore, the implementation of the first option may differ for a picture compared to further pictures. This is because a picture may have a color sampling format in which each color component is sampled evenly, and a further picture may have a different color sampling format in which multiple color components are sampled differently.
[0017] According to one embodiment, a block-based decoder / encoder is configured to select a partitioning scheme from a set of partitioning schemes, the set of partitioning schemes including a further partitioning scheme in which a picture is partitioned for a first color component using first partitioning information in a data stream, and the picture is partitioned for a second color component using second partitioning information present / signaled in a data stream separate from the first partitioning information. The block-based decoder / encoder is configured to partition still further pictures for multiple color components using the partitioning scheme or the further partitioning scheme.
[0018] According to one embodiment, the block-based decoder / encoder is configured to select from first and second options in response to signaling present / signaled in the data stream for each intra-predicted second color component block when the residual coding color transform mode is signaled to be deactivated for each intra-predicted second color component block in the data stream, and to select by inferring that the first option would be selected when the residual coding color transform mode is signaled to be activated for each intra-predicted second color component block in the data stream.
[0019] One embodiment relates to a method for block-based decoding / encoding, including dividing a picture of a multiple color component and color sampling format into blocks, in which each color component is equally sampled, using a partitioning scheme in which the picture is divided equally for each color component. The method further includes, for each intra-predicted first color component block of the picture, deriving a sample value vector of reference samples within the block, calculating a matrix-vector product between the sample value vector and a prediction matrix associated with a respective matrix-based intra-prediction mode to obtain a prediction vector, and predicting samples within the block based on the prediction vector, selecting one from a first set of intra-prediction modes, including a matrix-based intra-prediction mode according to each mode in which the interior of the block is predicted. The method further includes decoding / encoding the second color component of the picture on a block-by-block basis by intra-predicting a given second color component block of the picture using the matrix-based intra-prediction mode selected for the co-located intra-predicted first color component block.
[0020] The method as described above is based on the same considerations as the encoder / decoder described above, and the method may by completed with all the features and functionality also described with respect to the encoder / decoder.
[0021] One embodiment relates to a data stream having pictures or video encoded therein using the encoding method described herein.
[0022] An embodiment relates to a computer program having a program code for performing the methods described herein, when the computer program runs on a computer.
[0023] The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention.In the following description, various embodiments of the invention are described with reference to the following drawings: [Brief explanation of the drawings]
[0024] [Figure 1] 1 illustrates an embodiment for encoding into a data stream. [Figure 2] 1 illustrates an embodiment of an encoder. [Figure 3] 1 illustrates an embodiment of picture reconstruction. [Figure 4] 1 illustrates an embodiment of a decoder. [Figure 5] 1 shows a schematic diagram of prediction of a block for encoding and / or decoding according to one embodiment; [Figure 6] 1 illustrates a matrix operation for prediction of a block for encoding and / or decoding according to one embodiment. [Figure 7.1] 1 illustrates prediction of a block with a reduced sample value vector according to one embodiment. [Figure 7.2] 1 illustrates prediction of a block using sample interpolation according to one embodiment. [Figure 7.3] 10 illustrates prediction of a block with a reduced sample value vector in which only some boundary samples are averaged, according to one embodiment. [Figure 7.4] 10 illustrates prediction of a block with a reduced sample value vector in which groups of four boundary samples are averaged, according to one embodiment. [Figure 8] 1 illustrates a matrix operation performed by an apparatus according to one embodiment. [Figure 9-1] 4 illustrates detailed matrix operations performed by an apparatus according to one embodiment. [Figure 9-2] 4 illustrates detailed matrix operations performed by an apparatus according to one embodiment. [Figure 10] 10 illustrates detailed matrix operations performed by the device using offset and scaling parameters, according to one embodiment. [Figure 11] 10 illustrates detailed matrix operations performed by the device using offset and scaling parameters, according to one embodiment. [Figure 12] 1 illustrates one embodiment of a prediction of a given second color component block of a picture. [Figure 13] 1 illustrates one embodiment of different predictions for a given second color component block of a picture. DETAILED DESCRIPTION OF THE INVENTION
[0025] Equal or equivalent elements or elements with equal or equivalent functionality are represented in the following description by equal or equivalent reference signs, even if they occur in different figures.
[0026] In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in less detailed block diagram form to avoid obscuring the embodiments of the present invention. In addition, features of different embodiments described later in this specification may be combined with each other unless specifically stated otherwise.
[0027] 1. Introduction In the following, different inventive examples, embodiments, and aspects are described, at least some of which refer to methods and / or apparatuses for, inter alia, video coding and / or performing block-based prediction, e.g., for video and / or virtual reality applications, e.g., by neighboring sample downscaling and / or using linear or affine transformations for video delivery optimization (broadcast, streaming, file playback, etc.).
[0028] Additionally, examples, embodiments, and aspects may refer to High Efficiency Video Coding (HEVC) or later, and further embodiments, examples, and aspects are defined by the accompanying claims.
[0029] It should be noted that any embodiment, example, and aspect defined by the claims can be supplemented by any of the details (features and functions) described in the following sections.
[0030] Also, the embodiments, examples, and aspects described in the following sections can be used individually or can be supplemented by any of the features of another section or any of the features included in the claims.
[0031] It should also be noted that the individual examples, embodiments, and aspects described herein can be used individually or in combination, and thus details can be added to each of the individual aspects described above without adding details to another one of the examples, embodiments, and aspects described above.
[0032] It should also be noted that this disclosure explicitly or implicitly describes features of decoding and / or encoding systems and / or methods.
[0033] Furthermore, the features and functions disclosed herein in relation to the method can also be used in the device. Furthermore, any feature and function disclosed herein in relation to the device can also be used in the corresponding method. In other words, the method disclosed herein can be supplemented by any of the features and functions described in relation to the device.
[0034] Additionally, any of the features and functionality described herein may be implemented in hardware or software, or using both hardware and software, as described in the "Implementation Alternatives" section.
[0035] Also, some of the features highlighted in brackets (“(…)” or "[…]") may be considered optional in some examples, embodiments, or aspects.
[0036] 2 Encoder and decoder Below, various examples are described that can assist in achieving more efficient compression when using block-based prediction. In some examples, high compression efficiency is achieved by using a series of intra-prediction modes. Intra-prediction modes may be provided in addition to, or exclusively with, other heuristically designed intra-prediction modes, for example. Even other examples utilize both of the specialties described herein. However, as a variation of these embodiments, intra-prediction may be changed to inter-prediction by using reference samples in another picture instead.
[0037] To facilitate understanding of the following examples of this application, the description begins by presenting possible compatible encoders and decoders onto which the outlined examples of this application can then be built. Figure 1 shows an apparatus for block-wise encoding a picture 10 into a data stream 12. The apparatus is indicated using reference numeral 14 and may be a still picture encoder or a video encoder. In other words, picture 10 may be the current picture from a video 16 when encoder 14 is configured to encode video 16 containing picture 10 into data stream 12, or encoder 14 may exclusively encode picture 10 into data stream 12.
[0038] As described above, the encoder 14 performs block-wise or block-based encoding. To this end, the encoder 14 subdivides the picture 10 into blocks, which are the units at which the encoder 14 encodes the picture 10 into the data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are provided in more detail below. In general, the subdivision may ultimately result in blocks 18 of a fixed size, such as an array of blocks arranged in rows and columns, or blocks 18 of different block sizes, such as by using hierarchical multi-tree subdivision by starting with the entire picture area of the picture 10 or by pre-partitioning the picture 10 into an array of tree blocks; these examples should not be considered as excluding other possible ways of subdividing the picture 10 into blocks 18.
[0039] Furthermore, encoder 14 is a predictive encoder configured to predictively encode picture 10 into data stream 12. For a particular block 18, this means that encoder 14 determines a prediction for block 18 and encodes into data stream 12 the prediction residual, i.e., the prediction error, whereby the prediction deviates from the actual picture content in block 18.
[0040] The encoder 14 may support different prediction modes for deriving a prediction signal for a particular block 18. In the following example, the key prediction mode is an intra-prediction mode, whereby adjacent, already-encoded samples within the block 18 are spatially predicted from adjacent, already-encoded samples of the picture 10. The encoding of the picture 10 into the data stream 12, and the corresponding decoding procedure thereafter, may be based on a particular encoding order 20 defined among the blocks 18. For example, the encoding order 20 may traverse the blocks 18 in a raster scan order, e.g., from top to bottom row by row, while traversing each row from left to right. In the case of a hierarchical multi-tree-based subdivision, a raster scan ordering may be applied within each hierarchical level, or a depth-first traversal order may be applied, i.e., leaf nodes within a block at a particular hierarchical level may precede blocks at the same hierarchical level that have the same parent block, according to the encoding order 20. Depending on the encoding order 20, adjacent, already-encoded samples of the block 18 may typically be located on one or more edges of the block 18. In the example presented here, for example, the neighboring already coded samples of block 18 are located above and to the left of block 18 .
[0041] Intra-prediction modes need not be the only modes supported by encoder 14. If encoder 14 is, for example, a video encoder, encoder 14 may also support inter-prediction modes in which blocks 18 are temporally predicted from previously coded pictures of video 16. Such inter-prediction modes may be motion-compensated prediction modes in which a motion vector is signaled for such blocks 18, indicating the relative spatial offset of the portion from which the prediction signal for block 18 will be derived as a replica. Additionally or alternatively, other non-intra-prediction modes may also be available, such as inter-prediction modes in the case where encoder 14 is a multiview encoder, or non-predictive modes in which the interior of block 18 is coded as is, i.e., without any prediction.
[0042] Before beginning the description of this application by focusing on intra-prediction modes, we will describe a more specific example of a possible block-based encoder, i.e., a possible implementation of encoder 14 as described with respect to FIG. 2, and then present two corresponding examples of decoders compatible with FIGS. 1 and 2, respectively.
[0043] 2 illustrates a possible implementation of the encoder 14 of FIG. 1, i.e., an implementation in which the encoder is configured to use transform coding to encode the prediction residual. However, this is primarily an example, and the present application is not limited to that classification of prediction residual coding. According to FIG. 2, the encoder 14 includes a subtractor 22 configured to subtract a corresponding prediction signal 24 from an inbound signal, i.e., a picture 10, or, on a block-by-block basis, a current block 18, to obtain a prediction residual signal 26 that is encoded into the data stream 12 by a prediction residual encoder 28. The prediction residual encoder 28 includes a lossy encoding stage 28a and a lossless encoding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes samples of the prediction residual signal 26. As already mentioned above, this embodiment uses transform coding of the prediction residual signal 26, and accordingly the lossy coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such a spectrally decomposed prediction residual 26 by quantization by the quantizer 30 performed on the transformed coefficients representing the residual signal 26. The transform may be a DCT, DST, FFT, Hadamard transform, etc. The transformed and quantized prediction residual signal 34 is then subjected to lossless coding by the lossless coding stage 28b, which is an entropy coder that entropy codes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner that is also usable in the decoder, i.e. taking into account the coding loss in the quantizer 30. To this end, the prediction residual reconstruction stage 36 includes a dequantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs an inverse transform to the transform performed by the transformer 32, such as the inverse of a spectral decomposition, such as the inverse of any of the specific transform examples mentioned above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal as output by the inverse transformer 40 and the prediction signal 24 to output a reconstructed signal, i.e., reconstructed samples.This output is fed to a predictor 44 of the encoder 14, which then determines a prediction signal 24 based thereon, that is, a predictor 44 that supports all prediction modes already discussed above with respect to Figure 1. Figure 2 also illustrates that, if the encoder 14 is a video encoder, the encoder 14 may also include an in-loop filter 46 that filters the fully reconstructed picture that, after being filtered, forms the reference picture for the predictor 44 for the inter-predicted blocks.
[0044] As already mentioned above, the encoder 14 operates on a block basis. For the purposes of the following description, the block-based operation of interest is the operation of subdividing the picture 10 into blocks, in which an intra-prediction mode is selected from a set or plurality of intra-prediction modes supported by the predictor 44 or the encoder 14, respectively, and the selected intra-prediction mode is individually executed. However, other classifications of blocks into which the picture 10 may be subdivided may also exist. For example, the above-mentioned decision of whether the picture 10 is inter-coded or intra-coded may be made at a block granularity or unit deviating from the block 18. For example, the inter / intra mode decision may be performed at the coding block level into which the picture 10 is subdivided, and each coding block is subdivided into prediction blocks. Each prediction block resulting from encoding a block for which it has been determined that intra-prediction is to be used is subdivided into an intra-prediction mode decision. For this purpose, a decision is made for each of these prediction blocks as to which supported intra-prediction mode should be used for the respective prediction block. These predictive blocks form block 18, the block of interest here. Predictive blocks within coding blocks associated with inter-prediction are treated differently by predictor 44. They are inter-predicted from a reference picture by determining a motion vector and replicating a prediction signal for this block from the location in the reference picture pointed to by the motion vector. Another block subdivision involves subdivision into transform blocks in the units where transformation by transformer 32 and inverse transformer 40 is performed. Transformed blocks may be, for example, the result of further subdivision of coding blocks. Essentially, the examples shown herein should not be considered limiting, and other examples may exist. For completeness only, it should be noted that the subdivision into coding blocks may use, for example, multi-tree subdivision, and that prediction blocks and / or transform blocks may also be obtained by further subdividing coding blocks using multi-tree subdivision.
[0045] A decoder 54 or device for block-wise decoding compatible with the encoder 14 of FIG. 1 is described in FIG. 3. This decoder 54 performs the opposite of the encoder 14, i.e., it decodes the picture 10 from the data stream 12 in a block-wise manner and, to this end, supports multiple intra-prediction modes. The decoder 54 may, for example, include a residual provider 156. All other possibilities discussed above with respect to FIG. 1 are also valid for the decoder 54. To this end, the decoder 54 may be a still picture decoder or a video decoder, and all prediction modes and predictability are also supported by the decoder 54. The difference between the encoder 14 and the decoder 54 lies primarily in the fact that the encoder 14 chooses or selects coding decisions according to some optimization, such as to minimize some cost function, which may depend on the coding rate and / or coding distortion. One of those coding options or parameters may involve selecting the intra-prediction mode to be used for the current block 18 among the available or supported intra-prediction modes. The selected intra-prediction mode may then be signaled by the encoder 14 of the current block 18 in the data stream 12, with the decoder 54 remaking the selection using this signaling in the data stream 12 for the block 18. Similarly, the subdivision of the picture 10 into blocks 18 may be the subject of optimization within the encoder 14, and corresponding subdivision information may be conveyed in the data stream 12 by the decoder 54, which recovers the subdivision of the picture 10 into blocks 18 based on the subdivision information. To summarize the above, the decoder 54 may be a predictive decoder that operates on a block basis and in addition to intra-prediction modes, the decoder 54 may support other prediction modes, such as inter-prediction modes, for example, if the decoder 54 is a video decoder. When decoding, the decoder 54 may also use the coding order 20 discussed with reference to FIG. 1, and since this coding order 20 is adhered to in both the encoder 14 and the decoder 54, the same neighboring samples are available for the current block 18 in both the encoder 14 and the decoder 54.Therefore, in order to avoid unnecessary repetitions, the description of the modes of operation of the encoder 14, as far as it relates to the subdivision of the picture 10 into blocks, for example as far as it relates to prediction and as far as it relates to the coding of prediction residuals, may also apply to the decoder 54. The difference lies in the fact that the encoder 14 chooses, by optimization, some coding options or coding parameters and then signals in or inserts into the data stream 12 the coding parameters derived from the data stream 12 by the decoder 54 in order to perform the prediction, subdivision, etc. again.
[0046] FIG. 4 shows a possible implementation of the decoder 54 of FIG. 3, i.e., an adaptation to the implementation of the encoder 14 of FIG. 1 as shown in FIG. 2. Because many elements of the encoder 54 of FIG. 4 are identical to those in the corresponding encoder of FIG. 2, the same reference numerals with apostrophes are used in FIG. 4 to indicate those elements. In particular, the adder 42′, the optional in-loop filter 46′, and the predictor 44′ are connected in the prediction loop in the same way as they are in the encoder of FIG. 2. The reconstructed, i.e., quantized and retransformed, prediction residual signal applied to the adder 42′ is derived by a sequence of an entropy decoder 56, which reverses the entropy coding of the entropy encoder 28b, followed by a residual signal reconstruction stage 36′ consisting of a dequantizer 38′ and an inverse transformer 40′, just as in the encoding side. The output of the decoder is a reconstruction of the picture 10. The reconstruction of picture 10 may be available at the output of adder 42' or alternatively directly at the output of in-loop filter 46'. To improve picture quality, some post-filtering may be arranged at the output of the decoder to subject the reconstruction of picture 10 to some post-filtering, but this option is not described in FIG.
[0047] Again, with respect to Figure 4, the explanations brought out above with respect to Figure 2 should also be valid for Figure 4, except that the encoder only performs optimization tasks and related decisions regarding coding options. However, all explanations regarding block subdivision, prediction, dequantization, and retransformation are also valid for the decoder 54 of Figure 4.
[0048] 3. ALWIP (Affine Linear Weighted Intra Predictor) Although ALWIP is not always necessary to implement the techniques discussed herein, some non-limiting examples of ALWIP are discussed herein.
[0049] This application relates, inter alia, to the concept of improved block-based prediction modes for block-wise picture coding usable in video codecs such as HEVC or successors of HEVC. The prediction modes may be intra-prediction modes, although in theory the concepts described herein can also be transferred to inter-prediction modes where the reference samples are part of another picture.
[0050] A block-based prediction concept is required that allows for efficient implementation, including hardware-friendly implementation.
[0051] This object is achieved by the subject matter of the independent claims of the present application.
[0052] Intra-prediction modes are widely used in picture and video coding. In video coding, intra-prediction modes compete with other prediction modes, such as inter-prediction modes, such as motion-compensated prediction modes. In intra-prediction modes, a current block is predicted based on neighboring samples, i.e., samples that have already been coded as far as the encoder side is concerned and decoded as far as the decoder side is concerned. A prediction residual is transmitted in the data stream of the current block, and neighboring sample values are extrapolated to the current block to form a prediction signal for the current block. The better the prediction signal, the lower the prediction residual, and therefore the fewer bits required to code the prediction residual.
[0053] To be effective, several aspects need to be taken into account to form an effective framework for intra prediction in a block-wise picture coding environment. For example, the more intra prediction modes a codec supports, the higher the side information rate consumption for informing the decoder of the selection. On the other hand, the set of supported intra prediction modes must be able to provide a good prediction signal, i.e., a prediction signal that results in a low prediction residual.
[0054] As a comparative embodiment or basic example, we disclose below an apparatus (encoder or decoder) for block-wise decoding of pictures from a data stream, which supports at least one intra-prediction mode in which a prediction signal for a block of a given size of the picture is determined by applying a first template of neighboring samples of the current block to an affine linear predictor, called an affine linear weighted intra-predictor (ALWIP).
[0055] The apparatus may have at least one of the following characteristics (the same may apply to a method or another technology, e.g., implemented in a non-transitory storage unit that stores instructions, which when executed by a processor, causes the processor to perform the method and / or acts as an apparatus, etc.):
[0056] 3.1 Predictors may complement other predictors. The intra-prediction modes, which may form the subject of implementation improvements described further below, may complement other intra-prediction modes of the codec. Thus, they can complement the DC, planar, or angular prediction modes defined in the HEVC codec, respectively, and in the JEM reference software. Hereinafter, the latter three types of intra-prediction modes will be referred to as traditional intra-prediction modes. Therefore, for a given block in an intra-mode, a flag must be analyzed by the decoder to indicate whether one of the intra-prediction modes supported by the device is used.
[0057] 3.2 Multiple Proposed Prediction Modes A device can include multiple ALWIP modes, so if the decoder knows that one of the ALWIP modes supported by the device will be used, the decoder needs to parse additional information that indicates which of the ALWIP modes supported by the device will be used.
[0058] The signaling of supported modes may have the property that some ALWIP modes require fewer bins to encode than other ALWIP modes. Which of these modes require fewer bins and which require more bins can depend on information that can be extracted from the already decoded bitstream or that can be pre-corrected.
[0059] 4. Some Aspects 2 shows a decoder 54 for decoding a picture from data stream 12. Decoder 54 may be configured to decode a given block 18 of the picture. In particular, predictor 44 may be configured to map a set of P neighboring samples that neighbor the given block 18 to a set of Q prediction values for the samples of the given block using a linear or affine-linear transform (e.g., ALWIP).
[0060] As shown in FIG. 5, a given block 18 contains Q values to be predicted (which become "predicted values" at the end of the calculation). If block 18 has M rows and N columns, then Q=M·N. The Q values of block 18 may be in the spatial domain (e.g., pixels) or the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of block 18 may be predicted based on P values obtained from neighboring blocks 17a-17c, which are generally adjacent to block 18. The P values of neighboring blocks 17a-17c may be located closest to (e.g., adjacent to) block 18. The P values of neighboring blocks 17a-17c have already been processed and predicted. The P values are referred to as values of portions 17'a-17'c to distinguish them from the block of which they are a part (in some cases, 17'b is not used).
[0061] As shown in FIG. 6, to perform the prediction, it is possible to operate on a first vector 17P having P entries (each entry is associated with a specific position in the neighboring portions 17′a to 17′c), a second vector 18Q having Q entries (each entry is associated with a specific position in the block 18), and a mapping matrix 17M (each row is associated with a specific position in the block 18 and each column is associated with a specific position in the neighboring portions 17′a to 17′c). The mapping matrix 17M thus performs a prediction of the P values of the neighboring portions 17′a to 17′c to the values of the block 18 according to a predetermined mode. The entries in the mapping matrix 17M can therefore be understood as weighting coefficients. In the following text, the symbols 17a to 17c are used instead of 17′a to 17′c to refer to the neighboring portions of the boundary.
[0062] Several conventional modes are known in the art, such as DC mode, planar mode, and 65 directional prediction modes. For example, 67 modes may be known.
[0063] However, it has been noted that various modes, referred to herein as linear or affine-linear transformations, may also be used. A linear or affine-linear transformation includes P·Q weighting factors, at least ¼P·Q of which are non-zero, and for each of the Q predictors, a set of P weighting factors associated with that predictor. When the series is arranged one above the other in raster scan order between the samples of a given block, it forms an envelope that is nonlinear in all directions.
[0064] It is possible to map P positions of neighboring values 17'a-17'c (templates), Q positions of neighboring samples 17'a-17'c, and P*Q weighting factors of matrix 17M. A plane is an example of the envelope of a sequence for DC transformation (a plane for DC transformation). Because the envelope is clearly planar, it is excluded by the definition of linear or affine-linear transformation (ALWIP). Another example is a matrix resulting in the emulation of an angular mode, whose envelope is excluded from the definition of ALWIP and, in simple terms, resembles a hill extending diagonally from top to bottom along a direction in the P / Q plane. While the planar mode and the 65 directional prediction modes have different envelopes, the envelope is linear in at least one direction, i.e., in all directions for the illustrated DC mode, and in the hill direction for the angular mode.
[0065] Conversely, the envelope of a linear or affine transform is not linear in all directions. It is understood that such types of transforms may, in some circumstances, be optimal for performing predictions for block 18. Note that it is preferable that at least one-quarter of the weighting factors are different from zero (i.e., at least 25% of the P*Q weighting factors are different from zero).
[0066] The weighting factors may be unrelated to one another, according to normal mapping rules. Thus, the matrix 17M may be such that the values of its entries have no obvious discernible relationship. For example, the weighting factors cannot be described by any analytical or differential function.
[0067] In the example, the ALWIP transformation is such that the average of the maximum cross-correlation values between a first set of weighting factors associated with each predictor and a second set of weighting factors associated with predictors other than the respective predictor, or an inverted version of the latter set, may be higher or lower than a predetermined threshold (e.g., 0.2, 0.3, 0.35, or 0.1, e.g., a threshold in the range of 0.05 to 0.035). For example, for each combination (i1, i2) of rows in the ALWIP matrix 17M, the cross-correlation may be calculated by multiplying the P value in the i1 row by the P value in the i2 row. For each obtained cross-correlation, the maximum value may be obtained. Thus, the average (mean) may be obtained for the entire matrix 17M (i.e., the maximum cross-correlation values for all combinations are averaged). The threshold may then be, for example, 0.2, 0.3, 0.35, or 0.1, e.g., a threshold in the range of 0.05 to 0.035.
[0068] The P adjacent samples of blocks 17a-17c may lie along a one-dimensional path extending along the boundary (e.g., 18c, 18a) of a given block 18. For each of the Q predicted values of a given block 18, the set of P weighting factors associated with each predicted value may be ordered to traverse the one-dimensional path in a predetermined direction (e.g., left to right, top to bottom, etc.).
[0069] In an example, the ALWIP matrix 17M may be non-diagonal or non-block diagonal.
[0070] An example of an ALWIP matrix 17M for predicting a 4x4 block 18 from four adjacent samples that have already been predicted may be: { {37,59,77,28}, {32,92,85,25}, {31,69,100,24}, {33,36,106,29}, {24,49,104,48}, {24,21,94,59}, {29,0,80,72}, {35,2,66,84}, {32,13,35,99}, {39,11,34,103}, {45,21,34,106}, {51,24,40,105}, {50,28,43,101}, {56,32,49,101}, {61,31,53,102}, {61,32,54,100} }.
[0071] (Here, {37, 59, 77, 28} is the first row, {32, 92, 85, 25} is the second row, and {61, 32, 54, 100} is the 16th row of matrix 17M. Matrix 17M has dimensions 16x4 and includes 64 weighting factors (as a result of 16*4=64). This is because matrix 17M has dimensions QxP, where Q=M*N is the number of samples in block 18 to be predicted (block 18 is a 4x4 block), and P is the number of samples in the already predicted samples. Here, M=4, N=4, Q=16 (as a result of M*N=4*4=16), and P=4. The matrix is non-diagonal and non-block diagonal, and is not described by any particular rule.
[0072] As can be seen, less than a quarter of the weighting factors are zero (in the matrix above, one weighting factor out of 64 is zero). The envelope formed by these values, when placed one above the other in raster scan order, forms a nonlinear envelope in all directions.
[0073] Although the above description has been primarily discussed with reference to a decoder (eg, decoder 54), the same may be implemented in an encoder (eg, encoder 14).
[0074] In some examples, for each block size (within the set of block sizes), the ALWIP transforms of the intra prediction modes in the second set of intra prediction modes for the respective block size are different from each other. Additionally or alternatively, the cardinality of the second set of intra prediction modes for block sizes in the set of block sizes can match, but the associated linear or affine-linear transforms of the intra prediction modes in the second set of intra prediction modes for different block sizes can be made non-interchangeable with each other by scaling.
[0075] In some examples, an ALWIP transformation may be defined to have "nothing in common" with a conventional transformation (e.g., an ALWIP transformation may have "nothing in common" with a corresponding conventional transformation even if it is mapped via one of the above mappings).
[0076] In some examples, the ALWIP mode is used for both the luma and chroma components, while in other examples, the ALWIP mode is used for the luma component but not for the chroma component.
[0077] 5 Affine linear weighted intra prediction mode with encoder acceleration (e.g., test CE3-1.2.1)
[0078] 5.1 Description of the method or apparatus The Affine Linear Weighted Intra Prediction (ALWIP) mode tested in CE3-1.2.1 can be the same as that proposed in JVET-L0199 under test CE3-2.2.2, except for the following changes:
[0079] · Multiple Reference Line (MRL) intra prediction, especially harmonization of encoder estimation and signaling, i.e. MRL is not combined with ALWIP and transmission of MRL indices is restricted to non-ALWIP blocks.
[0080] Subsampling is now mandatory for all blocks WxH ≥ 32x32 (previously it was optional for 32x32). Therefore, the additional test and transmission of the subsampling flag in the encoder has been removed.
[0081] ALWIP for 64xN and Nx64 blocks (N≦32) is added by downsampling to 32xN and Nx32 respectively and applying the corresponding ALWIP mode.
[0082] Additionally, test CE3-1.2.1 includes the following encoder-optimizations for ALWIP:
[0083] Combined mode estimation: Conventional and ALWIP modes use a shared Hadamard candidate list for complete RD estimation, i.e., ALWIP mode candidates are added to the same list as conventional (and MRL) mode candidates based on their Hadamard cost.
[0084] EMT Intrafast and PB Intrafast are supported in the combined mode list, with additional optimizations to reduce the number of complete RD checks.
[0085] Only the MPMs of the available left and top blocks are added to the list for the complete RD estimation of ALWIP, following the same approach as in the conventional mode.
[0086] 5.2 Complexity Assessment Test CE3-1.2.1 required up to 12 multiplications per sample to generate the predicted signal, excluding calculations involving the discrete cosine transform. Furthermore, a total of 136,492 parameters, each 16 bits long, were required. This corresponds to 0.273 MB of memory.
[0087] 5.3 Experimental results The test evaluation was performed under the common test conditions JVET-J1010 [2] for Intra-only (AI) and Random Access (RA) configurations using VTM software version 3.0.1. The corresponding simulations were run on an Intel Xeon cluster (E5-2697Av4, AVX2 on, Turbo Boost off) with Linux OS and GCC 7.2.1 compiler.
[0088] [Table 1] [Table 2]
[0089] 5.4 Affine Linear Weighted Intra Prediction with Complexity Reduction (e.g., test CE3-1.2.2) The technique tested in CE2 is related to the "Affine Linear Intra Prediction" described in JVET-L0199 [1], but simplifies it in terms of memory requirements and computational complexity as follows:
[0090] There can be only three different sets of prediction matrices (e.g., S0, S1, S2, see also below) and bias vectors (e.g., to provide offset values) that cover all block shapes. As a result, the number of parameters is reduced to 14400 10-bit values, which is less memory than would be stored in a 128x128 CTU.
[0091] The input and output sizes of the predictor are further reduced. Furthermore, instead of transforming the boundaries through a DCT, averaging or downsampling can be performed on the boundary samples, and the generation of the prediction signal can use linear interpolation instead of an inverse DCT. Therefore, generating the prediction signal can require up to four multiplications per sample.
[0092] 6. Examples Here we describe how to perform several predictions (e.g., as shown in Figure 6) using ALWIP prediction.
[0093] In principle, referring to Figure 6, multiplications of Q*P samples of the QxP ALWIP prediction matrix 17M by P samples of the Px1 neighboring vector 17P should be performed to obtain Q=M*N values of the predicted MxN block 18. Thus, in general, at least P=M+N value multiplications are required to obtain each of the Q=M*N values of the predicted MxN block 18.
[0094] These multiplications have a highly undesirable effect: the dimension P of the boundary vector 17P generally depends on the number M+N of boundary samples (bins or pixels) 17a, 17c adjacent (e.g., nearby) to the M×N block 18 to be predicted. This means that if the size of the block 18 to be predicted is large, the number M+N of boundary pixels (17a, 17c) will be correspondingly large, and therefore the dimension P=M+N of the Px1 boundary vector, the length of each column of the Q×P ALWIP prediction matrix 17M, and thus the number of required multipliers (typically Q=M*N=W*H, where W (width) is another symbol for N, H (height) is another symbol for M, and P=M+N=H+W if the boundary vector is formed by only one row and / or one column of samples) will be large.
[0095] This problem is generally exacerbated by the fact that in microprocessor-based systems (or other digital processing systems), multiplication is generally a power-consuming operation. It can be imagined that a large number of multiplications performed on a large number of samples of a large number of blocks results in a waste of computing power, which is generally undesirable.
[0096] It is therefore desirable to reduce the number of multiplications Q*P required to predict an MxN block 18 .
[0097] It is understood that by intelligently selecting an operation that is easier to process instead of multiplication, the computational power required for each intra prediction of each predicted block 18 can be reduced in some way.
[0098] In particular, referring to FIGS. 7.1 to 7.4, an encoder or decoder can use a plurality of adjacent samples (e.g., 17a, 17c) to reduce the plurality of adjacent samples (e.g., by averaging or downsampling) (e.g., in step 811), and compare with the plurality of adjacent samples to obtain a set of reduced sample values with a lower number of samples. Then, (e.g., in step 812), the reduced set of sample values is subjected to a linear transformation or an affine linear transformation to obtain a predicted value of a predetermined sample of a predetermined block, thereby predicting a predetermined block (e.g., 18) of a picture.
[0099] In some cases, the decoder or encoder can also derive a predicted value of a further sample of a predetermined block based on a predetermined sample and predicted values of a plurality of adjacent samples, for example, by interpolation. Thus, an upsampling strategy can be obtained.
[0100] In the example, it is possible to perform some averaging on the samples of the boundary 17 to reach the reduced set 102 of samples with a reduced number of samples (FIGS. 7.1 to 7.4) (e.g., in step 811) (at least one of the samples of the reduced number of samples 102 may be the average of two samples of the original boundary samples or the selection of the original boundary samples). For example, if the original boundary has samples of P = M + N, the reduced set of samples can have P red <as being P, M red <M and N red <at least one of N, P red = M red + N red Thereby, the boundary vector 17P actually used for prediction (e.g., in step 812b) does not have a Px1 entry, Pred <P, which is P red has x1 entries. Similarly, the ALWIP prediction matrix 17M selected for prediction does not have a QxP dimension and has at least P red <P(M red <M and N red for at least one of N) has a reduced number of elements in the QxP red (or Q red xP red see below) matrix.
[0101] In some examples (e.g., Figures 7.2, 7.3), the block obtained by ALWIP (at step 812) is
Number
Number
Number
Number
Number
Number
[0102] These techniques allow matrix multiplication with reduced multiplication number (Q red *P red or Q*P red ), it may be advantageous because both the initial reduction (e.g., averaging or downsampling) and the final conversion (e.g., interpolation) may be performed by reducing (or avoiding) multiplications. For example, downsampling, averaging, and / or interpolation may be performed (e.g., in steps 811 and / or 813) by employing binary operations that do not require computational power, such as additions and shifts.
[0103] Also, addition is a very simple operation that can be easily performed without much computational effort.
[0104] This shifting operation can be used, for example, for averaging two boundary samples and / or for interpolating two samples (support values) of the downscaled prediction block (or taken from the boundary) to obtain the final prediction block. (Two sample values are needed for the interpolation: there are always two predetermined values within a block, but to interpolate samples along the left and top boundaries of the block, as in Figure 7.2, there is only one predetermined value, so the boundary sample is used as the support value for the interpolation.)
[0105] A two-step procedure may be used: First, the values of the two samples are summed. The sum is then halved (for example, by a right shift).
[0106] Alternatively, the following is possible: First, each of the samples is halved (eg, by left shifting). Next, sum the values of the two half samples.
[0107] Since it is only necessary to select one sample amount for a group of samples (e.g., adjacent samples to each other), simpler operations can be performed when downsampling (e.g., at step 811).
[0108] Therefore, it is possible to define techniques (s) for reducing the number of multiplications performed here. Some of these techniques may be based, among other things, on at least one of the following principles. Even if the actually predicted size of block 18 is MxN, the block is reduced (in at least one of the two dimensions) to a reduced size Q red xP red having an ALWIP matrix (
Number
Number
Number
number
number
number
[0109] According to the example shown in Figure 7.1, a 4x4 block 18 (M=4, N=4, Q=M*N=16) is predicted, and the neighbors 17 of samples 17a (vertical matrix containing 4 predicted samples) and 17c (horizontal row containing 4 predicted samples) have already been predicted in the previous iteration (neighbors 17a and 17c may be collectively denoted 17). A priori, by using the formulas shown in Figure 5, the prediction matrix 17M should be a QxP=16x8 matrix (since Q=M*N=4*4 and P=M+N=4+4=8), and the boundary vector 17P should have dimensions 8x1 (since P=8). However, this would result in the need to perform 8 multiplications for each of the 16 samples of the 4x4 block 18 being predicted, leading to the need to perform 16*8=128 multiplications in total. (Note that the average number of multiplications per sample is a good estimate of the computational complexity. Conventional intra prediction requires 4 multiplications per sample, which increases the associated computational effort. Therefore, it is possible to use this as an ALWIP upper bound that ensures that the complexity is reasonable and does not exceed that of conventional intra prediction.)
[0110] Nevertheless, by using this technique, in step 811, the number of neighboring samples 17a and 17c of the predicted block 18 is reduced from P to P redIt is understood that it is possible to reduce to . In particular, it is possible to average adjacent boundary samples (17a, 17c) (e.g., at 100 in Figure 7.1) to obtain a reduced boundary 102 having two horizontal rows and two vertical columns. Thus, it is understood that the operation as block 18 is a 2x2 block (the reduced boundary is formed by the average value). Alternatively, it is possible to perform downsampling. Thus, two samples are selected for row 17c and two samples are selected for column 17a. Thus, the horizontal row 17c is processed as having two samples (e.g., averaged samples) instead of having four original samples, and the vertical column 17a, which originally has four samples, is processed as having two samples (e.g., averaged samples). After subdividing row 17c and column 17a into groups 110 of two samples each, it can also be understood that one single sample is maintained (e.g., the average of the samples in group 110 or a simple selection between the samples in group 110). Thus, a so-called reduced set 102 of sample values is obtained by a set 102 having only four samples (M red =2, N red =2, P red =M red +N red =4, P red <P).
[0111] It is understood that operations (such as averaging or downsampling 100) can be performed at the processor level without performing too many multiplications. The averaging or downsampling 100 performed in step 811 can be easily obtained by simple and non-power-consuming operations such as addition and shift.
[0112] At this point, it is understood that the reduced set of sample values 102 can be subjected to a linear or affine-linear (ALWIP) transform 19 (e.g., using a prediction matrix such as matrix 17M of FIG. 5 ). In this case, the ALWIP transform 19 directly maps the four samples 102 to sample values 104 in block 18. In this case, no interpolation is required.
[0113] In this case, the dimension of the ALWIP matrix 17M is QxP red = 16x4, which follows from the fact that all Q = 16 samples of the predicted block 18 are obtained directly by ALWIP multiplication (no interpolation required).
[0114] Therefore, in step 812a, the dimension Q×P red An appropriate ALWIP matrix 17M is selected having A. The selection may be based, for example, at least in part, on signaling from data stream 12. The selected ALWIP matrix 17M may have A k where k can be understood as an index that can be signaled in the data stream 12 (in some cases the matrix
number
[0115] In step 812b, the selected QxP red ALWIP matrix 17M(A k (also shown as P red A multiplication is performed between the x1 boundary vector 17P.
[0116] In step 812c, an offset value (e.g., b k ) can be added to all the obtained values 104 of the vector 18Q obtained by ALWIP, for example. k Or in some cases
number
[0117] Therefore, the comparison with and without the present technology is now resumed. Without this technology, a block to be predicted 18 with dimensions M=4, N=4; Q=M*N=4*4=16 values to be predicted, P = M + N = 4 + 4 = 8 boundary samples, P=8 multiplications of each of the Q=16 values to be predicted, P*Q=8*16=128 total multiplications, With this technology, a block to be predicted 18 with dimensions M=4, N=4; Finally, Q=M*N=4*4=16 values to be predicted, Reduced dimension of boundary vector: P red =M red +N red =2+2=4; For each of the Q=16 values to be predicted by ALWIP, P red = 4 multiplications, P red *Q=4*16=64 total multiplications (half of 128!) The ratio between the number of multiplications and the number of final values obtained is P red *Q / Q=4, i.e. half of the P=8 multiplication for each sample to be predicted!
[0118] As can be seen, it is possible to obtain suitable values in step 812 by relying on easy and computationally low power operations such as averaging (possibly adding and / or shifting and / or downsampling).
[0119] Referring to Figure 7.2, the block 18 to be predicted is here an 8x8 block of 64 samples (M=8, N=8). Now, a priori, the prediction matrix 17M must have a size of QxP=64x16 (Q=M*N=8*8=64, Q=64 because M=8 and N=8, and P=M+N=8+8=16). Thus, a priori, P=16 multiplications are required for each of the Q=64 samples of the 8x8 block 18 to be predicted, arriving at 64*16=1024k multiplications for the entire 8x8 block 18!
[0120] However, as can be seen in Figure 7.2, instead of using all 16 samples of the boundary, a method 820 can be provided in which only 8 values are used (e.g., 4 in the horizontal boundary row 17c and 4 in the vertical boundary column 17a between the original samples of the boundary). From the boundary row 17c, 4 samples can be used instead of 8 (e.g., they can be a 2x2 average and / or a 1 of 2 sample selection). Thus, the boundary vector is not a Px1 = 16x1 vector, but a P red x1=8x1 vector only (P red =M red +N red =4+4). Instead of the original P=16 samples, P red It is understood that it is possible to select or average (for example, two by two) the samples of horizontal row 17c and vertical column 17a so as to have only Q = 8 boundary values, forming a reduced set 102 of sample values. This reduced set 102 makes it possible to obtain a reduced version of block 18, which is Q red =M red *N red = 4*4 = 16 samples (instead of Q = M*N = 8*8 = 64) of size Mred xN red The ALWIP matrix can be applied to predict a block with ∑ = 4x4. The reduced version of block 18 includes the samples shown in grey in scheme 106 of Figure 7.2. The samples shown in grey squares (including samples 118' and 118") are the Q obtained in step 812 of interest. red = 16 values, which is obtained by applying a linear transformation 19 in object step 812. After obtaining the values of the 4x4 reduced block, it is possible to obtain the values of the remaining samples (shown as white samples in scheme 106), for example by interpolation.
[0121] As with method 810 in Figure 7.1, method 820 uses the remaining QQ of the MxN=8x8 block 18 to be predicted. red The method may further include a step 813 of deriving predicted values for the remaining QQ = 64 - 16 = 48 samples (white squares), for example by interpolation. red = 64 - 16 = 48 samples are obtained directly by interpolation (e.g., the interpolation can also use the values of the boundary samples) red = 16 samples. As can be seen in Figure 7.2, samples 118' and 118'' were obtained in step 812 (as shown by the gray squares), while sample 108' (which is intermediate between samples 118' and 118'' and shown by the white square) is obtained in step 813 by interpolation between samples 118' and 118''. It is understood that interpolation can also be obtained by operations similar to those for averaging, such as shifting and adding. Thus, in Figure 7.2, value 108' can generally be determined as the intermediate value (may be the average) between the value of sample 118' and the value of sample 118''.
[0122] Interpolation may also be performed to arrive at the final version of the MxN=8x8 block 18 in step 813 based on the multiple sample values shown at 104 .
[0123] Therefore, the comparison between using and not using this technology is Without this technology, The prediction target block 18 has dimensions M = 8, N = 8, and Q = M * N = 8 * 8 = 64 samples within the prediction target block 18, P = M + N = 8 + 8 = 16 samples within the boundary 17, For each of the Q = 64 values to be predicted, P = 16 multiplications, The total number of multiplications is P * Q = 16 * 64 = 1028 multiplications. The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 16. With this technology, The prediction target block 18 has dimensions M = 8 and N = 8. The last Q = M * N = 8 * 8 = 64 values to be predicted. However, the Q used red xP red The ALWIP matrix has P red = M red + N red Q red = M red * N red M red = 4, N red = 4, and P red < P within the boundary, where P red = M red + N red = 4 + 4 = 8 samples, For the Q of the 4x4 reduced block to be predicted, red For each of the 16 values, P red = 8 multiplications (formed by the gray squares in scheme 106), P red * Q red = 8 * 16 = 128 multiplications in total (much less than 1024!) The ratio of the number of multiplications to the number of final values obtained is P red * Q red / Q = 128 / 64 = 2 (much less than 16 obtained without this technology!).
[0124] Therefore, the technology presented here has an 8-fold lower power requirement than the previous one.
[0125] FIG. 7.3 shows another example where the block 18 to be predicted is a rectangular 4×8 block (M = 8, N = 4) with Q = 4 * 8 = 32 samples to be predicted (which can be based on method 820). The boundary 17 is formed by a horizontal row 17c with N = 8 samples and a vertical column 17a with M = 4 samples. Therefore, a priori, the boundary vector 17P has dimension Px1 = 12x1, but the prediction ALWIP matrix needs to be a QxP = 32x12 matrix, and thus Q * P = 32 * 12 = 384 multiplications are required.
[0126] However, for example, it is possible to average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row with only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as it is (e.g., without averaging). In total, the dimension of the reduced boundary is P red = 8, and P red < P. Therefore, the boundary vector 17P has dimension P red x1 = 8x1. The ALWIP prediction matrix 17M becomes a matrix with dimension M * N red * P red = 4 * 4 * 8 = 64. The 4×4 reduced block (formed by the gray columns of the schema 107) directly obtained in the target step 812 has size Q red = M * N red = 4 * 4 = 16 samples (instead of Q = 4 * 8 = 32 of the original 4x8 block 18 to be predicted). When the reduced 4×4 block is obtained by ALWIP, an offset value b k can be added (step 812c), and interpolation can be performed in step 813. As can be seen in step 813 of FIG. 7.3, the reduced 4×4 block is expanded to the 4×8 block 18, and the value 108’ not obtained in step 812 is obtained in step 813 by interpolating the values 118’ and 118’’ (gray squares).
[0127] Therefore, the comparison between using and not using this technology is Without this technology, Block 18 to be predicted, having dimensions M = 4, N = 8 Q = M * N = 4 * 8 = 32 values to be predicted P = M + N = 4 + 8 = 12 samples within the boundary P multiplications for each of the Q = 32 values to be predicted Total number of P * Q = 12 * 32 = 384 multiplications The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 12. With this technology, Block 18 to be predicted, having dimensions M = 4, N = 8 Finally, Q = M * N = 4 * 8 = 32 values to be predicted. However, Q red × P red = 16x8 ALWIP matrix can be used, M = 4, N red = 4, Q red = M * N red = 16, P red = M + N red = 4 + 4 = 8, and Pred < P, P samples within the boundary red = M red + N red = 4 + 4 = 8 samples Q of the reduced block to be predicted red = P multiplications for each of the 16 values red = 8 multiplications Q red * P red = 16 * 8 = 128 total multiplications (less than 384!) The ratio of the number of multiplications to the number of れる final values obtained is Pred * Qred / Q = 128 / 32 = 4 (much less than the れる 12 obtained without this technology!).
[0128] Therefore, using this technology reduces the computational effort to one-third.
[0129] Figure 7.4 shows the case of block 18, which is predicted with dimensions MxN=16x16 and has Q=M*N=16*16=256 values predicted at the end, and P=M+N=16+16=32 boundary samples. This results in a prediction matrix with dimensions QxP=256x32, which means 256*32=8192 multiplications!
[0130] However, by applying method 820, it is possible to reduce the number of boundary samples in step 811 (e.g., by averaging or downsampling) from, for example, 32 to 8, so that, for example, for each group 120 of four consecutive samples in row 17a, one single sample (e.g., one selected from the four samples or the average of the samples) remains, and for each group of four consecutive samples in column 17c, one single sample (e.g., one selected from the four samples or the average of the samples) remains.
[0131] Here, the ALWIP matrix 17M is Q red xP red = 64x8 matrix. This means that P red This is due to the fact that =8 was chosen (by using 8 averaged or selected samples from the 32 boundary samples) and the fact that the reduced block to be predicted in step 812 is an 8x8 block (in scheme 109, the grey squares are 64).
[0132] Thus, once the 64 samples of the reduced 8x8 block are obtained in step 812, the remaining QQ of the block 18 to be predicted are obtained in step 813. red =256-64=192 values 104 can be derived.
[0133] In this case, it has been chosen to use all samples in boundary column 17a and only substitute samples in boundary row 17c to perform the interpolation. Other choices may also be made.
[0134] In this method, the ratio of the number of multiplications to the number of finally obtained values is Q red *P red / Q = 8 * 64 / 256 = 2, which is much less than the 32 multiplications for each value without using this technology!
[0135] Therefore, the comparison between using and not using this technology is Without this technology, Block 18 to be predicted, having dimensions M = 16, N = 16, Q = M * N = 16 * 16 = 256 values to be predicted, P = M + N = 16 + 16 = 32 samples within the boundary, 32 multiplications for each of the Q = 256 values to be predicted, Total number of multiplications: P * Q = 32 * 256 = 8192, Ratio of the number of multiplications to the number of finally obtained values is P * Q / Q = 32. With this technology, Block 18 to be predicted, having dimensions M = 16, N = 16, Finally, Q = M * N = 16 * 16 = 256 values to be predicted, However, the ALWIP matrix Q used red ×P red = 64x8, where the samples M predicted by ALWIP red = 4, N red = 4, Q red = 8 * 8 = 64, and P red = M red +N red = 4 + 4 = 8, P red < P, P within the boundary red = M red +N red = 4 + 4 = 8 samples, Reduced block Q to be predicted red For each of the Q = 64 values, P red = 8 multiplications, Q red *P red = 64 * 4 = 256 total multiplications (less than 8192!) The ratio between the number of multiplications and the number of final values obtained is P red *Q red / Q=8*64 / 256=2 (much less than the 32 you would get without this technique!).
[0136] Therefore, the computational power required for this technology is 16 times less than that of conventional technologies!
[0137] Thus, a given block (18) of a picture can be predicted using a plurality of neighboring samples (17) by shrinking a plurality of neighboring samples (100, 813) to obtain a reduced set (102) of sample values having a smaller number of samples compared to the plurality of neighboring samples (17), and subjecting the reduced set (102) to a linear or affine-linear transformation (19, 17M) to obtain (812) a predicted value of a given sample (104, 118', 188") of the given block (18).
[0138] In particular, the reduction (100, 813) can be performed by downsampling a number of adjacent samples to obtain a reduced set (102) of sample values with a smaller number of samples compared to the number of adjacent samples (17).
[0139] Alternatively, the reduction (100, 813) can be performed by averaging multiple adjacent samples to obtain a reduced set (102) of sample values with a smaller number of samples compared to the multiple adjacent samples (17).
[0140] Furthermore, it is possible to derive (813) a predicted value for a further sample (108, 108') of a given block (18) based on the predicted values of the given sample (104, 118', 118'') and a number of adjacent samples (17) by interpolation.
[0141] The plurality of adjacent samples (17a, 17c) may extend in one dimension along two sides of the given block (18) (e.g., to the right and downward in Figures 7.1-7.4). The given samples (e.g., those obtained by ALWIP in step 812) may also be arranged in rows and columns, and along at least one of the rows and columns, the given samples may be located at every nth position from the given samples 112 adjacent to the two sides of the given block 18.
[0142] A support value (118) for one of a plurality of neighboring positions (118) aligned with each of at least one of the rows and columns can be determined for each of the at least one of the rows and columns based on the plurality of neighboring samples (17). It is also possible to derive a predicted value 118 for a further sample (108, 108') of the given block (18) by interpolation based on the predicted value of the given sample (104, 118', 118'') and the support values of the neighboring samples (118) aligned with each of the at least one of the rows and columns.
[0143] A given sample (104) may be located every nth position along a row from adjacent samples (112) on two sides of a given block 18, and a given sample may be located every mth position along a column from adjacent samples (112) on two sides of a given block 18, where n, m > 1. In some cases, n = m (e.g., in Figures 7.2 and 7.3, samples 104, 118', 118'' obtained directly by ALWIP in 812 and shown as gray boxes are alternated along rows and columns with samples 108, 108' obtained subsequently in step 813).
[0144] It may be possible to perform the determination of the support values along at least one of the rows (17c) and columns (17a), for example, by downsampling or averaging (122) a group (120) of adjacent samples within a plurality of adjacent samples that includes, for each support value, the adjacent sample (118) for which the respective support value is determined. Thus, in Figure 7.4, in step 813, the value of sample 119 may be obtained by using the values of a given sample 118''' (previously obtained in step 812) and the adjacent samples 118 as support values.
[0145] The plurality of adjacent samples may extend in one dimension along two sides of a given block 18. It may be possible to perform the reduction 811 by grouping the plurality of adjacent samples 17 into one or more groups 110 of consecutive adjacent samples and performing downsampling or averaging on each of the one or more groups 110 of adjacent samples having two or more adjacent samples.
[0146] In the example, a linear or affine linear transformation is red *Q red or P red *Q may include weighting factors, P red is the number of sample values (102) in the reduced set of sample values, and Q red or Q is the number of given samples in a given block (18). red *Q red or 1 / 4P red *Q weighting coefficients are non-zero weight values. P red *Q red or P red *Q weighting factors are Q or Q red For each of the given samples, a set of P associated with each given sample is red The weighting coefficients may include a series of weighting coefficients that, when arranged one above the other in a raster scan order between given samples of a given block (18), form an envelope that is nonlinear in all directions.red *Q or P red *Q red The weighting factors may be independent of one another through a regular mapping rule. The average of the maximum value of the cross-correlation between a first series of weighting factors associated with each given sample and a second series of weighting factors associated with given samples other than each given sample, or the inverse of the latter series, is lower than a predetermined threshold value, even though the maximum value is higher. The predetermined threshold value may be 0.3 [or, in some cases, 0.2 or 0.1]. P red The adjacent samples (17) may be located along a one-dimensional path extending along two sides of a given block (18), and may be Q or Q red For each of the given samples, a set of P associated with each given sample is red The weighting factors are ordered to traverse a one-dimensional path in a given direction.
[0147] 6.1 Method and Apparatus Description
number
number
[0148] The generation of a predicted signal (eg, the value of a complete block 18) can be based on at least some of the following three steps: 1. Of the boundary samples 17, samples 102 (e.g., 4 samples when W=H=4 and / or 8 samples in other cases) may be extracted by averaging or downsampling (e.g., step 811). 2. A matrix-vector multiplication followed by an offset addition can be performed with the averaged samples (or samples remaining from downsampling) as input, and the result can be a reduced prediction signal for the set of subsampled samples in the original block (e.g., step 812). 3. Predictions at the remaining positions may be generated from the predictions on the subsampled set, for example by upsampling, for example by linear interpolation (eg, step 813).
[0149] Because of step 1 (811) and / or step 3 (813), the total number of multiplications required to compute the matrix-vector product is always
number
[0150] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b k ) is a set of matrices (e.g., a set of three), e.g.
number
[0151] In some cases, the set
number
number
number
[0152] In some cases, the set
number
number
number
number
number
[0153] Additionally, or alternatively, the set
Number
Number
Number
[0154] This set of these matrices and offset vectors or a subset of matrices and offset vectors may be used for all other block shapes.
[0155] 6.2 Averaging or Downsampling of Borders Here, features regarding step 811 are provided.
[0156] As described above, boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0157] In the first step, the input boundary
number
number
number
number
number
number
[0158] For a 4x4 block,
number
number
number
[0159] In all other cases (e.g., for blocks with a wither width or height different from 4), the block width W
number
number
number
[0160] In yet another case, the boundary can be downsampled (e.g., by selecting one particular boundary sample from a group of boundary samples) to arrive at a reduction in the number of samples.
number
number
number
[0161] Two reduced boundaries
number
number
number
[0162] where, if the mode (or the number of matrices in the set of matrices) is
number
[0163] For mode −17, which corresponds to a transposed mode of mode ≧18, it is possible to define
number
[0164] Therefore, according to a specific state (one state: mode < 18; another state: mode ≥ 18), we can compare the predicted values of the output vectors with different scan orders (e.g., one scan order).
number
number
[0165] Other strategies can be implemented. In other examples, the mode index "mode" does not necessarily lie in the range 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 does not necessarily have to have 18 matrices (thus, instead of a formula such as mode ≥ 18, the number of matrices in each set S0, S1, and S2, respectively, can be used).
number
[0166] The mode and transpose information is not necessarily stored and / or transmitted as one combined mode index "mode." In some examples, it may be explicitly signaled as a transpose flag and matrix index (0-15 for S0, 0-7 for S1, 0-5 for S2).
[0167] In some cases, a combination of a transpose flag and a matrix index may be interpreted as a set index. For example, there may be one bit that acts as a transpose flag and several bits that indicate matrix indexes, collectively referred to as a "set index."
[0168] 6.3 Generation of reduced prediction signal by matrix-vector multiplication Now, characteristics regarding step 812 are provided.
[0169] Reduced Input Vector
number
number
number
number
number
[0170] Shrinkage prediction signal
number
number
number
number
[0171]
number
number
number
number
[0172] Matrix A and
number
number
number
number
number
number
number
number
[0173] Other strategies can be implemented. In other examples, the mode index "mode" does not necessarily lie in the range 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, and S2 does not necessarily have to have 18 matrices (thus, instead of a formula like mode<18, it is the number of matrices in each set of matrices S0, S1, and S2, respectively).
number
[0174] 6.4 Linear Interpolation to Generate the Final Prediction Signal Now, characteristics regarding step 812 are provided.
[0175] For the interpolation of subsampled predictions in large blocks, a second version of the averaged boundary may be required:
number
number
number
number
number
number
number
[0176]
number
number
number
[0177] The linear interpolation may be given as follows (other examples are possible):
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0178] Here is an example of interpolation where the first interpolation (horizontal or vertical) uses the scaled down boundary samples, and the second interpolation (vertical or horizontal) uses the original boundary samples. Depending on the block size, only the second interpolation may be needed, or no interpolation may be needed. If both horizontal and vertical interpolation are needed, the order depends on the width and height of the block.
[0179] However, different techniques can be implemented, for example the original boundary samples can be used for both the first and second interpolation, and the order can be fixed, for example first horizontal, then vertical (or alternatively first vertical, then horizontal).
[0180] Therefore, the interpolation order (horizontal / vertical) and use of reduced / original boundary samples may be changed.
[0181] 6.5 Explaining the complete ALWIP process example The whole process of averaging, matrix-vector multiplication, and linear interpolation is illustrated for different shapes in Figures 7.1 to 7.4. It should be noted that the remaining shapes can be treated as one of the illustrated cases. Given a 1.4x4 block, ALWIP can take two averages along each axis of the boundary using the technique in Figure 7.1. The resulting four input samples are then input to a matrix-vector multiplication. The matrix is
number
number
number
number
number
[0182] 6.6 Evaluation of the number and complexity of parameters required The required parameters for all possible proposed intra prediction modes are:
number
[0183] 6.7 Signaling the proposed intra-prediction mode For luma blocks, for example, 35 ALWIP modes are proposed (other numbers of modes can also be used). For each coding unit (CU) in intra mode, a flag is transmitted in the bitstream indicating whether the ALWIP mode should be applied to the corresponding prediction unit (PU). The signaling of the latter index can be coordinated with the MRL in the same way as the first CE test. If the ALWIP mode is applied, the
number
[0184] Here, the derivation of MPM may be performed using the intra modes of the above and left PUs as follows:
number
number
number
number
number
number
number
number
[0185] The above PU is available, belongs to the same CTU as the current PU, is in intra mode, and is in conventional intra prediction mode.
number
number
[0186] In all other cases,
number
number
[0187] Finally, three fixed default lists
number
number
[0188] 6.8 Adaptive MPM List Derivation for Conventional Luma and Chroma Intra Prediction Modes The proposed ALWIP mode can be harmonized with conventional MPM-based coding of intra prediction modes as follows: The luma and chroma MPM list derivation process of conventional intra prediction modes is based on a fixed table
number
number
number
[0189] For Luma MPM list derivation, ALWIP mode
number
number
[0190] 7 Efficient implementation The above example will be briefly summarized as it can form the basis for further extensions of the embodiments described below.
[0191] To predict a given block 18 of a picture 10, a method using multiple neighboring samples 17a,c is used.
[0192] A reduction 100 by averaging multiple neighboring samples is performed to obtain a reduced set of sample values 102, which has a smaller number of samples compared to the multiple neighboring samples. This reduction is optional in embodiments herein and generates a so-called sample value vector, as described below. The reduced set of sample values is subjected to a linear or affine-linear transformation 19 to obtain a predicted value for a given sample 104 of a given block. This transformation is obtained by machine learning (ML) and is shown below using a matrix A and an offset vector b, which should be implemented efficiently.
[0193] Through interpolation, predicted values of further samples 108 of a given block are derived based on predicted values of the given sample and multiple adjacent samples. Theoretically, the result of the affine / linear transformation can be related to non-full-pel sample positions of the block 18, so that all samples of the block 18 can be obtained by interpolation according to an alternative embodiment. Interpolation may not even be necessary at all.
[0194] The plurality of neighboring samples may extend one-dimensionally along two sides of the predetermined block, with the predetermined samples arranged in rows and columns, and the predetermined samples may be arranged along at least one of the rows and columns, with the predetermined samples being positioned at every n-th position from the sample (112) of the predetermined sample adjacent to the two sides of the predetermined block. Based on the plurality of neighboring samples, a support value for one of the plurality of neighboring positions (118) may be determined for each of the rows and / or columns, which is aligned with the respective row and / or column. By interpolation, predicted values for further samples 108 of the predetermined block may be derived based on the predicted value for the predetermined sample and the support values of the neighboring samples aligned with the respective row and / or column. The predetermined sample may be positioned along the rows at every n-th position from the sample (112) of the predetermined sample adjacent to the two sides of the predetermined block, and the predetermined sample may be positioned along the columns at every m-th position from the sample (112) of the predetermined sample adjacent to the two sides of the predetermined block, where n and m > 1. n may equal m. Along at least one of the rows and columns, the determination of the support values may be performed by averaging 122 groups 120 of adjacent samples within the plurality of adjacent samples, including the adjacent sample 118 for which the respective support value is determined. The plurality of adjacent samples may extend one-dimensionally along two sides of the predetermined block, and the reduction may be performed by grouping the plurality of adjacent samples into one or more groups 110 of contiguous adjacent samples and performing the averaging for each of the one or more groups of adjacent samples having three or more adjacent samples.
[0195] For a given block, a prediction residual may be transmitted in the data stream, from which it is derived in the decoder, which uses the prediction residual and predicted values of a given sample to reconstruct the given block, and in the encoder, the prediction residual is coded into the data stream.
[0196] The picture may be subdivided into multiple blocks of different block sizes, including the predetermined block. Next, a linear or affine-linear transformation of block 18 is selected according to the width W and height H of the predetermined block, such that the linear or affine-linear transformation selected for the predetermined block is selected from the first set of linear or affine-linear transformations as long as the width W and height H of the predetermined block are within a first set of width / height pairs, and is selected from the second set of linear or affine-linear transformations as long as the width W and height H of the predetermined block are within a second set of width / height pairs that are different from the first set of width / height pairs. Similarly, it will be apparent later that the affine / linear transformation is represented by other parameters, namely, a weight C, and optionally offset and scale parameters.
[0197] The decoder and encoder may be configured to subdivide a picture into a plurality of blocks of different block sizes, including a predetermined block, and select a linear or affine-linear transformation according to a width W and a height H of the predetermined block, such that the selected linear or affine-linear transformation for the predetermined block is: As long as the width W and height H of a given block are within the first set of width / height pairs, a second set of linear or affine-linear transformations, so long as the width W and height H of a given block are within a second set of width / height pairs that are different from the first set of width / height pairs; and A third set of linear or affine-linear transformations is selected from the third set, so long as the width W and height H of a given block are within a third set of one or more width / height pairs that are different from the first and second sets of width / height pairs.
[0198] The third set of one or more width / height pairs includes only one width / height pair W', H', and each linear or affine-linear transformation in the first set of linear or affine-linear transformations is for transforming the N' sample values to a W'*H' predicted value in a W'xH' array of sample locations.
[0199] Each of the first and second sets of width / height pairs is W p H p First width / height vs. W not equal p , H p And, H q =W p and W q =H p The second width / height pair W is q , H q and
[0200] Each of the first and second sets of width / height pairs is connected to a third width / height pair W p is H p and W p is H p is equal to H p >H q is.
[0201] For a given block, a set index may be transmitted in the data stream that indicates which linear or affine-linear transform from a given set of linear or affine-linear transforms should be selected for the block 18.
[0202] The plurality of adjacent samples may extend one-dimensionally along two sides of the predetermined block, and the reduction may be performed by, for a first subset of the plurality of adjacent samples adjacent to a first side of the predetermined block, grouping the first subset into a first group 110 of one or more consecutive adjacent samples, and for a second subset of the plurality of adjacent samples adjacent to a second side of the predetermined block, grouping the second subset into a second group 110 of one or more consecutive adjacent samples, and performing averaging for each of the first and second groups of one or more adjacent samples having three or more adjacent samples to obtain a first sample value from the first group and a second sample value for the second group. Then, a linear or affine-linear transformation can be selected from a predetermined set of linear or affine-linear transformations depending on the set index, such that two different states of the set index result in the selection of one of the linear or affine-linear transformations of the predetermined set of linear or affine-linear transformations. For the set index assuming a first of two different states in the form of a first vector, the reduced set of sample values can be subjected to the predetermined linear or affine-linear transformation to generate an output vector of predicted values. The predicted values of the output vector can be distributed to predetermined samples of a predetermined block and of a predetermined block along the first scanning order. For the set index assuming a second of two different states in the form of a second vector, the first and second vectors differ such that a component input by one of the first sample values of the first vector is input by one of the second sample values of the second vector, and a component input by one of the second sample values of the first vector is input by one of the first sample values of the second vector, to generate an output vector of predicted values. The predicted values of the output vector can be distributed to predetermined samples of a predetermined block transposed with respect to the first scanning order along the second scanning order.
[0203] Each linear or affine linear transform in the first set of linear or affine linear transforms may be for transforming N1 sample values into w1*h1 predicted values for a w1xh1 array of sample locations, and each linear or affine linear transform in the first set of linear or affine linear transforms may be for transforming N2 sample values into w2*h2 predicted values for a w2xh2 array of sample locations, where for a first predetermined one of the first set of width / height pairs, w1 may exceed the width of the first predetermined width / height pair or h1 may exceed the height of the first predetermined width / height pair, and for a second predetermined one of the first set of width / height pairs, w1 may not exceed the width of the second predetermined width / height pair and h1 may not exceed the height of the second predetermined width / height pair. The averaging (100) of the plurality of adjacent samples to obtain a reduced set of sample values (102) may be performed such that the reduced set of sample values 102 has N1 sample values if the given block is of a first predetermined width / height pair and if the given block is of a second predetermined pair, and subjecting the reduced set of sample values to a selected linear or affine-linear transformation may be performed along the width dimension if w1 exceeds the width of one width / height pair, or along the height dimension if h1 exceeds the height of one width / height pair if the given block is of the first predetermined width / height pair, or if the given block is of the second predetermined width / height pair, the selected linear or affine-linear transformation may be performed entirely using only a first sub-portion of the selected linear or affine-linear transformation associated with subsampling the w1 x h1 array of sample locations.
[0204] Each linear or affine linear transformation in the first set of linear or affine linear transformations may be for transforming N1 sample values into w1*h1 predicted values for a w1xh1 array of sample locations where w1=h1, and each linear or affine linear transformation in the first set of linear or affine linear transformations is for transforming N2 sample values into w2*h2 predicted values for a w2xh2 array of sample locations where w2=h2.
[0205] All of the above-described embodiments are merely exemplary in that they may form the basis for embodiments described hereinafter. That is, the concepts and details above are useful for understanding the following embodiments and serve as a repository for possible extensions and modifications of the embodiments described hereinafter. In particular, many of the details described above are optional, such as the averaging of adjacent samples, the fact that adjacent samples are used as reference samples, etc.
[0206] More generally, the embodiments described herein assume that a prediction signal on a rectangular block is generated from already reconstructed samples, such as an intra prediction signal on a rectangular block being generated from already reconstructed samples in the left and top neighboring parts of the block. The generation of the prediction signal is based on the following steps: 1. However, among the reference samples, called boundary samples here, samples can be extracted by averaging, excluding the possibility of transferring the description to reference samples located elsewhere. Here, averaging is performed on both the left and top boundary samples of the block, or only on one of the boundary samples on either side. If averaging is not performed on one side, the sample on that side remains unchanged. 2. An optional addition of an offset is followed by a matrix-vector multiplication, where the input vector is either the concatenation of the averaged boundary sample to the left of the block if averaging was applied only to the left, or the concatenation of the original boundary sample to the left of the block and the averaged boundary sample above the block if averaging was applied only to the top side of the block, or the concatenation of the averaged boundary sample to the left of the block and the averaged boundary sample above the block if averaging was applied on both sides of the block. Again, alternatives exist, including those that do not use averaging at all. 3. The result of the matrix-vector multiplication and optional offset addition may optionally be a downscaled prediction of a subsampled set of samples in the original block. Predictions at the remaining positions may be generated from the predictions on the subsampled set by linear interpolation.
[0207] The matrix-vector product calculation in step 2 should preferably be performed in integer arithmetic. Thus,
number
number
number
number
number
number
number
number
number
number
number
number
number
[0208] 8 illustrates an improved ALWIP prediction. Samples of a given block can be predicted based on a first matrix-vector product between a matrix A 1100 derived by some machine learning-based training algorithm and a sample value vector 400. Optionally, an offset b 1110 can be added. To achieve an integer or fixed-point approximation of this first matrix-vector product, the sample value vector can undergo a reversible linear transform 403 to determine a further vector 402. A second matrix-vector product between a further matrix B 1200 and the further vector 402 can be equal to the result of the first matrix-vector product.
[0209] Due to the characteristics of the further vector 402, the second matrix-vector product may be an integer approximated by a matrix-vector product 404 between the predetermined prediction matrix C 405, the further vector 402, and the further offset 408. The further vector 402 and the further offset 408 may consist of integer or fixed-point values. For example, all components of the further offset are the same. The predetermined prediction matrix 405 may be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the predetermined prediction matrix 405 and the further vector 402 may be understood as a prediction vector 406.
[0210] Below we provide more details about this integer approximation.
[0211] Possible solution according to Example I: Subtracting and adding average values Expressions that can be used in the above scenarios
number
number
number
number
[0212]
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0213]
number
number
number
number
[0214] Alternatively, the predetermined value 1400 is a default value or a value signaled in the data stream in which the picture is encoded.
[0215] For example, a predetermined value of 1400 may be advantageous if the samples of a given block have small deviations from the expected values.
[0216] According to one embodiment, the device is configured to include a plurality of reversible linear transforms 403, each of which is associated with one component of the further vector 402. Furthermore, the device is configured to, for example, select a predetermined component 1500 from among the components of the sample value vector 400, and to use the reversible linear transform 403 as the predetermined reversible linear transform from among the plurality of reversible linear transforms associated with the predetermined component 1500. This is due to, for example, the different positions of the rows of the reversible linear transform 403 corresponding to the predetermined component depending on the i0-th row, i.e. the position of the predetermined component in the further vector. For example, if the first component of the further vector 402, i.e. y1, is the predetermined component, then the i o The th row replaces the first row of the reversible linear transformation.
[0217] 9b, the column 412 of the predetermined prediction matrix 405 corresponding to the predetermined component 1500 of the further vector 402, i.e., the matrix component 414 of the predetermined prediction matrix C 405 in the i0-th column, is, for example, all zero. In this case, the apparatus is configured to calculate the matrix-vector product 404 by performing a multiplication by calculating a matrix-vector product 407 between the reduced prediction matrix C′ 405 resulting from the predetermined prediction matrix C 405 by retaining the column 412 and the yet further vector 410 resulting from the further vector 402 by retaining the predetermined component 1500, as shown in FIG. 9c. Thus, the prediction vector 406 can be calculated with fewer multiplications.
[0218] 8, 9b and 9c, the apparatus may be configured to, when predicting samples of a predetermined block based on the prediction vector 406, calculate, for each component of the prediction vector 406, the sum of the respective component and a, i.e., a predetermined value 1400. This sum can be represented by the sum of the prediction vector 406 and the vector 409, in which all components of the vector 409 are equal to the predetermined value 1400, as shown in Figures 8 and 9c. Alternatively, the sum can be represented by the sum of the prediction vector 406 and a matrix-vector product 1310 between an integer matrix M 1300 and the further vector 402, as shown in Figure 9b, in which the matrix component of the integer matrix 1300 is one in the column of the integer matrix 1300 corresponding to the predetermined component 1500 of the further vector 402, i.e., the i0-th column, and all other components are, for example, zero.
[0219] The result of the sum of the given prediction matrix 405 and the integer matrix 1300 is, for example, equal to or approximates the further matrix 1200 shown in FIG.
[0220] In other words, the matrix obtained by summing each matrix element of the predetermined prediction matrix C 405 in the i0-th column 412 corresponding to the predetermined element 1500 of the further vector 402, i.e., the i0-th column, with the reversible linear transform 403 (i.e., matrix B), i.e., the further matrix B 1200, corresponds to a quantized version of the machine learning prediction matrix A 1100, as shown in, for example, Figures 8, 9a, and 9b. As shown in Figure 9b, summing each matrix element of the predetermined prediction matrix C 405 in the i0-th column 412 with 1 can correspond to the sum of the predetermined prediction matrix 405 and the integer matrix 1300. As shown in Figure 8, the machine learning prediction matrix A 1100 can be equal to the result of multiplying the further matrix 1200 with the reversible linear transform 403. This
number
[0221] Matrix multiplication using only integer arithmetic For a low complexity implementation (in terms of the complexity of adding and multiplying scalar values, and the storage required for the entries of the participation matrices), it is desirable to perform the matrix multiplication 404 using only integer arithmetic.
[0222]
number
number
number
number
[0223] The matrix vector product 404 with a matrix of size m×n, i.e., a given prediction matrix 405, can be performed as shown in this pseudocode, where <<, >> are arithmetic binary left and right shift operations, and +, - and * operate on integer values only. (1) final_offset=1<<(right_shift_result-1); for i in 0…m-1 { accumulator=0 for j in 0…n-1 { accumulator:=accumulator + y[j]*C[i,j] } z[i]=(accumulator+final_offset)>>right_shift_result; }
[0224] Here, array C, i.e., the predetermined prediction matrix 405, stores fixed point numbers, e.g., as integers. The final addition of final_offset and right shift operation output by right_shift_result reduces precision by rounding to obtain the required fixed point format at the output.
[0225] To be able to extend the range of real values representable by integers in C, two additional
number
number
number
number
number
number
[0226] In other words, the device may calculate predicted parameters, e.g.
number
[0227] According to one embodiment, the prediction parameters include weights each associated with a corresponding matrix element of a prediction matrix. In other words, a given prediction matrix is, for example, replaced or represented by the prediction parameters. The weights are, for example, integer and / or fixed-point values.
[0228] According to one embodiment, the prediction parameters are determined by one or more scaling factors, e.g., the value scale i,j each of which is a weight associated with one or more corresponding matrix elements of the predetermined prediction matrix 405, e.g.
number
number
[0229]
number
number
number
[0230]
number
[0231] A wide range of embodiments resulting from the solution The above solution implies the following embodiments: 1. A prediction method as described in Section I, wherein in Step 2 of Section I, the following is done for integer approximations of the matrix-vector products involved:
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0232] That is, according to an embodiment of the present application, the encoder and decoder operate as follows to predict a given block 18 of picture 10, with reference to Figure 8 in combination with one of Figures 7.1 to 7.4: For the prediction, multiple reference samples are used. As outlined above, the embodiment of the present application is not limited to intra-coding, and therefore the reference samples are not limited to being adjacent samples, i.e., samples of picture 10 adjacent to block 18. In particular, the reference samples are not limited to those located along the outer edge of block 18, such as samples abutting the outer edge of the block. However, this situation is certainly an embodiment of the present application.
[0233] To perform the prediction, a sample value vector 400 is formed from reference samples, such as reference samples 17a and 17c. Possible formations are described above. The formation may involve averaging, which may reduce the number of samples 102 or the number of components of the vector 400 compared to the reference samples 17 that contribute to the formation. The formation may also depend in some way on the dimensions or size, such as the width and height, of the blocks 18, as described above.
[0234] It is this vector 400 that must be subjected to an affine or linear transformation to obtain the prediction of block 18. A different nomenclature is used above. The aim is to perform the prediction by applying vector 400 to matrix A by a matrix-vector product within which, using the latest, a summation with an offset vector b is performed. The offset vector b is optional. The affine or linear transformation determined by A or A and b can be determined by the encoder and decoder, or more precisely, for the prediction based on the size and dimensions of block 18, as already mentioned above.
[0235] However, to achieve the computational efficiency improvements outlined above or to make prediction more effective from an implementation perspective, the affine or linear transformation is quantized, and the encoder and decoder, or their predictors, use C and T, which represent the quantized version of the affine transformation and are applied in the manner described above, to represent and perform the linear or affine transformation. In particular, instead of directly applying vector 400 to matrix A, the predictors of the encoder and decoder apply vector 402, obtained from sample value vector 400, by subjecting it to a mapping through a predetermined invertible linear transformation T. The transformation T used here is the same, or at least the same for different affine / linear transformations, as long as vector 400 is the same size, i.e., independent of the block dimensions, i.e., width and height. In the above, vector 402 is denoted y. The exact matrix for performing the affine / linear transformation, determined by machine learning, was B. However, instead of executing B exactly, the predictions in the encoder and decoder are made by its approximation or a quantized version. In particular, the representation is done by appropriately representing C in the manner outlined above, and C+M represents a quantized version of B.
[0236] Therefore, prediction in the encoder and decoder is further performed by calculating a matrix-vector product 404 between the vector 402 and a predetermined prediction matrix C, which is appropriately represented and stored in the encoder and decoder in the manner described above. The vector 406 resulting from this matrix-vector product is then used to predict the samples 104 of the block 18. As described above, for the prediction, each component of the vector 406 may be subjected to a summation with a parameter a, as indicated at 408, to compensate for the corresponding definition of C. An optional summation of the vector 406 with an offset vector b may also be involved in deriving the prediction of the block 18 based on the vector 406. As described above, each component of the vector 406, and thus each component of the sum of the vector 406, the vector of all a's indicated at 408, and the optional vector b, directly corresponds to a sample 104 of the block 18 and may therefore indicate a predicted value of the sample. It is also possible that only a subset of the samples 104 of the block are predicted in this way, with the remaining samples of the block 18, such as 108, being derived by interpolation.
[0237] As mentioned above, there are different embodiments for setting α. For example, it may be the arithmetic mean of the components of vector 400. In this case, see, for example, FIG. 9. The invertible linear transform T may be as shown in FIG. 9a. i0 denotes a predetermined component of the sample value vector and vector 402, respectively, and is replaced by a. However, as mentioned above, there are other possibilities. However, as far as the representation of C is concerned, it has also been mentioned above that it may be embodied differently. For example, the matrix-vector product 404 may end up in its actual calculation as a smaller matrix-vector product having a lower dimension, see, for example, FIG. 9c. In particular, as mentioned above, due to the definition of C, its entire i0-th column 412 may be zero, and therefore the actual calculation of product 404 may be
number
[0238] The weights of C or C', i.e., the elements of this matrix, can be represented and stored in a fixed-point representation. However, these weights 414 can also be stored in a manner associated with different scales and / or offsets, as described above. The scale and offset can be defined for the entire matrix C, i.e., equal for all weights 414 in matrix C or matrix C', or constant or equal for all weights 414 in the same row or column of matrices C and C', respectively. In this regard, FIG. 10 illustrates that the matrix-vector product calculation, i.e., the result of the product, can actually be performed slightly differently, i.e., by shifting the multiplication with the scale(s) toward vector 402 or 404, thereby reducing the number of further multiplications that must be performed. FIG. 11 illustrates the use of one scale and one offset for all weights 414 in C or C', as in calculation (2) above.
[0239] 8. Special Case Embodiments Using Block-Based Intra Prediction Mode for All Channels All of the above descriptions shall be considered optional implementation details of the embodiments described herein. Note that hereinafter, the term matrix-based intra prediction (MIP) is used to denote an intra prediction mode, such as a block-based intra prediction mode, which may be embodied by or equivalent to the ALWIP described above, as described with respect to Figures 5-11.
[0240] In this document, for 4:4:4 chroma format and single tree, on chroma intra blocks whose chroma intra mode is direct mode (DM) mode and whose luma intra mode is MIP mode, we propose that the chroma intra prediction signal is generated using this MIP mode.
[0241] In direct mode, the chroma intra prediction mode is derived from the luma intra prediction mode. For example, the chroma intra prediction mode is equal to the luma intra prediction mode. An exception is when the luma intra prediction mode is the MIP mode; the chroma intra prediction mode is the MIP mode only when the color sampling format is 4:4:4; in the single-tree case and all other cases, the chroma intra prediction mode is the planar mode. The 4:4:4 color sampling format represents a color sampling format according to which each color component is sampled evenly.
[0242] 8.1 Description of the proposed method In the current VTM, MIP is only used for the luma component [3]. If the intra mode of a chroma intra block is direct mode (DM) and the intra mode of the co-located luma block is MIP mode, the chroma block must use planar mode to generate the intra prediction signal. The main reason for treating DM mode in this way in the MIP case is that in the 4:2:0 case or dual-tree case, the co-located luma block may have a different shape than the color blocks. Therefore, since MIP mode is not applicable to all block shapes, the MIP mode of the co-located luma block may not be applicable to the chroma block in that case.
[0243] On the other hand, when the chroma format is 4:4:4 and a single tree is used, it can be said that the above luma MIP mode cannot be applied to the chroma component. Therefore, in this case, when the chroma intra mode is the DM mode on an intra block and the luma intra mode is the MIP mode, it is proposed that the chroma intra prediction signal be generated in this MIP mode.
[0244] This proposed change is particularly relevant when adaptive color transformation (ACT) is enabled, since the chroma mode is assumed to be DM mode when ACT is used for a particular block. Specifically for ACT, a strong correlation of predicted signals across all channels is beneficial, and such correlation can be enhanced if the same intra-prediction mode is used for all three components of a block. The experimental results reported below support this view, in that the proposed change significantly impacts camera-captured content in RGB format, which is claimed to be the intersection for which MIP and ACT were primarily designed.
[0245] 8.2 Implementation of the proposed method 12 illustrates one embodiment of a prediction of a given second color component block 182 of picture 10. A block-based decoder, a block-based encoder, and / or an apparatus for predicting picture 10 may be configured to perform this prediction.
[0246] According to one embodiment, a division scheme 11' is used in which the picture 10 is divided equally for each color component 101, 102, e.g., into a first color component block 18' 11 ~18' 1n and the second color component block 18' 21 ~18' 2n1, a picture 10 has a plurality of color components 101, 102 and a color sampling format 11, in which each color component 101, 102 is sampled equally into blocks. According to the color sampling format 11, each color component 101, 102 is, for example, sampled equally, and according to the division scheme 11', each color component 101, 102 is, for example, divided equally. Thus, for example, according to the division scheme 11', a first color component block 18' 11 ~18' 1n is the second color component block 18' 21 ~18' 2n and according to the color sampling format 11, the first color component block 18' 11 ~18' 1n is the second color component block 18' 21 ~18' 2n The sampling is indicated by a sample point within a given second color component block 182 and the co-located intra-predicted first color component block 181.
[0247] The first color component 101 of the picture 10 is an intra-predicted first color component block 18' of the picture 10. 11 ~18' 1n For each of the blocks, the block interior 18 derives a sample value vector 514 of the reference sample 17, and the block interior 18 is adjacent to the block interior 18. To obtain the prediction vector 514, the sample value vector 514 and each matrix-based intra prediction mode 5101 to 510 are used. m The matrix-based intra prediction modes (MIP modes) 5101-510 are used to predict the block interior 18 by calculating a matrix-vector product 512 between the block interior 18 and a prediction matrix 516 associated with the block interior 18 and predicting the block interior 18 samples based on the prediction vector 514. mThe sample value vector 514 has the characteristics and / or functionality described with respect to the sample value vector 400 in FIG. 6 or with respect to the further vector 402 in FIGS. 8-11. In other words, the matrix-based intra predictions 5101-510 m is performed, for example, as described with respect to one of FIGS.
[0248] Optionally, the number of components of the prediction vector 518 is less than the number of samples in the block interior 18. In this case, the samples in the block interior 18 may be predicted based on the prediction vector 518 by interpolating samples based on the components of the prediction vector 518 assigned to supporting sample positions in the block interior 18.
[0249] Optionally, matrix-based intra-prediction modes 5101-510 comprise the first set of intra-prediction modes 508. m is selected from the set of disjoint subsets of matrix-based intra-prediction modes according to the block dimension of each intra-predicted first color component block 181 to be a subset of matrix-based intra-prediction modes. The prediction matrices 516 associated with the set of disjoint subsets of matrix-based intra-prediction modes are, for example, machine-learned, and the prediction matrices 516 of one subset of matrix-based intra-prediction modes are equal in size, while the prediction matrices 516 of two subsets of matrix-based intra-prediction modes selected for different block sizes are different in size. In other words, the prediction matrices of one subset are equal in size, but the prediction matrices of different subsets have different sizes.
[0250] Optionally, matrix-based intra-prediction modes 5101-510 comprise the first set of intra-prediction modes 508. m The prediction matrices 516 associated with the s are of equal size to each other and are machine learned.
[0251] Optionally, matrix-based intra-prediction modes 5101 to 5102 comprise the first set of intra-prediction modes 508. m The intra prediction modes of the first set 508 of intra prediction modes other than 501 include a DC mode 506, a planar mode 504, and a directional mode 5001-5001.
[0252] The second color component 102 of picture 10 is decoded and / or coded block-wise by intra-predicting a given second color component block 182 of picture 10 using a matrix-based intra-prediction mode selected for a co-located intra-predicted first color component block 181. Thus, the decoder and / or encoder selects one of MIP modes 5101-510 for at least one intra-coded second / third color component block. m This is done, for example, when the initial prediction mode of a given second color component block 182 is direct mode.
[0253] In direct mode, the intra-prediction mode of a given second color component block 182 is derived, for example, from the intra-prediction mode of the co-located intra-predicted first color component block 181. For example, the intra-prediction mode of a given second color component block 182 is equal to the intra-prediction mode of the co-located intra-predicted first color component block 181. In other words, if the intra-prediction mode of the co-located intra-predicted first color component block 181 is one of the MIP modes 5101-510, m If so, the predetermined second color component block 182 also has the same MIP mode 5101-510 m is predicted by
[0254] According to one embodiment, if no index indicating a prediction mode is signaled for a given second color component block 182 in the data stream, direct mode can be inferred, otherwise the mode signaled in the data stream is used for predicting the given second color component block 182.
[0255] In other words, the second color component block 18' of the picture 10 21 ~18' 2n For each of the co-located intra-predicted first color component blocks 18′, one can be selected from a first option and a second option. 11 -18' 1n The intra prediction mode selected for the matrix-based intra prediction mode 5101 to 510 m If so, then each second color component block 18' 21 -18' 2n , the intra prediction mode for the co-located intra predicted first color component block 18' 11 ~18' 1n Each second color component block 18' is equal to 21 ~18' 2n , the intra prediction mode for the co-located intra predicted first color component block 18' 11 ~18' 1n The co-located intra-predicted first color component block 18' is derived based on the intra-prediction mode selected for the co-located intra-predicted first color component block 18'. 11 ~18' 1n The intra prediction mode selected for the image is a planar intra prediction mode 504, a DC intra prediction mode 506, and / or a directional intra prediction mode 5001-5002. l As shown in the matrix-based intra prediction modes 5101 to 510, m If the second color component block 18' is one of the other intra-prediction modes other than 21 ~18' 2n The intra prediction mode of the co-located intra predicted first color component block 18'11 ~18' 1n The first option may represent a direct mode. According to a second option, the first option may represent a direct mode. 21 ~18' 2n The intra prediction modes for the respective second color component blocks 18' 21 ~18' 2n Optionally, the intra-mode index is selected based on an intra-mode index present in the data stream for the co-located intra-predicted first color component block 18′ for selection from a first set of intra-prediction modes 508. 11 ~18' 1n In addition to the further intra-mode index present in the data stream for each second color component block 18', 21 ~18' 2n The intra-mode index present in the data stream for each second color component block 18' 21 ~18' 2n , indicating a prediction mode from, for example, the first set of intra-prediction modes 508 to be selected for prediction of .
[0256] According to one embodiment, when the 4:4:4 color sampling format 11 of a picture 10 and the prediction mode of a given second color component block 182 is direct mode, an apparatus for predicting a given block of a picture, a block-based decoder and / or a block-based encoder may predict the given second color component block 182 in the same MIP mode 5101-5102 as the co-located intra-predicted first color component block 181. mThe partitioning scheme 11′ is configured to use a co-located intra-predicted first color component block 181 and a predetermined second color component block 182, which represent predetermined blocks associated with different color components, i.e., the first color component and the second color component. The co-located intra-predicted first color component block 181 and the predetermined second color component block 182, for example, have the same geometric characteristics and the same spatial positioning in the picture 10. In other words, the first color component coding tree, for example, the luma coding tree, is equal to the second color component coding tree, for example, the chroma coding tree. For example, a single tree is used. The picture 10 is divided equally with respect to each color component 101, 102. The partitioning scheme 11′, for example, defines single-tree processing of the picture 10.
[0257] According to one embodiment, dual tree processing of picture 10 is also possible. Dual tree processing can be associated with a further partitioning scheme 11′, where picture 10 is partitioned with respect to a first color component 101 using first partitioning information in a data stream, and picture 10 is partitioned with respect to a second color component 2 using second partitioning information present in a data stream separate from the first partitioning information.
[0258] If not a dual-tree or single-tree, the device, block-based decoder and / or block-based encoder for predicting a given block of a picture may, for example, be configured to predict the given block of a picture in a manner that the prediction mode of the co-located intra-predicted first color component block 181 is MIP mode 5101-5102. mand the prediction mode of a given second color component block 182 is direct mode, the planar intra prediction mode 504 is configured to be used. However, if the prediction mode of the co-located intra-predicted first color component block 181 is not MIP mode, for example, if the prediction mode of the co-located intra-predicted first color component block 181 is DC intra prediction mode 506, planar intra prediction mode 504, or directional intra prediction mode 5001-5001 (e.g., angular intra prediction mode), and the prediction mode of a given second color component block 182 is direct mode, the prediction mode of the given second color component block 182 is equal to the prediction mode of the co-located intra-predicted first color component block 181. This may also be applicable when using a single tree but not using a 4:4:4 color sampling format 11, for example, when using a 4:1:1 color sampling format 11, a 4:2:2 color sampling format 11, or a 4:2:0 color sampling format 11.
[0259] The color sampling format 11 can be expressed as x:y:z, where the first number x refers to, for example, the size of the color component block 18′, and the next two numbers y and z both refer to the second and / or third color component samples. These, i.e., y and z, are both relative to the first number and define horizontal and vertical sampling, respectively. A signal with 4:4:4 is not compressed (and therefore not subsampled) and carries the first color component and additional color component data, e.g., the second and / or third color component samples, i.e., the entire chroma samples. In a 4x2 array of pixels, 4:2:2 has half the chroma samples of 4:4:4, while 4:2:0 and 4:1:1 have one-quarter of the available chroma information. A 4:2:2 signal has half the sampling rate in the horizontal direction but maintains full sampling in the vertical direction. On the other hand, 4:2:0 samples the color only from half the pixels in the first row and completely ignores the second row of samples, while 4:1:1 samples the color from only one pixel in the first row and one pixel in the second row of samples, i.e. a 4:1:1 signal has a quarter sampling rate horizontally but maintains full sampling vertically.
[0260] According to one embodiment, if the prediction mode of the co-located intra-predicted first color component block 181 is not a MIP mode 5101-5101, such as, for example, a DC intra-prediction mode 506, a planar intra-prediction mode 504, or a directional intra-prediction mode 5001-5001, and if the prediction mode of a given second color component block 182 is a direct mode, the device for predicting a given block of a picture, the block-based decoder and / or the block-based encoder is configured to use the prediction mode of the co-located intra-predicted first color component block 181 as the prediction mode of the given second color component block, independently of the color sampling format 11 and / or the partitioning scheme of the picture 10 for each color component, i.e., using a single tree or a dual tree.
[0261] In DC intra prediction mode 506, for example, a single value, a quasi-DC value, is derived based on neighboring samples 17 spatially adjacent to a given block 18, for example, a co-located intra predicted first color component block 181 and / or a given second color component block 181, and this single DC value is attributed to all samples of the given block 18 to obtain an intra prediction signal.
[0262] In planar intra prediction mode 504, for example, a two-dimensional linear function defined by a horizontal gradient, a vertical gradient, and an offset is derived based on neighboring samples 17 spatially adjacent to a given block 18, for example, a co-located intra-predicted first color component block 181 and / or a given second color component block 182, and this linear function defines the predicted sample values of the given block 18.
[0263] In directional intra-prediction modes 5001-5001, such as angular intra-prediction modes, reference samples 17 neighboring a given block 18, e.g., a co-located intra-predicted first color component block 181 and / or a given second color component block 182, are used to fill the given block 18 to obtain an intra-predicted signal for the given block 18. In particular, reference samples 17 located along the boundary of the given block 18, such as along the top and left edges of the given block 18, represent picture content that is extrapolated or copied along a predetermined direction 502 into the interior of the given block 18. Prior to extrapolation or copying, the picture content represented by the neighboring samples 17 may be subject to interpolation filtering, or in other words, may be derived from the neighboring samples 17 by interpolation filtering. Angular intra-prediction modes 5001-500 l The intra prediction directions 502 of the angle intra prediction modes 5001 to 5002 are different from each other. l may have associated indices, and the angular intra prediction modes 5001 to 500 lThe index association to the angular intra prediction modes 5001 to 5002 is performed according to the associated mode index. l may be such that the direction 502 rotates monotonically clockwise or counterclockwise when ordering.
[0264] FIG. 13 shows in more detail the prediction of a given second color component block 182 of picture 10 depending on different partitioning schemes and color sampling formats.
[0265] Prediction is shown for a picture 10a, another picture 10b, and yet a further picture 10c, with different conditions set for the different pictures.
[0266] According to one embodiment, an apparatus, such as a block-based decoder, a block-based encoder, etc. and / or an apparatus for predicting a picture, includes or has access to a set of partitioning schemes 11' to select from and / or obtain a partitioning scheme for a picture 10. The set of partitioning schemes 11' includes a partitioning scheme 11'1, i.e., a first partitioning scheme in which the picture 10 is partitioned equally with respect to each color component 101, 102, and a second partitioning scheme in which the picture 10 is partitioned with respect to the first color component 101 using first partitioning information 11'2 in the data stream 12, and the picture 10 is partitioned with respect to the first color component 101 using first partitioning information 11'3. 2a The second division information 11' exists in the data stream 12 separately from the 2b The first partitioning scheme 11'1 may represent a single-tree processing of the picture 10, and the further partitioning scheme 11'2 may represent a dual-tree processing of the picture.
[0267] In a further division scheme 11'2, the first color component 101 is divided differently from the second color component 102. One of the color components 101, 102 can be divided finely, while the other color component 101, 102 can be divided coarsely. 21 ~18'24 The boundary of the first color component block 18' 11 ~18' 116 It is possible that the second color component block 18' is not located at the same location as the boundary of the second color component block 18'. 21 ~18' 24 The boundary of 2alternative Or as shown in further division of picture 10b, first color component block 18' 11 ~18' 116 It is possible to pass through the interior of the block.
[0268] For picture 10a, the same predictions as described with respect to Figure 12 can be applied, with a first division scheme 11'1 selected from the set of division schemes 11' and a color sampling format being used in which each color component 10a1, 10a2 is sampled equally.
[0269] 13 is configured to select one of the first and second options for each intra-predicted second color component block of picture 10a. In the first option, the intra-prediction mode selected for co-located intra-predicted first color component block 18a1 is one of matrix-based intra-prediction modes 5101-510, as shown in FIG. m
[0046] If the intra-prediction mode for each intra-predicted second color component block 18a2 is one of
[0046] (i.e., for a given second color component block 18a2), the intra-prediction mode for each intra-predicted second color component block 18a2 is derived based on the intra-prediction mode selected for the co-located intra-predicted first color component block 18a1, such that the intra-prediction mode for each intra-predicted second color component block 18a2 is equal to the intra-prediction mode selected for the co-located intra-predicted first color component block 18a1. In a second option, the intra-prediction mode for each intra-predicted second color component block 18a2 is selected based on an intra-mode index 509 present in the data stream for each intra-predicted second color component block 18a2. The intra-mode index 509 is present in the data stream in addition to a further intra-mode index 507 present in the data stream for the co-located intra-predicted first color component block 18a1 for selection from the first set of intra-prediction modes.
[0270] Optionally, the device is configured to divide the further picture 10b of multiple color components 10b1, 10b2 and color sampling format into further blocks using a further division scheme 11'2, in which each color component is sampled equally. Thus, the first color component 10b1 has the same color sampling as the second color component 10b2, but the block sizes are different. As shown in Figure 13, a further block 18b2 of a given second color component of the further picture 10b has, for example, one-quarter the size of a further block 18b1 of the first color component at the same location in the further picture 10b. The further picture 10b is, for example, sampled in a 4:4:4 color sampling format.
[0271] When using the further division scheme 11'2, the further block 18b1 of the first color component that is co-located with the further block 18b2 of a given second color component is determined, for example, based on a pixel that is located in both blocks 18b1 and 18b2. According to one embodiment, this pixel is located in the upper left corner, the upper right corner, the lower left corner, the lower right corner and / or in the middle of the further block 18b2 of the given second color component.
[0272] The first color component 10b1 of the further picture 10 is decoded in units of further blocks by selecting one from the first set of intra prediction modes for each further block of the intra predicted first color component of the further picture 10b.
[0273] For each further block of the intra-predicted second color component further picture 10b, one of the first and second options can be selected. According to the first option, the intra-prediction mode for each intra-predicted second color component further block, e.g., a given second color component further block 18b2, is the matrix-based intra-prediction mode 5101-5102, and the intra-prediction mode selected for the co-located first color component further block 18b1 is the matrix-based intra-prediction mode 5101-5102. m , the intra-prediction mode selected for the further block 18b1 of the co-located first color component is one of the intra-prediction modes 501-502, for example, the intra-prediction mode is the planar intra-prediction mode 504, the DC intra-prediction mode 506, or the directional intra-prediction mode 5001-5001, and the matrix-based intra-prediction mode 5101-5102 is one of the intra-prediction modes 5101-5103. m, then the intra-prediction mode of each intra-predicted second-color component further block 18b2 is equal to the intra-prediction mode selected for the co-located first-color component further block 18b1. According to a second option, the intra-prediction mode for each intra-predicted second-color component block 18b2 is selected based on an intra-mode index 5092 present in the data stream 12 for the respective intra-predicted second-color component block 18b2. The intra-mode index 5092 is present in the data stream 12 in addition to a further intra-mode index 5072 present in the data stream for the co-located intra-predicted first-color component block 18b1, for example, for selection from the first set of intra-prediction modes.
[0274] Optionally, the device is configured to split the still further picture 10c in multiple color components and different color sampling formats, in which the multiple color components are sampled differently. As shown in FIG. 13, a first color component 10c1 is sampled differently from a second color component 10c2. The sampling is indicated by sample points. According to the embodiment shown in FIG. 13, the still further picture is sampled according to a 4:2:1 color sampling format. However, it is clear that other color sampling formats than the 4:4:4 color sampling format can also be used. This is shown once for the use of a first partitioning scheme 11'1 and once for the use of a further partitioning scheme 11'2. A subsequent prediction of the still further picture 10c can be performed independently of the partitioning scheme 11'.
[0275] The first color component 10c1 of the yet further picture 10c is decoded in units of yet further blocks by selecting one from the first set of intra prediction modes for each of the yet further blocks of the intra predicted first color component of the yet further picture 10c.
[0276] Further, the apparatus is configured to select, for example, for each intra-predicted second color component, still further blocks of the still further picture 10c from the first option and the second option, whereby the intra-prediction mode for each intra-predicted second color component further block 18c2 is the matrix-based intra-prediction mode 5101-5102, and the intra-prediction mode for each intra-predicted second color component further block 18c2 is the matrix-based intra-prediction mode 5101-5103. m If the intra prediction mode selected for the yet further block 18c1 of the co-located first color component is one of the planar intra prediction mode 504, the intra prediction mode selected for the yet further block 18c1 of the co-located first color component is equal to the planar intra prediction mode 504. If the intra prediction mode selected for the yet further block 18c1 of the co-located first color component is one of the planar intra prediction mode 504, the DC intra prediction mode 506, or the directional intra prediction modes 5001-5002, the intra prediction mode selected for the yet further block 18c1 of the co-located first color component is equal to the planar intra prediction mode 504, the DC intra prediction mode 506, or the directional intra prediction modes 5001-5002. l Matrix-based intra prediction modes 5101 to 510 m , then the intra-prediction mode for each intra-predicted yet further block 18c2 of the second color component is equal to the intra-prediction mode selected for the co-located yet further block 18c1 of the first color component. According to a second option, the intra-prediction mode for each intra-predicted yet further block 18c2 of the second color component is selected based on an intra-mode index 5093 present in the data stream for each intra-predicted yet further block 18c2 of the second color component. Optionally, the intra-mode index 5093 is present in the data stream 12 in addition to an intra-mode index 5073 present in the data stream for the co-located yet further block 18c1 of the intra-predicted first color component for selection from the first set of intra-prediction modes.
[0277] According to one embodiment, if the residual coding color transform mode is signaled to be deactivated, e.g., ACT (adaptive color transform) is deactivated, for blocks 18a2, 18b2, and / or 18c2 in data stream 12, the selection between the first and second options depends on signaling present in data stream 12 for the respective intra-predicted second color component block 18a2, the respective intra-predicted second color component further block 18b2, and / or the respective intra-predicted second color component yet further block 18c2. If the residual coding color transform mode is signaled to be activated, e.g., ACT (adaptive color transform) is activated, for blocks 18a2, 18b2, and / or 18c2 in data stream 12, no signaling is required. In this case, i.e., the residual coding color transform mode is signaled to be activated for blocks 18a2, 18b2 and / or 18c2, and, for example, when ACT (adaptive color transform) is activated, the device may be configured to infer that the first option should be selected.
[0278] When ACT is used on a given block, for example, the chroma mode is assumed to be DM mode, i.e., the first option. Specifically for ACT, it can be argued that a strong correlation of prediction signals across all channels is beneficial, and that such correlation can be enhanced if the same intra-prediction mode is used for all three components of the block. The experimental results reported below support this view, in that the proposed changes have a significant impact on camera-captured content in RGB format, which is claimed to be the intersection for which MIP and ACT were primarily designed.
[0279] 8.3 Experimental results In this section, experimental results are reported for 4:4:4 according to typical test conditions. Tables 1 and 2 report the results of the proposed changes compared to the VTM-7.0 anchor for AI and RA configurations, respectively. The corresponding simulations were performed on an Intel Xeon cluster (E5-2697A v4, AVX2 on, Turbo Boost off) with Linux OS and GCC 7.2.1 compiler. Results are reported for RGB and YUV according to CE-8 conditions. For single-tree off, the proposed changes do not change the bitstream, so single-tree is always enabled.
[0280] [Table 3] [Table 4] [Table 5] [Table 6]
[0281] 8.4 Conclusion This document proposes enabling MIP for all three channels in the case of 4:4:4 content and a single tree. In this case, if the luma component uses MIP mode on an intra block and the chroma intra mode is DM mode, it is proposed to generate a chroma intra prediction signal in this MIP mode. It is proposed that the technique described in this document be adopted in the next working draft of VVC.
[0282] 9 References [1] P. Helle et al., “Non-linear weighted intra prediction”, JVET-L0199, Macao, China, October 2018. [2] F. Bossen, J. Boyce, K. Suehring, X. Li, V. Seregin, “JVET common test conditions and software reference configurations for SDR video”, JVET-K1010, Ljubljana, SI, July 2018. [3] B.Bross, J.Chen, S.Liu, Y.-K.Wang,Verstatile Video Coding (Draft 7),Document JVET-P2001,Version 14 ,Geneva,Switzerland,October 2019 […]Common test conditions for 4:4:4
[0283] Further Embodiments and Examples Generally, the examples may be implemented as a computer program product having program instructions operable to perform one of the methods when the computer program product is executed on a computer. The program instructions may be stored on, for example, a machine-readable medium.
[0284] Another example comprises the computer program stored on a machine-readable medium for performing one of the methods described herein.
[0285] In other words, an example of a method is therefore a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
[0286] A further example of a method is therefore a data carrier medium (or digital storage medium, or computer readable medium) comprising recorded thereon a computer program for performing one of the methods described herein. The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-transitory, rather than an intangible and transitory signal.
[0287] A further example of a method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. A data stream or a sequence of signals may be transmitted via a data communication connection, for example via the Internet.
[0288] Further examples include a processing means, such as a computer, or a programmable logic device, for performing one of the methods described herein.
[0289] A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0290] Further examples include an apparatus or system that transfers (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0291] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods may be performed by any suitable hardware apparatus.
[0292] The above examples are merely illustrative to illustrate the principles described above. It is understood that modifications and variations of the arrangements and details described herein will become apparent. It is therefore the intention to be limited only by the scope of the appended claims and not by the specific details presented by way of description and illustration of the examples herein.
[0293] Equal or equivalent elements or elements with equal or equivalent functionality are represented in the following description by equal or equivalent reference signs, even if they occur in different figures.
Claims
1. 1. A video decoder comprising one or more processors, the one or more processors comprising: determining that the picture is in a 4:4:4 color sampling format; determining a matrix-based intra-prediction (MIP) mode based at least in part on an index signaled in the data stream; decoding luma blocks of the picture using the MIP mode; For a chroma block of the picture, determining whether the coding tree for the chroma block is a single tree; selecting an intra prediction mode for decoding the chroma block, in response to determining that the coding tree for the chroma block is a single tree, the selected intra prediction mode is an MIP mode; and in response to determining that the coding tree for the chroma block is not a single tree, the selected intra prediction mode is a planar intra prediction mode. selecting the intra-prediction mode; decoding the chroma block using the selected intra-prediction mode; and a video decoder configured to:
2. The video decoder of claim 1 , wherein the single tree indicates that the coding tree for the chroma block is the same as the coding tree for the luma block.
3. 10. The video decoder of claim 1, wherein the one or more processors are further configured to determine a MIP mode from a set of MIP modes that depend at least in part on a dimension of the luma block.
4. To decode the luma block, the one or more processors:
10. The video decoder of claim 1, further configured to multiply a vector based on at least a portion of above and left neighboring samples of the luma block by a matrix corresponding to a MIP mode to generate a prediction of the luma block.
5. determining that the picture is in a 4:4:4 color sampling format; determining a matrix-based intra-prediction (MIP) mode based at least in part on an index signaled in the data stream; decoding luma blocks of the picture using the MIP mode; For a chroma block of the picture, determining whether the coding tree for the chroma block is a single tree; selecting an intra prediction mode for decoding the chroma block, in response to determining that the coding tree for the chroma block is a single tree, the selected intra prediction mode is an MIP mode; and in response to determining that the coding tree for the chroma block is not a single tree, the selected intra prediction mode is a planar intra prediction mode. selecting the intra-prediction mode; decoding the chroma block using the selected intra-prediction mode; and 1. A video decoding method comprising:
6. The method of claim 5 , wherein the single tree indicates that the coding tree for the chroma block is the same as the coding tree for the luma block.
7. The method of claim 5 , further comprising determining a MIP mode from a set of MIP modes that depends at least in part on a dimension of the luma block.
8. Decoding the luma block comprises:
6. The method of claim 5, comprising multiplying a vector based on at least a portion of neighboring samples above and to the left of the luma block by a matrix corresponding to a MIP mode to generate a prediction of the luma block.
9. A non-transitory processor-readable medium having stored thereon instructions for causing at least one processor to perform the method of claim 5.
10. 1. A video encoder comprising one or more processors, the one or more processors comprising: determining that the picture is in a 4:4:4 color sampling format; signaling, in a data stream, an index indicating a matrix-based intra-prediction (MIP) mode for coding luma blocks of the picture; encoding the luma block using the MIP mode; For a chroma block of the picture, determining whether the coding tree for the chroma block is a single tree; selecting an intra prediction mode for encoding the chroma block, in response to determining that the coding tree for the chroma block is a single tree, the selected intra prediction mode is an MIP mode; and in response to determining that the coding tree for the chroma block is not a single tree, the selected intra prediction mode is a planar intra prediction mode. selecting the intra-prediction mode; encoding the chroma block using the selected intra-prediction mode; and a video encoder configured to:
11. The video encoder of claim 10 , wherein the single tree indicates that the coding tree for the chroma block is the same as the coding tree for the luma block.
12. The video encoder of claim 10 , wherein the one or more processors are further configured to determine a MIP mode from a set of MIP modes that depend at least in part on a dimension of the luma block.
13. To encode the luma block, the one or more processors:
11. The video encoder of claim 10, further configured to multiply a vector based on at least a portion of above and left neighboring samples of the luma block by a matrix corresponding to a MIP mode to generate a prediction of the luma block.
14. determining that the picture is in a 4:4:4 color sampling format; signaling, in a data stream, an index indicating a matrix-based intra-prediction (MIP) mode for coding luma blocks of the picture; encoding the luma block using the MIP mode; For a chroma block of the picture, determining whether the coding tree for the chroma block is a single tree; selecting an intra prediction mode for encoding the chroma block, in response to determining that the coding tree for the chroma block is a single tree, the selected intra prediction mode is an MIP mode; and in response to determining that the coding tree for the chroma block is not a single tree, the selected intra prediction mode is a planar intra prediction mode. selecting the intra-prediction mode; encoding the chroma block using the selected intra-prediction mode; and A video encoding method comprising:
15. The method of claim 14 , wherein the single tree indicates that the coding tree for the chroma block is the same as the coding tree for the luma block.
16. The method of claim 14 , further comprising determining a MIP mode from a set of MIP modes that depends at least in part on a dimension of the luma block.
17. encoding the luma block 15. The method of claim 14, comprising multiplying a vector based on at least a portion of neighboring samples above and to the left of the luma block by a matrix corresponding to a MIP mode to generate a prediction of the luma block.
18. 15. A non-transitory processor-readable medium having stored thereon instructions for causing at least one processor to perform the method of claim 14.
Citation Information
Patent Citations
Image processing method and device and storage medium
CN110708559A
Moving image coding apparatus and moving image coding method
JP2010045853A
Video encoding method and system using adaptive color conversion
JP2017005688A
Data encoding and decoding
JP2018129858A
Method and device for intra-predictive encoding / decoding of a coding unit comprising picture data, wherein the intra-predictive encoding relies on a prediction tree and a transform tree
JP2019511153A