Encoding using matrix-based intra prediction and quadratic transformation
By selecting the subset of the quadratic transformation and the matrix vector product, the encoding efficiency of the intra prediction mode based on the matrix is optimized, the problem of excessive memory demand is solved, and more efficient encoding is achieved.
Patent Information
- Application Number
- CN202080049521.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-25
- Filing Date
- 2020-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-06-23
AI Technical Summary
In the prior art, the combination of matrix-based intra prediction mode and secondary transformation leads to excessive memory requirements, which increases the cost of encoding efficiency.
By selecting a subset of the quadratic transformation, the encoding efficiency is optimized by optimizing the encoding efficiency by combining matrix vector product and cascade transformation of the quadratic transformation.
It effectively reduces storage memory requirements, improves encoding efficiency, and reduces encoding costs.
Smart Images

Figure CN114073081B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of matrix-based intra prediction and secondary transformation. Background Art
[0002] For traditional intra prediction modes such as planar mode, DC mode, and angular mode, the non-separable secondary transform (LFNST) is a tool for transforming the prediction residuals corresponding to these intra prediction modes. Here, a set S of transform sets is given such that each traditional intra prediction mode is associated with one of these transform sets. Then, at the decoder, it can be extracted from the bitstream whether to apply LFNST to a given block. If this is the case, depending on the intra prediction mode used on the current block, one of the transform sets in the set S is given, and if the transform set consists of more than one transform, it can be extracted from the bitstream which transform T in the set to use. Then, at the decoder, the transform T is applied as a secondary transform, which means applying the transform T to a subset of the residual transform coefficients of the separable primary transform. However, the above secondary transform is only defined a priori for traditional intra prediction modes.
[0003] Therefore, it is desirable to provide concepts for more efficient picture coding and / or video coding to support secondary transformation for matrix-based intra prediction (MIP), i.e., block-based intra prediction.
[0004] This is achieved by the subject matter of the independent claims of this application.
[0005] Other embodiments according to the invention are defined by the subject matter of the dependent claims of this application. Summary of the Invention
[0006] According to a first aspect of the present invention, the inventors of the present application have realized that a problem encountered when attempting to associate a secondary transform with a matrix-based intra prediction mode stems from the fact that providing a specific secondary transform for each MIP mode may be too costly in terms of the memory requirements for storing additional transforms. According to a first aspect of the present application, this difficulty is overcome by selecting a subset of one or more secondary transforms from a set of secondary transforms, the set of secondary transforms including transforms associated with matrix-based intra prediction modes and non-matrix-based intra prediction modes. The secondary transforms in the set of secondary transforms can be defined for one or more prediction modes, which reduces the memory capacity required for the set of secondary transforms. The transform defined for the planar intra prediction mode and / or the transform defined for the DC intra prediction mode can also be used for the selection of matrix-based intra prediction modes. Although the cost of the bitstream and thus the signaling may increase due to the additional syntax elements required for the blocks associated with the matrix-based intra prediction mode to indicate the use of the secondary transform, the coding efficiency can be improved through a specific selection of the subset of secondary transforms for the matrix-based intra prediction mode.
[0007] Accordingly, in a first aspect of the present application, an apparatus for decoding a predetermined block of a picture using intra prediction, i.e., a decoder, is configured to select a predetermined intra prediction mode from a plurality of intra prediction modes based on a data stream, the plurality of intra prediction modes including a first set of intra prediction modes and a second set of matrix-based intra prediction modes. The first set of intra prediction modes includes a DC intra prediction mode, an angular prediction mode, and optionally a planar intra prediction mode. If a matrix-based intra prediction mode from the second set is selected as the predetermined intra prediction mode, the decoder is configured to obtain a prediction vector using a matrix-vector product between: a vector derived from reference samples in the neighborhood of the predetermined block, and a prediction matrix associated with the corresponding matrix-based intra prediction mode, and the decoder is configured to predict samples of the predetermined block based on the prediction vector. The decoder is configured to: derive a prediction signal for the predetermined block using the predetermined intra prediction mode, and select one or more subsets of secondary transforms from a set of secondary transforms in a manner depending on the predetermined intra prediction mode, such that the subset is non-empty in the case where the predetermined intra prediction mode is included in the first set of intra prediction modes and in the case where the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. The first set and the second set define the intra prediction modes available for the secondary transforms. Thus, for a predetermined intra prediction mode selected from the first set or from the second set, the decoder is configured to select one or more subsets of secondary transforms from a set of secondary transforms specifically associated with the selected predetermined intra prediction mode. Additionally, the decoder is configured to: in the case where the predetermined intra prediction mode is included in the first set of intra prediction modes and in the case where the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes, derive a transformed version of the prediction residual of the predetermined block from the data stream, the transformed version of the prediction residual of the predetermined block being related to the spatial domain version of the prediction residual of the predetermined block via a transform defined by a cascade of a primary transform and a predetermined secondary transform applied to a subset of the coefficients of the primary transform. The primary transform is, for example, default set, and the predetermined secondary transform is, for example, selected by the decoder from the subset of secondary transforms. The decoder may be configured to select the predetermined secondary transform from the subset of secondary transforms by deriving a secondary transform indication syntax element from the data stream. The decoder is configured to reconstruct the predetermined block using the prediction signal and the prediction residual of the predetermined block.
[0008] According to a first aspect of the present application, a device for encoding a predetermined block of a picture using intra prediction in parallel with a decoder, i.e., an encoder, is configured to select a predetermined intra prediction mode from a plurality of intra prediction modes, the plurality of intra prediction modes including a first set of intra prediction modes and a second set of matrix-based intra prediction modes, the first set of intra prediction modes including a DC intra prediction mode, an angular prediction mode, and optionally a planar intra prediction mode, and according to each matrix-based intra prediction mode in the second set of matrix-based intra prediction modes, obtaining a prediction vector using a matrix-vector product between: a vector derived from reference samples in a neighborhood of the predetermined block, and a prediction matrix associated with the corresponding matrix-based intra prediction mode, and predicting samples of the predetermined block based on the prediction vector. The encoder is configured to signal the predetermined intra prediction mode in a data stream and derive a prediction signal for the predetermined block using the predetermined intra prediction mode. Additionally, the encoder is configured to select one or more subsets of secondary transforms from a set of secondary transforms in a manner depending on the predetermined intra prediction mode such that the subset is non-empty in a case where the predetermined intra prediction mode is included in the first set of intra prediction modes and in a case where the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. The encoder is configured to encode a transformed version of the prediction residual of the predetermined block into the data stream in a case where the predetermined intra prediction mode is included in the first set of intra prediction modes and in a case where the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes, the transformed version of the prediction residual of the predetermined block being related to a spatial-domain version of the prediction residual of the predetermined block via a transform defined by a cascade of a primary transform and a predetermined secondary transform applied to a subset of coefficients of the primary transform from the subset of secondary transforms. The predetermined block can be reconstructed using the prediction signal and the prediction residual of the predetermined block.
[0009] According to an embodiment, the decoder / encoder is configured to: select one or more subsets of secondary transforms from a set of secondary transforms in a manner depending on the predetermined intra prediction mode such that each secondary transform in the set of secondary transforms is included in one or more subsets of secondary transforms selected for at least one intra prediction mode within the first set and the second set of intra prediction modes. For one or more intra prediction modes within the first set or the second set, the corresponding selected subset can include all secondary transforms in the set of secondary transforms.
[0010] According to an embodiment, the decoder / encoder is configured to: select a subset of one or more secondary transforms from a set of secondary transforms in a manner that depends on a predetermined intra prediction mode, such that each secondary transform in each subset of secondary transforms selected for any matrix-based intra prediction mode is included in a subset of secondary transforms selected for at least one intra prediction mode within a first set that does not belong to an angular prediction mode. The subset selected for a matrix-based intra prediction mode may include secondary transforms associated with one or more non-angular intra prediction modes (e.g., DC intra prediction mode and / or planar intra prediction mode) within the first set. Thus, a secondary transform in the set of secondary transforms may be part of more than one subset of one or more secondary transforms. The set of secondary transforms does not necessarily have to include additional secondary transforms that are only available for blocks having a predetermined intra prediction mode that is a matrix-based intra prediction mode. For a predetermined block having a predetermined intra prediction mode that is a matrix-based intra prediction mode, the decoder / encoder is configured to select the same secondary transform from the set of secondary transforms as the predetermined intra prediction mode that is a non-angular prediction mode within the first set. The subset selected for a matrix-based intra prediction mode may be equal to the subset selected for intra prediction modes within the first set that do not belong to an angular prediction mode, or may include some of the secondary transforms in the subset selected for intra prediction modes within the first set that do not belong to an angular prediction mode, or may include some or all of the secondary transforms in two or more subsets selected for intra prediction modes within the first set that do not belong to an angular prediction mode.
[0011] The method for encoding or decoding is based on the same considerations as the apparatus for encoding or decoding described above. In this way, these methods can be accomplished with all the features and functions also described with respect to the apparatus for encoding and / or decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are not necessarily to scale, and instead generally focus on illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0013] Figure 1 An embodiment of encoding into a data stream is shown;
[0014] Figure 2 An embodiment of an encoder is shown;
[0015] Figure 3 An embodiment of picture reconstruction is shown;
[0016] Figure 4 An embodiment of a decoder is shown;
[0017] Figure 5Shows a schematic diagram for predicting a block for encoding and / or decoding according to an embodiment;
[0018] Figure 6 Shows a matrix operation for predicting a block for encoding and / or decoding according to an embodiment;
[0019] Figure 7.1 Shows predicting a block using a reduced sample value vector according to an embodiment;
[0020] Figure 7.2 Shows predicting a block using interpolation of samples according to an embodiment;
[0021] Figure 7.3 Shows predicting a block using a reduced sample value vector according to an embodiment, where only some boundary samples are averaged;
[0022] Figure 7.4 Shows predicting a block using a reduced sample value vector according to an embodiment, where a group of four boundary samples are averaged;
[0023] Figure 8 Shows a matrix operation performed by a device according to an embodiment;
[0024] Figures 9a to 9c Shows a detailed matrix operation performed by a device according to an embodiment;
[0025] Figure 10 Shows a detailed matrix operation performed by a device using offset and scale parameters according to an embodiment;
[0026] Figure 11 Shows a detailed matrix operation performed by a device using offset and scale parameters according to different embodiments;
[0027] Figure 12 Shows a schematic diagram of details of intra prediction of a predetermined block using a prediction mode in a most probable mode list according to an embodiment;
[0028] Figure 13 Shows a schematic diagram of decoding a predetermined block using a secondary transform according to an embodiment;
[0029] Figure 14 Shows a schematic diagram of applying a primary transform and a secondary transform according to an embodiment;
[0030] Figure 15 Shows a schematic diagram of selecting a subset of secondary transforms according to an embodiment;
[0031] Figure 16 Shows a schematic diagram of a predetermined block having a non - zero transform domain region according to an embodiment;
[0032] Figure 17 A block diagram showing a method for decoding a predetermined block according to an embodiment;
[0033] Figure 18 A block diagram showing a method for encoding a predetermined block according to an embodiment; and
[0034] Figures 19a to 19d Showing a syntax element portion of a data stream. DETAILED DESCRIPTION
[0035] Even if reference numerals appear in different figures, in the following description, the same or equivalent elements or elements having the same or equivalent functions are denoted by the same or equivalent reference numerals.
[0036] In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the present invention. However, it will be clear to those skilled in the art that the embodiments of the present invention can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than specifically, to avoid obscuring the embodiments of the present invention. Additionally, unless otherwise specifically indicated, the features of the different embodiments described below can be combined with each other.
[0037] 1 Introduction
[0038] Hereinafter, different inventive examples, embodiments, and aspects will be described. At least some of these examples, embodiments, and aspects particularly relate to methods and / or apparatuses for the following: for video coding, and / or for performing intra prediction using, for example, a linear or affine transformation with adjacent sample reduction, and / or optimizing video delivery (e.g., broadcast, streaming, file playback, etc.), for example, for video applications and / or for virtual reality applications.
[0039] In addition, the examples, embodiments, and aspects may relate to High Efficiency Video Coding (HEVC) or successors. Additionally, other embodiments, examples, and aspects will be defined by the appended claims.
[0040] It should be noted that any embodiment, example, and aspect defined by the claims can be supplemented by any details (features and functions) described in the following sections.
[0041] In addition, the embodiments, examples, and aspects described in the following sections can be used alone and can also be supplemented by any feature in another section or by any feature included in the claims.
[0042] In addition, it should be noted that the individuals, examples, embodiments, and aspects described herein can be used alone or in combination. Therefore, details can be added to each of the individual aspects without adding details to another of the examples, embodiments, and aspects.
[0043] It should also be noted that the present disclosure explicitly or implicitly describes features of decoding and / or encoding systems and / or methods.
[0044] Furthermore, the method-related features and functions disclosed herein can also be used for a device. In addition, any features and functions regarding a device disclosed herein can also be used for the corresponding method. In other words, the methods disclosed herein can be supplemented by any features and functions described regarding a device.
[0045] In addition, as will be described in the "Implementation Alternatives" section, any features and functions described herein can be implemented in hardware, or in software, or using a combination of hardware and software.
[0046] In addition, in some examples, embodiments, or aspects, any features described in parentheses ("(...)") or "[...]" can be considered optional.
[0047] 2 Encoder, Decoder
[0048] Hereinafter, various examples are described that can help achieve more efficient compression when using block-based prediction. Some examples achieve high compression efficiency by consuming a set of intra prediction modes. The latter can be added to other intra prediction modes designed heuristically, for example, or can be provided specifically. Other examples even utilize the two special cases just discussed. However, as a variant of these embodiments, intra prediction can be transformed into inter prediction by using reference samples in another picture.
[0049] To facilitate understanding of the following examples of the present application, the description begins with a presentation of a possible encoder and decoder into which the examples outlined subsequently in the present application can be incorporated. Figure 1 A device for encoding picture 10 block by block into data stream 12 is shown. The device is indicated by reference numeral 14 and can be a still picture encoder or a video encoder. In other words, when encoder 14 is configured to encode video 16 including picture 10 into data stream 12, picture 10 can be the current picture in video 16, or encoder 14 can simply encode picture 10 into data stream 12.
[0050] As described above, the encoder 14 performs encoding in a block-by-block manner or based on blocks. To this end, the encoder 14 subdivides the picture 10 into blocks, and the encoder 14 encodes the picture 10 into the data stream 12 in units of these blocks. Examples of possible subdivisions of the picture 10 into blocks 18 are elaborated in more detail below. Generally, it can be finally subdivided into blocks 18 of a fixed size (e.g., a block array arranged by rows and columns), or, for example, subdivided into blocks 18 of different block sizes by using a hierarchical multi-tree subdivision, which starts from the entire picture area of the picture or from a pre-partitioning of the picture 10 into a tree block array. These examples should not be regarded as excluding other possible ways of subdividing the picture 10 into blocks 18.
[0051] In addition, the encoder 14 is a predictive encoder, which is configured to predictively encode the picture 10 into the data stream 12. For a certain block 18, this means that the encoder 14 determines the prediction signal of the block 18 and encodes the prediction residual (i.e., the prediction error by which the prediction signal deviates from the actual picture content within the block 18) into the data stream 12.
[0052] The encoder 14 can support different prediction modes in order to derive the prediction signal of a certain block 18. The prediction mode, which is an important part in the following examples, is the intra prediction mode. According to the intra prediction mode, the inside of the block 18 is spatially predicted based on the adjacent encoded samples of the picture 10. Encoding the picture 10 into the data stream 12 and thus the corresponding decoding process can be based on a certain encoding order 20 defined between the blocks 18. For example, the encoding order 20 can traverse the blocks 18 in a raster scan order, for example, row by row from top to bottom and column by column from left to right within each row. In the case of a hierarchical multi-tree subdivision, a raster scan sorting can be applied within each level, where a depth-first traversal order can be applied, that is, the leaf nodes within a block at a certain level can be before the blocks at the same level that have the same parent block according to the encoding order 20. Depending on the encoding order 20, the adjacent encoded samples of the block 18 can generally be located on one or more sides of the block 18. In the case of the examples presented herein, for example, the adjacent encoded samples of the block 18 are located at the top and left of the block 18.
[0053] The intra prediction mode may not be the only mode supported by encoder 14. In the case where encoder 14 is a video encoder, for example, encoder 14 may also support an inter prediction mode, according to which a block 18 is predicted temporally based on a previously encoded picture of video 16. Such an inter prediction mode may be a motion compensated prediction mode, according to which, for such a block 18, a motion vector is signaled, the motion vector indicating a relative spatial offset of a portion of the predicted signal of block 18 that is to be derived as a copy. Additionally or alternatively, other non-intra prediction modes may also be available, such as an inter prediction mode in the case where encoder 14 is a multi-view encoder, or a non-prediction mode, according to which the interior of block 18 is encoded as is (i.e., without any prediction).
[0054] Before focusing the description of the present application on the intra prediction mode, with respect to Figure 2 a more specific example of a possible block-based encoder is described, i.e., a possible implementation for encoder 14, and then two corresponding examples of decoders respectively adapted to Figure 1 and Figure 2 are presented.
[0055] Figure 2 Illustrated is Figure 1 a possible implementation of encoder 14, i.e., an implementation in which the encoder is configured to encode prediction residuals using transform coding, although this is only an example and the present application is not limited to such prediction residual coding. According to Figure 2, the encoder 14 includes a subtractor 22 configured to subtract a corresponding prediction signal 24 from an inbound signal (i.e., picture 10 or, in the case of block-based, the current block 18) to obtain a prediction residual signal 26, which is then encoded by a prediction residual encoder 28 into the data stream 12. The prediction residual encoder 28 consists of a lossy coding stage 28a and a lossless coding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes the samples of the prediction residual signal 26. As already described above, this example uses transform coding of the prediction residual signal 26, and thus, the lossy coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform out such spectrally decomposed prediction residual 26, where the quantization of the quantizer 30 occurs on the transformed coefficients representing the residual signal 26. The transform can be a DCT, DST, FFT, Hadamard transform, etc. The lossless coding stage 28b then losslessly encodes the transformed and quantized prediction residual signal 34, and the lossless coding stage 28b is an entropy encoder that entropy encodes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner also available at the decoder (i.e., taking into account that the coding loss is caused by the quantizer 30). To this end, the prediction residual reconstruction stage 36 includes an inverse quantizer 38, followed by an inverse transformer 40. The inverse quantizer 38 performs the inverse operation of the quantization by the quantizer 30, and the inverse transformer 40 performs the inverse transform with respect to the transform performed by the transformer 32, e.g., the inverse operation of spectral decomposition, e.g., the inverse operation of any of the above specific transform examples. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transformer 40 to the prediction signal 24 to output a reconstructed signal, i.e., reconstructed samples. This output is fed into a predictor 44 of the encoder 14, which then determines the prediction signal 24 based on this output. The predictor 44 supports all the prediction modes discussed above with respect to Figure 1 those already discussed. Figure 2 It is also shown that in the case where the encoder 14 is a video encoder, the encoder 14 may further include a loop filter 46 that filters the fully reconstructed pictures, which, after being filtered, form reference pictures for the predictor 44 with respect to inter-frame prediction blocks.
[0056] As described above, the encoder 14 operates on a block basis. For the following description, the block basis of interest is to divide the picture 10 into blocks for which an intra prediction mode is selected from the set or plurality of intra prediction modes supported by the predictor 44 or the encoder 14, respectively, and the selected intra prediction mode is performed individually. However, there may also be other types of blocks into which the picture 10 is divided. For example, the above decision as to whether the picture 10 is inter-coded or intra-coded can be made at a granularity deviating from that of the block 18 or in units of blocks deviating from the block 18. For example, the inter / intra mode decision can be performed at the level of the coding blocks into which the picture 10 is divided, and each coding block is divided into prediction blocks. For a coding block for which intra prediction has been determined to be used, an intra prediction mode decision is made for each of the prediction blocks into which it is divided. To this end, for each of these prediction blocks, it is determined which of the supported intra prediction modes is applied to the corresponding prediction block. These prediction blocks will form the block 18 of interest here. The predictor 44 will process the prediction blocks within the coding blocks associated with inter prediction differently. These blocks are inter-predicted from the reference picture by determining a motion vector and copying the prediction signal of the block from the position in the reference picture pointed to by the motion vector. Another block subdivision relates to the subdivision of transform blocks, and the transformer 32 and the inverse transformer 40 perform the transform on a transform block basis. For example, a transform block can be the result of a further subdivision of a coding block. Of course, the examples set forth herein should not be considered limiting, and there are other examples. For completeness only, it should be noted that the subdivision to coding blocks can, for example, use a quadtree subdivision, and the coding blocks can also be further subdivided using a quadtree subdivision to obtain prediction blocks and / or transform blocks.
[0057] Figure 3 depicted in is suitable for Figure 1 the decoder 54 or apparatus for block-by-block decoding of the encoder 14. The decoder 54 operates in a manner opposite to that of the encoder 14, i.e., it decodes the picture 10 from the data stream 12 in a block-by-block manner and, to this end, supports a plurality of intra prediction modes. For example, the decoder 54 can include a residual provider 156. As described above with respect to Figure 1All other possibilities discussed are also valid for decoder 54. To this end, decoder 54 can be a still picture decoder or a video decoder, and decoder 54 also supports all prediction modes and prediction possibilities. The difference between encoder 14 and decoder 54 mainly lies in that encoder 14 makes or selects encoding decisions based on some optimizations, such as minimizing some cost functions that may depend on the coding rate and / or coding distortion. One of these coding options or coding parameters can involve selecting the intra prediction mode to be used for the current block 18 from the available or supported intra prediction modes. Then, encoder 14 can signal the selected intra prediction mode for the current block 18 within data stream 12, and decoder 54 uses this signaling for block 18 in data stream 12 to redo the selection. Similarly, the subdivision of picture 10 into blocks can be optimized within encoder 14, and the corresponding subdivision information is passed within data stream 12, and decoder 54 restores the subdivision of picture 10 into blocks based on the subdivision information. In summary, decoder 54 can be a prediction decoder based on block operations, and in addition to the intra prediction mode, decoder 54 can support other prediction modes, such as the inter prediction mode in the case where decoder 54 is a video decoder. In decoding, decoder 54 can also use the coding order 20 Figure 1 discussed, and since this coding order 20 is adhered to at both encoder 14 and decoder 54, the same neighboring samples can be used for the current block 18 at both encoder 14 and decoder 54. Therefore, to avoid unnecessary repetition, the description of the operating mode of encoder 14 should also apply to decoder 54 in terms of the subdivision of picture 10 into blocks, such as in terms of prediction and in terms of the coding of prediction residuals. The difference is that encoder 14 selects some coding options or coding parameters and signals through optimization, signals these coding parameters in data stream 12 or inserts these coding parameters into data stream 12, and then decoder 54 derives these coding parameters from data stream 12 for re - prediction, subdivision, etc.
[0058] Figure 4 shows Figure 3 a possible implementation of decoder 54, that is, an implementation suitable for Figure 1 the implementation of encoder 14 (as Figure 2 shown). Since Figure 4 many elements of encoder 54 are the same as the elements that appear in the Figure 2 corresponding encoder, the same reference numerals with an apostrophe are used in Figure 4 to indicate these elements. Specifically, adder 42', optional loop filter 46' and predictor 44' are in the same manner as they are in Figure 2is connected to the prediction loop in the same way as in the encoder. The reconstructed (i.e., inverse quantized and inverse transformed) prediction residual signal applied to the adder 42' is derived by a sequence of an entropy decoder 56 followed by a residual signal reconstruction stage 36'. The entropy decoder 56 reverses the entropy encoding of the entropy encoder 28b. Similar to the encoding side, the residual signal reconstruction stage 36' consists of an inverse quantizer 38' and an inverse transformer 40'. The output of the decoder is Figure 10 a reconstruction. The reconstruction of picture 10 can be directly available at the output of the adder 42' or, alternatively, at the output of the loop filter 46'. Some post-filtering can be arranged at the output of the decoder to perform some post-filtering on the reconstruction of picture 10 in order to improve the picture quality, but Figure 4 this option is not depicted in
[0059] Similarly, with respect to Figure 4 , except that only the encoder performs the optimization tasks and the relevant decisions regarding the coding options, the above description presented with respect to Figure 2 should also apply to Figure 4 . However, all the descriptions regarding block subdivision, prediction, inverse quantization, and inverse transformation also apply to the decoder 54 of Figure 4 .
[0060] 3 ALWIP (Affine Linear Weighted Intra Prediction)
[0061] Some non-limiting examples regarding ALWIP are discussed herein, even though ALWIP is not always required to embody the techniques discussed here.
[0062] This application particularly relates to an improved block-based prediction mode concept for per-block picture coding that can be used, for example, in a video codec such as HEVC or any successor to HEVC. The prediction mode can be an intra prediction mode, but in theory, the concepts described herein can also be transferred to an inter prediction mode where the reference samples are part of another picture.
[0063] A block-based prediction concept that allows for an efficient implementation, such as a hardware-friendly implementation, is sought.
[0064] This object is achieved by the subject matter of the independent claims of this application.
[0065] Intra prediction modes are widely used in picture coding and video coding. In video coding, intra prediction modes compete with other prediction modes such as inter prediction modes (e.g., motion compensation prediction modes). In intra prediction modes, the current block is predicted based on neighboring samples, i.e., samples that have been encoded on the encoder side and decoded on the decoder side. The neighboring sample values are extrapolated into the current block in order to form a prediction signal for the current block, and the prediction residual is transmitted in the data stream for the current block. The better the prediction signal, the lower the prediction residual, and thus, the fewer bits required to encode the prediction residual.
[0066] To be effective, several aspects should be considered in order to form an effective framework for intra prediction in a per-block picture coding environment. For example, the larger the number of intra prediction modes supported by the codec, the greater the rate consumption of the side information for notifying this selection signal to the decoder. On the other hand, the set of supported intra prediction modes should be able to provide a good prediction signal, i.e., a prediction signal that results in a low prediction residual.
[0067] Hereinafter, as a comparative example or a basic example, a device (encoder or decoder) for decoding pictures block by block from a data stream is disclosed, which supports at least one intra prediction mode. According to this intra prediction mode, the intra prediction signal of a block of a predetermined size of a picture is determined by applying a first template of samples adjacent to the current block to an affine linear predictor, which will be referred to as an affine linear weighted intra predictor (ALWIP) hereinafter.
[0068] The device may have at least one of the following attributes (which may also apply to a method or another technology, e.g., implemented in a non-transitory storage unit storing instructions, which when executed by a processor, cause the processor to implement the method and / or operate as a device):
[0069] 3.1 The predictor can be complementary to other predictors
[0070] The intra prediction modes that can form the subject of the implementation improvements described further below can be complementary to other intra prediction modes of the codec. Thus, they can be complementary to the DC prediction mode, the planar prediction mode, or the angular prediction mode defined in the HEVC codec (the corresponding JEM reference software). The latter three intra prediction modes of the intra prediction modes shall be referred to as traditional intra prediction modes from now on. Thus, for a given block in the intra mode, the decoder needs to parse a flag that indicates whether one of the intra prediction modes supported by the device is to be used.
[0071] 3.2 More than one proposed prediction mode
[0072] The apparatus may include more than one ALWIP mode. Thus, in the case where the decoder knows to use one of the ALWIP modes supported by the apparatus, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the apparatus is to be used.
[0073] The signaling of the supported modes may have the following property: the encoding of some ALWIP modes may require fewer bins than other ALWIP modes. Which of these modes require fewer bins and which require more bins may depend on information extractable from the already decoded bitstream, or may be fixed in advance.
[0074] 4 Some aspects
[0075] Figure 2 A decoder 54 for decoding a picture from a data stream 12 is shown. The decoder 54 may be configured to decode a predetermined block 18 of the picture. Specifically, the predictor 44 may be configured to map a set of P neighboring samples in the neighborhood of the predetermined block 18 to a set of Q predicted values of the samples of the predetermined block using a linear or affine linear transform [e.g., ALWIP].
[0076] As Figure 5 shown, the predetermined block 18 includes Q values to be predicted (which will be "predicted values" at the end of the operation). If the block 18 has M rows and N columns, then Q = M·N. The Q values of the block 18 may be in the spatial domain (e.g., pixels) or in the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of the block 18 may be predicted based on P values taken from neighboring blocks 17a to 17c that are typically adjacent to the block 18. The P values of the neighboring blocks 17a to 17c may be in the position closest to the block 18 (e.g., adjacent). The P values of the neighboring blocks 17a to 17c have been processed and predicted. The P values are indicated as the values in the portions 17'a to 17'c to distinguish them from the blocks of which they are a part (in some examples, 17'b is not used).
[0077] As Figure 6As shown, for performing prediction, it is possible to operate with a first vector 17P having P entries (each entry associated with a specific position in adjacent parts 17'a to 17'c), a second vector 18Q having Q entries (each entry associated with a specific position in block 18), and a mapping matrix 17M (each row associated with a specific position in block 18 and each column associated with a specific position in adjacent parts 17'a to 17'c). Thus, the mapping matrix 17M performs predicting P values of adjacent parts 17'a to 17'c as values of block 18 according to a predetermined pattern. Thus, the entries in the mapping matrix 17M can be understood as weighting factors. In the following paragraphs, we will use labels 17a to 17c instead of 17'a to 17'c to refer to the adjacent parts of the boundary.
[0078] In the art, several traditional patterns are known, such as the DC pattern, the planar pattern, and 65 directional prediction patterns. For example, 67 patterns may be known.
[0079] However, it has been noted that different patterns can also be used, here referred to as linear or affine linear transformations. The linear or affine linear transformation includes P·Q weighting factors, where at least 1 / 4 P·Q weighting factors are non-zero weighting values, and for each of the Q predicted values, the non-zero weighting value includes a series of P weighting factors related to the corresponding predicted value. When this series is arranged one by one according to the raster scan order between samples of a predetermined block, it forms an envelope that is non-linear omnidirectionally.
[0080] It is possible to map the P positions of adjacent values 17'a to 17'c (templates), the Q positions of adjacent samples 17'a to 17'c, and map at the values of the P*Q weighting factors of matrix 17M. The plane is an example of the envelope of the sequence for DC transformation (is the plane for DC transformation). The envelope is clearly planar and is thus excluded by the definition of linear or affine linear transformation (ALWIP). Another example is the matrix that produces the simulation of the angular pattern: the envelope would be excluded from the ALWIP definition and, frankly, it looks like a hill sloping from top to bottom along a direction in the P / Q plane. The planar pattern and the 65 directional prediction patterns will have different envelopes, however these envelopes will be linear in at least one direction, i.e., for the exemplary DC in all directions, for the angular pattern in the hill direction for example.
[0081] In contrast, the envelope of the linear or affine transformation will not be linearly omnidirectional. It has been understood that in some cases, this type of transformation can be optimal for performing the prediction of block 18. It has been noted that preferably, at least 1 / 4 of the weighting factors are different from zero (i.e., at least 25% of the P*Q weighting factors are different from 0).
[0082] According to any conventional mapping rules, the weighting factors can be independent of each other. Thus, the matrix 17M can be such that the values of its entries do not have an obvious recognizable relationship. For example, the weighting factors cannot be described by any analytical function or difference function.
[0083] In an example, the ALWIP transform enables the mean value of the maximum of the cross-correlation between the following two to be less than a predetermined threshold (e.g., 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold within the range between 0.05 and 0.035): a first series of weighting factors related to the corresponding predicted values, a second series of weighting factors related to predicted values other than the corresponding predicted values, or the reverse version of the latter series (whichever results in a higher maximum). For example, for each pair (i1, i2) of rows in the ALWIP matrix 17M, the cross-correlation can be calculated by multiplying the P values in the i1-th row by the P values in the i2-th row. For each obtained cross-correlation, the maximum value can be obtained. Thus, the mean value (average value) of the entire matrix 17M can be obtained (i.e., averaging the maximum values of the cross-correlations in all combinations). Thereafter, the threshold can be, for example, 0.2 or 0.3 or 0.35 or 0.1, e.g., a threshold within the range between 0.05 and 0.035.
[0084] The P adjacent samples of blocks 17a to 17c can be located along a one-dimensional path that extends along the border of a predetermined block 18 (e.g., 18c, 18a). For each of the Q predicted values of the predetermined block 18, a series of P weighting factors related to the corresponding predicted value can be sorted in a manner that traverses the one-dimensional path along a predetermined direction (e.g., from left to right, from top to bottom, etc.).
[0085] In an example, the ALWIP matrix 17M can be non-diagonal or non-block diagonal.
[0086] An example of the ALWIP matrix 17M for predicting a 4×4 block 18 from 4 already predicted adjacent samples can be:
[0087] {
[0088] {37,59,77,28},
[0089] {32,92,85,25},
[0090] {31,69,100,24},
[0091] {33,36,106,29},
[0092] {24,49,104,48},
[0093] {24,21,94,59},
[0094] {29, 0, 80, 72},
[0095] {35, 2, 66, 84},
[0096] {32, 13, 35, 99},
[0097] {39, 11, 34, 103},
[0098] {45, 21, 34, 106},
[0099] {51, 24, 40, 105},
[0100] {50, 28, 43, 101},
[0101] {56, 32, 49, 101},
[0102] {61, 31, 53, 102},
[0103] {61, 32, 54, 100}
[0104] }。
[0105] (Here, {37, 59, 77, 28} is the first row of matrix 17M; {32, 92, 85, 25} is the second row of matrix 17M; and {61, 32, 54, 100} is the 16th row of matrix 17M.) Matrix 17M has a size of 16×4 and includes 64 weighting factors (as a result of 16*4 = 64). This is because the size of matrix 17M is Q×P, where Q = M*N, i.e., the number of samples of the block 18 to be predicted (block 18 is a 4×4 block), and P is the number of samples of the predicted samples. Here, M = 4, N = 4, Q = 16 (as a result of M*N = 4*4 = 16), and P = 4. The matrix is non-diagonal and non-block diagonal and has no specific rules to describe.
[0106] It can be seen that less than 1 / 4 of the weighting factors are 0 (in the case of the matrix shown above, one of the 64 weighting factors is zero). When arranged one by one according to the raster scan order, the envelope formed by these values forms an omnidirectional non-linear envelope.
[0107] Even though the above description is mainly discussed with reference to a decoder (e.g., decoder 54), the same operations can also be performed at an encoder (e.g., encoder 14).
[0108] In some examples, for each block size in the set of block sizes, the ALWIP transforms of the intra prediction modes within the second set of intra prediction modes for the corresponding block size are different from each other. Additionally or alternatively, the cardinality of the second set of intra prediction modes for the block sizes in the set of block sizes may coincide, but the associated linear or affine-linear transforms of the intra prediction modes within the second set of intra prediction modes for different block sizes cannot be transformed into each other by scaling.
[0109] In some examples, the ALWIP transforms may be defined in such a way that they "have nothing in common" with traditional transforms (e.g., the ALWIP transforms may "have nothing" in common with the corresponding traditional transforms, although they are mapped via one of the above mappings).
[0110] In an example, the ALWIP mode is used for both the luminance component and the chrominance component, but in other examples, the ALWIP mode is used for the luminance component and not for the chrominance component.
[0111] 5 Affine-linear weighted intra prediction mode with encoder acceleration (e.g., test CE3-1.2.1)
[0112] 5.1 Description of the method or apparatus
[0113] Except for the following changes, the affine-linear weighted intra prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as that proposed in JVET-L0199 under test CE3-2.2.2:
[0114] · Coordinate with multi-reference line (MRL) intra prediction, especially encoder estimation and signaling, i.e., MRL is not combined with ALWIP, and the transmission of the MRL index is limited to non-ALWIP blocks.
[0115] · For all blocks with W×H≥32×32, subsampling is now mandatory (previously it was optional for 32×32); thus, the additional tests and the transmission of the subsampling flag at the encoder have been removed.
[0116] · ALWIP for 64×N blocks and N×64 blocks (N≤32) has been added by downsampling to 32×N and N×32 respectively and applying the corresponding ALWIP mode.
[0117] In addition, test CE3-1.2.1 includes the following encoder optimizations for ALWIP:
[0118] · Combined mode estimation: The traditional mode and the ALWIP mode use a shared Hadamard candidate list for full RD estimation, i.e., ALWIP mode candidates are added to the same list as traditional (and MRL) mode candidates based on the Hadamard cost.
[0119] · The combined mode list supports fast intra EMT and fast intra PB, with additional optimizations to reduce the number of full RD checks.
[0120] · Following the same approach as the traditional mode, only the MPMs of the available left and above blocks are added to the list for the full RD estimation of ALWIP.
[0121] 5.2 Complexity Evaluation
[0122] In Test CE3-1.2.1, excluding the calculations for calling the discrete cosine transform, each sample requires at most 12 multiplications to generate the prediction signal. Additionally, a total of 136492 parameters are required, with each parameter being 16 bits. This corresponds to 0.273 megabytes of memory.
[0123] 5.3 Experimental Results
[0124] The tests are evaluated according to the common test conditions JVET-J1010[2], and for the intra-only (AI) and random access (RA) configurations, the VTM software version 3.0.1 is used. The corresponding simulations are carried out on an Intel Xeon cluster (E5-2697A v4, AVX2 enabled, Intel Turbo Boost Technology turned off) with a Linux operating system and a GCC7.2.1 compiler.
[0125] Table 1. Results of CE3-1.2.1 for VTM AI Configuration
[0126] Y U V Encoding time Decoding time Class A1 -2,08% -1,68% -1,60% 155% 104% Class A2 -1,18% -0,90% -0,84% 153% 103% Class B -1,18% -0,84% -0,83% 155% 104% Class C -0,94% -0,63% -0,76% 148% 106% Class E -1,71% -1,28% -1,21% 154% 106% Total -1,36% -1,02% -1,01% 153% 105% Class D -0,99% -0,61% -0,76% 145% 107% Class F (optional) -1,38% -1,23% -1,04% 147% 104%
[0127] Table 2. Results of CE3-1.2.1 for VTM RA Configuration
[0128] Y U V Encoding time Decoding time Class A1 -1,25% -1,80% -1,95% 113% 100% Class A2 -0,68% -0,54% -0,21% 111% 100% Class B -0,82% -0,72% -0,97% 113% 100% Class C -0,70% -0,79% -0,82% 113% 99% Class E Total -0,85% -0,92% -0,98% 113% 100% Class D -0,65% -1,06% -0,51% 113% 102% Class F (optional) -1,07% -1,04% -0,96% 117% 99%
[0129] 5.4 Complexity-Reduced Affine Linear Weighted Intra Prediction (e.g., Test CE3-1.2.2)
[0130] The technique tested in CE2 is related to the "affine linear intra prediction" described in JVET-L0199[1], but it is simplified in terms of memory requirements and computational complexity:
[0131] · There can be only three different sets of prediction matrices (e.g., S0, S1, S2, see below) and bias vectors covering all block shapes (e.g., to provide offset values). As a result, the number of parameters is reduced to 14400 10-bit values, which is less than the storage amount stored in a 128×128 CTU.
[0132] · The input size and output size of the predictor are further reduced. Additionally, instead of transforming the boundaries via DCT, averaging or downsampling can be performed on the boundary samples, and the generation of the prediction signal can use linear interpolation instead of the inverse DCT. Thus, at most four multiplications are required for each sample to generate the prediction signal.
[0133] 6. Example:
[0134] How to perform some predictions with ALWIP prediction is discussed here (e.g., as Figure 6 shown).
[0135] In principle, referring to Figure 6 , to obtain the Q = M*N values of the M×N block 18 to be predicted, the Q*P samples of the Q×PALWIP prediction matrix 17M should be multiplied with the P samples of the P×1 adjacent vector 17P. Thus, generally, to obtain each of the Q = M*N values of the M×N block 18 to be predicted, at least P = M+N multiplications are required.
[0136] These multiplications have a very adverse effect. The size P of the boundary vector 17P generally depends on the number M+N of boundary samples (binary bins or pixels) 17a, 17c adjacent (e.g., neighboring) to the M×N block 18 to be predicted. This means that if the size of the block 18 to be predicted is large, the number M+N of boundary pixels (17a, 17c) is correspondingly large, thus increasing the size P = M+N of the P×1 boundary vector 17P and the length of each row of the Q×P ALWIP prediction matrix 17M, and thus also increasing the necessary number of multiplications (generally, Q = M*N = W*H, where W (width) is another symbol for N, and H (height) is another symbol for M; in the case where the boundary vector consists of only one row of samples and / or one column of samples, P is P = M+N = H+W).
[0137] Generally, the following fact exacerbates this problem: in a microprocessor-based system (or other digital processing system), multiplication is generally a power-consuming operation. It can be imagined that performing a large number of multiplications on a very large number of samples of a large number of blocks will result in a waste of computing power, which is usually not desirable.
[0138] Therefore, it is preferred to reduce the number of multiplications Q*P required for predicting the M×N block 18.
[0139] It has been understood that by intelligently selecting alternative multiplications and operations that are easier to process, the computing power required for each intra-frame prediction of each block 18 to be predicted can be reduced in some way.
[0140] Specifically, referring to Figures 7.1 to 7.4, it has been understood that an encoder or a decoder can perform prediction on a predetermined block (e.g., 18) of a picture by using a plurality of adjacent samples (e.g., 17a, 17c) through the following operations:
[0141] Reduce (e.g., in step 811) (e.g., by averaging or downsampling) a plurality of adjacent samples (e.g., 17a, 17c) to obtain a reduced set of sample values, where the number of samples in the reduced set of sample values is smaller than that of the plurality of adjacent samples,
[0142] Perform (e.g., in step 812) a linear or affine-linear transformation on the reduced set of sample values to obtain a predicted value of a predetermined sample of the predetermined block.
[0143] In some cases, a decoder or an encoder can also derive predicted values of other samples of the predetermined block based on the predicted value of the predetermined sample and the plurality of adjacent samples, for example, by interpolation. Thus, an upsampling strategy can be obtained.
[0144] In an example, it is possible to perform (e.g., in step 811) some averaging on the samples of the boundary 17 in order to obtain a reduced set 102 of samples with a reduced number of samples (at least one of the samples of the samples with a reduced number 102 can be the average of two samples of the original boundary samples, or selected from the original boundary samples) ( Figures 7.1 to 7.4 ). For example, if the original boundary has P = M + N samples, the reduced set of samples can have P red = M red + N red samples, where M red < M and N red < N, and at least one of them satisfies such that P red < P. Therefore, the boundary vector 17P actually used for prediction (e.g., in step 812b) will not have P × 1 entries, but will have P red × 1 entries, where P red < P. Similarly, the ALWIP prediction matrix 17M selected for prediction will not have a Q × P size, but the number of elements of the matrix is reduced to Q × P red (or Q red × P red , see below), at least because P red < P (by means of M red < M and N red < N, and at least one of them).
[0145] In some examples (e.g., Figure 7.2 , Figure 7.3 ), if the block obtained by ALWIP (in step 812) has a size of M′ red × N′ redThe reduced block, where M′ red <M and / or N′ red <N (i.e., the number of samples directly predicted by ALWIP is less than the number of samples of the block 18 to be actually predicted), and it is even possible to further reduce the number of multiplications. Therefore, set Q red = M′ red * N′ red , and this will obtain the ALWIP prediction by using Q red * P red multiplications instead of Q * P red multiplications (where Q red * P red <Q * P red <Q * P). This multiplication will predict the reduced block of size M′ red ×N′ red . Although it will be possible to perform (e.g., in subsequent step 813) upsampling from the reduced M′ red ×N′ red predicted block to the final M×N predicted block (e.g., obtained by interpolation).
[0146] Although matrix multiplication involves a reduction in the number of multiplications (Q red * P red or Q * P red ), both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or even avoiding) multiplications, and these techniques can be advantageous. For example, downsampling, averaging, and / or interpolation can be performed (e.g., in steps 811 and / or 813) by adopting binary operations such as addition and shift that have no power requirements for computing.
[0147] In addition, addition is a very easy operation that can be easily performed without a large amount of computational work.
[0148] This shift operation can be used, for example, to average two boundary samples and / or to interpolate two samples (support values) of the reduced predicted block (or taken from the boundary) to obtain the final predicted block. (For interpolation, two sample values are required. Inside the block, we always have two predetermined values, but to interpolate samples along the left and upper borders of the block, we only have one predetermined value, as Figure 7.2 shown, so we use the boundary samples as the support values for interpolation.)
[0149] A two-step process can be used, for example:
[0150] First, sum the values of the two samples;
[0151] Then, halve the sum value (e.g., by right shift).
[0152] Alternatively, it is possible to:
[0153] First, halve each of the samples in the sample (e.g., by shifting left);
[0154] Then, sum the values of the two halved samples.
[0155] When sampling is performed currently (e.g., at step 811), since only one sample needs to be selected from a set of samples (e.g., samples adjacent to each other), easier operations can be performed.
[0156] Therefore, it is now possible to define techniques for reducing the number of multiplications to be performed. Some of these techniques can be based in particular on at least one of the following principles:
[0157] Even if the size of the block 18 to be actually predicted is M×N, the block can be reduced (in at least one of the two dimensions) and an ALWIP matrix with a reduced size of Q red ×P red can be applied (where Q red = M′ red * N′ red , P red = N red + M red , where M′ red < M and / or N′ red < N and / or M red < M and / or N red < N). Therefore, the boundary vector 17P will have a size of P red ×1, meaning that there will be only P red < P multiplications (where P red = M red + N red and P = M + N).
[0158] P red ×1 boundary vector 17P can be easily obtained from the original boundary 17, for example:
[0159] By downsampling (e.g., by only selecting some samples of the boundary); and / or
[0160] By averaging multiple samples of the boundary (which can be easily obtained by addition and shifting without multiplication).
[0161] Additionally or alternatively, instead of predicting all Q = M * N values of the block 18 to be predicted by multiplication, it is possible to predict only the reduced block with the reduced size (e.g., Q red = M′ red * N′ red, where M′ red <M and / or N′ red <N). The remaining samples of block 18 to be predicted will use, for example, Q red samples as the support values for the remaining Q - Q red values to be obtained by interpolation.
[0162] According to Figure 7.1 the example shown, block 18 to be predicted is 4×4 (M = 4, N = 4, Q = M*N = 16) and the neighborhoods 17 of sample 17a (vertical column with 4 predicted samples) and 17c (horizontal row with 4 predicted samples) have been predicted in the previous iteration (neighborhoods 17a and 17c can be jointly indicated by 17). A priori, by using Figure 6 the equation shown, prediction matrix 17M should be a Q×P = 16×8 matrix (by means of Q = M*N = 4*4 and P = M + N = 4 + 4 = 8), and the size of boundary vector 17P should be 8×1 (by means of P = 8). However, this would result in the need to perform 8 multiplications for each of the 16 samples of the 4×4 block 18 to be predicted, resulting in a total of 16*8 = 128 multiplications. (Note that the average number of multiplications per sample is a good assessment of the computational complexity. For traditional intra prediction, each sample requires four multiplications, which increases the computational workload involved. Therefore, it can be used as an upper limit for ALWIP to ensure reasonable complexity and not exceed the complexity of traditional intra prediction.)
[0163] Nevertheless, it has been understood that by using the present technique, the number of samples 17a and 17c adjacent to block 18 to be predicted can be reduced from P to P red <P. Specifically, it has been understood that the boundary samples (17a, 17c) adjacent to each other can be averaged (e.g., at Figure 7.1 100 in), to obtain a reduced boundary 102 with two horizontal rows and two vertical columns, thus operating on block 18 as a 2×2 block (the reduced boundary is formed by the averaged values). Alternatively, downsampling can be performed, so two samples are selected for row 17c and two samples are selected for column 17a. Thus, horizontal row 17c, instead of having four original samples, is processed as having two samples (e.g., averaged samples), while vertical column 17a, which originally had four samples, is processed as having two samples (e.g., averaged samples). It can also be understood that after subdividing row 17c and column 17a into groups 110 each having two samples, one single sample is maintained (e.g., the average of the samples in group 110 or a simple selection between the samples in group 110). Thus, by means of only having four samples (M red = 2, N red = 2, P red= M red + N red = 4, where P red <a set 102 of P), obtaining a so-called reduced set 102 of sample values.
[0164] It can be understood that it is possible to perform operations (e.g., averaging or downsampling 100) without performing too many multiplications at the processor level: the averaging or downsampling 100 performed at step 811 can be obtained simply by direct and computationally non-power-consuming operations such as addition and shifting.
[0165] It has been understood that at this point, it is possible to (e.g., use a prediction matrix such as Figure 6 a matrix 17M of ) perform a linear or affine-linear (ALWIP) transformation 19 on the reduced set 102 of sample values. In this case, the ALWIP transformation 19 directly maps four samples 102 to the sample values 104 of block 18. Interpolation is not required in the current case.
[0166] In this case, the size of the ALWIP matrix 17M is Q × P red = 16 × 4: This follows the fact that all Q = 16 samples of the block 18 to be predicted are obtained directly by ALWIP multiplication (no interpolation is required).
[0167] Thus, at step 812a, a suitable ALWIP matrix 17M of size Q × P red is selected. This selection can be at least partially based on, for example, signaling from the data stream 12. The selected ALWIP matrix 17M can also be indicated by A k , where k can be understood as an index, which can be signaled in the data stream 12 (in some cases, the matrix is also indicated as see below). The selection can be performed according to the following scheme: for each size (e.g., the height / width pair of the block 18 to be predicted), an ALWIP matrix 17M is selected from one of, for example, three sets S0, S1, S2 of matrices (each of the three sets S0, S1, S2 can group multiple ALWIP matrices 17M of the same size, and the ALWIP matrix selected for prediction will be one of them).
[0168] At step 812b, the selected Q × P red ALWIP matrix 17M (also indicated as A k ) is multiplied by the P red × 1 boundary vector 17P.
[0169] At step 812c, an offset value (e.g., b k)Added to all the obtained values 104 of vector 18Q obtained, for example, through ALWIP. The offset value (b k or in some cases also with indicated, see below) can be associated with a specifically selected ALWIP matrix (A k ), and can be based on an index (for example, the index can be signaled in the data stream 12).
[0170] Therefore, here we resume the comparison between using this technology and not using this technology:
[0171] In the case of not using this technology:
[0172] The size of the block 18 to be predicted is M = 4, N = 4;
[0173] Q = M * N = 4 * 4 = 16 values to be predicted;
[0174] P = M + N = 4 + 4 = 8 boundary samples
[0175] For each of the Q = 16 values to be predicted, P = 8 multiplications,
[0176] In total, P * Q = 8 * 16 = 128 multiplications;
[0177] In the case of using this technology, we have:
[0178] The size of the block 18 to be predicted is M = 4, N = 4;
[0179] Finally, Q = M * N = 4 * 4 = 16 values to be predicted;
[0180] Reduced size of the boundary vector: P red = M red + N red = 2 + 2 = 4;
[0181] For each of the Q = 16 values to be predicted by ALWIP, P red = 4 multiplications,
[0182] In total, P red * Q = 4 * 16 = 64 multiplications (half of 128!)
[0183] The ratio between the number of multiplications and the number of final values to be obtained is P red * Q / Q = 4, that is, half of P = 8 multiplications for each sample to be predicted!
[0184] It can be understood that by relying on direct and computationally non-power-demanding operations such as averaging (and, in case, addition and / or shifting and / or downsampling), appropriate values can be obtained in step 812.
[0185] Reference Figure 7.2 , the block 18 to be predicted here is an 8×8 block of 64 samples (M = 8, N = 8). Here, a priori, the size of the prediction matrix 17M should be Q×P = 64×16 (Q = 64, by virtue of Q = M*N = 8*8 = 64, M = 8 and N = 8, and by virtue of P = M + N = 8 + 8 = 16). Therefore, a priori, for each of the Q = 64 samples of the 8×8 block 18 to be predicted, P = 16 multiplications will be required, amounting to 64*16 = 1024 multiplications for the entire 8×8 block 18!
[0186] However, as Figure 7.2 can be seen, a method 820 can be provided according to which, instead of using all 16 samples of the boundary, only 8 values are used (for example, 4 values in the horizontal boundary row 17c between the original samples of the boundary and 4 values in the vertical boundary column 17a). From the boundary row 17c, 4 samples can be used instead of 8 samples (for example, they can be the average of a pair of samples and / or one sample selected from two samples). Therefore, the boundary vector is not a P×1 = 16×1 vector, but only a P red ×1 = 8×1 vector (P red = M red + N red = 4 + 4). It has been understood that the samples of the horizontal row 17c and the samples of the vertical column 17a can be selected or averaged (for example, in pairs) to have only P red = 8 boundary values instead of the original P = 16 samples, thereby forming a reduced set 102 of sample values. This reduced set 102 will allow a reduced version of the block 18 to be obtained, which has Q red = M red * N red = 4*4 = 16 samples (instead of Q = M*N = 8*8 = 64). The ALWIP matrix can be applied to predict a block of size M red × N red = 4×4. The reduced version of the block 18 includes the samples indicated in gray in the Figure 7.2 scheme 106: the samples indicated by the gray squares (including sample 118' and sample 118”) form a 4×4 reduced block, which has Q red= 16 values. A 4×4 reduced block has been obtained by applying a linear transformation 19 in step 812. After obtaining the values of the 4×4 reduced block, the values of the remaining samples (the samples indicated by white samples in scenario 106) can be obtained, for example, by interpolation.
[0187] Relative to Figure 7.1 method 810, method 820 may additionally include step 813: Deriving the predicted values of the remaining Q-Q of the M×N = 8×8 block 18 to be predicted, for example, by interpolation red = 64 - 16 = 48 samples (white squares). The remaining Q-Q red = 64 - 16 = 48 samples can be obtained based on Q red = 16 samples directly obtained by interpolation (for example, interpolation can also utilize boundary samples). As can be seen from Figure 7.2 Although samples 118' and 118'' have been obtained in step 812 (as shown by the gray squares), sample 108' (in the middle between sample 118' and sample 118'' and indicated by a white square) is obtained by interpolation between sample 118' and 118'' in step 813. It has been understood that interpolation can also be obtained by operations similar to those for averaging, such as shifting and adding. Therefore, in Figure 7.2 value 108' can generally be determined as the intermediate value (which can be the average value) between the value of sample 118' and the value of sample 118''.
[0188] By performing interpolation, in step 813, the final version of the M×N = 8×8 block 18 can also be obtained based on the multiple sample values indicated in 104.
[0189] Therefore, the comparison between using this technology and not using this technology is:
[0190] Without using this technology:
[0191] The size of the block 18 to be predicted is M = 8, N = 8, and there are Q = M*N = 8*8 = 64 samples in the block 18 to be predicted;
[0192] P = M + N = 8 + 8 = 16 samples in the boundary 17;
[0193] For each of the Q = 64 values to be predicted, P = 16 multiplications,
[0194] A total of P*Q = 16*64 = 1028 multiplications
[0195] The ratio between the number of multiplications and the number of final values to be obtained is P*Q / Q = 16
[0196] When using this technology:
[0197] The size of block 18 to be predicted is M = 8, N = 8;
[0198] Finally, Q = M * N = 8 * 8 = 64 values are to be predicted;
[0199] but Q red ×P red ALWIP matrix is to be used, where P red = M red + N red , Q red = M red * N red , M red = 4, N red = 4
[0200] P in the boundary red = M red + N red = 4 + 4 = 8 samples, where P red <P For each of the Q of the 4×4 reduced block to be predicted (formed by the gray squares in Scenario 106) red = 16 values, P red = 8 multiplications, for a total of P red * Q red = 8 * 16 = 128 multiplications (much smaller than 1024!) The ratio between the number of multiplications and the number of final values to be obtained is P red * Q red / Q = 128 / 64 = 2 (much smaller than 16 obtained without using this technology!)
[0201] Therefore, the power requirement of the technology presented here is 1 / 8 of the prior art
[0202] Figure 7.3 Another example is shown (which may be based on Method 820), where the block 18 to be predicted is a rectangular 4×8 block (M = 8, N = 4), with Q = 4 * 8 = 32 samples to be predicted. The boundary 17 is formed by a horizontal row 17c with N = 8 samples and a vertical column 17a with M = 4 samples. Thus, a priori, the size of the boundary vector 17P will be P×1 = 12×1, and the predicted ALWIP matrix should be a Q×P = 32×12 matrix, thus requiring Q * P = 32 * 12 = 384 multiplications
[0203] However, it is possible to, for example, average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row with only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as is (e.g., not averaged). Overall, the size of the reduced boundary will be Pred = 8, where P red < P. Therefore, the boundary vector 17P will have the dimension P red ×1 = 8×1. The ALWIP prediction matrix 17M will be of dimension M*N red *P red = 4*4*8 = 64. The 4×4 reduced block directly obtained when performing step 812 (formed by the gray columns in scenario 107) will have a size of Q red = M*N red = 4*4 = 16 samples (instead of Q = 4*8 = 32 for the original 4×8 block 18 to be predicted). Once the reduced 4×4 block is obtained through ALWIP, the offset value b k (step 812c) can be added and interpolation is performed in step 813. As can be seen from Figure 7.3 step 813 in, the reduced 4×4 block is expanded to a 4×8 block 18, where the value 108' not obtained in step 812 is obtained in step 813 by interpolating the values 118' and 118” (gray squares) obtained in step 812.
[0204] Therefore, the comparison between using this technology and not using this technology is as follows:
[0205] Without using this technology:
[0206] The size of the block 18 to be predicted is M = 4, N = 8
[0207] Q = M*N = 4*8 = 32 values to be predicted;
[0208] There are P = M + N = 4 + 8 = 12 samples in the boundary;
[0209] For each of the Q = 32 values to be predicted, P = 12 multiplications,
[0210] A total of P*Q = 12*32 = 384 multiplications
[0211] The ratio between the number of multiplications and the number of final values to be obtained is P*Q / Q = 12
[0212] When using this technology:
[0213] The size of the block 18 to be predicted is M = 4, N = 8
[0214] Finally, Q = M*N = 4*8 = 32 values are to be predicted;
[0215] However, Q red ×P red = 16×8 ALSIP matrix can be used, where M = 4, N red= 4, Q red = M * N red = 16, P red = M + N red = 4 + 4 = 8
[0216] P in the boundary red = M + N red = 4 + 4 = 8 samples, where P red < P
[0217] For Q of the reduced block to be predicted red = for each of the 16 values, P red = 8 multiplications,
[0218] In total Q red * P red = 16 * 8 = 16 * 8 = 128 multiplications (less than 384!)
[0219] The ratio between the number of multiplications and the number of final values to be obtained is P red * Q red / Q = 128 / 32 = 4 (much less than 12 obtained without using this technology!)
[0220] Therefore, by using this technology, the computational workload is reduced to one - third
[0221] Figure 7.4 It shows the case where the block 18 to be predicted has dimensions M × N = 16 × 16 and finally has Q = M * N = 16 * 16 = 256 values to be predicted, and this block has P = M + N = 16 + 16 = 32 boundary samples. This will result in a prediction matrix of size Q × P = 256 × 32, which means 256 * 32 = 8192 multiplications!
[0222] However, by applying method 820, it is possible to reduce (e.g., by averaging or down - sampling) the number of boundary samples at step 811, for example, from 32 to 8: For each group 120 of four consecutive samples in row 17a, one single sample is retained (e.g., selected from the four samples, or the average of the samples). Additionally, for each group of four consecutive samples in column 17c, one single sample is retained (e.g., selected from the four samples, or the average of the samples).
[0223] Here, the ALWIP matrix 17M is Q red × P red = 64 × 8 matrix: This is because: P has been selected red= 8 (by using 8 averaged or selected samples out of 32 samples from the boundary), and the reduced block to be predicted in step 812 is an 8×8 block (in scenario 109, the grey squares are 64).
[0224] Thus, once the 64 samples of the reduced 8×8 block are obtained in step 812, the remaining Q - Q of the block 18 to be predicted can be derived in step 813 red = 256 - 64 = 192 values 104.
[0225] In this case, to perform interpolation, all samples of the boundary column 17a have been selected and only alternative samples in the boundary row 17c are used. Other selections can be made.
[0226] While in the case of using this method, the ratio between the number of multiplications and the number of finally obtained values is Q red *P red / Q = 8 * 64 / 256 = 2, which is much less than 32 multiplications for each value in the case of not using this technology!
[0227] The comparison between using this technology and not using this technology is:
[0228] In the case of not using this technology:
[0229] The size of the block 18 to be predicted is M = 16, N = 16
[0230] Q = M * N = 16 * 16 = 256 values to be predicted;
[0231] P = M + N = 16 + 16 = 32 samples in the boundary;
[0232] For each of the Q = 256 values to be predicted, P = 32 multiplications,
[0233] In total, P * Q = 32 * 256 = 8192 multiplications;
[0234] The ratio between the number of multiplications and the number of finally obtained values is P * Q / Q = 32
[0235] In the case of using this technology:
[0236] The size of the block 18 to be predicted is M = 16, N = 16
[0237] Finally, Q = M * N = 16 * 16 = 256 values to be predicted;
[0238] But Q red ×P red = 64 × 8 ALWIP matrix, where M red = 4, Nred = 4, to predict Q through ALWIP red = 8 * 8 = 64 samples, P red = M red + N red = 4 + 4 = 8
[0239] P in the boundary red = M red + N red = 4 + 4 = 8 samples, where P red < P
[0240] For Q of the reduced block to be predicted red = For each of the 64 values, P red = 8 multiplications,
[0241] In total Q red * P red = 64 * 4 = 16 * 8 = 256 multiplications (less than 8192!)
[0242] The ratio between the number of multiplications and the number of final values to be obtained is P red * Q red / Q = 8 * 64 / 256 = 2 (much smaller than the 32 obtained without using this technology).
[0243] Therefore, the computing power required by this technology is 1 / 16 of the traditional technology!
[0244] Therefore, it is possible to use multiple adjacent samples (17) to predict a predetermined block (18) of a picture through the following operations:
[0245] Reduce (100, 813) multiple adjacent samples (17) to obtain a reduced set (102) of sample values that is fewer in number of samples compared to the multiple adjacent samples (17),
[0246] Perform a linear or affine linear transformation (19, 17M) on the reduced set of sample values to obtain the predicted values of the predetermined samples (104, 118', 188”) of the predetermined block (18).
[0247] Specifically, the reduction (100, 813) can be performed by downsampling multiple adjacent samples to obtain a reduced set (102) of sample values that is fewer in number of samples compared to the multiple adjacent samples (17).
[0248] Alternatively, the reduction (100, 813) can be performed by averaging multiple adjacent samples to obtain a reduced set (102) of sample values that is fewer in number of samples compared to the multiple adjacent samples (17).
[0249] In addition, based on the predicted values of predetermined samples (104, 118', 118") and multiple adjacent samples (17), the predicted values of other samples (108, 108') of a predetermined block (18) can be derived (813) by interpolation.
[0250] Multiple adjacent samples (17a, 17c) can extend one-dimensionally along both sides of the predetermined block (18) (e.g., to the right and downward in Figures 7.1 to 7.4 ). The predetermined samples (e.g., the predetermined samples obtained by ALWIP in step 812) can also be arranged in rows and columns and along at least one of the rows and columns, and the predetermined samples can be located at every nth position starting from the samples (112) adjacent to both sides of the predetermined block 18 in the predetermined samples 112.
[0251] Based on multiple adjacent samples (17), for each of at least one of the rows and columns, a support value (118) of one adjacent position (118) among multiple adjacent positions can be determined, and the support value (118) is aligned with the corresponding one of at least one of the rows and columns. It is also possible to derive the predicted values 118 of other samples (108, 108') of the predetermined block (18) by interpolation based on the predicted values of the predetermined samples (104, 118', 118") and the support values of the adjacent samples (118) aligned with at least one of the rows and columns.
[0252] The predetermined sample (104) can be located at every nth position along the row starting from the samples (112) adjacent to both sides of the predetermined block 18, and the predetermined sample is located at every mth position along the column starting from the samples adjacent to both sides of the predetermined block (18) in the predetermined samples (112), where n, m > 1. In some cases, n = m (e.g., in Figure 7.2 and Figure 7.3 where the samples 104, 118', 118" directly obtained by ALWIP at 812 and indicated by gray squares alternate with the samples 108, 108' subsequently obtained in step 813 along the rows and columns).
[0253] Along at least one of the row (17c) and the column (17a), for each support value, the determination of the support value can be performed, for example, by downsampling or averaging (122) a set (120) of adjacent samples within the multiple adjacent samples, and the set of adjacent samples includes the adjacent sample (118) for which the corresponding support value is determined. Thus, in Figure 7.4 in step 813, the value of the sample 119 can be obtained by using the value of the predetermined sample 118''' (previously obtained in step 812) and the value of the adjacent sample 118 as support values.
[0254] Multiple adjacent samples can extend one-dimensionally along both sides of a predetermined block (18). Reduction (811) can be performed by: grouping multiple adjacent samples (17) into one or more groups (110) of consecutive adjacent samples and downsampling or averaging each adjacent sample in the one or more groups (110) of adjacent samples, the one or more groups (110) of adjacent samples having two or more adjacent samples.
[0255] In an example, the linear or affine linear transformation can include P red *Q red or P red *Q weighting factors, where P red is the number of sample values (102) within a reduced set of sample values, and Q red or Q is the number of predetermined samples within the predetermined block (18). At least 1 / 4P red *Q red or 1 / 4P red *Q weighting factors are non-zero weight values. For each of the Q or Q red predetermined samples, P red *Q red or P red *Q weighting factors can include a series of P red weighting factors related to the respective predetermined sample, where when the series is arranged one after another in the predetermined samples of the predetermined block (18) according to a raster scan order, it forms an envelope of omnidirectional non-linearity. P red *Q or P red *Q red weighting factors can be uncorrelated with each other via any conventional mapping rule. The mean of the maximum values of the cross-correlation between the following two is less than a predetermined threshold: a first series of weighting factors related to a respective predetermined sample, a second series of weighting factors related to a predetermined sample other than the respective predetermined sample, or a reverse version of the latter series (whichever results in a higher maximum value). The predetermined threshold can be 0.3 [or in some cases 0.2 or 0.1]. P red adjacent samples (17) can be located along a one-dimensional path extending along both sides of the predetermined block (18), and for the Q or Q red predetermined samples, a series of P red weighting factors related to the respective predetermined sample are sorted in a manner that traverses the one-dimensional path in a predetermined direction.
[0256] 6.1 Description of the method and apparatus
[0257] To predict a sample of a rectangular block with width W (also denoted as N) and height H (also denoted as M), affine linear weighted intra prediction (ALWIP) can take as input one H - sized reconstructed adjacent boundary sample on the left side of the block and one W - sized reconstructed adjacent boundary sample above the block. If the reconstructed samples are not available, they can be generated as in traditional intra prediction.
[0258] The generation of the prediction signal (e.g., the value for the complete block 18) can be based on at least some of the following three steps:
[0259] 1. Among the boundary samples 17, samples 102 can be extracted by averaging or downsampling (e.g., four samples in the case of W = H = 4 and / or eight samples in other cases) (e.g., step 811).
[0260] 2. The averaged samples (or the remaining samples after downsampling) can be used as input to perform matrix - vector multiplication and then perform the addition of an offset. The result can be a reduced prediction signal on a subsampled set of samples in the original block (e.g., step 812).
[0261] 3. The prediction signal at the remaining positions can be generated, for example, by upsampling (e.g., by linear interpolation) based on the prediction signal on the subsampled set (e.g., step 813).
[0262] Due to step 1 (811) and / or step 3 (813), the total number of multiplications required to compute the matrix - vector product can be such that it is always less than or equal to 4 * W * H. Additionally, the averaging operation on the boundary and the linear interpolation of the reduced prediction signal are performed only by using addition and shift. In other words, in the example, each sample in the ALWIP mode requires at most four multiplications.
[0263] In some examples, the matrix (e.g., 17M) and the offset vector (e.g., b k ) for generating the prediction signal can be taken from a set of matrices (e.g., three sets) that can be stored in the storage units of the decoder and the encoder, such as S0, S1, S2.
[0264] In some examples, the set S0 can include (e.g., contain) n0 (e.g., n0 = 16 or n0 = 18 or another number) matrices i ∈ {0,…,n0 - 1}, each of which can have 16 rows and 4 columns and 18 offset vectors i ∈ {0,...,n0 - 1} (each offset vector has a size of 16) to perform the technique according to Figure 7.1 The matrices and offset vectors of this set are used for the 4×4 - sized block 18. Once the boundary vectors have been reduced to P red= 4 vectors (such as Figure 7.1 in step 811), it is possible to directly map the P of the reduced set 102 of samples red = 4 samples directly to the Q = 16 samples of the 4×4 block 18 to be predicted.
[0265] In some examples, the set S1 may include (e.g., contain) n1 (e.g., n1 = 8 or n1 = 18 or another number) matrices i ∈ {0, …, n1 - 1}, and each of these matrices may have 16 rows and 8 columns and 18 offset vectors i ∈ {0, …, n1 - 1} (each offset vector having a size of 16) to perform according to Figure 7.2 or Figure 7.3 of the technology. The matrices and offset vectors of the set S1 can be used for blocks of sizes 4×8, 4×16, 4×32, 4×64, 16×4, 32×4, 64×4, 8×4, and 8×8. Additionally, it can also be used for blocks of size WxH, where max(W, H) > 4 and min(W, H) = 4, i.e., for blocks of size 4×16 or 16×4, 4×32 or 32×4, and 4×64 or 64×4. The 16×8 matrix refers to a reduced version of the block 18, which is a 4×4 block, as obtained in Figure 7.2 and Figure 7.3 obtained.
[0266] Additionally or alternatively, the set S2 may include (e.g., contain) n2 (e.g., n2 = 6 or n2 = 18 or another number) matrices i ∈ {0,..., n2 - 1}, and each of these matrices may have 64 rows and 8 columns and 18 offset vectors i ∈ {0, …, n2 - 1} (the offset vectors having a size of 64). The 64×8 matrix refers to a reduced version of the block 18, which is an 8×8 block, e.g., as obtained in Figure 7.4 obtained. The matrices and offset vectors of this set can be used for blocks of sizes 8×16, 8×32, 8×64, 16×8, 16×16, 16×32, 16×64, 32×8, 32×16, 32×32, 32×64, 64×8, 64×16, 64×32, 64×64.
[0267] The matrices and offset vectors of this set or a part of these matrices and offset vectors can be used for all other block shapes.
[0268] 6.2 Averaging or downsampling of the boundaries
[0269] Here, the features of step 811 are provided.
[0270] As explained above, the boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0271] In a first step, the input boundaries bdry top (e.g., 17c) and bdry left (e.g., 17a) can be reduced to smaller boundaries and to obtain the reduced set 102. Here, in the case of a 4×4 block, and both consist of 2 samples, and in other cases, both consist of 4 samples.
[0272] In the case of a 4×4 block, it is possible to define:
[0273]
[0274]
[0275] And similarly define Thus, is the average value obtained, for example, using a shift operation.
[0276] In all other cases (e.g., for a block with a width or height not equal to 4), if the block width W is W = 4 * 2 k , then for 0 ≤ i < 4, it is defined that:
[0277]
[0278] And similarly define
[0279] In some other cases, it is possible to downsample the boundary (e.g., by selecting a specific boundary sample from a set of boundary samples) to obtain a reduced number of samples. For example, it is possible to select from bdry top [0] and bdry top [1] And it is possible to select from bdry top [2] and bdry top [3] It is also possible to similarly define
[0280] Two reduced boundaries and can be cascaded to the reduced boundary vector bdry red (associated with the reduced set 102), also denoted as 17P. The reduced boundary vector bdryred For a block of shape 4×4 ( Figure 7.1 example), the block size can be 4 (P red = 4), and for blocks of all other shapes ( Figures 7.2 to 7.4 example), the block size can be 8 (P red = 8).
[0281] Here, if mode < 18 (or the number of matrices in the set of matrices), then it is possible to define:
[0282]
[0283] If mode ≥ 18, which corresponds to the transposed mode of mode - 17, then it is possible to define:
[0284]
[0285] Therefore, depending on the specific state (one state: mode < 18; the other state: mode ≥ 18), it is possible to assign the predicted values of the output vector along different scan orders (e.g., one scan order: another scan order: ).
[0286] Other strategies can be executed. In other examples, the mode index "mode" does not have to be in the range of 0 to 35 (other ranges can be defined). Additionally, each of the three sets S0, S1, S2 does not have to have 18 matrices (thus, instead of an expression like mode ≥ 18, it is possible to use mode ≥ n0, n1, n2, which are the number of matrices in each set of matrices S0, S1, S2 respectively). Moreover, each set can have a different number of matrices (e.g., it can be: S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0287] The mode and transpose information do not have to be stored and / or transmitted as a combined mode index "mode": in some examples, it is possible to explicitly signal the transpose flag and matrix indices (S0 is 0 - 15, S1 is 0 - 7, and S2 is 0 - 5).
[0288] In some cases, the combination of the transpose flag and matrix indices can be interpreted as a set index. For example, there can be one bit serving as the transpose flag and some bits indicating the matrix indices, collectively referred to as the "set index".
[0289] 5.4 Generating a Reduced Prediction Signal through Matrix - Vector Multiplication
[0290] Here, the features regarding step 812 are provided.
[0291] In reducing the input vector bdry red (boundary vector 17P), a reduced prediction signal pred red can be generated. The latter signal can be a signal on a downsampled block of width W red and height H red . Here, W red and H red can be defined as:
[0292] If max(W, H) ≤ 8, then W red = 4, H red = 4,
[0293] Otherwise, W red = min(W, 8), H red = min(H, 8).
[0294] The reduced prediction signal pred can be calculated by computing a matrix-vector product and adding an offset red :
[0295] pred red = A · bdry red + b.
[0296] Here, A is a matrix (e.g., prediction matrix 17M), which can have W red * H red rows and 4 columns if W = H = 4, and 8 columns in all other cases, and b is a vector of size W red * H red .
[0297] If W = H = 4, A can have 4 columns and 16 rows, and thus, in this case, each sample may require 4 multiplications to compute pred red . In all other cases, A can have 8 columns, and it can be verified that in these cases 8 * W red * H red ≤ 4 * W * H, i.e., in these cases, each sample requires at most 4 multiplications to compute pred red .
[0298] The matrix A and the vector b can be taken from one of the following sets S0, S1, S2. Define the index idx = idx(W,H) by setting idx(W,H): If W = H = 4, then set idx(W,H) = 0; if max(W,H) = 8, then set idx(W,H) = 1, and in all other cases set idx(W,H) = 2. Additionally, if mode < 18, then m can be made equal to mode, otherwise make m = mode - 17. Then, if idx ≤ 1 or idx = 2 and min(W,H) > 4, then and in the case where idx = 2 and min(W,H) = 4, make A the matrix produced by removing every other row of which, in the case where W = 4, corresponds to the odd x - coordinates in the down - sampling block, or in the case where H = 4, corresponds to the odd y - coordinates in the down - sampling block. If mode ≥ 18, replace the reduced prediction signal with its transposed signal. In an alternative example, different strategies can be executed. For example, instead of reducing the size of a larger matrix (“removing”), use a smaller matrix S1 (idx = 1) with W red = 4 and H red = 4. That is, such a block is now assigned to S1 instead of S2.
[0299] Other strategies can be executed. In other examples, the mode index “mode” does not have to be in the range 0 to 35 (other ranges can be defined). Additionally, each of the three sets S0, S1, S2 does not have to have 18 matrices (so, instead of an expression like mode < 18, one could use mode < n0,n1,n2, which are the number of matrices in each set of matrices S0, S1, S2 respectively). Additionally, each set can have a different number of matrices (for example, it could be: S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0300] 6.4 Linear interpolation to generate the final prediction signal
[0301] Here, the features regarding step 812 are provided.
[0302] Interpolation of the subsampled prediction signal may require a second version of the border to be averaged on large blocks. That is, if min(W,H)>8 and W ≥ H, write W = 8*2 l , and for 0 ≤ i < 8, define:
[0303]
[0304] If min(W,H)>8 and H > W, define similarly
[0305] Additionally or alternatively, it may be "difficult to downsample", where is equal to:
[0306]
[0307] Furthermore, it can be defined similarly
[0308] At the sample positions removed when generating pred red , the final prediction signal can be generated by linear interpolation from pred red (e.g., step 813 in the example of Figures 7.2 to 7.4 ). In some examples, if W = H = 4 (e.g., the example of Figure 7.1 ), then this linear interpolation may be unnecessary.
[0309] The linear interpolation can be given as follows (although other examples are possible). Assume W ≥ H. Then, if H > H red , then vertical upsampling of pred red can be performed. In this case, pred red can extend one row to the top as follows. If W = 8, then pred red can have width W red = 4 and can be extended to the top by the averaged boundary signal , e.g., as defined above. If W > 8, then pred red has width W red = 8 and it is extended to the top by the averaged boundary signal , e.g., as defined above. For the first row of pred red , it can be obtained that pred red [x][-1]. Then, the signal red on a block with width W red and height 2*H can be given as follows:
[0310]
[0311]
[0312] where 0 ≤ x < W red and 0 ≤ y < H red . The subsequent process can be performed k times until 2 k *H red= H. Therefore, if H = 8 or H = 16, it can be performed at most once. If H = 32, it can be performed twice. If H = 64, it can be performed three times. Next, a horizontal upsampling operation can be applied to the result of the vertical upsampling. The subsequent upsampling operations can use all the boundaries to the left of the prediction signal. Finally, if H > W, a similar process can be performed by first upsampling in the horizontal direction (if necessary) and then in the vertical direction.
[0313] This is an example of interpolation using decimated boundary samples for the first interpolation (horizontally or vertically) and original boundary samples for the second interpolation (vertically or horizontally). Depending on the block size, only the second interpolation may be required or no interpolation may be needed. If both horizontal and vertical interpolations are required, the order depends on the width and height of the block.
[0314] However, different techniques can be implemented: for example, the original boundary samples can be used for both the first and second interpolations, and the order can be fixed, such as first horizontal and then vertical (in other cases, first vertical and then horizontal).
[0315] Therefore, the interpolation order (horizontal / vertical) and the use of decimated boundary samples / original boundary samples can be changed.
[0316] 6.5 Description of an example of the entire ALWIP process
[0317] For Figures 7.1 to 7.4 the different shapes in, the entire process of averaging, matrix-vector multiplication, and linear interpolation is shown. Note that the remaining shapes will be treated as one of the cases depicted.
[0318] 1. Given a 4×4 block, ALWIP can take two averages along each axis of the boundary by using Figure 7.1 the technique of. The resulting four input samples enter the matrix-vector multiplication. The matrix is taken from the set S0. After adding the offset, this can produce 16 final prediction samples. No linear interpolation is required to generate the prediction signal. Therefore, a total of (4*16) / (4*4) = 4 multiplications are performed per sample. See, for example, Figure 7.1 .
[0319] 2. Given an 8×8 block, ALWIP can take four averages along each axis of the boundary. By using Figure 7.2 the technique of, the resulting eight input samples enter the matrix-vector multiplication. The matrix is taken from the set S1. This produces 16 samples at the odd positions of the prediction block. Therefore, a total of (8*16) / (8*8) = 2 multiplications are performed per sample. After adding the offset, these samples can be interpolated, for example, by using vertical interpolation of the top boundary, and, for example, by using horizontal interpolation of the left boundary. See, for example,Figure 7.2 。
[0320] 3. Given an 8×4 block, ALWIP can take four averages along the horizontal axis of the boundary by using Figure 7.3 's technique and take four original boundary values on the left boundary. The resulting eight input samples enter matrix-vector multiplication. The matrix is taken from the set S1. This produces 16 samples at the odd horizontal positions of the prediction block and at each vertical position. Therefore, each sample performs a total of (8 * 16) / (8 * 4) = 4 multiplications. For example, after adding an offset, these samples will be horizontally interpolated by using the left boundary. See, for example, Figure 7.3 。
[0321] The transposed case is handled accordingly.
[0322] 4. Given a 16×16 block, ALWIP can take four averages along each axis of the boundary. By using Figure 7.2 's technique, the resulting eight input samples enter matrix-vector multiplication. The matrix is taken from the set S2. This produces 64 samples at the odd positions of the prediction block. Therefore, each sample performs a total of (8 * 64) / (16 * 16) = 2 multiplications. After adding an offset, these samples are vertically interpolated by using the top boundary and horizontally interpolated by using the left boundary. See, for example, Figure 7.2 。See, for example, Figure 7.4 。
[0323] For larger shapes, the process may be basically the same, and it is easy to check whether the number of multiplications per sample is less than two.
[0324] For a W×8 block, since the samples are given at the odd horizontal positions and at each vertical position, only horizontal interpolation is required. Therefore, in these cases, each sample performs at most (8 * 64) / (16 * 8) = 4 multiplications.
[0325] Finally, for a W×4 block, where W > 8, let A k be the matrix produced by removing each row corresponding to the odd entries along the horizontal axis of the downsampled block. Therefore, the output size can be 32, and again, only horizontal interpolation remains to be performed. Each sample can perform at most (8 * 32) / (16 * 4) = 4 multiplications.
[0326] The transposed case can be handled accordingly.
[0327] 6.6 Required number of parameters and complexity assessment
[0328] The parameters required for all possible intra prediction modes can be covered by matrices and offset vectors belonging to sets S0, S1, and S2. All matrix coefficients and offset vectors can be stored as 10-bit values. Thus, according to the above description, the proposed method may require a total of 14,400 parameters, each with a precision of 10 bits. This corresponds to 0.018 megabytes of memory. It should be noted that currently, a CTU of size 128×128 in the standard 4:2:0 chroma subsampling consists of 24,576 values, each value being 10 bits. Therefore, the memory requirement of the proposed intra prediction tool does not exceed that of the current picture reference tool adopted in the previous meeting. Additionally, it should be noted that due to the PDPC tool or the 4-tap interpolation filter for the angular prediction mode with fractional angular positions, the traditional intra prediction mode requires four multiplications per sample. Thus, in terms of operation complexity, the proposed method does not exceed the traditional intra prediction mode.
[0329] 6.7 Signaling of the Proposed Intra Prediction Mode
[0330] For example, for a luminance block, 35 ALWIP modes are proposed (other numbers of modes can be used). For each coding unit (CU) in the intra mode, a flag indicating whether the ALWIP mode is to be applied to the corresponding prediction unit (PU) is sent in the bitstream. The signaling of the latter index can be coordinated with the MRL in the same way as in the first CE test. If the ALWIP mode is to be applied, an MPM list with 3 MPMS can be used to signal the index predmode of the ALWIP mode.
[0331] Here, the intra modes of the upper PU and the left PU can be used as follows to perform the derivation of the MPM. There can be tables, for example, three fixed tables map_angular_to_alwip idx , idx ∈ {0, 1, 2}, which can assign an ALWIP mode to each traditional intra prediction mode predmode Angular :
[0332] predmode ALWIP = map_angular_to_alwip idx [predmode Angular .
[0333] For each PU with width W and height H, define and index
[0334] idx(PU) = idx(W, H) ∈ {0, 1, 2}
[0335] It indicates from which one of the three sets the ALWIP parameters will be obtained, as described in Section 4 above. If the upper prediction unit PU above is available, belongs to the same CTU as the current PU and is in the intra mode, then if idx(PU) = idx(PU above ) and if ALWIP is applied to the PU in the ALWIP mode above , then make:
[0336]
[0337] If the upper PU is available, belongs to the same CTU as the current PU and is in the intra mode, and if the traditional intra prediction mode is applied to the upper PU, then make:
[0338]
[0339] In all other cases, make:
[0340]
[0341] It means that this mode is not available. Derive the mode in the same way but without restricting that the left PU needs to belong to the same CTU as the current PU:
[0342]
[0343] Finally, three fixed default lists list idx are provided, idx ∈ {0, 1, 2}, each of which contains three different ALWIP modes. In the default list list idx(PU) and the modes and , three different MPMs are constructed by replacing -1 with the default value and eliminating duplicates.
[0344] The embodiments described herein are not limited by the above signaling of the proposed intra prediction mode. According to an alternative embodiment, there is no MPM and / or mapping table for MIP(ALWIP).
[0345] 6.8 Derivation of MPM Lists Applicable to Traditional Luminance Intra Prediction Mode and Chrominance Intra Prediction Mode
[0346] The proposed ALWIP mode can be coordinated with the MPM-based coding of the traditional intra prediction mode as follows. The process of deriving the luminance MPM list and the chrominance MPM list of the traditional intra prediction mode can use the fixed table map_alwip_to_angular idx, where idx ∈ {0, 1, 2}, map the ALWIP mode predmode on the given PU ALWIP to one of the traditional intra prediction modes:
[0347] predmode Angular =
[0348] map_alwip_to_angular idx(PU) [predmode ALWIP .
[0349] For luma MPM list derivation, whenever an adjacent luma block using the ALWIP mode predmode ALWIP is encountered, the block can be processed as if it were using the traditional intra prediction mode predmode Angular . For chroma MPM list derivation, whenever the current luma block uses the LWIP mode, the ALWIP mode can be converted to a traditional intra prediction mode using the same mapping.
[0350] Obviously, the ALWIP mode can also be coordinated with the traditional intra prediction mode without using the MPM and / or the mapping table. For example, for chroma blocks, whenever the current luma block uses the ALWIP mode, the ALWIP mode can be mapped to the planar intra prediction mode.
[0351] 7. Efficient embodiments
[0352] Let us briefly summarize the above examples as they may form the basis for other extended embodiments described below.
[0353] To predict a predetermined block 18 of Picture 10, multiple adjacent samples 17a, 17c are used.
[0354] Reduction 100 has been performed by averaging multiple adjacent samples to obtain a reduced set 102 of sample values that is fewer in the number of samples compared to the multiple adjacent samples. This reduction is optional in the embodiments herein and results in a so-called sample value vector described below. A linear or affine-linear transformation 19 is performed on the reduced set of sample values to obtain the predicted value of the predetermined samples 104 of the predetermined block. This transformation is later indicated using matrix A and offset vector b, has been obtained by machine learning (ML), and should be an effectively preformed implementation.
[0355] Through interpolation, the predicted value of other samples 108 of a predetermined block is derived based on the predicted value of a predetermined sample and multiple adjacent samples. It should be noted that theoretically, the result of an affine / linear transformation can be associated with the non-full pixel sample positions of block 18, such that all samples of block 18 can be obtained through interpolation according to an alternative embodiment. It is also possible that interpolation may not be required at all.
[0356] The multiple adjacent samples can extend one-dimensionally along two sides of the predetermined block. The predetermined sample is arranged by rows and columns and along at least one of the rows and columns, where the predetermined sample can be located at every nth position starting from the samples (112) adjacent to the two sides of the predetermined block among the predetermined samples. Based on the multiple adjacent samples, for each of at least one of the rows and columns, a support value of one adjacent position (118) among the multiple adjacent positions can be determined. The support value is aligned with the corresponding one of at least one of the rows and columns, and through interpolation, the predicted value of other samples 108 of the predetermined block can be derived based on the predicted value of the predetermined sample and the support values of the adjacent samples aligned with at least one of the rows and columns. The predetermined sample can be located at every nth position along the row starting from the samples 112 adjacent to the two sides of the predetermined block 18 of the predetermined sample, and the predetermined sample can be located at every mth position along the column starting from the samples 112 adjacent to the two sides of the predetermined block, where n, m > 1. It is possible that n = m. Along at least one of the rows and columns, for each support value, the support value can be determined by averaging (122) a set of adjacent samples 120 within the multiple adjacent samples. The set of adjacent samples 120 includes the adjacent sample 118 for which the corresponding support value is determined. The multiple adjacent samples can extend one-dimensionally along two sides of the predetermined block, and the reduction can be completed by grouping the multiple adjacent samples into one or more groups 110 of consecutive adjacent samples and performing averaging on each adjacent sample in the one or more groups of adjacent samples. The one or more groups of adjacent samples have more than two adjacent samples.
[0357] For a predetermined block, a prediction residual can be transmitted in the data stream. It can be derived from the data stream at the decoder, and the predetermined block is reconstructed using the prediction residual and the predicted value of the predetermined sample. At the encoder, the prediction residual is encoded into the data stream at the encoder.
[0358] The picture can be subdivided into multiple blocks of different block sizes, and the multiple blocks include a predetermined block. Then, a linear or affine-linear transformation for block 18 can be selected depending on the width W and height H of the block, such that the linear or affine-linear transformation selected for the predetermined block is selected from the following: a first set of linear or affine-linear transformations, provided that the width W and height H of the predetermined block are within a first set of width / height pairs; a second set of linear or affine-linear transformations, provided that the width W and height H of the predetermined block are within a second set of width / height pairs, the second set of width / height pairs being disjoint from the first set of width / height pairs. Also, as will become clear later, the affine / linear transformation is represented by other parameters, namely the weights C and optionally offset and scale parameters.
[0359] The decoder and encoder can be configured to subdivide a picture into multiple blocks of different block sizes, which include a predetermined block, and to select a linear or affine-linear transformation depending on the width W and height H of the predetermined block, such that the linear or affine-linear transformation selected for the predetermined block is selected from the following:
[0360] a first set of linear or affine-linear transformations, provided that the width W and height H of the predetermined block are within a first set of width / height pairs,
[0361] a second set of linear or affine-linear transformations, provided that the width W and height H of the predetermined block are within a second set of width / height pairs that is disjoint from the first set of width / height pairs, and
[0362] a third set of linear or affine-linear transformations, provided that the width W and height H of the predetermined block are within a third set of one or more width / height pairs, the third set of width / height pairs being disjoint from the first set of width / height pairs and the second set of width / height pairs.
[0363] The third set of one or more width / height pairs includes only one width / height pair W', H', and each linear or affine-linear transformation within the first set of linear or affine-linear transformations is used to transform N' sample values into W'*H' predicted values in a W'xH' array of sample positions.
[0364] Each of the first and second sets of width / height pairs may include a first width / height pair W p , H p , where W p is not equal to H p , and includes a second width / height pair W q , H q , where H q =W p and W q =H p .
[0365] Each of the first set and the second set of width / height pairs may additionally include a third width / height pair W p 、H p ,where W p is equal to H p and H p >H q 。
[0366] For a predetermined block, a set index may be transmitted in the data stream, which indicates which linear or affine linear transformation from a predetermined set of linear or affine linear transformations is to be selected for block 18.
[0367] Multiple adjacent samples may extend one-dimensionally along both sides of a predetermined block, and the reduction may be accomplished by: for a first subset of multiple adjacent samples adjacent to a first side of the predetermined block, grouping the first subset into one or more first groups 110 of consecutive adjacent samples; and for a second subset of multiple adjacent samples adjacent to a second side of the predetermined block, grouping the second subset into one or more second groups 110 of consecutive adjacent samples; and performing an average on each of the one or more first groups of adjacent samples and the one or more second groups of adjacent samples, each group having more than two adjacent samples, so as to obtain a first sample value from the first group and a second sample value of the second group. Then, a linear or affine linear transformation may be selected depending on the set index in the predetermined set of linear or affine linear transformations, such that two different states of the set index result in the selection of one of the linear or affine linear transformations in the predetermined set of linear or affine linear transformations. In the case where the set index assumes the first state of the two different states in the form of a first vector, a predetermined linear or affine linear transformation may be performed on the reduced set of sample values to produce an output vector of predicted values, and the predicted values of the output vector may be assigned to predetermined samples of the predetermined block along a first scan order. And in the case where the second state of the two different states is assumed in the form of a second vector, the first vector and the second vector are different, such that a component filled with one of the first sample values in the first vector is filled with one of the second sample values in the second vector, and a component filled with one of the second sample values in the first vector is filled with one of the first sample values in the second vector, so as to produce an output vector of predicted values, and the predicted values of the output vector may be assigned to predetermined samples of the predetermined block transposed relative to the first scan order along a second scan order.
[0368] For a w1×h1 array of sample positions, each linear or affine linear transformation within a first set of linear or affine linear transformations can be used to transform N1 sample values into w1*h1 predicted values, and for a w2×h2 array of sample positions, each linear or affine linear transformation within a second set of linear or affine linear transformations is used to transform N2 sample values into w2*h2 predicted values, where for a first predetermined width / height pair in a first set of width / height pairs, w1 can exceed the width of the first predetermined width / height pair, or h1 can exceed the height of the first predetermined width / height pair, and for a second predetermined width in the first set of width / height pairs, w1 cannot exceed the width of the second predetermined width / height pair and h1 cannot exceed the height of the second predetermined width / height pair. Then reduction (100) can be performed by averaging a plurality of adjacent samples to obtain a reduced set (102) of sample values such that: if a predetermined block is a predetermined block of the first predetermined width / height pair and if the predetermined block is a predetermined block of the second predetermined width / height pair, the reduced set 102 of sample values has N1 sample values, and a selected linear or affine linear transformation can be performed on the reduced set of sample values by only using a first sub - portion of the selected linear or affine linear transformation: if w1 exceeds the width of a width / height pair, the first sub - portion is related to subsampling of the w1×h1 array of sample positions along the width dimension, or if h1 exceeds the height of a width / height pair, the first sub - portion is related to subsampling of the w1×h1 array of sample positions along the height dimension, and if the predetermined block is the first predetermined width / height pair, the selected linear or affine linear transformation is performed completely on the reduced set of sample values.
[0369] For a w1×h1 array of sample positions (w1 = h1), each linear or affine linear transformation within a first set of linear or affine linear transformations can be used to transform N1 sample values into w1*h1 predicted values, and for a w2×h2 array of sample positions (w2 = h2), each linear or affine linear transformation within a second set of linear or affine linear transformations is used to transform N2 sample values into w2*h2 predicted values.
[0370] All of the above embodiments are illustrative only, as they can form the basis of the embodiments described below. That is, the above concepts and details are applied to understand the following embodiments and serve as a repository for possible extensions and modifications of the embodiments described below. Specifically, many of the above details are optional, such as the averaging of adjacent samples, the fact that adjacent samples are used as reference samples, etc.
[0371] More generally, the embodiments described herein assume that the prediction signals on a rectangular block are generated from already reconstructed samples. For example, the intra-prediction signals on a rectangular block are generated from adjacent reconstructed samples to the left and above the block. The generation of the prediction signals is based on the following steps.
[0372] 1. In the reference samples, now referred to as boundary samples, but without excluding the possibility of transferring the description to reference samples located elsewhere, samples can be extracted by taking an average. Here, the boundary samples on both the left and above the block or only the boundary samples on one of the two sides are averaged. If one side is not averaged, the samples on that side remain unchanged.
[0373] 2. Perform a matrix-vector multiplication, optionally followed by adding an offset. Wherein, if averaging is only applied on the left side, the input vector of the matrix-vector multiplication is the concatenation of the averaged boundary samples on the left side of the block and the original boundary samples above the block; or if averaging is only applied on the upper side, the input vector of the matrix-vector multiplication is the concatenation of the original boundary samples on the left side of the block and the averaged boundary samples above the block; or if averaging is applied on both sides of the block, the input vector of the matrix-vector multiplication is the concatenation of the averaged boundary samples on the left side of the block and the averaged boundary samples above the block. Similarly, there will be alternatives, such as an alternative of not using averaging at all.
[0374] 3. The result of the matrix-vector multiplication and the optional offset addition can optionally be a reduced prediction signal on a subsampled set of samples in the original block. The prediction signals at the remaining positions can be generated from the prediction signal on the subsampled set by linear interpolation.
[0375] The calculation of the matrix-vector product in step 2 is preferably performed with integer arithmetic. Thus, if x = (x1, …, x n ) represents the input of the matrix-vector product, that is, x represents the concatenation of the (averaged) boundary samples on the left and above the block, then in x, the (reduced) prediction signal calculated in step 2 should be calculated only using shifts, addition of offset vectors, and multiplication by integers. Ideally, the prediction signal in step 2 will be given as Ax + b, where b is an offset vector that may be zero, and where A is derived by some machine learning-based training algorithm. However, such a training algorithm typically only produces a matrix A = A float . Therefore, there is a problem of specifying integer operations in the foregoing sense such that these integer operations approximate the expression A float x well. Here, it is important to mention that these integer operations do not have to be chosen such that they approximate the expression A float x under the assumption that the vector x is uniformly distributed, but usually the input of this expression A floatThe vector x to be approximated is a (mean) boundary sample from a natural video signal, where some correlations between the components of x can be expected. i between them.
[0376] Figure 8 An improved ALWIP prediction is shown. Samples of a predetermined block can be predicted based on a first matrix-vector product between a matrix A 1100 derived by some machine learning-based training algorithm and a sample value vector x 400. Optionally, an offset b 1110 can be added. To achieve an integer approximation or a fixed-point approximation of this first matrix-vector product, a reversible linear transformation 403 can be performed on the sample value vector to determine another vector 402. A second matrix-vector product between another matrix B 1200 and the other vector 402 can be equal to the result of the first matrix-vector product.
[0377] Due to the characteristics of the other vector 402, the second matrix-vector product can be an integer approximated by a matrix-vector product 404 between a predetermined prediction matrix C 405 and the other vector 402 plus another offset 408. The other vector 402 and the other offset 408 can consist of integer or fixed-point values. All components of the other offset are, for example, the same. The predetermined measurement matrix 405 can be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the predetermined prediction matrix 405 and the other vector 402 can be understood as a prediction vector 406.
[0378] More details about this integer approximation are provided below.
[0379] Possible solution according to Example I: Subtracting and adding the mean
[0380] An expression A that can be used for the above situation float One possible combination for the integer approximation of x is to replace the i0-th component of x (i.e., the sample value vector 400) with the mean of the components of x, mean(x) (i.e., the predetermined value 1400) (i.e., the predetermined component 1500), and subtract this mean from all other components. In other words, as Figure 9a shown, the reversible linear transformation 403 is defined such that: the predetermined component 1500 of the other vector 402 becomes a, and each of the other components in the other vector 402 except the predetermined component 1500 is equal to the corresponding component of the sample value vector 400 minus a, where a is the predetermined value 1400, which is, for example, the average of the components of the sample value vector 400, such as the arithmetic mean or the weighted average. This operation on the input is given by the reversible transformation T403, which has an obvious integer implementation, especially if the size n of x is a power of 2.
[0381] Since A float =(Afloat T -1 )T, if such a transformation is performed on the input x, an integer approximation of the matrix-vector product By must be found, where B = (A float T -1 ) and y = Tx. Since the matrix-vector product A float x represents a prediction of a rectangular block (i.e., a predetermined block), and since x400 is included in the (e.g., averaged) boundary samples of the block, it should be expected that in the case where all sample values of x are equal, i.e., for all i, x i = mean(x), each sample value in the prediction signal A float x should be close to mean(x) or exactly equal to mean(x). This means that it should be expected that the i0-th column (i.e., the column corresponding to the predetermined component of B) is very close to or equal to the column consisting only of 1s. Therefore, if M(i0), i.e., the integer matrix 1300, is the matrix: whose i0-th row consists of 1s and all its other rows are zero, written as By = Cy + M(i0), where C = B - M(i0), it should be expected that the i0-th row 412 of C (i.e., the predetermined prediction matrix 405) has small entries or is zero, as Figure 9b shown. In addition, since the components of x are correlated, it can be expected that for each i ≠ i0, the i-th component y i = x i - mean(x) component generally has a much smaller absolute value than the i-th component of x. Since the matrix M(i0)1300 is an integer matrix, if an integer approximation of Cy is given, an integer approximation of By is achieved, and from the above discussion, it can be expected that the quantization error resulting from quantizing each entry of C405 in a suitable manner corresponds to A float x and only slightly affects the error in the quantization result of By.
[0382] The predetermined value 1400 does not have to be the mean mean(x). The following alternative definition of the predetermined value 1400 can also be used to achieve the integer approximation of the expression A float x described herein:
[0383] In another possible combination of the integer approximation of the expression A float x, the i0-th component of x remains unchanged, and the same value is subtracted from all other components i.e., for each i ≠ i0, In other words, the preset value 1400 can be the component in the sample value vector 400 corresponding to the preset component 1500.
[0384] Alternatively, the predetermined value 1400 is a default value or a value signaled in the data stream into which the picture is encoded.
[0385] The predetermined value 1400 is equal to 2, for example. bitdepth-1 . In this case, another vector 402 can be defined by y0 = 2 bitdepth-1 and y i = x i - x0 (for i > 0).
[0386] Alternatively, the predetermined component 1500 becomes a constant minus the predetermined value 1400. The constant is equal to 2, for example. bitdepth-1 . According to an embodiment, the predetermined component 1500 of another vector y402 is equal to 2 bitdepth-1 minus the component in the sample value vector 400 corresponding to the predetermined component 1500 and all other components of another vector 402 are equal to the corresponding components of the sample value vector 400 minus the component in the sample value vector 400 corresponding to the predetermined component 1500.
[0387] For example, it is advantageous if the predetermined value 1400 has a small deviation from the predicted value of the samples of the predetermined block.
[0388] According to an embodiment, the apparatus 1000 is configured to include a plurality of reversible linear transforms 403, where each reversible linear transform is associated with one component of another vector 402. Further, the apparatus is configured to select, for example, the predetermined component 1500 from the components of the sample value vector 400, and use the reversible linear transform 403 associated with the predetermined component 1500 among the plurality of reversible linear transforms as the predetermined reversible linear transform. This is because, for example, the different positions of the i0-th row (i.e., the row of the reversible linear transform 403 corresponding to the predetermined component) depend on the position of the predetermined component in the other vector. If, for example, the first component of another vector 402 (i.e., y1) is the predetermined component, the i0-th row replaces the first row of the reversible linear transform.
[0389] As Figure 9b shown, the matrix components 414 of the predetermined prediction matrix C 405 within the column 412 (i.e., the i0-th column) of the predetermined prediction matrix 405 are all zero, for example, and the matrix component 414 corresponds to the predetermined component 1500 of another vector 402. In this case, the apparatus is configured to perform multiplication to calculate the matrix-vector product 404 by, for example: calculating the matrix-vector product 407 between the reduced prediction matrix C'405 obtained by ignoring the column 412 from the predetermined prediction matrix C 405 and another vector 410 obtained by ignoring the predetermined component 1500 from another vector 412, as Figure 9c shown. Thus, the prediction vector 406 can be calculated with fewer multiplications.
[0390] As Figure 8 , Figure 9b and Figure 9c shown, the apparatus 1000 may be configured to: when predicting samples of a predetermined block based on the prediction vector 406, for each component of the prediction vector 406, calculate the sum of the corresponding component and a (i.e., the predetermined value 1400). This sum may be represented by the sum of the prediction vector 406 and the vector 409, where all components of the vector 409 are equal to the predetermined value 1400, as Figure 8 and Figure 9c shown. Alternatively, this sum may be represented by the sum of the matrix-vector product 1310 between the prediction vector 406 and the integer matrix M 1300 and another vector 402, as Figure 9b shown, where the matrix component in the integer matrix 1300 corresponding to the predetermined component 1500 of the other vector 402 is 1 within a column of the integer matrix 1300 (i.e., the i0-th column), and all other components are, for example, zero.
[0391] The result of the sum of the predetermined prediction matrix C 405 and the integer matrix 1300 is equal to or approximate to, for example Figure 8 another matrix B 1200 shown in
[0392] In other words, another matrix B 1200 generated by summing each matrix component in the predetermined prediction matrix C 405 corresponding to the predetermined component 1500 of the other vector 402 within the column 412 (i.e., the i0-th column) of the predetermined prediction matrix 405 with 1 (i.e., matrix B) and multiplying by the invertible linear transformation 403 corresponds to a quantized version of, for example, the machine learning prediction matrix A 1100, as Figure 8 , Figure 9a and Figure 9b shown. The sum of each matrix component in the predetermined prediction matrix C 405 within the i0-th column 412 with 1 may correspond to the sum of the predetermined prediction matrix 405 and the integer matrix 1300, as Figure 9b shown. As Figure 8 shown, the machine learning prediction matrix A 1100 may be equal to the result of multiplying another matrix 1200 by the invertible linear transformation 403. This is because A·x = BT·yT -1 . The predetermined prediction matrix 405 is, for example, a quantized matrix, an integer matrix, and / or a fixed-point matrix, whereby a quantized version of the machine learning prediction matrix A 1100 can be achieved.
[0393] Matrix multiplication using only integer operations
[0394] For low-complexity implementations (in terms of the complexity of adding and multiplying scalar values, and in terms of the storage required for the entries of the matrices involved), it is desirable to perform matrix multiplication 404 using only integer arithmetic.
[0395] To compute an approximation of z = Cy, namely:
[0396]
[0397] According to an embodiment, using only integer operations, the real-valued C i,j must be mapped to integer values This can be done, for example, by uniform scalar quantization or by taking into account specific correlations between the values of y i The integer values represent, for example, fixed-point numbers, each of which can be stored with a fixed number of bits n_bits, e.g., n_bits = 8.
[0398] The matrix-vector product 404 with a matrix of size m × n (i.e., the predetermined prediction matrix 405) can then be performed as shown in this pseudocode, where <<, >> are arithmetic binary left shift and right shift operations, and +, -, and * operate only on integer values. (1)
[0400]
[0401] Here, the array C, i.e., the predetermined prediction matrix 405, stores fixed-point numbers as integers. The final addition of final_offset and the right shift operation of right_shift_result reduce the precision by rounding to obtain the required fixed-point format at the output.
[0402] To allow for an increased range of real values represented by integers in C, two additional matrices offset i,j and scale i,j , as Figure 10 and Figure 11 shown in the embodiments of j such that each coefficient b i,j of y
[0403]
[0404] in the following matrix-vector product is given by
[0405]
[0406] offset i,j and scale i,jThe value itself is an integer value. For example, these integers can represent fixed-point numbers, and each fixed-point number can be stored with a fixed number of bits (e.g., 8 bits) or, for example, the same number of bits n_bits as the number of bits used to store the value of the value.
[0407] In other words, the apparatus 1000 is configured to use prediction parameters (e.g., integer values and the value offset i,j and scale i,j ) to represent a predetermined prediction matrix 405, and to calculate the matrix-vector product 404 by performing multiplication and summation on the components of another vector 402, the prediction parameters, and the resulting intermediate results, where the absolute value of the prediction parameter can be represented by an n-bit fixed-point number representation, where n is equal to or less than 14, or alternatively equal to or less than 10, or alternatively equal to or less than 8. For example, the components of another vector 402 are multiplied by the prediction parameters to produce a product as an intermediate result, and then the intermediate results are summed or form the addends of the sum.
[0408] According to an embodiment, the prediction parameters include weights, each weight being associated with a corresponding matrix component of the prediction matrix. In other words, the predetermined prediction matrix is replaced or represented by the prediction parameters, for example. The weights are, for example, integer values and / or fixed-point values.
[0409] According to an embodiment, the prediction parameters further include one or more scaling factors, such as the value scale i,j , each scaling factor being associated with one or more corresponding matrix components of the predetermined prediction matrix 405 to scale the weights associated with one or more corresponding matrix components of the predetermined prediction matrix 405, such as integer values Additionally or alternatively, the prediction parameters include one or more offsets, such as the value offset i,j , each offset being associated with one or more corresponding matrix components of the predetermined prediction matrix 405 to offset the weights associated with one or more corresponding matrix components of the predetermined prediction matrix 405, such as integer values
[0410] To reduce the storage amount required for offset i,j and scale i,j , their values can be chosen to be constant for a particular set of indices i, j. For example, their entries can be constant for each column, or they can be constant for each row, or they can be constant for all i, j, as Figure 10 shown.
[0411] For example, in a preferred embodiment, offset i,j and scale i,jAll values of a matrix for a prediction mode are constant, as Figure 11 shown in. Thus, when there are K prediction modes, where k = 0..K-1, only a single value o k and a single value s k are needed to compute the prediction for mode k.
[0412] According to an embodiment, for all matrix-based intra prediction modes, offset i,j and / or scale i,j is constant, i.e., the same. Additionally or alternatively, for all block sizes, offset i,j and / or scale i,j may be constant, i.e., the same.
[0413] In the case where o k represents the offset and s k represents the scale, the calculation in (1) can be modified to: (2)
[0415]
[0416] Extended embodiments resulting from this solution
[0417] The above solution implies the following embodiments:
[0418] 1. A prediction method as in Part I, wherein in step 2 of Part I, the following operations are performed on the integer approximation of the matrix-vector product involved: In the (averaged) boundary samples x = (x1,..., x n ), for a fixed i0 (where 1 ≤ i0 ≤ n), the vector y = (y1,..., y n ) is computed, where y i = x i - mean(x) (for i ≠ i0) and where and where mean(x) represents the mean of x. The vector y is then used as the input to the integer implementation of the matrix-vector product Cy such that the (downsampled) prediction signal pred from step 2 of Part I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents the signal at each sample position in the domain of the (downsampled) prediction signal, which is equal to mean(x). (See, for example, Figure 9b )
[0419] 2. A prediction method as in Part I, wherein in step 2 of Part I, the following operations are performed on the integer approximation of the matrix-vector product involved: In the (averaged) boundary samples x = (x1,..., x n) For a fixed i0 where 1 ≤ i0 ≤ n, compute the vector y = (y1, …, y n-1 ), where y i = x i - mean(x) for i < i0, and where y i = x i+1 - mean(x) for i ≥ i0, and where mean(x) represents the mean of x. The vector y is then used as the input to the matrix-vector product Cy (the integer implementation thereof), such that the (downsampled) predicted signal pred from step 2 of Part I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents the signal at each sample position in the domain of the (downsampled) predicted signal, which is equal to mean(x). (See, for example Figure 9c )
[0420] 3. A prediction method as in Part I, wherein the integer implementation of the matrix-vector product Cy is given by using the coefficients in the matrix-vector product z i = ∑ j b i,j *y j . (See, for example ) Figure 10 )
[0421] 4. A prediction method as in Part I, wherein step 2 uses one of K matrices such that multiple prediction patterns can be computed, each prediction pattern using a different matrix where k = 0…K - 1, where the integer implementation of the matrix-vector product C k y is given by using the coefficients in (the matrix-vector product z i = ∑ j b i,j *y j ). (See, for example ) Figure 11 )
[0422] That is, according to an embodiment of the present application, the encoder and decoder operate as follows to predict the predetermined block 18 of picture 10, see Figure 8 . For prediction, multiple reference samples are used. As outlined above, embodiments of the present application will not be limited to intra coding, and thus, the reference samples will not be limited to neighboring samples, that is, samples in picture 10 adjacent to block 18. Specifically, the reference samples will not be limited to samples arranged along the outer edge of block 18, such as samples adjacent to the outer edge of the block. However, this case is of course an embodiment of the present application.
[0423] To perform the prediction, a sample value vector 400 is formed from reference samples such as reference sample 17a and reference sample 17c. The possible formation has been described above. The formation can involve averaging, whereby the number of samples 102 or the number of components of the vector 400 is reduced compared to the reference samples contributing to the formation. As described above, the formation can also depend in some way on the size or dimensions of block 18, such as its width and height.
[0424] The vector 400 should be subjected to an affine or linear transformation in order to obtain the prediction of block 18. Different nomenclatures have been used above. Using the most recent nomenclature, the aim is to perform the prediction by applying the vector 400 to matrix A via a matrix-vector product when performing the summation with the offset vector b. The offset vector b is optional. The affine or linear transformation determined by A, or A and b, can be determined by the encoder and decoder, or more precisely, for the prediction based on the size and dimensions of block 18 as already described above, the affine or linear transformation is determined by the encoder and decoder.
[0425] However, in order to achieve the computational efficiency improvement outlined above or to make the prediction more effective in terms of implementation, the affine or linear transformation has been quantized, and the encoder and decoder or their predictors use C and T mentioned above in order to represent and perform the linear or affine transformation, where C and T applied in the above manner represent the quantized version of the affine transformation. Specifically, the predictors in the encoder and decoder do not directly apply the vector 400 to matrix A, but instead apply the vector 402 generated from the sample value vector 400 to matrix A in a way that is mapped via a predetermined invertible linear transformation T. As long as the vector 400 has the same size (i.e., does not depend on the dimensions of the block, i.e., width and height) or is at least the same for different affine / linear transformations, the transformation T used here can be the same. Above, the vector 402 was represented as y. To perform the affine / linear transformation determined by machine learning, the exact matrix would be B. However, the prediction in the encoder and decoder is not exactly performed as B, but rather through its approximate version or quantized version. Specifically, the representation is performed via C appropriately represented in the manner outlined above, where C + M represents the quantized version of B.
[0426] Thus, prediction in the encoder and decoder is further performed by computing the matrix-vector product 404 between the vector 402 and a predetermined prediction matrix C that is appropriately represented and stored at the encoder and decoder in the above-described manner. The vector 406 resulting from this matrix-vector product is then used to predict the samples 104 of block 18. As described above, for prediction, as indicated at 408, each component of the vector 406 can be summed with a parameter a in order to compensate for the corresponding definition of C. Deriving the prediction of block 18 based on the vector 406 can also involve an optional summation of the vector 406 with an offset vector b. As described above, each component of the vector 406, and accordingly each component of the summation of the vector 406, the vector of all a' indicated at 408, and the optional vector b, can directly correspond to the samples 104 of block 18 and thus indicate the predicted values of the samples. It is also possible that only a subset of the samples 104 of the block are predicted in this way, while the remaining samples of block 18, such as 108, are derived by interpolation.
[0427] As described above, there are different embodiments for setting a. For example, it may be the arithmetic mean of the components of the vector 400. For this case, see Figure 9a . The invertible linear transformation T 403 can be as Figure 9a indicated. i0 are respectively the sample value vector and the predetermined component of the vector 402, which are replaced with a. However, as indicated above, there are other possibilities. However, in terms of the representation of C, it has also been indicated above that C can be embodied in different ways. For example, the matrix-vector product 404 can end with the actual calculation of a smaller matrix-vector product of lower dimension in its actual calculation. Specifically, as indicated above, due to the definition of C, the entire i0-th column 412 of C can become 0, such that the actual calculation of the product 404 can be performed with a reduced version of the vector 402, which is produced by omitting the component (i.e., multiplying this reduced vector 410 with a reduced matrix C' produced from C by omitting the i0-th column 412) to produce the vector 402.
[0428] The weights of C or the weights of C' (i.e., the components of this matrix) can be represented and stored in fixed-point representation. However, as described above, these weights 414 can also be stored in a manner related to different scaling and / or offset. Scaling and offset can be defined for the entire matrix C, i.e., equal for all weights 414 of the matrix C or matrix C', or can be defined in such a way that the weights 414 of the same column or the same row of the matrix C and matrix C' are respectively constant or equal. In this regard, Figure 10 it is shown that the calculation of the matrix-vector product (i.e., the result of the product) can actually be performed slightly differently, i.e., for example, by bringing the multiplication with the scaling towards the vector 402 or vector 410, thereby reducing the number of multiplications that have to be further performed. Figure 11Shows the case of using one scaling and one offset for all weights 414 of C or C', as performed in the above calculation (2).
[0429] According to an embodiment, the apparatus for predicting a predetermined block of a picture described herein may be configured to use matrix-based intra-sample prediction, which includes the following features:
[0430] The apparatus is configured to form a sample value vector pTemp[x]400 from a plurality of reference samples 17. Assuming pTemp[x] is 2*boundarySize, pTemp[x] can be filled with the following samples, for example, by direct copying or by subsampling or pooling: adjacent samples redT[x] (where x = 0..boundarySize–1) located at the top of the predetermined block, followed by adjacent samples redL[x] (where x = 0..boundarySize–1) located to the left of the predetermined block (for example, in the case of isTransposed = 0), or vice versa in the case of a transpose process (for example, in the case of isTransposed = 1).
[0431] Derive input values p[x], where x = 0..inSize-1, that is, the apparatus is configured to derive another vector p[x] from the sample value vector pTemp[x] by mapping the sample value vector pTemp[x] to a vector p[x] through a predetermined invertible linear transformation, or more specifically, a predetermined invertible affine linear transformation, as follows:
[0432] – If mipSizeId is equal to 2, the following applies:
[0433] p[x] = pTemp[x+1] - pTemp[0]
[0434] – Otherwise (mipSizeId is less than 2), the following applies:
[0435] p[0] = (1<<(BitDepth-1)) - pTemp[0]
[0436] p[x] = pTemp[x] - pTemp[0] for x = 1..inSize-1
[0437] Here, the variable mipSizeId indicates the size of the predetermined block. That is, according to this embodiment, the invertible transformation used to derive another vector from the sample value vector depends on the size of the predetermined block. The dependency can be given as follows:
[0438] mipSizeId boundarySize predSize 0 2 4 1 4 4 2 4 8
[0439] where predSize indicates the number of predicted samples in a predetermined block, and inSize = (2*boundarySize)-(mipSizeId == 2)? 1 : 0, 2*bondarySize indicates the size of the sample value vector and is related to inSize (i.e., the size of another vector). More precisely, inSize indicates the number of those components of the other vector that actually participate in the calculation. For a smaller block size, inSize is as large as the size of the sample value vector, and for a larger block size, inSize is one component smaller. In the former case, one component can be ignored, i.e., the component corresponding to a predetermined component of the other vector, as in the matrix-vector product to be calculated later, the contribution of the corresponding vector component will anyway result in zero, and thus, no actual calculation is required. In the case of an alternative embodiment, the dependence on the block size can be ignored, in which alternative embodiment only one of two alternatives is inevitably used, i.e., independent of the block size (the option corresponding to mipSizeId is less than 2, or the option corresponding to mipSizeId is equal to 2).
[0440] In other words, the predetermined invertible linear transformation is defined such that a predetermined component of another vector p becomes a, while all other components correspond to the components of the sample value vector minus a, where for example a = pTemp[0]. In the case where the first option corresponds to mipSizeId equal to 2, this is easily seen, and further consider only the components of the other vector formed in a different way. That is, in the case of the first option, the other vector is actually {p[0…inSize]; pTemp[0]}, where pTemp[0] is a, and the actual calculation part (i.e., the result of the multiplication) of the matrix-vector multiplication that produces the matrix-vector product is limited to the inSize components of the other vector and the corresponding columns of the matrix, since the matrix has zero columns that do not need to be calculated. In the other cases corresponding to mipSizeId less than 2, a = pTemp[0] is selected as all components of the other vector except p[0] (i.e., each of the other components p[x] (where x = 1..inSize - 1) of the other vector p except the predetermined component p[0]) is equal to the corresponding component of the sample value vector pTemp[x] minus a, but p[0] is selected as a constant minus a. Then the matrix-vector product is calculated. The constant is the mean of representable values, i.e., 2 x-1(i.e., 1 << (BitDepth - 1)), where x represents the bit depth of the computational representation used. It should be noted that if p[0] is selected as pTemp[0], the computed product will simply deviate from the product computed using p[0] as indicated above (p[0] = (1 << (BitDepth - 1)) - pTemp[0]) by a constant vector, which can be considered when predicting the internal block based on the product, i.e., the prediction vector. The value a is thus a predetermined value, e.g., pTemp[0]. The predetermined value pTemp[0] is, for example, the component in the sample value vector pTemp corresponding to the predetermined component p[0] in this case. It can be an adjacent sample at the top or left side of the predetermined block, closest to the upper left corner of the predetermined block.
[0441] For the intra-sample prediction process according to predModeIntra (e.g., specifying the intra prediction mode), the apparatus is configured to apply, for example, the following steps, and at least perform the first step:
[0442] 1. The intra prediction samples based on matrix predMip[x][y] (where x = 0..predSize – 1, y = 0..predSize - 1) are derived as follows:
[0443] – Set the variable modeId to be equal to predModeIntra.
[0444] – Derive the weighted matrix mWeight[x][y] (where x = 0..inSize - 1, y = 0..predSize * predSize - 1) by calling the MIP weighted matrix derivation process with mipSizeId and modeId as inputs.
[0445] – The intra prediction samples based on matrix predMip[x][y] (where x = 0..predSize - 1, y = 0..predSize - 1) are derived as follows:
[0446]
[0447]
[0448] In other words, the apparatus is configured to compute a matrix-vector product between another vector p[i] or, in the case where mipSizeId is equal to 2, {p[i]; pTemp[0]} and a predetermined prediction matrix mWeight or, in the case where mipSizeId is less than 2, a prediction matrix mWeight having additional zero-weight lines corresponding to the omitted components of p, in order to obtain a prediction vector, which has been assigned here to an array of block positions {x, y} distributed inside a predetermined block in order to produce an array predMip[x][y]. The prediction vector will respectively correspond to the concatenation of the rows or the columns of predMip[x][y].
[0449] According to an embodiment, or according to a different interpretation, only the component is understood as the prediction vector, and the apparatus is configured to: when predicting the samples of a predetermined block based on the prediction vector, for each component of the prediction vector, compute the sum of the corresponding component and a (e.g., pTemp[0]).
[0450] The apparatus may optionally be configured to: when predicting the samples of a predetermined block based on a prediction vector such as predMip or perform the following additional steps.
[0451] 2. Clip the in-frame prediction samples predMip[x][y] (where x = 0..predSize-1, y = 0..predSize-1) based on a matrix as follows:
[0452] predMip[x][y] = Clip1(predMip[x][y])
[0453] 3. When isTransposed is equal to true, transpose the predSize×predSize array predMip[x][y] (where x = 0..predSize-1, y = 0..predSize-1) as follows:
[0454] predTemp[y][x] = predMip[x][y]
[0455] predMip = predTemp
[0456] 4. Derive the predicted samples predSamples[x][y] (where x = 0..nTbW-1, y = 0..nTbH-1) as follows:
[0457] – If the specified transform block width nTbW is greater than predSize or the specified transform block height nTbH is greater than predSize, then the input block size predSize, matrix-based intra prediction samples predMip[x][y] (where x = 0..predSize-1, y = 0..predSize-1), transform block width nTbW, transform block height nTbH, top reference samples refT[x] (where x = 0..nTbW-1), and left reference samples refL[y] (where y = 0..nTbH-1) are used as inputs to call the MIP prediction upsampling process, and the output is the predicted sample array predSamples.
[0458] – Otherwise, set predSamples[x][y] (where x = 0..nTbW-1, y = 0..nTbH-1) to be equal to predMip[x][y].
[0459] In other words, the apparatus is configured to predict the samples predSamples of a predetermined block based on the prediction vector predMip.
[0460] 8. Use block / matrix-based intra prediction mode and other intra prediction modes
[0461] The following description presents again the possibility of combining block / matrix-based prediction with other intra prediction modes. It represents another presentation of the possibility, based on which the embodiments described in the subsequent parts can be embodied.
[0462] Note that hereinafter, the term block-based intra prediction is used to denote an intra prediction mode that can be embodied by or is equivalent to the above ALWIP.
[0463] Thus, hereinafter regarding Figure 12The described embodiments relate to decoders and encoders that support intra prediction for decoding / encoding a predetermined block 18, where different intra prediction modes are supported. There is an angular intra prediction mode 500, according to which reference samples 17 adjacent to the predetermined block 18 are used to fill the predetermined block 18 in order to obtain an intra prediction signal for the predetermined block 18. Specifically, the reference samples 17 arranged along the boundary of the predetermined block 18 (e.g., along the upper edge and the left edge of the predetermined block 18) represent picture content that is extrapolated or copied into the interior of the predetermined block 18 along a predetermined direction 502. Before extrapolation or copying, the picture content represented by the adjacent samples 17 can be subjected to interpolation filtering, or in other words, the picture content represented by the adjacent samples 17 can be derived from the adjacent samples 17 by means of interpolation filtering. The angular intra prediction modes 500 differ from each other in the intra prediction direction 502. Each angular intra prediction mode 500 can have an associated index, where the association of the index with the angular intra prediction mode 500 can be such that: when the angular intra prediction modes 500 are sorted according to the associated mode index, the direction 500 rotates monotonically clockwise or counterclockwise.
[0464] There can also be non-angular intra prediction modes. For example, in Figure 12 504 shows a planar intra prediction mode that is optionally included in the set 508. According to this planar intra prediction mode, a two-dimensional linear function defined by a horizontal slope, a vertical slope, and an offset is derived based on the adjacent samples 17, and the predicted sample values of the predetermined block 18 are defined by this linear function. The horizontal slope, the vertical slope, and the offset are derived based on the adjacent samples 17. According to an embodiment, the first set 508 of intra prediction modes includes the planar intra prediction mode 504.
[0465] 506 shows a specific non-angular intra prediction mode included in the set 508, the DC mode. Here, a value, i.e., a quasi-DC value, is derived based on the adjacent samples 17, and this one DC value is attributed to all samples of the predetermined block 18 in order to obtain an intra prediction signal. Although two examples of non-intra prediction modes are shown, there can be only one or more than two examples.
[0466] The intra prediction modes 500, 504, and 506 form a set 508 of intra prediction modes supported by the encoder and decoder, which compete with the block-based intra prediction modes (the example above uses the abbreviation ALWIP) typically indicated by reference numeral 510 in the sense of rate / distortion optimization. As described above, according to these block-based intra prediction modes 510, a matrix-vector product 520 is performed between a vector 514 derived from neighboring samples 17 on one hand and a predetermined prediction matrix 516 on the other hand. The result of the multiplication 520 is a prediction vector 518 for predicting the samples of a predetermined block 18. The block-based intra prediction modes 510 differ from each other in the prediction matrix 516 associated with the respective mode.
[0467] Thus, in short, the encoder and decoder according to the embodiments described herein include a set 508 of intra prediction modes (i.e., a first set of intra prediction modes) and a set 520 of block-based intra prediction modes (i.e., a second set of matrix-based intra prediction modes), and these two sets compete with each other.
[0468] According to an embodiment of the present application, intra prediction is used to encode / decode a predetermined block 18 in the following manner. Specifically, first, a set selection syntax element 522 selects whether to use any mode in the set 508 of intra prediction modes or any mode in the set 520 of block-based intra prediction modes to predict the predetermined block 18. If the set selection syntax element indicates to use any mode in the set 508 (i.e., the first set of intra prediction modes) to predict the predetermined block 18, then based on the intra prediction modes used for the neighboring blocks adjacent to block 18 that have been predicted (exemplarily indicated at 524 and 526), a list 528 of the most likely candidates in the set 508 is constructed / formed at the decoder and encoder. The neighboring blocks 524 and 526 can be determined in a predetermined manner relative to the position of the predetermined block 18, for example, by determining those neighboring blocks that cover certain neighboring samples of block 18 (e.g., the sample above the sample in the upper left corner of block 18), and the block 526 that contains the sample to the left of the just-mentioned corner sample. Of course, this is only an example. This also applies to the number of neighboring blocks used for mode prediction, which is not limited to two for all embodiments. More than two or only one can be used. If any of these blocks 524 and 526 is missing, the default intra prediction mode can be used by default as an alternative to the intra prediction mode of the missing neighboring block. This also applies if any of the blocks 524 and 526 has been encoded / decoded using an inter prediction mode (e.g., prediction by motion compensation).
[0469] The construction of the list of modes in set 508 (i.e., the list 528 of the most likely intra prediction modes) is as follows. The list length of list 528, i.e., the number of the most likely modes, can be fixed by default. This length can be 4 as shown in Figure 12 or can be different from it, for example, 5 or 6. The latter case applies to the specific examples described below. The index in the data stream to be described later can indicate one mode in list 528 to be used for the predetermined block 18. Indexing is performed along the list order or sorting 530, where the list index is, for example, variable length coded such that the length of the index increases monotonically along the order 530. Therefore, it is worth first filling list 528 only with the most likely modes in set 508 and placing the modes that are more likely to be suitable for block 18 upstream of the modes with lower likelihood along the order 530. The modes in list 528 are derived based on the modes used for blocks 524 and 526 (i.e., the neighboring blocks adjacent to the predetermined block 18). If any of blocks 524 and 526 has been intra predicted using the block-based mode 510 in set 520, the aforementioned mapping from such "ALWIP" or block-based mode 510 to the modes within set 508 (the non-ALWIP modes we mentioned) is used. The latter mapping can, for example, map most (i.e., more than half) of the block-based mode 510 to the DC mode 506 (or either the DC 506 or the planar mode 504).
[0470] According to an embodiment, the list 528 of the most likely intra prediction modes is filled with the planar intra prediction mode 504 in a manner independent of the intra prediction mode used when predicting neighboring blocks. Thus, for example, depending on the intra prediction mode used for predicting neighboring blocks 524 and 526, only the DC intra prediction mode 506 and the angular intra prediction mode 500 are filled in list 528. For example, independent of the intra prediction mode used when predicting neighboring blocks 524 and 526, the planar intra prediction mode 504 is located at the first position of the list 528 of the most likely intra prediction modes.
[0471] In the manner exemplarily shown in more detail below, the list construction of the list 528 of most probable intra prediction modes proceeds as follows: such that if the adjacent blocks 524 and 526 have already been pre - predicted by any angular intra prediction mode 500, then there is no DC intra prediction mode 506 in the list 528. If one of the adjacent blocks 524 or 526 is predicted by any angular intra prediction mode 500 and / or if both of the adjacent blocks 524 and the adjacent block 526 are predicted by any angular intra prediction mode 500, then the DC intra prediction mode 506 is not in the list 528 of most probable intra prediction modes. According to the embodiments set forth below in this document, for example, the list 528 is filled with the DC mode 506 only when the following is true for all adjacent blocks 524 and 526: the adjacent blocks 524 and 526 have been encoded using either of the non - angular intra prediction modes 504 and 506, or the adjacent blocks 524 and 526 have been intra - predicted using any block - based intra prediction mode 510, which is mapped to either of the non - angular intra prediction modes 504 and 506 by the aforementioned mapping from the block - based intra prediction mode 510 to a mode within the set 508. Only in such a case, the DC intra prediction mode 506 is located in the list 528. In such a case, it can be positioned in the order 530 before any angular intra prediction mode 500, as can be seen from subsequent examples.
[0472] In other words, the list 528 of most probable intra prediction modes is filled with the DC intra prediction mode 506 only when, for each of the adjacent blocks 524 and 526, the corresponding adjacent block is predicted using at least one of the non - angular intra prediction modes 504 and 506 within a first set 508 (which includes the DC intra prediction mode 506), or using any of the block - based intra prediction modes 510, which is mapped to any of the at least one non - angular intra prediction modes 500 by a mapping from a second set 520 of block - based intra prediction modes 510 to the intra prediction modes within the first set 508 (which is used to form the list 528 of most probable intra prediction modes). The DC intra prediction mode 506 is, for example, located before any angular intra prediction mode 500 in the list 528 of most probable intra prediction modes.
[0473] Accordingly, continuing the description of how to encode the predetermined block 18 into the data stream 12, if the set selection syntax element 522 indicates that the predetermined block 18 is to be encoded by any mode in the first set 508, the data stream 12 optionally includes an MPM syntax element 532 indicating whether the intra prediction mode to be used for the predetermined block 18 is within the list 528. If "yes", the data stream 12 includes an MPM list index 534 pointing to the list 528. The MPM list index 534 indicates, by indexing the list 528 along the order 530, the mode in the list 528 to be used for the predetermined block 18, i.e., the predetermined intra prediction mode. However, if the mode in the set 508 is not within the list 528, as indicated by the MPM syntax element 532, the data stream 12 includes another syntax element 536 for the block 18 to indicate which mode in the set 508 is to be used for the block 18, i.e., the predetermined intra prediction mode. The another syntax element 536 can indicate the mode in a way that differentiates only between those modes in the set 508 that are not included in the list 528.
[0474] In other words, the apparatus for decoding the predetermined block 18 is configured, for example, to derive the MPM syntax element 532 from the data stream. If the set selection syntax element 522 indicates that a mode in the first set 508 of intra prediction modes is to be used to predict the predetermined block 18, the MPM syntax element 532 indicates whether the predetermined intra prediction mode in the first set 508 of intra prediction modes is within the list 528 of the most probable intra prediction modes. If the MPM syntax element 532 indicates that the predetermined intra prediction mode in the first set 508 of intra prediction modes is within the list 528 of the most probable intra prediction modes, the apparatus is configured, for example, to perform the formation of the list 528 of the most probable intra prediction modes based on the intra prediction modes used when predicting the neighboring blocks 524, 526 adjacent to the predetermined block 100, and to perform the derivation of the MPM list index 534 from the data stream 12. The MPM list index 534 points to the predetermined intra prediction mode in the list 528 of the most probable intra prediction modes. If the MPM syntax element 532 from the data stream 12 indicates that the predetermined intra prediction mode in the first set 508 of intra prediction modes is not within the list 528 of the most probable intra prediction modes, the apparatus is configured to derive another list index 536 from the data stream. The another list index 536 indicates the predetermined intra prediction mode in the first set of intra prediction modes. Accordingly, based on the MPM syntax element 532, the data stream 12 includes either the MPM list index 534 or another list index 536 for the prediction of the predetermined block 18.
[0475] By removing the cases in which the list 528 includes the DC intra prediction mode 506, the following advantages are achieved. Specifically, the inventors of the present application have found that since the DC intra prediction mode 506 in the set 508 competes with the block-based intra prediction mode 510 anyway, using the DC intra prediction mode 506 in the set 508 to encode / decode the predetermined block 18 "consumes" valuable list positions in the list 528, which will have a negative impact on the encoding efficiency. Any mode indicated by the syntax element 522 (i.e., the set selection syntax element) in the intra prediction modes of the set 508 should be used to encode / decode the predetermined block 18. Therefore, "consuming" the list positions in the list 528 with such a DC intra prediction mode 506 in the set 508 will increase the likelihood of the following situation: the intra prediction mode ultimately to be used for the predetermined block 18 (i.e., the predetermined intra prediction mode) is not in the list 528, such that the syntax element 536 (i.e., another list index) needs to be transmitted in the data stream 12.
[0476] Specifically, since the syntax element 522 has already indicated for the block 18 whether any mode within the set 508 or any mode among the block-based modes 510 in the set 520 should be used to predict the block 18, it seems that if the syntax element 522 indicates that the mode within the set 508 is preferred for the block 18 and thus the block-based mode 510 is not used for the block 18, the likelihood that the DC prediction mode 506 in the set 508 can be applied to the block 18 is relatively low, such that its occurrence in the list 528 should be limited to a very limited set of cascades of the modes for the adjacent blocks 524 and 526, i.e., the cascades described above.
[0477] In other cases, i.e., in cases where the set selection syntax element 522 indicates that any one of the block-based intra prediction modes 510 is to be used to predict the predetermined block 18, encoding the block 18 into the data stream 12 and decoding from the data stream can be performed in the manner described above. For this purpose, an index can be used to index the selected one of the block-based intra prediction modes 510 to be used in the set 520 (i.e., the second set of block-based intra prediction modes), or to indicate which one of the block-based intra prediction modes in the set 520 is to be used. Another MPM syntax element 538 can indicate whether the indexing is done via the index 540 (i.e., via another MPM list index), the index 540 indicating the block-based intra prediction mode 510 in the list 542 of the most likely block-based intra prediction modes 510 to be used for the block 18, i.e., by indexing along the list order 544, or whether the block-based intra prediction mode 510 to be used for the block 18 is indicated by another syntax element 546 (i.e., by yet another list index), the latter syntax element 546 being able to distinguish, for example, only between those modes 510 within the set 520 that are not already included in the list 542. The list construction of the list 542 can be performed based on the modes used to predict the blocks 524 and 526. If either of the blocks 524 and 526 is not available because it is outside the picture or because inter prediction is in progress, a default intra prediction mode, such as one of the intra prediction modes in the set 508, can be used instead. For each of the blocks 524 and 526, in cases where intra prediction has been performed using the modes in the set 508 instead of the set 520, the aforementioned mapping from the modes in the set 508 to the modes in the set 520 is used to obtain the intra prediction mode 510 for the corresponding block (i.e., the predetermined block 18), i.e., the predetermined block-based intra prediction mode, and the list 542 is constructed based on the block-based intra prediction modes obtained for the blocks 524 and 526.
[0478] According to an embodiment, a device for decoding a predetermined block 18 is configured to derive from a data stream 12 another MPM syntax element 532, the another MPM syntax element 532 indicating whether, if a set selection syntax element 522 indicates that the predetermined block 18 will not be predicted using one of the modes in a first set 508 of intra prediction modes, a predetermined block-based intra prediction mode in a second set 520 of block-based intra prediction modes 510 is within a list 542 of most likely block-based intra prediction modes. If the another MPM syntax element 538 indicates that the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes 510 is within the list 542 of most likely block-based intra prediction modes, the device is configured, for example, to form the list 542 of most likely block-based intra prediction modes based on the intra prediction modes used when predicting using adjacent blocks 524, 526 adjacent to the predetermined block 18, and to derive from the data stream 12 another MPM list index 540 that points to the predetermined block-based intra prediction mode in the list 542 of most likely block-based intra prediction modes. If the another MPM syntax element 538 indicates that the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes is not within the list 542 of most likely block-based intra prediction modes, the device is configured to derive from the data stream 12 yet another list index 546 that indicates the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes. Thus, based on the another MPM syntax element 538, the data stream 12 includes either the another MPM list index 540 or the yet another list index 546 for the prediction of the predetermined block 18.
[0479] Although the another MPM syntax element 538, the another MPM list index 540, and the yet another list index 546 are represented in the data stream 12 in Figure 12 as being parallel to the MPM syntax element 532, the MPM list index 534, and the another list index 536, it is clear that the data stream 12 includes either the another MPM syntax element 538 and an index associated with the another MPM syntax element (e.g., the another MPM list index 540 or the yet another list index 546), or includes the MPM syntax element 532 and an index associated with the MPM syntax element (e.g., the MPM list index 534 or the another list index 536). Which of the syntax element and the index the data stream 12 includes depends, for example, on the set selection syntax element 522.
[0480] An example of writing the syntax element portion of the data stream 12 in pseudocode can be as Figures 19a to 19d shown, where the reference numerals indicate which syntax elements correspond to the syntax elements discussed above.
[0481] The list construction of list 528 can be defined as follows: where candIntraPredModeA / B indicates the intra prediction mode used for prediction of either block 524 or block 526, for example A for block 524 and B for block 526, or indicates which mode in set 508 the intra prediction mode maps to in the case where the corresponding block 524 or 526 has been intra predicted using any of the block-based intra prediction modes 510. INTRA_DC is used to indicate mode 506, and the angular mode 500 is indicated by INTRA_ANGULAR# where the number (#) sorts the angular modes as exemplified above (i.e., in a manner such that the angular direction 502 decreases or increases monotonically with increasing number). The sorting between the modes within set 508 can be defined as in the subsequent table, where the INTRA_PLANAR table indicates mode 504.
[0482] Note that in the above example, the index 534 is actually distributed over the syntax elements 534' and 534": the former 534' is specific to the first position in list 528 in order 530, where according to the example, the INTRA_PLANAR mode 504 is inevitably located at this first position. The latter 534" points to any of the subsequent positions in list 528, where as described, the DC mode 506 is only included in the specific cases described.
[0483] In addition, in the above example, in the case where the syntax element 522 indicates the use of any of the patterns within the set 508, the data stream includes some other syntax elements that parameterize, in some way, the intra prediction modes within the set 500. For example, the syntax element 600 parameterizes or alters the region in which the reference sample 17 is located, based on which the patterns in the set 508 perform intra prediction of the interior of the block 18, for example, in terms of the distance from the outer periphery of the block 18. Additionally or alternatively, the syntax element 602 parameterizes or alters whether the patterns in the set 508 use the reference sample 17 to perform intra prediction of the interior of the block 18 globally or block - by - block, or whether to perform intra prediction by slices or portions into which the block 18 is subdivided, and the slice or portion performs intra prediction in turn, such that the prediction residuals encoded into the data stream for one portion can be used to supplement new reference samples for intra prediction of subsequent portions. The latter - encoded portion controlled by the syntax element is available (and the corresponding syntax element may be present in the data stream) only if the syntax element 600 has a predetermined state corresponding, for example, to the region in which the reference sample 17 is located being adjacent to the block 18. This portion can be defined by subdividing the block along a predetermined direction, for example, horizontally, resulting in a portion as high as the block 18, or vertically, resulting in a portion as wide as the block 18. If the activation of the partitioning is signaled, the syntax element 604 may be present in the data stream, which controls which partitioning direction to use. It can be seen that the positions reserved for the INTRA_PLANAR mode in the list 528 may be available only in the case where the mode is parameterized in some way by the parameterizing syntax elements just mentioned, for example, only if the syntax element 600 has a predetermined state corresponding to, for example, the region in which the reference sample 17 is located being adjacent to the block 18, and / or the per - portion intra prediction mode is not signaled as active by the syntax element 602.
[0484] All syntax elements shown in the table and not specifically mentioned above are optional and are not further discussed herein.
[0485] – If candIntraPredModeB is equal to candIntraPredModeA and candIntraPredModeA is greater than INTRA_DC, then candModeList[x], x = 0..4 is derived as follows:
[0486] candModeList[0] = candIntraPredModeA
[0487] candModeList[1] = 2 + ((candIntraPredModeA + 61) % 64)
[0488] candModeList[2] = 2 + ((candIntraPredModeA - 1) % 64)
[0489] candModeList[3] = 2 + ((candIntraPredModeA + 60) % 64)
[0490] candModeList[4] = 2 + (candIntraPredModeA % 64)
[0491] - Otherwise, if candIntraPredModeB is not equal to candIntraPredModeA and candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC, the following applies:
[0492] - The variables minAB and maxAB are derived as follows:
[0493] minAB = Min(candIntraPredModeA, candIntraPredModeB)
[0494] maxAB = Max(candIntraPredModeA, candIntraPredModeB)
[0495] - If both candIntraPredModeA and candIntraPredModeB are greater than INTRA_DC, candModeList[x], x = 0..4 are derived as follows:
[0496] candModeList[0] = candIntraPredModeA
[0497] candModeList[1] = candIntraPredModeB
[0498] - If maxAB - minAB equals 1, the following applies:
[0499] candModeList[2] = 2 + ((minAB + 61) % 64)
[0500] candModeList[3] = 2 + ((maxAB - 1) % 64)
[0501] candModeList[4] = 2 + ((minAB + 60) % 64)
[0502] – Otherwise, if maxAB – minAB is greater than or equal to 62, then the following applies:
[0503] candModeList[2] = 2 + ((minAB - 1) % 64)
[0504] candModeList[3] = 2 + ((maxAB + 61) % 64)
[0505] candModeList[4] = 2 + (minAB % 64)
[0506] – Otherwise, if maxAB – minAB equals 2, then the following applies:
[0507] candModeList[2] = 2 + ((minAB - 1) % 64)
[0508] candModeList[3] = 2 + ((minAB + 61) % 64)
[0509] candModeList[4] = 2 + ((maxAB - 1) % 64)
[0510] – Otherwise, the following applies:
[0511] candModeList[2] = 2 + ((minAB + 61) % 64)
[0512] candModeList[3] = 2 + ((minAB - 1) % 64)(8 - 36)
[0513] candModeList[4] = 2 + (((maxAB + 61)) % 64)
[0514] – Otherwise (candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC), then candModeList[x], x = 0..4 is derived as follows:
[0515] candModeList[0] = maxAB
[0516] candModeList[1] = 2 + ((maxAB + 61) % 64)
[0517] candModeList[2] = 2 + ((maxAB - 1) % 64)(8 - 41)
[0518] candModeList[3] = 2 + ((maxAB + 60) % 64)
[0519] candModeList[4] = 2 + (maxAB % 64)
[0520] – Otherwise, the following applies:
[0521] candModeList[0] = INTRA_DC
[0522] candModeList[1] = INTRA_ANGULAR50(
[0523] candModeList[2] = INTRA_ANGULAR18
[0524] candModeList[3] = INTRA_ANGULAR46
[0525] candModeList[4] = INTRA_ANGULAR54
[0526] Intra prediction mode Associated name 0 INTRA_PLANAR 1 INTRA_DC 2..66 INTRA_ANGULAR2..INTRA_ANGULAR66
[0527] 9. Embodiments using block / matrix-based intra prediction modes and other intra prediction modes and using quadratic transformation
[0528] The following description presents embodiments for combining block / matrix-based prediction with other intra prediction modes and encoding prediction residuals using quadratic transformation. The above presentation of matrix-based intra prediction (ALWIP) and its possibility of combination with other intra prediction modes is to be used as an example for implementing the embodiments described below. For example, in Figure 12 all details regarding the construction of the MPM list that includes the DC mode restriction are optional.
[0529] As described above, matrix-based intra prediction (MIP), also referred to herein as block-based intra prediction and ALWIP, generates an intra prediction signal on a rectangular block by performing matrix-vector multiplication, where the output of the matrix-vector multiplication can be regarded as the prediction signal for the downsampled block, and where the downsampled boundary samples can include the input to the matrix-vector multiplication. If the output is regarded as the prediction signal for the downsampled block, then this prediction signal needs to go through an upsampling (or linear interpolation) stage before the final prediction signal is obtained.
[0530] On the other hand, for traditional intra prediction modes such as planar mode 504, DC mode 506, and angular mode 500 (also represented as modes in set 508 in the above description), the non-separable quadratic transform (LFNST) is a tool for transforming the prediction residuals corresponding to these intra prediction modes. Here, a set S of transform sets is given such that each traditional intra prediction mode is associated with one of these transform sets. Then, at the decoder, it can be extracted from the bitstream whether to apply LFNST to a given block. If this is the case, depending on the intra prediction mode used on the current block, one of the transform sets in set S is given, and if the transform set consists of more than one transform, it can be extracted from the bitstream which transform T in the set is to be used. Then, at the decoder, transform T is applied as a quadratic transform Ts, which means it is applied to a subset 622 of the residual transform coefficients 620 of the separable primary transform Tp, for example, as Figure 14 shown.
[0531] The problem is that the above quadratic transform Ts is defined a priori only for traditional intra prediction modes. Providing a specific quadratic transform Ts for each MIP mode 510 may be too costly in terms of the memory requirements for storing additional transforms.
[0532] Figure 13 A decoder that solves this problem is shown. The decoder decodes a predetermined block 18 of a picture using intra prediction. According to an embodiment, the encoder includes features and / or functions parallel to the decoder.
[0533] The decoder / encoder is configured to select 602 a predetermined intra prediction mode 604 from a plurality of intra prediction modes 600, the plurality of intra prediction modes 600 including a first set 508 of intra prediction modes and a second set 520 of matrix-based intra prediction modes 510. This intra mode selection 602 is performed by the decoder based on the data stream 12, where the encoder is configured to signal the predetermined intra prediction mode 604 in the data stream 12. The intra mode selection 602 can be performed as described with respect to Figure 12 described.
[0534] The first set 508 of intra prediction modes includes a DC intra prediction mode 506, an angular prediction mode 500, and an optional planar intra prediction mode 504. In the case where a predetermined intra prediction mode 604 is a matrix-based intra prediction mode 510 in a second set 520, the decoder / encoder is configured to obtain a prediction vector 518 using a matrix-vector product 512 between: a vector 514 derived from reference samples 17 in a neighborhood of a predetermined block 18, and a prediction matrix 516 associated with the corresponding matrix-based intra prediction mode 510, and to predict samples of the predetermined block 18 based on the prediction vector 518. Using the matrix-based intra prediction mode 510 as the predetermined intra prediction mode 604 to predict the predetermined block 18 may be performed by a decoder / encoder according to an Figures 6 to 11 embodiment. The decoder / encoder is configured to derive a prediction signal 606 for the predetermined block 18 using the predetermined intra prediction mode 604.
[0535] The decoder / encoder is configured to: select 608 one or more quadratic transforms T (1) –T (N) from a set 612 of quadratic transforms Ts (e.g., Ts s (i1) -T s (in) (where i1 ranges from 1 to N and i n ranges from i1 to N) of a subset 610 in a manner depending on the predetermined intra prediction mode 604 such that the subset 610 is non-empty in the case where the predetermined intra prediction mode 604 is included in the first set 508 of intra prediction modes and in the case where the predetermined intra prediction mode 604 is included in the second set 520 of matrix-based intra prediction modes.
[0536] According to an embodiment, the decoder / encoder is configured to: select 608 the subset 610 such that each quadratic transform Ts in the set 612 of quadratic transforms Ts is included in a subset 610 of one or more quadratic transforms selected for at least one intra prediction mode within the first set 508 and the second set 520 of intra prediction modes. Thus, the subset 610 may be equal to the set 612 of quadratic transforms. Such a subset may be selected for one or more matrix-based intra prediction modes 510. For one or more matrix-based intra prediction modes 510, it may be possible to select all quadratic transforms Ts in the set 612 of quadratic transforms, whereby the subset 610 of the one or more matrix-based intra prediction modes 510 includes quadratic transforms Ts that may be selected for intra prediction modes within the first set 508 of intra prediction modes. Thus, for at least one of the matrix-based intra prediction modes 510, no specific additional quadratic transforms are required in the set 612 of quadratic transforms.
[0537] According to an embodiment, the decoder / encoder is configured to select the 608 subset 610 in such a way that each subset 610 of the secondary transform Ts selected for any matrix-based intra prediction mode 510 matrix contains the subset of the secondary transform Ts selected for at least one intra prediction mode within the first set 508 that does not belong to the angular prediction mode 500, e.g., 610 DC and / or 610 planar . Figure 15 Indicates different possible subsets of the secondary transform Ts selected for any matrix-based intra prediction mode 510. The decoder / encoder can be configured to: for each matrix-based intra prediction mode 510, select the first union 611 of the subsets 610 of the secondary transform for the matrix-based intra prediction mode 510 matrix of the subset 610.
[0538] As Figure 15 shown, the subset 610 that can be selected for one or more matrix-based intra prediction modes 510 can be equal to the subset, e.g., 610 DC1 or 610 DC2 , that can be selected for the DC intra prediction mode 506, or can be equal to the subset, e.g., 610 planar1 or 610p lanar2 , that can be selected for the planar intra prediction mode 504. The subset 610 that can be selected for one or more matrix-based intra prediction modes 510 may only include one or more secondary transforms (e.g., Ts DC ) of the subset of the secondary transform Ts that can be selected for only one intra prediction mode within the first set 508 that does not belong to the angular prediction mode 500 (e.g., 610 (ax) to Ts (ay) ), as indicated by the subset 610 matrix4 .
[0539] The subset 610 that can be selected for the matrix-based intra prediction mode 510 can contain one or more secondary transforms of two or more subsets (e.g., 610 DC1 and 610 DC2 ) that can be selected for the DC intra prediction mode 506, as indicated by the subset 610 matrix1 , or can contain one or more secondary transforms of two or more subsets (e.g., 610 planar1 and 610 planar2 ) that can be selected for the planar intra prediction mode 504, as indicated by the subset 610 matrix3 .
[0540] Another possible subset 610 that can be selected for the matrix-based intra prediction mode 510 can include one or more subsets that can be selected for the DC intra prediction mode 506 (e.g., one or more among 610 DC2 ) of the secondary transforms, and one or more subsets that can be selected for the planar intra prediction mode 504 (e.g., one or more among 610 planar1 ) of the secondary transforms, as indicated by the subset 610 matrix2 .
[0541] According to an embodiment, the decoder / encoder is configured to select 608 the subset 610 in such a way that a first union 611 of the subset 610 of the secondary transforms selected for the matrix-based intra prediction mode 510 matrix and a second union 611 of the subset 610 of the secondary transforms selected for all the angular intra prediction modes angular has an empty intersection. A third union 611 angular includes all the subsets 610 DC selected for the DC intra prediction mode 506, and a fourth union 611 DC includes all the subsets 610 planar selected for the planar intra prediction mode 504. For example, since downsampling is optionally applied to reduce the prediction signal 606, the MIP mode 510 is as non-directional as the planar mode 504 and the DC mode 506, and thus their prediction residuals 618 have more statistical similarity with the DC mode 506 and the planar mode 504 compared to the angular mode 500. planar
[0542] In addition, as Figure 13 shown, the decoder is configured to derive 614 from the data stream a transformed version 616 of the prediction residual of a predetermined block 18 encoded by the encoder into the data stream, the transformed version 616 of the prediction residual of the predetermined block 18 being related to the spatial domain version of the prediction residual of the predetermined block 18 via a transform T, and the transform T being defined by a cascade of a primary transform Tp and a predetermined secondary transform Ts in the subset 610 of the secondary transforms. As Figure 14 shown, in the case where the predetermined intra prediction mode is included in a first set 508 of the intra prediction modes, and in the case where the predetermined intra prediction mode is included in a second set 520 of the matrix-based intra prediction modes 510, the encoder can be configured to apply the transform T to a subset 622 of the coefficients 620 of the primary transform Tp. The decoder can be configured to use the inverse T -1 of the transform T to obtain the spatial domain version 618 of the prediction residual of the predetermined block 18. The primary transform is, for example, a separable 2D transform, and the secondary transform is, for example, a non-separable 2D transform.
[0543] The decoder is configured to reconstruct 624 the predetermined block 18 using the prediction signal 606 and the prediction residual 618 of the predetermined block 18.
[0544] If a subset 610 of one or more secondary transforms Ts contains more than one secondary transform Ts, the decoder may be configured to: depending on the secondary transform indication syntax element transmitted in the data stream 12 for the predetermined block, select a predetermined secondary transform Ts from the subset of one or more secondary transforms 610. The secondary transform indication syntax element may be an index pointing to the selected subset 610 among one or more secondary transforms. In this case, the encoder may be configured to transmit the secondary transform indication syntax element in the data stream 12.
[0545] According to an embodiment, the decoder is configured to: if the size of the predetermined block 18 meets a predetermined criterion, infer that the transform version 616 of the prediction residual of the predetermined block 18 is the primary transform Tp via which it is related to the spatial domain version 618 of the prediction residual of the predetermined block. Otherwise, the transform T may be a concatenation of the primary transform Tp and a predetermined secondary transform Ts. The decoder makes available a second set 520 of 510 matrix-based intra prediction modes for the selection 602 of the predetermined intra prediction mode 604, regardless of whether the size of the predetermined block 18 meets the predetermined criterion. The predetermined criterion for the size of the predetermined block 18 may be related only to the subset selection 608 and not to the intra mode selection 602. Although LFNST (i.e., subset selection 608) is not available, there are block sizes for which, for example, MIP / ALWIP 510 is available. For blocks 18 of such a size, it is not necessary to transmit and read the secondary transform indication syntax element in the data stream. For example, if the size is below a predetermined threshold, the predetermined criterion is met. For small predetermined blocks 18, this may not be beneficial, meaning that for the latter shape, the additional signaling cost of signaling whether to apply LFNST to the block in the MIP mode is on average higher than the gain obtained by allowing the LFNST transform for the MIP mode.
[0546] According to an embodiment, the decoder is configured to read a non-zero region indication transmitted in the data stream 12 for the corresponding predetermined block 18, the non-zero region indication being used to indicate a non-zero transform domain region 623 within the transform version 616 of the prediction residual of the predetermined block 18, as Figure 16As shown. All non-zero coefficients are only located in the non-zero transform domain region 623. For example, the non-zero region indication is transmitted in the data stream 12 by the encoder. The decoder / encoder is configured to decode the coefficients within the non-zero transform domain region 623 from the data stream 12 / encode the coefficients within the non-zero transform domain region 623 into the data stream 12. The LP syntax element can be used as the non-zero region indication. The last non-zero coefficient position along the scan path from the DC coefficient position to the opposite (or highest frequency) coefficient position is indicated by the LP syntax element in this document. The LP syntax element can be used as a measure of the expected count of the non-zero coefficients within the non-zero transform domain region 623.
[0547] According to an embodiment, the decoder is configured to: depending on that the extension and / or position of the non-zero transform domain region 623 meet a first predetermined criterion, and / or the number of non-zero coefficients within the non-zero transform domain region 623 meets a second predetermined criterion, infer that the transform version 616 of the prediction residual of the predetermined block 18 is the main transform Tp via the transform T related to the spatial domain version 618 of the prediction residual of the predetermined block 18.
[0548] For example, the first predetermined criterion is such that if the following conditions are met, the first predetermined criterion is also met: the non-zero transform domain region 623 does not only cover the coefficients of the main transform Tp that are a subset 622 of the coefficients obtained by cascading the application of the secondary transform Tp. This is based on the idea that the secondary transform Ts should cover all non-zero coefficients among the transform coefficients of the main transform Tp. In the case where the non-zero coefficients are outside the subset 622 of the coefficients of the main transform Tp, applying the predetermined secondary transform may not be beneficial, so the decoder infers that the transform T is the main transform Tp. The decoder can perform subset selection 608 and apply the transform T defined by the cascade of the main transform Tp and the predetermined secondary transform Ts applied to the subset 622 of the coefficients of the main transform Tp within the subset 610 of the secondary transform to obtain the spatial domain version 618 of the prediction residual of the predetermined block 18 when the non-zero transform domain region 623 is completely located within the subset 622 of the coefficients of the main transform Tp obtained by cascading the application of the secondary transform Tp, see for example Figure 16 .
[0549] For example, if the number of non-zero coefficients within the non-zero transform domain region 623 is below a predetermined threshold, the second predetermined criterion is met. This is based on the idea that if the number of non-zero coefficients within the non-zero transform domain region 623 is below a predetermined threshold, there is no need to further reduce the number of non-zero coefficients through the secondary transform Ts. In the case where the primary transform Tp has a small number of non-zero coefficients, applying an additional predetermined secondary transform may not be beneficial, and thus the decoder infers that the transform T is the primary transform Tp. In the case where the number of non-zero coefficients is below a predetermined threshold, the additional signaling cost of the predetermined secondary transform will be higher than the increased coding efficiency achieved by the predetermined secondary transform. The decoder may perform subset selection 608 and apply the transform T defined by the cascade of the primary transform Tp and a subset of the coefficients of the primary transform Tp to which the subset 610 of the secondary transform Ts is applied, to obtain a spatial domain version 618 of the prediction residual of the predetermined block 18 when the number of non-zero coefficients within the non-zero transform domain region 623 is equal to or exceeds the predetermined threshold.
[0550] According to an embodiment, the decoder / encoder is configured to derive from the data stream 12 a set selection syntax element 522 / to encode the set selection syntax element 522 into the data stream 12, the set selection syntax element 522 indicating whether to use one of the intra prediction modes in the first set 508 of intra prediction modes to predict the predetermined block 18, e.g., as described regarding Figure 12 If the set selection syntax element 522 indicates to use one of the intra prediction modes in the first set 508 of intra prediction modes to predict the predetermined block 18, the decoder / encoder is configured to form a list 542 of the most likely intra prediction modes based on the intra prediction modes used to predict the neighboring blocks 524, 526 adjacent to the predetermined block 18, and to derive from the data stream 12 an MPM list index 540 pointing to a predetermined intra prediction mode in the list 542 of the most likely intra prediction modes, or to signal the MPM list index 540 pointing to a predetermined intra prediction mode in the list 542 of the most likely intra prediction modes into the data stream 12. If the set selection syntax element 522 indicates not to use one of the intra prediction modes in the first set 508 of intra prediction modes to predict the predetermined block 18, the decoder is configured to derive another index 540 and / or 546 from the data stream, or the encoder is configured to encode another index 540 and / or 546 into the data stream, the another index 540 and / or 546 indicating a predetermined intra prediction mode 604 in the second set 520 of matrix-based intra prediction modes 510.
[0551] Regarding Figure 13 The described decoder and encoder may include other features and / or functions as described regarding Figure 12 described.
[0552] When the prediction mode 604 within a predetermined frame of a predetermined block 18 is a matrix-based intra prediction mode 510 in the second set 520, the apparatus for decoding the predetermined block 18 (i.e., the decoder according to Figure 13 ), and / or the apparatus for encoding the predetermined block 18 (i.e., the encoder according to Figure 13 ) may include one or more of the following features.
[0553] According to an embodiment, the apparatus is configured to form a sample value vector based on a plurality of reference samples 17 (e.g., the sample value vector 400 described with respect to Figure 6 to one of the embodiments of FIG. 9), and derive a vector 514 from the sample value vector such that the sample value vector is mapped onto the vector 514 by a predetermined invertible linear transformation. In this case, the vector 514 can be understood as another vector. For example, the vector 514 is determined and / or defined as described for another vector 402 with respect to Figures 8 to 11 one of the embodiments.
[0554] According to an embodiment, the apparatus is configured to form a sample value vector based on a plurality of reference samples 17 by adopting, for each component of the sample value vector, one of the plurality of reference samples as the corresponding component of the reference sample, and / or by averaging two or more components of the sample value vector to obtain the corresponding component of the sample value vector.
[0555] The plurality of reference samples 17 are arranged, for example, along the outer edge of the predetermined block 18 within the picture.
[0556] The invertible linear transformation is defined, for example, such that a predetermined component of the vector 514 (e.g., of another vector) becomes a, and each of the other components of the vector 514 except the predetermined component is equal to the corresponding component of the sample value vector minus a. The value a is, for example, a predetermined value 1400.
[0557] According to an embodiment, the predetermined value 1400 is one of the following: the average value (e.g., arithmetic mean or weighted average) of the components of the sample value vector, a default value, a value signaled in the data stream into which the picture is encoded, the component corresponding to a preset component in the sample value vector.
[0558] The invertible linear transformation is defined, for example, such that a predetermined component of the vector 514 (e.g., of another vector) becomes a, and each of the other components of the vector 514 except the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the arithmetic mean of the components of the sample value vector.
[0559] A reversible linear transformation is defined, for example, such that a predetermined component of vector 514 (e.g., of another vector) becomes a, and each of the other components of vector 514 other than the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the component of the sample value vector corresponding to the predetermined component. The apparatus is configured, for example, to include a plurality of reversible linear transformations, each of the plurality of reversible linear transformations being associated with one component of vector 514, to select a predetermined component from the components of the sample value vector, and to use the reversible linear transformation associated with the predetermined component among the plurality of reversible linear transformations as the predetermined reversible linear transformation.
[0560] According to an embodiment, matrix components of prediction matrix 516 corresponding to a predetermined component of vector 514 (e.g., of another vector) within a column of prediction matrix 516 are all zero. The apparatus is configured to calculate matrix-vector product 512 between: a reduced prediction matrix generated from prediction matrix 516 by removing a column, and another vector 410 generated from vector 514 by removing the predetermined component, by performing multiplication to calculate matrix-vector product 512, as Figure 9c shown.
[0561] According to an embodiment, the apparatus is configured to, when predicting samples of predetermined block 18 based on prediction vector 518, calculate the sum of the corresponding component and a for each component of prediction vector 518.
[0562] By summing 1 with each matrix element of prediction matrix 516 corresponding to a predetermined component of vector 514 (e.g., of another vector) within a column of prediction matrix 516 (i.e., the sum of matrix C 405 and matrix M 1300, as Figure 9a shown, resulting in Figure 8 matrix B in), which, when multiplied by the reversible linear transformation, corresponds to, for example, a quantized version of a machine learning prediction matrix (i.e., Figure 8 prediction matrix A 1100 shown in).
[0563] According to an embodiment, the apparatus is configured to calculate matrix-vector product 512 using fixed-point arithmetic operations.
[0564] According to an embodiment, the apparatus is configured to calculate matrix-vector product 512 without using floating-point arithmetic operations.
[0565] According to an embodiment, the apparatus is configured to store a fixed-point representation of prediction matrix 516.
[0566] According to an embodiment, the apparatus is configured to: represent a prediction matrix 516 using prediction parameters, and calculate the matrix-vector product 512 by performing multiplication and addition on the components of a vector 514 (e.g., another vector), the prediction parameters, and the resulting intermediate results, where the absolute value of the prediction parameters can be represented by a fixed-point number of n bits, where n is equal to or less than 14, or alternatively equal to or less than 10, or alternatively equal to or less than 8. This can be performed similarly to that described in Figure 10 or Figure 11 or as described in Figure 10 or Figure 11 as described therein.
[0567] The prediction parameters include, for example, weights, each weight being associated with a corresponding matrix component of the prediction matrix 516.
[0568] The prediction parameters further include, for example: one or more scaling factors, each of the one or more scaling factors being associated with one or more corresponding matrix components in the prediction matrix 516 for scaling the weights associated with the one or more corresponding matrix components of the prediction matrix 516; and / or one or more offsets, each of the one or more offsets being associated with one or more corresponding matrix components in the prediction matrix 516 for offsetting the weights associated with the one or more corresponding matrix components of the prediction matrix 516.
[0569] According to an embodiment, the apparatus is configured to: use interpolation to calculate at least one sample position of a predetermined block 18 based on a prediction vector 518 when predicting samples of the predetermined block 18, each component of the prediction vector being associated with a corresponding position within the predetermined block 18.
[0570] Regarding Figure 13 the decoder and encoder described may include other features and / or functions described in one or more embodiments of Figures 6 to 11 as described.
[0571] Accordingly, the solution provided by the present invention is to associate a specific set 160 of the transforms in the given transform set S 612 with each MIP mode 510, where the specific set 160 of the transforms in the given transform set S 612 is initially defined for conventional intra prediction modes (i.e., the intra prediction modes within the first set 508). A specific way of doing this is that all MIP modes 510 use the LFNST transform, i.e., the quadratic transform initially designed for the planar mode 504 and the DC mode 506. For example, since downsampling is optionally applied to the reduced prediction signal, the MIP modes 510 are as non-directional as the planar mode 504 and the DC mode 506, and thus, compared with the angular mode 500, their prediction residuals have more statistical similarity with the DC mode 506 and the planar mode 504.
[0572] A possible outcome is that allowing the LFNST, i.e., allowing the use of the quadratic transform for MIP 510, may be beneficial in terms of coding efficiency for some block shapes in the foregoing manner, while it may not be beneficial for other block shapes, which means that for the latter shapes, the cost of additionally signaling whether to apply the LFNST to block 18 in the MIP mode 510 on average is higher than the gain obtained by allowing the LFNST transform for the MIP mode 510. Accordingly, in an embodiment of the present invention, the foregoing combination of the LFNST and the MIP 510 is only allowed for a subset of all those block shapes for which such a combination is possible in principle.
[0573] Figure 17Disclosed is a method 6000 for decoding a predetermined block (18) of a picture using intra prediction, comprising: selecting (602) a predetermined intra prediction mode (604) from a plurality (600) of intra prediction modes based on a data stream, the plurality (600) of intra prediction modes including a first set (508) of intra prediction modes and a second set (520) of matrix-based intra prediction modes (510), the first set (508) of intra prediction modes including a DC intra prediction mode (506) and angular prediction modes (500) and optionally a planar intra prediction mode (504), for each matrix-based intra prediction mode (510) in the second set (520) of matrix-based intra prediction modes (510), obtaining a prediction vector (518) using a matrix-vector product (512) between: a vector (514) derived from reference samples (17) in a neighborhood of the predetermined block, and a prediction matrix (516) associated with the corresponding matrix-based intra prediction mode, predicting samples of the predetermined block based on the prediction vector (518). Deriving 6100 a prediction signal (606) for the predetermined block using the predetermined intra prediction mode, and selecting (608) one or more subsets (610) of secondary transforms from a set (612) of secondary transforms in a manner depending on the predetermined intra prediction mode, such that the subset (610) is non-empty in the case where the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510). The method 6000 includes: in the case where the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510), deriving (614) a transformed version (616) of the prediction residual of the predetermined block (18) from the data stream, the transformed version (616) of the prediction residual of the predetermined block (18) being related to a spatial-domain version (618) of the prediction residual of the predetermined block (18) via a transform (T), the transform (T) being defined by a cascade of a primary transform (Tp) and a predetermined secondary transform (Ts) applied to a subset (622) of coefficients (620) of the primary transform in a subset (610) of secondary transforms. Additionally, the method 6000 includes reconstructing (624) the predetermined block using the prediction signal and the prediction residual of the predetermined block.
[0574] Figure 18A method 7000 for encoding a predetermined block (18) of a picture using intra prediction is shown, including: selecting (602) a predetermined intra prediction mode (604) from a plurality (600) of intra prediction modes, the plurality (600) of intra prediction modes including a first set (508) of intra prediction modes and a second set (520) of matrix-based intra prediction modes (510), the first set (508) of intra prediction modes including a DC intra prediction mode (506) and angular prediction modes (500), and for each matrix-based intra prediction mode (510) in the second set (520) of matrix-based intra prediction modes (510), obtaining a prediction vector (518) using a matrix-vector product (512) between: a vector (514) derived from reference samples (17) in the neighborhood of the predetermined block (18), and a prediction matrix (516) associated with the corresponding matrix-based intra prediction mode (510), and predicting samples of the predetermined block (18) based on the prediction vector (518). Additionally, method 7000 includes: signaling 7100 the predetermined intra prediction mode (604) in a data stream, deriving 7200 a prediction signal (606) for the predetermined block (18) using the predetermined intra prediction mode, and selecting (608) one or more subsets (610) of secondary transforms from a set (612) of secondary transforms in a manner depending on the predetermined intra prediction mode such that the subset (610) is non-empty in the case where the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510). The method includes: in the case where the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and in the case where the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510), encoding (614) a transformed version (616) of the prediction residual of the predetermined block (18) into the data stream, the transformed version (616) of the prediction residual of the predetermined block (18) being related to a spatial-domain version (618) of the prediction residual of the predetermined block (18) via a transform (T), the transform (T) being defined by a cascade of a primary transform (Tp) and a predetermined secondary transform (Ts) applied to a subset (622) of coefficients (620) of the primary transform in a subset (610) of secondary transforms, wherein the predetermined block can be reconstructed (624) using the prediction signal (606) and the prediction residual of the predetermined block.
[0575] References
[0576] [1] P. Helle et al., “Non-linear weighted intra prediction”, JVET-L0199, Macau, China, October 2018.
[0577] [2] F. Bossen, J. Boyce, K. Suehring, X. Li, V. Seregin, “JVET common test conditions and software reference configurations for SDR video”, JVET-K1010, Ljubljana, Slovenia, July 2018.
[0578] Other embodiments and examples
[0579] Generally, an example can be implemented as a computer program product having program instructions that are operable to perform one of the methods when the computer program product is run on a computer. The program instructions can be stored, for example, on a machine-readable medium.
[0580] Other examples include a computer program stored on a machine-readable carrier for performing one of the methods described herein.
[0581] In other words, an example of a method is thus a computer program having program instructions for performing one of the methods described herein when the computer program is run on a computer.
[0582] Thus, another example of a method is a data carrier medium (or digital storage medium or computer-readable medium) on which a computer program is recorded for performing one of the methods described herein. The data carrier medium, digital storage medium or recording medium is tangible and / or non-transitory, rather than intangible and transitory signals.
[0583] Thus, another example of a method is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence can be transmitted, for example, via a data communication connection (e.g., via the Internet).
[0584] Another example includes a processing device, e.g., a computer or a programmable logic device, for performing one of the methods described herein.
[0585] Another example includes a computer on which a computer program is installed for performing one of the methods described herein.
[0586] Another example includes a device or system for transmitting a computer program (e.g., electronically or optically) to a receiver, the computer program for performing one of the methods described herein. The receiver can be, for example, a computer, a mobile device, a storage device, etc. The device or system can include, for example, a file server for transmitting the computer program to the receiver.
[0587] In some examples, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods can be performed by any suitable hardware device.
[0588] The above examples are illustrative only of the principles disclosed above. It should be understood that modifications and variations of the arrangements and details described herein will be apparent. Accordingly, it is intended to be limited by the scope of the appended claims rather than by the specific details given by way of description and explanation of the examples herein.
[0589] In the following description, like or equivalent elements or elements having like or equivalent functions are denoted by like or equivalent reference numerals (even if occurring in different figures).
Claims
1. An apparatus (54) for decoding a picture (10) from a data stream, the apparatus (54) having a processor and a memory storing instructions which, when executed by the processor, cause the processor to perform the following operations: For a block (18) of the picture, an intra prediction mode (604) is selected (602) from a first set (508) of intra prediction modes or a second set (520) of intra prediction modes based on indications included in the data stream, wherein The first set (508) of intra prediction modes includes a DC intra prediction mode (506), a planar intra prediction mode (504), and at least one angular prediction mode (500), and wherein the second set (520) of intra prediction modes includes at least one matrix-based intra prediction mode (510), and for each matrix-based intra prediction mode (510) in the at least one matrix-based intra prediction mode (510), a prediction vector (518) is obtained according to a matrix-vector product (512) between: a vector (514) derived from a plurality of reference samples (17) in a neighborhood of the block (18), and a prediction matrix (516) associated with the selected matrix-based intra prediction mode (510), and samples of the block (18) are predicted according to the selected matrix-based intra prediction mode (510). Derive a prediction signal (606) for the block (18) using the selected intra prediction mode (604). Select (608) a subset (610) of quadratic transforms from a set (612) of quadratic transforms based on the selected intra prediction mode (604) and the size of the block (18), the set (612) of quadratic transforms including a plurality of low-frequency non-separable quadratic transforms (LFNSTs), wherein the selected subset (610) of quadratic transforms includes at least one of the plurality of LFNSTs, and wherein the same subset (610) of quadratic transforms is selected for the planar intra prediction mode (504) and any matrix-based intra prediction mode in the at least one matrix-based intra prediction mode (510). Derive a prediction residual for the block (18) from the data stream (12). Transform the prediction residual using a transform (T) defined by a cascade of a primary transform (Tp) and a quadratic transform (Ts) selected from the subset (610) of quadratic transforms, the quadratic transform (Ts) corresponding to an LFNST. Reconstruct (624) the block (18) using the prediction signal (606) of the block (18) and the transformed prediction residual.
2. The apparatus (54) according to claim 1, configured to: select one or more subsets (610) of quadratic transforms from the set (612) of quadratic transforms such that each quadratic transform in the set (612) of quadratic transforms is included in the subset (610) of quadratic transforms selected for at least one intra prediction mode in the first set (508) and the second set (520) of intra prediction modes.
3. The apparatus (54) according to claim 1, configured to: select a subset (610) of one or more secondary transforms from the set (612) of secondary transforms such that each secondary transform in each subset (610) of secondary transforms selected for any matrix-based intra prediction mode (510) is included in a subset (610) of secondary transforms selected for at least one intra prediction mode within the first set (508) that does not belong to the angular prediction mode (500).
4. The apparatus (54) according to claim 1, configured to: select one or more subsets (610) of secondary transforms from the set (612) of secondary transforms such that a first union (611 matrix ) of the subsets (610) of secondary transforms selected for the matrix-based intra prediction mode (510) and a second union (611 angular ) of the subsets (610) of secondary transforms selected for all angular intra prediction modes (500) have an empty intersection.
5. The apparatus (54) according to claim 1, configured to: if the subset (610) of one or more secondary transforms contains more than one secondary transform, select the secondary transform (Ts) from the subset (610) of one or more secondary transforms depending on a secondary transform indication syntax element transmitted in the data stream (12) for the block (18).
6. The apparatus (54) according to claim 1, configured to: If the size of the block (18) meets the criteria, it is inferred that the transform through which the transformed prediction residual of the block (18) is correlated with the spatial domain version (618) of the prediction residual of the block (18) is the primary transform, where, the apparatus (54) enables a second set (520) of intra prediction modes to be used for the selection (602) of the intra prediction mode (604), regardless of whether the size of the block (18) meets the criterion.
7. The apparatus according to claim 6, wherein, if the size is below a threshold, the criterion is met.
8. The apparatus (54) according to claim 1, configured to: read a non-zero region indication transmitted in the data stream (12) for a corresponding block (18), the non-zero region indication being used to indicate a non-zero transform domain region (623) in which all non-zero coefficients are located only within a transformed version (616) of the prediction residual of the block (18), and decode the coefficients within the non-zero transform domain region (623) from the data stream (12), depending on the extension and / or position of the non-zero transform domain region (623) meeting a first criterion, and / or the number of non-zero coefficients within the non-zero transform domain region (623) meeting a second criterion, infer that the transform via which the transformed version (616) of the prediction residual of the block (18) is related to the spatial domain version (618) of the prediction residual of the block (18) is the primary transform.
9. The apparatus (54) according to claim 1, wherein, The primary transform is a separable 2D transform and the secondary transform is a non-separable 2D transform.
10. The apparatus (54) according to claim 1, configured to: derive a set selection syntax element (522) from the data stream (12), the set selection syntax element (522) indicating whether to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), the first set (508) of intra prediction modes including a DC intra prediction mode (506) and an angular prediction mode (500), if the set selection syntax element (522) indicates to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), A list (528) of most probable intra prediction modes MPM is then formed based on intra prediction modes for predicting neighboring blocks (524, 526) adjacent to the block (18). An MPM list index (534) is derived from the data stream (12), the MPM list index (534) indicating which most probable intra prediction mode in the list (528) of most probable intra prediction modes is to be selected as the intra prediction mode (604). If the set selection syntax element (522) indicates not to use an intra prediction mode from the first set (508) of intra prediction modes to predict the block (18), then another index (540; 546) is derived from the data stream (12), the another index (540; 546) indicating an intra prediction mode (604) from the second set (520) of intra prediction modes.
11. The apparatus (54) according to claim 10, configured to: If the set selection syntax element (522) indicates to use an intra prediction mode from the first set (508) of intra prediction modes to predict the block (18), then an MPM syntax element (532) is derived from the data stream (12), the MPM syntax element (532) indicating whether the intra prediction mode (604) from the first set (508) of intra prediction modes is within the list (528) of most probable intra prediction modes. If the MPM syntax element (532) indicates that the intra prediction mode (604) from the first set (508) of intra prediction modes is within the list (528) of most probable intra prediction modes, then perform: Form the list (528) of most probable intra prediction modes based on intra prediction modes for predicting neighboring blocks (524, 526) adjacent to the block (18); Derive the MPM list index (534) from the data stream (12), the MPM list index (534) pointing to the intra prediction mode (604) in the list (528) of most probable intra prediction modes. If the MPM syntax element (532) from the data stream (12) indicates that the intra prediction mode (604) from the first set (508) of intra prediction modes is not within the list (528) of most probable intra prediction modes, then another list index (536) is derived from the data stream (12), the another list index (536) indicating the intra prediction mode (604) from the first set (508) of intra prediction modes.
12. The apparatus (54) according to claim 10, configured to: If the set selection syntax element (522) indicates not to use an intra prediction mode from the first set (508) of intra prediction modes to predict the block (18), Another MPM syntax element (538) is derived from the data stream (12), and the another MPM syntax element (538) indicates whether at least one matrix-based intra prediction mode (510) in the second set (520) of the intra prediction modes is within a list (542) of the most probable intra prediction modes (510). If the another MPM syntax element (538) indicates that at least one matrix-based intra prediction mode (510) in the second set (520) of the intra prediction modes is within the list (542) of the most probable intra prediction modes (510), a list (542) of the most probable intra prediction modes (510) is formed based on an intra prediction mode for predicting an adjacent block (524, 526) adjacent to the block (18). Another MPM list index (540) is derived from the data stream (12), and the another MPM list index (540) points to the intra prediction mode (510) in the list (542) of the most probable matrix-based intra prediction modes (510). If the another MPM syntax element (538) indicates that at least one matrix-based intra prediction mode (510) in the second set (520) of the intra prediction modes is not within a list (542) of the most probable matrix-based intra prediction modes (510), yet another list index (546) is derived from the data stream (12), and the yet another list index (546) indicates a matrix-based intra prediction mode (510) in the second set (520) of the intra prediction modes (510).
13. The apparatus (54) according to claim 11, configured to form a list (528) of the most probable intra prediction modes based on an intra prediction mode for predicting an adjacent block (524, 526) adjacent to the block (18), such that: the list (528) is filled with the DC intra prediction mode (506) only when, for each of the adjacent blocks, any one of at least one non-angular intra prediction mode (504, 506) within a first set (508) including the DC intra prediction mode (506) is used to predict the corresponding adjacent block, or any one of the matrix-based intra prediction modes (510) is used to predict the corresponding adjacent block, and any one of the matrix-based intra prediction modes (510) is mapped to any one of the at least one non-angular intra prediction modes through a mapping from the second set (520) of the intra prediction modes (510) to the intra prediction modes within the first set (508) for forming the list (528) of the most probable intra prediction modes.
14. The apparatus (54) according to claim 11, configured to form a list (528) of the most probable intra prediction modes based on intra prediction modes for predicting adjacent blocks (524, 526) adjacent to the block (18), such that for each of the adjacent blocks, in the following cases: any one of at least one non - angular intra prediction mode (504, 506) within a first set (508) including a DC intra prediction mode (506) is used to predict the corresponding adjacent block, or any one of the intra prediction modes (510) is used to predict the corresponding adjacent block, any one of the intra prediction modes (510) is mapped to any one of the at least one non - angular intra prediction modes through a mapping from a second set (520) of the intra prediction modes (510) to the intra prediction modes within the first set (508) for forming the list (528) of the most probable intra prediction modes; The DC intra prediction mode (506) is before any angular intra prediction mode (500) in the list (528) of the most probable intra prediction modes.
15. The apparatus (54) according to claim 11, configured to form a list (528) of the most probable intra prediction modes based on intra prediction modes for predicting adjacent blocks (524, 526) adjacent to the block (18), such that: The list (528) is filled with the planar intra prediction mode (504) in a manner independent of the intra prediction mode used to predict the adjacent blocks.
16. The apparatus (54) according to claim 15, configured to form a list (528) of the most probable intra prediction modes based on the intra prediction modes used for predicting adjacent blocks (524, 526) adjacent to the block (18), such that: Independent of the intra prediction mode used to predict the adjacent blocks, the planar intra prediction mode (504) is at a first position in the list (528) of the most probable intra prediction modes.
17. The apparatus (54) according to claim 1, configured to: Form a sample value vector (400) according to the plurality of reference samples (17), Derive the vector (514) from the sample value vector (400) such that the sample value vector (400) is mapped to the vector (514) through a reversible linear transformation (403).
18. The apparatus (54) according to claim 17, wherein, The reversible linear transformation (403) is defined such that: Each component of the vector (514, 402) other than a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is a predetermined value (1400).
19. The apparatus (54) according to claim 18, wherein, The predetermined value (1400) is one of the following: The average value of the components of the sample value vector (400), such as an arithmetic mean or a weighted average, A default value, A value signaled in the data stream (12) into which the picture (10) is encoded, and The component in the sample value vector (400) corresponding to the component (1500).
20. The apparatus (54) according to claim 17, wherein, The invertible linear transformation (403) is defined such that: Each component among the other components of the vector (514, 402) except for a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is the arithmetic mean of the components of the sample value vector (400).
21. The apparatus (54) according to claim 17, wherein, The invertible linear transformation (403) is defined such that: Each component among the other components of the vector (514, 402) except for a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is the component in the sample value vector (400) corresponding to the component (1500), where the device (54) is configured to: include a plurality of invertible linear transformations (403), each of the plurality of invertible linear transformations (403) being associated with one component of the vector (514, 402), select the component (1500) from the components of the sample value vector (400), and use the invertible linear transformation (403) associated with the component (1500) among the plurality of invertible linear transformations (403) as the invertible linear transformation (403).
22. The apparatus (54) according to claim 18, wherein, All matrix components of the prediction matrix (516) corresponding to the component (1500) of the vector (514, 402) within the column (412) of the prediction matrix (516) are zero, and the device (54) is configured to: calculate the matrix-vector product (407) between: the reduced prediction matrix (405) obtained by removing the column (412) from the prediction matrix (516), and another vector (410) obtained by removing the component (1500) from the vector (514, 402), and perform multiplication to calculate the matrix-vector product (512).
23. The device (54) according to claim 18, configured to: when predicting the samples of the block (18) based on the prediction vector (518), for each component of the prediction vector (518), calculate the sum of the corresponding component and a.
24. The apparatus (54) according to claim 18, wherein, Sum each matrix component of the prediction matrix (516) corresponding to the component (1500) of the vector (514, 402) within the column (412) of the prediction matrix (516) with 1 to generate a matrix, and the multiplication of the matrix by the invertible linear transformation (403) corresponds to a quantized version of the machine learning prediction matrix (1100).
25. The device (54) according to claim 17, configured to: for each component of the sample value vector (400), form (100) the sample value vector (400) based on the plurality of reference samples (17) by: adopting one of the plurality of reference samples (17) as the corresponding component of the sample value vector (400), and / or Average two or more components of the sample value vector (400) to obtain corresponding components of the sample value vector (400).
26. The apparatus (54) according to claim 1, wherein, The plurality of reference samples (17) are arranged within the picture (10) along the outer edge of the block (18).
27. The apparatus (54) according to claim 1, configured to calculate the matrix-vector product (512) using fixed-point arithmetic operations.
28. The apparatus (54) according to claim 1, configured to calculate the matrix-vector product (512) without performing floating-point arithmetic operations.
29. The apparatus (54) according to claim 1, configured to store a fixed-point number representation of the prediction matrix (516).
30. The apparatus (54) according to claim 18, configured to: represent the prediction matrix (516) using prediction parameters, and calculate the matrix-vector product (512) by performing multiplications and additions on components of the vectors (514, 402), the prediction parameters, and intermediate results thereby obtained, wherein, The absolute value of the prediction parameter can be represented by a fixed-point number representation of n bits, where n is equal to or less than 14, or equal to or less than 10, or equal to or less than 8.
31. The apparatus (54) according to claim 30, wherein, The prediction parameter includes: Weights, each of the weights being associated with a corresponding matrix component of the prediction matrix (516).
32. The apparatus (54) according to claim 31, wherein, The prediction parameter further includes: One or more scaling factors, each of the one or more scaling factors being associated with one or more corresponding matrix components of the prediction matrix (516) for scaling the weights associated with one or more corresponding matrix components of the prediction matrix (516), and / or One or more offsets, each of the one or more offsets being associated with one or more corresponding matrix components of the prediction matrix (516) for offsetting the weights associated with one or more corresponding matrix components of the prediction matrix (516).
33. The apparatus (54) according to claim 1, configured to: when predicting samples of the block (18) based on the prediction vector (518), Calculate at least one sample position of the block (18) using interpolation based on the prediction vector (518), each component of the prediction vector (518) being associated with a corresponding position within the block (18).
34. An apparatus (14) for encoding a block (18) of a picture (10), the apparatus (14) having a processor and a memory storing instructions which, when executed by the processor, cause the processor to perform the following operations: Select (602) an intra prediction mode (604) from a first set (508) of intra prediction modes or a second set (520) of intra prediction modes, the first set (508) of intra prediction modes including a DC intra prediction mode (506), a planar intra prediction mode (508), and at least one angular prediction mode (500), the second set (520) of intra prediction modes including at least one matrix-based intra prediction mode (510), and for each matrix-based intra prediction mode (510) in the at least one matrix-based intra prediction mode (510), generate a prediction vector (518) using a matrix-vector product (512) between: a vector (514) derived from a plurality of reference samples (17) in a neighborhood of the block (18), and a prediction matrix (516) associated with the selected matrix-based intra prediction mode (510), and predict samples of the block (18) based on the selected matrix-based intra prediction mode (510). Signal the intra prediction mode (604) in a data stream (12). Derive a prediction signal (606) for the block (18) using the selected intra prediction mode (604). Select (608) a subset (610) of quadratic transforms from a set (612) of quadratic transforms based on the selected intra prediction mode (604) and the size of the block (18), the set (612) of quadratic transforms including a plurality of low-frequency non-separable quadratic transforms LFNST, wherein, A selected subset (610) of the secondary transforms includes at least one of the plurality of LFNSTs, and wherein the same subset (610) of secondary transforms is selected for the planar intra prediction mode (504) and any matrix-based intra prediction mode (510) among the at least one matrix-based intra prediction mode (510). Transform the prediction residual using a transform (T) defined by a cascade of a primary transform (Tp) and a secondary transform (Ts) selected from the subset (610) of secondary transforms, the secondary transform (Ts) corresponding to an LFNST, and Encode a transformed version (616) of the prediction residual of the block (18) into the data stream (12). wherein the block (18) can be reconstructed (624) using the prediction signal (606) of the block (18) and the prediction residual.
35. The apparatus (14) according to claim 34, configured to: select one or more subsets (610) of secondary transforms from the set (612) of secondary transforms such that each secondary transform in the set (612) of secondary transforms is included in a subset (610) of secondary transforms selected for at least one intra prediction mode among the intra prediction modes within the first set (508) and the second set (520).
36. The apparatus (14) according to claim 34, configured to: select one or more subsets (610) of secondary transforms from the set (612) of secondary transforms such that each secondary transform in each subset (610) of secondary transforms selected for any matrix-based intra prediction mode (510) is included in a subset (610) of secondary transforms selected for at least one intra prediction mode within the first set (508) that does not belong to the angular prediction mode (500).
37. The apparatus (14) according to claim 34, configured to: select one or more subsets (610) of secondary transforms from the set (612) of secondary transforms such that a first union (611 matrix of the subset (610) of secondary transforms selected for the matrix-based intra prediction mode (510) and a second union (611 angular ) of the subset (610) of secondary transforms selected for all angular intra prediction modes (500) has an empty intersection.
38. The apparatus (14) according to claim 34, configured to: if a subset (610) of one or more secondary transforms contains more than one secondary transform, select the secondary transform (Ts) from the subset (610) of one or more secondary transforms depending on a secondary transform indication syntax element transmitted in the data stream (12) for the block (18).
39. The apparatus (14) according to claim 34, configured such that: If the size of the block (18) meets the criteria, it is inferred that the transform via which the transformed prediction residual of the block (18) is related to the spatial domain version (618) of the prediction residual of the block (18) is the primary transform, where, the apparatus (14) enables the second set (520) of intra prediction modes to be used for the selection (602) of the intra prediction mode (604), regardless of whether the size of the block (18) meets the criterion.
40. The apparatus (14) according to claim 39, wherein the criterion is met if the size is below a threshold.
41. The apparatus (14) according to claim 34, configured to: transmit, in the data stream (12), a non-zero region indication for a respective block (18), the non-zero region indication being used to indicate a non-zero transform domain region (623) in which all non-zero coefficients in the transform version (616) are located, and encode the coefficients within the non-zero transform domain region (623) into the data stream (12), Among them, infer that the transform via which the transform version (616) of the prediction residual of the block (18) is related to the spatial domain version (618) of the prediction residual of the block (18) is the primary transform depending on the extension and / or position of the non-zero transform domain region (623) meeting a first criterion and / or the number of non-zero coefficients within the non-zero transform domain region (623) meeting a second criterion.
42. The apparatus (14) according to claim 34, wherein, The primary transform is a separable 2D transform and the secondary transform is a non-separable 2D transform.
43. The apparatus (14) according to claim 34, configured to: signal in the data stream (12) a set selection syntax element (522) that indicates whether to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), the first set (508) of intra prediction modes including a DC intra prediction mode (506) and an angular prediction mode (500), if the set selection syntax element (522) indicates to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), A list (528) of most probable intra prediction modes MPM is then formed based on the intra prediction modes used to predict neighboring blocks (524, 526) adjacent to the block (18). An MPM list index (534) is signaled in the data stream (12), the MPM list index (534) indicating which one of the most probable intra prediction modes in the list (528) is to be selected as the intra prediction mode (604). If the set selection syntax element (522) indicates not to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), then another index (540; 546) is signaled in the data stream (12), the another index (540; 546) indicating the intra prediction mode (604) in the second set (520) of intra prediction modes (510).
44. The apparatus (14) according to claim 43, configured to: If the set selection syntax element (522) indicates to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), then an MPM syntax element (532) is signaled in the data stream (12), the MPM syntax element (532) indicating whether the intra prediction mode (604) in the first set (508) of intra prediction modes is within the list (528) of most probable intra prediction modes. If the MPM syntax element (532) indicates that the intra prediction mode (604) in the first set (508) of intra prediction modes is within the list (528) of most probable intra prediction modes, then perform: Form the list (528) of most probable intra prediction modes based on the intra prediction modes used to predict neighboring blocks (524, 526) adjacent to the block (18); Signal the MPM list index (534) in the data stream (12), the MPM list index (534) pointing to the intra prediction mode (604) in the list (528) of most probable intra prediction modes. If the MPM syntax element (532) from the data stream (12) indicates that the intra prediction mode (604) in the first set (508) of intra prediction modes is not within the list (528) of most probable intra prediction modes, then another list index (536) is signaled in the data stream (12), the another list index (536) indicating the intra prediction mode (604) in the first set (508) of intra prediction modes.
45. The apparatus (14) according to claim 43, configured to: If the set selection syntax element (522) indicates not to use one of the intra prediction modes in the first set (508) of intra prediction modes to predict the block (18), Then, another MPM syntax element (538) is signaled in the data stream (12), and the another MPM syntax element (538) indicates whether at least one matrix-based intra prediction mode (510) among the matrix-based intra prediction modes (510) in the second set (520) of the intra prediction modes is within the list (542) of the most probable intra prediction modes (510). If the another MPM syntax element (538) indicates that at least one matrix-based intra prediction mode (510) among the matrix-based intra prediction modes (510) in the second set (520) of the intra prediction modes is within the list (542) of the most probable intra prediction modes (510), then, the list (542) of the most probable intra prediction modes (510) is formed based on the intra prediction modes used to predict the neighboring blocks (524, 526) adjacent to the block (18). Another MPM list index (540) is signaled in the data stream (12), and the another MPM list index (540) points to the intra prediction mode (510) in the list (542) of the most probable matrix-based intra prediction modes (510). If the another MPM syntax element (538) indicates that at least one matrix-based intra prediction mode (510) among the matrix-based intra prediction modes (510) in the second set (520) of the intra prediction modes is not within the list (542) of the most probable matrix-based intra prediction modes (510), then, yet another list index (546) is signaled in the data stream (12), and the yet another list index (546) indicates the matrix-based intra prediction modes (510) in the second set (520) of the intra prediction modes.
46. The apparatus (14) according to claim 44, configured to form the list (528) of the most probable intra prediction modes based on the intra prediction modes used to predict the neighboring blocks (524, 526) adjacent to the block (18), such that:[[]] The list (528) is filled with the DC intra prediction mode (506) only when, for each of the neighboring blocks, any one of the at least one non-angular intra prediction modes (504, 506) within the first set (508) including the DC intra prediction mode (506) is used to predict the corresponding neighboring block, or any one of the matrix-based intra prediction modes (510) is used to predict the corresponding neighboring block, and through the mapping from the matrix-based intra prediction modes (510) in the second set (520) of the intra prediction modes to any one of the at least one non-angular intra prediction modes within the first set (508) used to form the list (528) of the most probable intra prediction modes, any one of the matrix-based intra prediction modes (510) is mapped to any one of the at least one non-angular intra prediction modes.
47. The apparatus (14) according to claim 44, configured to form a list (528) of the most probable intra prediction modes based on intra prediction modes for predicting neighboring blocks (524, 526) adjacent to the block (18), such that for each of the neighboring blocks, in the following cases: any one of at least one non - angular intra prediction mode (504, 506) within the first set (508) including the DC intra prediction mode (506) is used to predict the corresponding neighboring block, or any one of the matrix - based intra prediction modes (510) is used to predict the corresponding neighboring block, any one of the matrix - based intra prediction modes (510) is mapped to any one of the at least one non - angular intra prediction modes through a mapping from the second set (520) of the intra prediction modes (510) to the intra prediction modes within the first set (508) for forming the list (528) of the most probable intra prediction modes; The DC intra prediction mode (506) is before any angular intra prediction mode (500) in the list (528) of the most probable intra prediction modes.
48. The apparatus (14) according to claim 43, configured to form a list (528) of the most probable intra prediction modes based on intra prediction modes for predicting neighboring blocks (524, 526) adjacent to the block (18), such that: The list (528) is filled with the planar intra prediction mode (504) in a manner independent of the intra prediction mode used to predict the neighboring blocks.
49. The apparatus (14) according to claim 48, configured to form a list (528) of the most probable intra prediction modes based on intra prediction modes for predicting neighboring blocks (524, 526) adjacent to the block (18), such that: Independent of the intra prediction mode used to predict the neighboring blocks, the planar intra prediction mode (504) is at the first position in the list (528) of the most probable intra prediction modes.
50. The apparatus (14) according to claim 34, configured to: Form a sample value vector (400) according to the plurality of reference samples (17), Derive the vectors (514, 402) from the sample value vector (400) such that the sample value vector (400) is mapped to the vectors (514, 402) through a predetermined invertible linear transformation (403).
51. The apparatus (14) according to claim 50, wherein, The invertible linear transformation (403) is defined such that: Each component of the vectors (514, 402) other than a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is a predetermined value (1400).
52. The apparatus (14) according to claim 51, wherein, The predetermined value (1400) is one of the following: The average value of the components of the sample value vector (400), such as the arithmetic mean or the weighted average, The default value, The value signaled in the data stream (12) into which the picture (10) is encoded, and The component in the sample value vector (400) corresponding to the component (1500).
53. The device (14) according to claim 50, wherein, The reversible linear transformation (403) is defined such that: Each component among the components of the vector (514, 402) other than a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is the arithmetic mean of the components of the sample value vector (400).
54. The device (14) according to claim 50, wherein, The reversible linear transformation (403) is defined such that: Each component among the components of the vector (514, 402) other than a certain component (1500) is equal to the corresponding component of the sample value vector (400) minus a, where a is the component in the sample value vector (400) corresponding to the component (1500), where the device (14) is configured to: include a plurality of reversible linear transformations (403), each of the plurality of reversible linear transformations (403) being associated with a component of the vector (514, 402), select the component (1500) from the components of the sample value vector (400), and use the reversible linear transformation (403) associated with the component (1500) among the plurality of reversible linear transformations (403) as the reversible linear transformation (403).
55. The device (14) according to claim 51, wherein, All matrix components of the prediction matrix (516) corresponding to the component (1500) of the vector (514, 402) within the column (412) of the prediction matrix (516) are zero, and the device (14) is configured to: calculate the matrix - vector product (407) between: the reduced prediction matrix (405) obtained by removing the column (412) from the prediction matrix (516), and another vector (410) obtained by removing the component (1500) from the vector (514, 402), and perform multiplication to calculate the matrix - vector product (512).
56. The device (14) according to claim 51, configured to: when predicting samples of the block (18) based on the prediction vector (518), for each component of the prediction vector (518), calculate the sum of the corresponding component and a.
57. The apparatus (14) according to claim 51, wherein, Sum each matrix component of the prediction matrix (516) corresponding to the component (1500) of the vector (514, 402) within the column (412) of the prediction matrix (516) with 1 to produce a matrix, and the multiplication of the matrix by the reversible linear transformation (403) corresponds to a quantized version of the machine - learning prediction matrix (1100).
58. The device (14) according to claim 50, configured to: for each component of the sample value vector (400), form (100) the sample value vector (400) based on the plurality of reference samples (17) by: adopting one of the plurality of reference samples (17) as the corresponding component of the sample value vector (400), and / or Average two or more components of the sample value vector (400) to obtain corresponding components of the sample value vector (400).
59. The apparatus (14) according to claim 34, wherein, The plurality of reference samples (17) are arranged within the picture (10) along the outer edge of the block (18).
60. The apparatus (14) according to claim 34, configured to calculate the matrix-vector product (512) using fixed-point arithmetic operations.
61. The apparatus (14) according to claim 34, configured to calculate the matrix-vector product (512) without performing floating-point arithmetic operations.
62. The apparatus (14) according to claim 34, configured to store a fixed-point representation of the prediction matrix (516).
63. The apparatus (14) according to claim 51, configured to: represent the prediction matrix (516) using prediction parameters, and calculate the matrix-vector product (512) by performing multiplications and additions on the components of the vectors (514, 402), the prediction parameters, and the resulting intermediate results, wherein, The absolute value of the prediction parameter can be represented by a fixed-point representation of n bits, where n is equal to or less than 14, or equal to or less than 10, or equal to or less than 8.
64. The apparatus (14) according to claim 63, wherein, The prediction parameter includes: Weights, each of the weights being associated with a corresponding matrix component of the prediction matrix (516).
65. The apparatus (14) according to claim 64, wherein, The prediction parameter further includes: One or more scaling factors, each of the one or more scaling factors being associated with one or more corresponding matrix components of the prediction matrix (516) for scaling the weights associated with the one or more corresponding matrix components of the prediction matrix (516), and / or One or more offsets, each of the one or more offsets being associated with one or more corresponding matrix components of the prediction matrix (516) for offsetting the weights associated with the one or more corresponding matrix components of the prediction matrix (516).
66. The apparatus (14) according to claim 34, configured to: when predicting samples of the block (18) based on the prediction vector (518), Calculate at least one sample position of the block (18) using interpolation based on the prediction vector (518), each component of the prediction vector (518) being associated with a corresponding position within the block (18).
67. A method (6000) for decoding a picture (10) from a data stream, comprising: For a block (18) of the picture, an intra prediction mode (604) is selected (602) from a first set (508) of intra prediction modes or a second set (520) of intra prediction modes based on indications included in the data stream, the first set (508) of intra prediction modes including a DC intra prediction mode (506), a planar intra prediction mode (504), and at least one angular prediction mode (500), the second set (520) of intra prediction modes including at least one matrix-based intra prediction mode (510), and for each matrix-based intra prediction mode (510) in the at least one matrix-based intra prediction mode (510), a prediction vector (518) is obtained based on a matrix-vector product (512) between: a vector (514) derived from a plurality of reference samples (17) in a neighborhood of the block (18), and a prediction matrix (516) associated with the selected matrix-based intra prediction mode (510), and samples of the block (18) are predicted based on the selected matrix-based intra prediction mode (510). A prediction signal (606) of the block (18) is derived using the selected predetermined intra prediction mode (604). A subset (610) of secondary transforms is selected (608) from a set (612) of secondary transforms based on the selected intra prediction mode (604) and the size of the block (18), the set (612) of secondary transforms including a plurality of low-frequency non-separable secondary transforms (LFNSTs), wherein the selected subset (610) of secondary transforms includes at least one of the plurality of LFNSTs, and wherein the same subset (610) of secondary transforms is selected for the planar intra prediction mode (504) and any matrix-based intra prediction mode in the at least one matrix-based intra prediction mode (510). A prediction residual of the block (18) is derived from the data stream (12). The prediction residual is transformed using a transform (T) defined by a cascade of a primary transform (Tp) and a secondary transform (Ts) selected from the subset (610) of secondary transforms, the secondary transform (Ts) corresponding to an LFNST. The block (18) is reconstructed (624) using the prediction signal (606) of the block (18) and the transformed prediction residual.
68. A method (7000) for encoding a block (18) of a picture (10), comprising: An intra prediction mode (604) is selected (602) from a first set (508) of intra prediction modes and a second set (520) of intra prediction modes. The first set (508) of intra prediction modes includes a DC intra prediction mode (506), a planar intra prediction mode (508), and at least one angular prediction mode (500). The second set (520) of intra prediction modes includes at least one matrix-based intra prediction mode (510). For each matrix-based intra prediction mode (510) in the at least one matrix-based intra prediction mode (510), a prediction vector (518) is obtained using a matrix-vector product (512) between: a vector (514) derived from a plurality of reference samples (17) in a neighborhood of the block (18); and a prediction matrix (516) associated with the selected matrix-based intra prediction mode (510). The samples of the block (18) are predicted based on the selected matrix-based intra prediction mode (510). The intra prediction mode (604) is signaled in the data stream (12); A prediction signal (606) for the block (18) is derived using the selected intra prediction mode (604); A subset (610) of quadratic transforms is selected (608) from a set (612) of quadratic transforms based on the selected intra prediction mode (604) and the size of the block (18). The set (612) of quadratic transforms includes a plurality of low-frequency non-separable quadratic transforms (LFNSTs). The selected subset (610) of quadratic transforms includes at least one of the plurality of LFNSTs, and the same subset (610) of quadratic transforms is selected for the planar intra prediction mode (504) and any matrix-based intra prediction mode (510) in the at least one matrix-based intra prediction mode (510). The prediction residuals are transformed using a transform (T) defined by a cascade of a primary transform (Tp) and a quadratic transform (Ts) selected from the subset (610) of quadratic transforms, where the quadratic transform (Ts) corresponds to an LFNST; and The transformed version (616) of the prediction residuals of the block (18) is encoded into the data stream (12). wherein the block (18) can be reconstructed (624) using the prediction signal (606) of the block (18) and the prediction residuals.
69. A computer program product having program code for performing the method according to any one of claims 67 or 68 when run on a computer.