Coding using matrix-based intra prediction and quadratic transform
By selecting a subset of secondary transforms and performing matrix-vector product processing, the problem of high storage requirements of matrix-based intra-frame prediction modes is solved, and coding efficiency and performance are improved.
Patent Information
- Application Number
- CN202511018792.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-25
- Filing Date
- 2020-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, the combination of matrix-based intra prediction mode and secondary transform has the problem of high memory requirements for storing additional transforms, resulting in reduced coding efficiency.
Transforms are defined for matrix-based intra prediction modes and non-matrix-based intra prediction modes by selecting one or more subsets of secondary transforms from a set of secondary transforms, reducing storage requirements, and processing prediction residuals by concatenation of matrix-vector products and secondary transforms.
It improves coding efficiency, reduces storage requirements, supports more efficient intra-frame prediction mode selection, and improves encoding and decoding performance.
Smart Images

Figure CN120751124A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application "Encoding using matrix-based intra-frame prediction and secondary transform" (application number: 202080049521.X) with an application date of June 23, 2020. Technical Field
[0002] The present application relates to the field of matrix-based intra prediction and secondary transform. Background Art
[0003] For traditional intra prediction modes such as planar mode, DC mode, and angular mode, a non-separable secondary transform (LFNST) is a tool used to transform the prediction residuals corresponding to these intra prediction modes. Here, a set S of transform sets is given, so that each traditional intra prediction mode is associated with one of these transform sets. Then, at the decoder, it is possible to extract from the bitstream whether LFNST is to be applied to a given block. If this is the case, one of the transform sets in the set S is given, depending on the intra prediction mode used on the current block, and if the transform set consists of more than one transform, it is possible to extract from the bitstream which transform T in the set is to be used. Then, at the decoder, transform T is applied as a secondary transform, which means that the transform T is applied to a subset of the residual transform coefficients of the separable primary transform. However, the above-mentioned secondary transform is defined a priori only for traditional intra prediction modes.
[0004] Therefore, it is desirable to provide concepts for more efficiently rendering picture coding and / or video coding to support secondary transforms for matrix-based intra prediction (MIP), ie, block-based intra prediction.
[0005] This is achieved by the subject matter of the independent claims of the present application.
[0006] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application. Summary of the Invention
[0007] According to a first aspect of the present invention, the inventors of the present application recognized that one problem encountered when attempting to associate a secondary transform with a matrix-based intra prediction mode stems from the fact that providing a specific secondary transform for each MIP mode may be prohibitively expensive in terms of memory requirements for storing the additional transforms. According to the first aspect of the present application, this difficulty is overcome by selecting one or more subsets of secondary transforms from a set of secondary transforms, the set of secondary transforms including transforms associated with both matrix-based intra prediction modes and non-matrix-based intra prediction modes. Secondary transforms in the set of secondary transforms can be defined for one or more prediction modes, which reduces the memory capacity required for the set of secondary transforms. Transforms defined for planar intra prediction mode and / or transforms defined for DC intra prediction mode can also be used for selection of matrix-based intra prediction modes. Although the additional syntax elements required for blocks associated with matrix-based intra prediction modes to indicate the use of secondary transforms may increase bitstream and, therefore, signaling costs, coding efficiency can be improved by specifically selecting the subset of secondary transforms for matrix-based intra prediction modes.
[0008] Therefore, according to a first aspect of the present application, a device for decoding a predetermined block of a picture using intra-frame prediction, i.e., a decoder, is configured to select a predetermined intra-frame prediction mode from a plurality of intra-frame prediction modes based on a data stream, the plurality of intra-frame prediction modes including a first set of intra-frame prediction modes and a second set of matrix-based intra-frame prediction modes. The first set of intra-frame prediction modes includes a DC intra-frame prediction mode and an angular prediction mode and an optional planar intra-frame prediction mode. If a matrix-based intra-frame prediction mode in the second set is selected as the predetermined intra-frame prediction mode, the decoder is configured to obtain a prediction vector using a matrix-vector product between the following two: a vector derived from a reference sample in a neighborhood of the predetermined block and a prediction matrix associated with the corresponding matrix-based intra-frame prediction mode, and the decoder is configured to predict samples of the predetermined block based on the prediction vector. The decoder is configured to derive a prediction signal for a predetermined block using a predetermined intra prediction mode, and to select one or more subsets of secondary transforms from a set of secondary transforms in a manner that depends on the predetermined intra prediction mode, such that the subset is non-empty if the predetermined intra prediction mode is included in a first set of intra prediction modes and if the predetermined intra prediction mode is included in a second set of matrix-based intra prediction modes. The first set and the second set define intra prediction modes available for the secondary transform. Thus, for a predetermined intra prediction mode selected from the first set or from the second set, the decoder is configured to select one or more subsets of secondary transforms from the set of secondary transforms that are specifically associated with the selected predetermined intra prediction mode. In addition, the decoder is configured to: derive, from the data stream, a transformed version of a prediction residual for a predetermined block, if the predetermined intra prediction mode is included in a first set of intra prediction modes, and if the predetermined intra prediction mode is included in a second set of matrix-based intra prediction modes, the transformed version of the prediction residual for the predetermined block being related to a spatial domain version of the prediction residual for the predetermined block via a transform defined by a concatenation of a primary transform and a predetermined secondary transform applied to a subset of coefficients of the primary transform in a subset of secondary transforms. The primary transform is, for example, set by default, and the predetermined secondary transform is, for example, selected by the decoder from the subset of secondary transforms. The decoder may be configured to select the predetermined secondary transform from the subset of secondary transforms by deriving a secondary transform indication syntax element from the data stream. The decoder is configured to reconstruct the predetermined block using the prediction signal and the prediction residual for the predetermined block.
[0009] According to a first aspect of the present application, an apparatus for encoding a predetermined block of a picture using intra-frame prediction, i.e., an encoder, is configured to, in parallel with a decoder, select a predetermined intra-frame prediction mode from a plurality of intra-frame prediction modes, the plurality of intra-frame prediction modes including a first set of intra-frame prediction modes and a second set of matrix-based intra-frame prediction modes, the first set of intra-frame prediction modes including a DC intra-frame prediction mode, an angular prediction mode, and an optional planar intra-frame prediction mode, and obtain a prediction vector based on each matrix-based intra-frame prediction mode in the second set of matrix-based intra-frame prediction modes using a matrix-vector product between the following two: a vector derived from reference samples in a neighborhood of the predetermined block and a prediction matrix associated with the corresponding matrix-based intra-frame prediction mode, and predicting samples of the predetermined block based on the prediction vector. The encoder is configured to signal the predetermined intra-frame prediction mode in a data stream and derive a prediction signal for the predetermined block using the predetermined intra-frame prediction mode. In addition, the encoder is configured to select one or more subsets of secondary transforms from the set of secondary transforms in a manner that depends on the predetermined intra-frame prediction mode, such that the subset is non-empty if the predetermined intra-frame prediction mode is included in the first set of intra-frame prediction modes and if the predetermined intra-frame prediction mode is included in the second set of matrix-based intra-frame prediction modes. The encoder is configured to encode a transformed version of the prediction residual of the predetermined block into the data stream if the predetermined intra-frame prediction mode is included in the first set of intra-frame prediction modes and if the predetermined intra-frame prediction mode is included in the second set of matrix-based intra-frame prediction modes, the transformed version of the prediction residual of the predetermined block being related to a spatial domain version of the prediction residual of the predetermined block via a transform defined by a concatenation of a primary transform and a predetermined secondary transform applied to a subset of coefficients of the primary transform in the subset of the secondary transforms. The predetermined block can be reconstructed using the prediction signal and the prediction residual of the predetermined block.
[0010] According to an embodiment, the decoder / encoder is configured to select a subset of one or more secondary transforms from a set of secondary transforms in a manner that depends on a predetermined intra-prediction mode, such that each secondary transform in the set of secondary transforms is included in the subset of one or more secondary transforms selected for at least one of the intra-prediction modes in the first set and the second set. For one or more intra-prediction modes in the first set or the second set, the corresponding selected subset may include all secondary transforms in the set of secondary transforms.
[0011] According to an embodiment, the decoder / encoder is configured to select one or more subsets of secondary transforms from a set of secondary transforms in a manner that depends on a predetermined intra prediction mode, such that each secondary transform in each subset of secondary transforms selected for any matrix-based intra prediction mode is included in the subset of secondary transforms selected for at least one intra prediction mode in the first set that is not an angular prediction mode. The subset selected for the matrix-based intra prediction mode may include secondary transforms associated with one or more non-angular intra prediction modes in the first set (e.g., DC intra prediction mode and / or planar intra prediction mode). Thus, the secondary transforms in the set of secondary transforms may be part of more than one subset of one or more secondary transforms. The set of secondary transforms does not necessarily include additional secondary transforms that are only applicable to blocks having the predetermined intra prediction mode as a matrix-based intra prediction mode. For a predetermined block having the predetermined intra prediction mode as a matrix-based intra prediction mode, the decoder / encoder is configured to select from the set of secondary transforms the same secondary transform as the predetermined intra prediction mode as a non-angular prediction mode in the first set. The subset selected for the matrix-based intra-frame prediction mode may be equal to the subset selected for the intra-frame prediction mode that does not belong to the angular prediction mode in the first set, or may include some secondary transforms in the subset selected for the intra-frame prediction mode that does not belong to the angular prediction mode in the first set, or may include some or all secondary transforms in two or more subsets selected for the intra-frame prediction mode that does not belong to the angular prediction mode in the first set.
[0012] The methods for encoding or decoding are based on the same considerations as the above-mentioned devices for encoding or decoding. In this way, these methods can be completed with all the features and functions also described with respect to the devices for encoding and / or decoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are not necessarily drawn to scale, emphasis instead generally being placed on illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0014] Figure 1 An embodiment of encoding into a data stream is shown;
[0015] Figure 2 An embodiment of an encoder is shown;
[0016] Figure 3 An embodiment of the reconstruction of a picture is shown;
[0017] Figure 4 An embodiment of a decoder is shown;
[0018] Figure 5A schematic diagram illustrating prediction of a block for encoding and / or decoding according to an embodiment is shown;
[0019] Figure 6 Matrix operations for predicting a block for encoding and / or decoding according to an embodiment are shown;
[0020] Figure 7.1 shows prediction of a block using a reduced sample value vector according to an embodiment;
[0021] Figure 7.2 shows prediction of a block using interpolation of samples according to an embodiment;
[0022] Figure 7.3 shows prediction of a block using a reduced sample value vector according to an embodiment, where only some boundary samples are averaged;
[0023] Figure 7.4 shows prediction of a block using a reduced sample value vector according to an embodiment, wherein groups of four boundary samples are averaged;
[0024] Figure 8 shows matrix operations performed by the apparatus according to an embodiment;
[0025] Figure 9a and Figure 9b shows detailed matrix operations performed by the apparatus according to an embodiment;
[0026] Figure 10 shows detailed matrix operations performed by a device using offset and scaling parameters according to an embodiment;
[0027] Figure 11 shows detailed matrix operations performed by a device using offset and scaling parameters according to various embodiments;
[0028] Figure 12 A schematic diagram illustrating details of performing intra-frame prediction on a predetermined block using a prediction mode in a most probable mode list according to an embodiment;
[0029] Figure 13 A schematic diagram illustrating decoding a predetermined block using a secondary transform according to an embodiment is shown;
[0030] Figure 14 shows a schematic diagram of applying a primary transform and a secondary transform according to an embodiment;
[0031] Figure 15 A schematic diagram illustrating selecting a subset of secondary transformations according to an embodiment is shown;
[0032] Figure 16 A schematic diagram illustrating a predetermined block having a non-zero transform domain region according to an embodiment;
[0033] Figure 17 A block diagram illustrating a method for decoding a predetermined block according to an embodiment is shown;
[0034] Figure 18 A block diagram showing a method for encoding a predetermined block according to an embodiment; and
[0035] Figures 19a to 19d The syntax elements portion of the data stream is shown. DETAILED DESCRIPTION
[0036] In the following description, the same or equivalent elements or elements having the same or equivalent functions are denoted by the same or equivalent reference numerals even if reference numerals appear in different drawings.
[0037] In the following description, a number of details are set forth to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid confusion regarding the embodiments of the present invention. Furthermore, unless specifically indicated otherwise, the features of the different embodiments described below may be combined with each other.
[0038] 1 Introduction
[0039] In the following, various inventive examples, embodiments, and aspects are described, at least some of which relate, inter alia, to methods and / or apparatus for video coding and / or for performing intra-frame prediction, for example using linear or affine transformations with neighboring sample reduction, and / or for optimizing video delivery (e.g., broadcast, streaming, file playback, etc.), for example, for video applications and / or for virtual reality applications.
[0040] Furthermore, examples, embodiments and aspects may relate to High Efficiency Video Coding (HEVC) or successors.Further, other embodiments, examples and aspects are defined by the following claims.
[0041] It should be noted that any embodiment, example and aspect defined by the claims may be supplemented by any of the details (features and functions) described in the following sections.
[0042] Furthermore, the embodiments, examples and aspects described in the following sections can be used alone and can also be supplemented by any features in another section or by any features included in the claims.
[0043] In addition, it should be noted that the individuals, examples, embodiments, and aspects described herein can be used alone or in combination. Therefore, details can be added to each of the individual aspects without adding details to another of the examples, embodiments, and aspects.
[0044] It should also be noted that this disclosure explicitly or implicitly describes features of decoding and / or encoding systems and / or methods.
[0045] Furthermore, any features and functions disclosed herein in relation to a method may also be used in a device. Furthermore, any features and functions disclosed herein in relation to a device may also be used in a corresponding method. In other words, any features and functions described in relation to a device may be supplemented by the methods disclosed herein.
[0046] Furthermore, as will be described in the "Implementation Alternatives" section, any features and functions described herein may be implemented in hardware or in software, or using a combination of hardware and software.
[0047] Furthermore, in some examples, embodiments or aspects, any features described within brackets ("(...)" or "[...]") may be considered optional.
[0048] 2Encoder and Decoder
[0049] Below, we describe various examples that can help achieve more efficient compression when using block-based prediction. Some examples achieve high compression efficiency by consuming a set of intra-prediction modes. This can be added to other intra-prediction modes, for example, using heuristics, or can be provided specifically. Other examples even utilize the two special cases just discussed. However, as a variation of these embodiments, intra-prediction can be converted to inter-prediction by using reference samples from another picture.
[0050] To facilitate understanding of the following examples of the present application, the description begins with the presentation of possible encoders and decoders into which the examples subsequently outlined in the present application can be built. Figure 1 An apparatus for encoding a picture 10 block by block into a data stream 12 is shown. The apparatus is indicated using reference numeral 14 and can be a still picture encoder or a video encoder. In other words, when the encoder 14 is configured to encode a video 16 including the picture 10 into the data stream 12, the picture 10 can be the current picture in the video 16, or the encoder 14 can only encode the picture 10 into the data stream 12.
[0051] As described above, the encoder 14 performs encoding in a block-by-block manner or on a block basis. To this end, the encoder 14 subdivides the picture 10 into blocks, which the encoder 14 encodes into the data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are described in more detail below. Generally, the subdivision into fixed-size blocks 18 (e.g., an array of blocks arranged in rows and columns) or into blocks 18 of different block sizes can be achieved, for example, by using a hierarchical multi-tree subdivision starting from the entire picture area of the picture or starting from a pre-partition of the picture 10 into an array of tree blocks, wherein these examples should not be considered to exclude other possible ways of subdividing the picture 10 into blocks 18.
[0052] Furthermore, the encoder 14 is a predictive encoder that is configured to predictively encode the picture 10 into the data stream 12. For a certain block 18, this means that the encoder 14 determines a prediction signal for the block 18 and encodes a prediction residual (i.e., a prediction error of the prediction signal from the actual picture content within the block 18) into the data stream 12.
[0053] Encoder 14 can support different prediction modes to derive a prediction signal for a block 18. The prediction mode that is important in the following example is intra-frame prediction mode, according to which the interior of block 18 is spatially predicted based on neighboring coded samples of picture 10. The encoding of picture 10 into data stream 12, and therefore the corresponding decoding process, can be based on a coding order 20 defined between blocks 18. For example, coding order 20 can traverse blocks 18 in a raster scan order, e.g., from top to bottom, row by row, and from left to right within each row. In the case of a hierarchical multi-tree subdivision, a raster scan order can be applied within each level, where a depth-first traversal order can be applied, i.e., leaf nodes within a block at a certain level can precede blocks at the same level with the same parent block according to coding order 20. Depending on coding order 20, the neighboring coded samples of block 18 can generally be located on one or more sides of block 18. In the example presented herein, for example, the neighboring coded samples of block 18 are located at the top and to the left of block 18.
[0054] The intra prediction mode may not be the only mode supported by the encoder 14. In the case where the encoder 14 is a video encoder, for example, the encoder 14 may also support an inter prediction mode, according to which the block 18 is temporally predicted based on previously encoded pictures of the video 16. Such an inter prediction mode may be a motion compensated prediction mode, according to which a motion vector is signaled for such a block 18, the motion vector indicating the relative spatial offset of the portion of the prediction signal for the block 18 to be derived as a copy. Additionally or alternatively, other non-intra prediction modes may also be available, such as the inter prediction mode in the case where the encoder 14 is a multi-view encoder, or a non-prediction mode, according to which the interior of the block 18 is encoded as is (i.e., without any prediction).
[0055] Before focusing the description of this application on intra prediction mode, Figure 2 A more specific example for a possible block-based encoder is described, i.e., a possible implementation for the encoder 14, and then the respective embodiments are presented. Figure 1 and Figure 2 Two corresponding examples of decoders for .
[0056] Figure 2 Shown Figure 1 A possible implementation of the encoder 14 is one in which the encoder is configured to encode the prediction residual using transform coding, although this is merely an example and the application is not limited to such prediction residual coding. Figure 2, the encoder 14 includes a subtractor 22 configured to subtract the corresponding prediction signal 24 from the incoming signal (i.e., the picture 10, or, in the case of a block-based scenario, the current block 18) to obtain a prediction residual signal 26, which is then encoded into the data stream 12 by a prediction residual encoder 28. The prediction residual encoder 28 consists of a lossy encoding stage 28a and a lossless encoding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30, which quantizes the samples of the prediction residual signal 26. As already described above, this example uses transform coding of the prediction residual signal 26, and therefore, the lossy encoding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform this spectrally decomposed prediction residual 26, wherein the quantization of the quantizer 30 occurs on the transformed coefficients representing the residual signal 26. The transform may be a DCT, a DST, an FFT, a Hadamard transform, etc. The transformed and quantized prediction residual signal 34 is then losslessly encoded by a lossless encoding stage 28b, which is an entropy encoder that entropy encodes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further comprises a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 in order to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner that is also usable at the decoder (i.e., taking into account the coding losses caused by the quantizer 30). To this end, the prediction residual reconstruction stage 36 comprises an inverse quantizer 38, which performs the inverse operation of the quantization of the quantizer 30, followed by an inverse transformer 40, which performs the inverse transform relative to the transform performed by the transformer 32, such as the inverse of the spectral decomposition, such as the inverse of any of the specific transform examples described above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transformer 40 to the prediction signal 24 to output a reconstructed signal, i.e., reconstructed samples. This output is fed into a predictor 44 of the encoder 14, which then determines the prediction signal 24 based on the output. The predictor 44 supports the above-mentioned Figure 1 All the forecasting models already discussed. Figure 2 It is also shown that in the case where the encoder 14 is a video encoder, the encoder 14 may also include a loop filter 46 that filters the fully reconstructed pictures which, after having been filtered, form the reference pictures for the predictor 44 with respect to the inter-prediction blocks.
[0057] As described above, encoder 14 operates on a block-based basis. For the purposes of the subsequent description, the block-based basis of interest is the subdivision of picture 10 into blocks, for which an intra-prediction mode is selected from a set or multiple intra-prediction modes supported by predictor 44 or encoder 14, respectively, and the selected intra-prediction mode is independently performed. However, other types of blocks into which picture 10 is subdivided are also possible. For example, the aforementioned determination of whether picture 10 is inter-coded or intra-coded can be made at the granularity of block 18 or in units of blocks offset from block 18. For example, the inter / intra mode decision can be made at the level of the coding blocks into which picture 10 is subdivided, with each coding block being subdivided into prediction blocks. For coding blocks for which intra-prediction has been determined, an intra-prediction mode decision is made for each prediction block. To this end, for each of these prediction blocks, a decision is made as to which supported intra-prediction mode should be applied to the corresponding prediction block. These prediction blocks will form the block 18 of interest herein. Predictor 44 treats prediction blocks within coding blocks associated with inter-prediction differently. These blocks are inter-predicted from the reference picture by determining a motion vector and copying the prediction signal for the block from the position in the reference picture pointed to by the motion vector. Another block subdivision involves the subdivision of the transform block, the transformer 32 and the inverse transformer 40 performing the transform in units of transform blocks. For example, the transform block can be the result of a further subdivision of the coding block. Of course, the examples set forth herein should not be considered limiting, and other examples exist. Just for the sake of completeness, it should be noted that the subdivision into coding blocks can, for example, use multi-tree subdivision, and that the coding block can also be further subdivided using multi-tree subdivision to obtain prediction blocks and / or transform blocks.
[0058] Figure 3 Described in Figure 1 The decoder 54 or means for block-by-block decoding of the encoder 14. The decoder 54 operates inversely to the encoder 14, i.e., it decodes the picture 10 from the data stream 12 in a block-by-block manner and supports multiple intra prediction modes for this purpose. For example, the decoder 54 may include a residual provider 156. Figure 1All other possibilities discussed are also valid for decoder 54. To this end, decoder 54 can be a still picture decoder or a video decoder, and decoder 54 also supports all prediction modes and prediction possibilities. The difference between encoder 14 and decoder 54 lies primarily in the fact that encoder 14 chooses or selects coding decisions based on some optimization, such as minimizing some cost function that may depend on the coding rate and / or coding distortion. One of these coding options or coding parameters may involve selecting an intra-prediction mode to be used for current block 18 from available or supported intra-prediction modes. The selected intra-prediction mode can then be signaled by encoder 14 within data stream 12 for current block 18, and decoder 54 uses this signaling for block 18 in data stream 12 to redo the selection. Similarly, the subdivision of picture 10 into block 18 can be optimized within encoder 14, with corresponding subdivision information delivered within data stream 12, and decoder 54 recovers the subdivision of picture 10 into block 18 based on the subdivision information. In summary, the decoder 54 can be a prediction decoder based on block operations, and in addition to the intra-frame prediction mode, the decoder 54 can support other prediction modes, such as the inter-frame prediction mode when the decoder 54 is a video decoder. Figure 1 The coding order 20 discussed is used, and since this coding order 20 is observed at both the encoder 14 and the decoder 54, the same adjacent samples can be used for the current block 18 at both the encoder 14 and the decoder 54. Therefore, in order to avoid unnecessary repetition, the description of the mode of operation of the encoder 14, with respect to the subdivision of the picture 10 into blocks, for example, with respect to prediction, and with respect to the coding of the prediction residual, should also apply to the decoder 54. The difference is that the encoder 14 selects some coding options or coding parameters and signals by optimization and signals or inserts these coding parameters into the data stream 12, which are then derived from the data stream 12 by the decoder 54 in order to re-perform the prediction, subdivision, etc.
[0059] Figure 4 Shown Figure 3 A possible implementation of the decoder 54 is, i.e., suitable for Figure 1 The implementation of encoder 14 (such as Figure 2 As shown in the following example, Figure 4 Many components of the encoder 54 are related to Figure 2 The same elements appear in the corresponding encoder, so Figure 4 The same reference numerals with primes are used in order to designate these elements. Specifically, the adder 42', the optional loop filter 46' and the predictor 44' are shown in the same manner as they are in FIG. Figure 2is connected to the prediction loop in the same way as in the encoder of . The reconstructed (i.e. dequantized and retransformed) prediction residual signal applied to the adder 42 ′ is derived from a sequence of an entropy decoder 56 that inverses the entropy coding of the entropy encoder 28 b followed by a residual signal reconstruction stage 36 ′ consisting of a dequantizer 38 ′ and an inverse transformer 40 ′, as in the case on the encoding side. The output of the decoder is Figure 10 The reconstruction of the picture 10 may be available directly at the output of the adder 42' or, alternatively, at the output of the loop filter 46'. Some post-filter may be arranged at the output of the decoder in order to perform some post-filtering on the reconstruction of the picture 10 in order to improve the picture quality, but Figure 4 This option is not depicted in .
[0060] Likewise, about Figure 4 , except that only the encoder performs optimization tasks and related decisions about encoding options, the above Figure 2 The description proposed should also Figure 4 However, all descriptions of block subdivision, prediction, inverse quantization and retransformation also apply to Figure 4 The decoder 54 is effective.
[0061] 3ALWIP (Affine Linear Weighted Intra Predictor)
[0062] This document discusses some non-limiting examples of ALWIP, even though ALWIP need not always embody the techniques discussed herein.
[0063] The present application relates in particular to an improved block-based prediction mode concept for block-by-block picture coding, such as can be used in video codecs such as HEVC or any successor of HEVC. The prediction mode can be an intra-prediction mode, but in theory the concepts described herein can also be transferred to inter-prediction mode, where the reference sample is part of another picture.
[0064] A block-based prediction concept is sought that allows for efficient implementation, e.g., hardware-friendly implementation.
[0065] This object is achieved by the subject-matter of the independent claims of the present application.
[0066] Intra-frame prediction mode is widely used in picture coding and video coding. In video coding, intra-frame prediction mode competes with other prediction modes such as inter-frame prediction mode (e.g., motion compensated prediction mode). In intra-frame prediction mode, the current block is predicted based on neighboring samples, that is, samples that have been encoded on the encoder side and samples that have been decoded on the decoder side. The neighboring sample values are extrapolated to the current block to form a prediction signal for the current block, and the prediction residual is transmitted in the data stream for the current block. The better the prediction signal, the lower the prediction residual, and therefore, the fewer bits are required to encode the prediction residual.
[0067] To be effective, several aspects should be considered to form an efficient framework for intra prediction in a block-by-block picture coding environment. For example, the greater the number of intra prediction modes supported by the codec, the greater the rate consumption of auxiliary information used to signal the selection to the decoder. On the other hand, the set of supported intra prediction modes should be able to provide a good prediction signal, that is, a prediction signal that results in low prediction residuals.
[0068] Hereinafter, as a comparative embodiment or basic example, a device (encoder or decoder) for decoding a picture block by block from a data stream is disclosed, wherein the device supports at least one intra-frame prediction mode, according to which an intra-frame prediction signal of a block of a predetermined size of the picture is determined by applying a first template of samples adjacent to the current block to an affine linear predictor. In the following, the affine linear predictor will be referred to as an affine linear weighted intra-frame predictor (ALWIP).
[0069] The apparatus may have at least one of the following properties (which may also apply to a method or another technique, such as being implemented in a non-transitory memory unit storing instructions that, when executed by a processor, cause the processor to implement the method and / or operate as an apparatus):
[0070] 3.1 Predictors can complement other predictors
[0071] The intra-frame prediction modes that may form the subject of the improvements described further below may be complementary to other intra-frame prediction modes of the codec. Thus, they may be complementary to the DC prediction mode, the planar prediction mode, or the angular prediction mode defined in the HEVC codec (corresponding JEM reference software). The latter three intra-frame prediction modes shall from now on be referred to as conventional intra-frame prediction modes. Therefore, for a given block in intra-frame mode, the decoder needs to parse a flag indicating whether one of the intra-frame prediction modes supported by the device is to be used.
[0072] 3.2 More than one proposed prediction model
[0073] The device may include more than one ALWIP mode. Therefore, in the case where the decoder knows that one of the ALWIP modes supported by the device is to be used, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the device is to be used.
[0074] The signaling of supported modes may have the following properties: the encoding of some ALWIP modes may require fewer binary bins than other ALWIP modes. Which of these modes requires fewer binary bins and which requires more binary bins may depend on information that can be extracted from the already decoded bitstream or may be fixed in advance.
[0075] 4 Some aspects
[0076] Figure 2 A decoder 54 is shown for decoding a picture from the data stream 12. The decoder 54 may be configured to decode a predetermined block 18 of the picture. Specifically, the predictor 44 may be configured to map a set of P neighboring samples in the neighborhood of the predetermined block 18 to a set of Q predicted values for the samples of the predetermined block using a linear or affine linear transform [e.g., ALWIP].
[0077] like Figure 5 As shown, a predetermined block 18 includes Q values to be predicted (which will be "predicted values" at the end of the operation). If block 18 has M rows and N columns, then Q=M·N. The Q values of block 18 can be in the spatial domain (e.g., pixels) or in the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of block 18 can be predicted based on P values taken from neighboring blocks 17a to 17c, which are generally adjacent to block 18. The P values of neighboring blocks 17a to 17c can be located closest to block 18 (e.g., adjacent). The P values of neighboring blocks 17a to 17c have already been processed and predicted. The P values are indicated as values in portions 17'a to 17'c to distinguish them from the blocks they are part of (in some examples, 17'b is not used).
[0078] like Figure 6As shown, to perform prediction, a first vector 17P having P entries (each entry is associated with a specific position in the neighboring parts 17'a to 17'c), a second vector 18Q having Q entries (each entry is associated with a specific position in the block 18), and a mapping matrix 17M (each row is associated with a specific position in the block 18, and each column is associated with a specific position in the neighboring parts 17'a to 17'c) can be used. Thus, the mapping matrix 17M performs a prediction of the P values of the neighboring parts 17'a to 17'c into values of the block 18 according to a predetermined pattern. Therefore, the entries in the mapping matrix 17M can be understood as weighting factors. In the following paragraphs, we will use the labels 17a to 17c instead of the labels 17'a to 17'c to refer to the neighboring parts of the boundary.
[0079] In the art, several conventional modes are known, such as DC mode, planar mode and 65 directional prediction modes. For example, 67 modes may be known.
[0080] However, it has been noted that a different mode can also be used, referred to herein as a linear or affine linear transform. The linear or affine linear transform comprises P·Q weighting factors, wherein at least 1 / 4 of the P·Q weighting factors are non-zero weighting values, and for each of the Q predicted values, the non-zero weighting value comprises a series of P weighting factors associated with the corresponding predicted value. This series, when arranged one after another according to a raster scan order between samples of a predetermined block, forms an omnidirectional nonlinear envelope.
[0081] It is possible to map P positions of adjacent values 17'a to 17'c (templates), Q positions of adjacent samples 17'a to 17'c, and to map at the values of P*Q weighting factors of the matrix 17M. The plane is an example of an envelope of a sequence for DC transformation (it is a plane for DC transformation). The envelope is obviously planar and is therefore excluded by the definition of linear or affine linear transformation (ALWIP). Another example is a matrix that produces a simulation of an angular pattern: the envelope would be excluded from the ALWIP definition, and frankly, it looks like a hill that slopes from top to bottom along the direction in the P / Q plane. The planar mode and the 65 directional prediction modes will have different envelopes, but these envelopes will be linear in at least one direction, that is, for example, all directions for the exemplary DC and for example, the hill direction for the angular mode.
[0082] In contrast, the envelope of a linear or affine transformation will not be linear in all directions. It is understood that in some cases this type of transformation may be optimal for performing the prediction of block 18. It is noted that preferably at least 1 / 4 of the weighting factors are different from zero (i.e., at least 25% of the P*Q weighting factors are different from 0).
[0083] According to any conventional mapping rules, the weighting factors may be unrelated to each other. Thus, the matrix 17M may be such that the values of its entries have no obvious identifiable relationship. For example, the weighting factors cannot be described by any analytical function or differential function.
[0084] In an example, the ALWIP transformation is performed such that the mean of the maximum values of the cross-correlations between: a first set of weighting factors associated with the corresponding predicted values, a second set of weighting factors associated with predicted values other than the corresponding predicted values, or an inverse version of the second set (whichever results in a higher maximum value) can be less than a predetermined threshold (e.g., 0.2, 0.3, 0.35, or 0.1, e.g., a threshold within a range between 0.05 and 0.035). For example, for each pair (i1, i2) of rows in the ALWIP matrix 17M, a cross-correlation can be calculated by multiplying the P values in the i1th row by the P values in the i2th row. For each obtained cross-correlation, a maximum value can be obtained. Thus, a mean (average) value for the entire matrix 17M can be obtained (i.e., averaging the maximum values of the cross-correlations in all combinations). The threshold value can then be, for example, 0.2, 0.3, 0.35, or 0.1, e.g., a threshold within a range between 0.05 and 0.035.
[0085] The P adjacent samples of blocks 17a to 17c may be located along a one-dimensional path that extends along a border (e.g., 18c, 18a) of the predetermined block 18. For each of the Q predicted values of the predetermined block 18, a series of P weighting factors associated with the corresponding predicted value may be ordered in a manner that traverses the one-dimensional path along a predetermined direction (e.g., from left to right, from top to bottom, etc.).
[0086] In an example, the ALWIP matrix 17M may be non-diagonal or non-block diagonal.
[0087] An example of an ALWIP matrix 17M for predicting a 4×4 block 18 from 4 already predicted neighboring samples may be:
[0088] {
[0089] {37,59,77,28},
[0090] {32,92,85,25},
[0091] {31,69,100,24},
[0092] {33,36,106,29},
[0093] {24,49,104,48},
[0094]
[0095] (Here, {37, 59, 77, 28} is the first row of matrix 17M; {32, 92, 85, 25} is the second row of matrix 17M; and {61, 32, 54, 100} is the 16th row of matrix 17M.) Matrix 17M has a size of 16×4 and includes 64 weighting factors (as a result of 16*4=64). This is because matrix 17M has a size of Q×P, where Q=M*N, i.e., the number of samples of block 18 to be predicted (block 18 is a 4×4 block), and P is the number of samples of predicted samples. Here, M=4, N=4, Q=16 (as a result of M*N=4*4=16), and P=4. The matrix is non-diagonal and non-block diagonal, and there is no particular rule to describe it.
[0096] It can be seen that less than 1 / 4 of the weighting factors are 0 (in the case of the matrix shown above, one of the 64 weighting factors is zero). When arranged one after another according to a raster scan order, the envelope formed by these values forms an omnidirectional nonlinear envelope.
[0097] Even though the above description is primarily discussed with reference to a decoder (eg, decoder 54), the same operations may be performed at an encoder (eg, encoder 14).
[0098] In some examples, for each block size (in the set of block sizes), ALWIP transforms of intra prediction modes within the second set of intra prediction modes for the corresponding block size are different from each other. Additionally or alternatively, the cardinality of the second set of intra prediction modes for the block sizes in the set of block sizes may coincide, but the associated linear or affine linear transforms of the intra prediction modes within the second set of intra prediction modes for different block sizes may not be convertible into each other by scaling.
[0099] In some examples, ALWIP transforms may be defined in such a way that they share “nothing” with legacy transforms (eg, an ALWIP conversion may share “nothing” with a corresponding legacy transform even though they are mapped via one of the mappings described above).
[0100] In an example, the ALWIP mode is used for both the luma component and the chroma components, but in other examples, the ALWIP mode is used for the luma component but not for the chroma components.
[0101] 5 Affine linearly weighted intra prediction mode with encoder acceleration (e.g., test CE3-1.2.1)
[0102] 5.1 Description of method or apparatus
[0103] The affine linear weighted intra prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as proposed in JVET-L0199 under test CE3-2.2.2 except for the following changes:
[0104] • Coordination with Multiple Reference Line (MRL) intra prediction, especially encoder estimation and signaling, i.e., MRL is not combined with ALWIP and transmission of MRL indices is restricted to non-ALWIP blocks.
[0105] - Subsampling is now mandatory for all blocks with WxH ≥ 32x32 (previously optional for 32x32); therefore, the additional test and sending of the subsampling flag at the encoder has been removed.
[0106] By downsampling to 32×N and N×32 respectively and applying the corresponding ALWIP mode,
[0107] ALWIP for 64xN blocks and Nx64 blocks (N≤32) has been added.
[0108] Additionally, test CE3-1.2.1 includes the following encoder optimizations for ALWIP:
[0109] • Combined mode estimation: Legacy and ALWIP modes use a shared Hadamard candidate list for full RD estimation, i.e., ALWIP mode candidates are added to the same list as legacy (and MRL) mode candidates based on Hadamard cost.
[0110] Combined mode lists support EMT Intra-Frame Fast and PB Intra-Frame Fast, with additional optimizations to reduce the number of full RD checks.
[0111] Following the same approach as the conventional mode, only the available MPMs of the left and above blocks are added to this list for full RD estimation of ALWIP.
[0112] 5.2 Complexity Evaluation
[0113] In test CE3-1.2.1, excluding the computation of the discrete cosine transform, a maximum of 12 multiplications are required per sample to generate the prediction signal. Furthermore, a total of 136,492 parameters are required, each of 16 bits. This corresponds to 0.273 megabytes of memory.
[0114] 5.3 Experimental Results
[0115] The tests were evaluated according to the common test conditions JVET-J1010 [2], using VTM software version 3.0.1 for intra-only (AI) and random access (RA) configurations. The corresponding simulations were performed on an Intel Xeon cluster (E5-2697A v4, AVX2 enabled, Intel Turbo Boost disabled) with the Linux operating system and the GCC 7.2.1 compiler.
[0116] Table 1. Results for CE3-1.2.1 configured for VTM AI
[0117] Y U V Coding Time Decoding time Category A1 -2,08% -1,68% -1,60% 155% 104% Category A2 -1,18% -0,90% -0,84% 153% 103% Category B -1,18% -0,84% -0,83% 155% 104% Category C -0,94% -0,63% -0,76% 148% 106% Category E -1,71% -1,28% -1,21% 154% 106% total -1,36% -1,02% -1,01% 153% 105% Class D -0,99% -0,61% -0,76% 145% 107% Class F (optional) -1,38% -1,23% -1,04% 147% 104%
[0118] Table 2. Results for CE3-1.2.1 configured for VTM RA
[0119] Y U V Coding Time Decoding time Category A1 -1,25% -1,80% -1,95% 113% 100% Category A2 -0,68% -0,54% -0,21% 111% 100% Category B -0,82% -0,72% -0,97% 113% 100% Category C -0,70% -0,79% -0,82% 113% 99% Category E total -0,85% -0,92% -0,98% 113% 100% Class D -0,65% -1,06% -0,51% 113% 102% Class F (optional) -1,07% -1,04% -0,96% 117% 99%
[0120] 5.4 Complexity Reduced Affine Linear Weighted Intra Prediction (e.g., Test CE3-1.2.2)
[0121] The techniques tested in CE2 are related to the “affine linear intra prediction” described in JVET-L0199 [1], but are simplified in terms of memory requirements and computational complexity:
[0122] There may be only three different sets of prediction matrices (e.g. S0, S1, S2, see below) and a bias vector (e.g. to provide offset values) covering all block shapes. As a result, the number of parameters is reduced to 14400 10-bit values, which is less than 128×
[0123] The amount of storage stored in 128CTU is less.
[0124] The input and output sizes of the predictor are further reduced. Furthermore, instead of transforming the boundaries via DCT, boundary samples can be averaged or downsampled, and the prediction signal can be generated using linear interpolation instead of inverse DCT. Consequently, a maximum of four multiplications are required per sample to generate the prediction signal.
[0125] 6. Example:
[0126] Here we discuss how to use ALWIP forecasting to perform some forecasting (e.g. Figure 6 shown).
[0127] In principle, reference Figure 6In order to obtain Q=M*N values of the M×N block 18 to be predicted, the Q*P samples of the Q×PALWIP prediction matrix 17M are multiplied by the P samples of the P×1 neighboring vector 17P. Therefore, in general, in order to obtain each of the Q=M*N values of the M×N block 18 to be predicted, at least P=M+N multiplications are required.
[0128] These multiplications have a very negative effect. The size P of the boundary vector 17P generally depends on the number M+N of boundary samples (binary bins or pixels) 17a, 17c adjacent to (e.g., neighboring) the M×N block 18 to be predicted. This means that if the size of the block 18 to be predicted is large, the number M+N of boundary pixels (17a, 17c) is correspondingly large, thus increasing the size P=M+N of the P×1 boundary vector 17P and the length of each row of the Q×P ALWIP prediction matrix 17M, and thus also increasing the number of necessary multiplications (in general, Q=M*N=W*H, where W (width) is the other sign of N and H (height) is the other sign of M; in the case where the boundary vector consists of only one row and / or column of samples, P is P=M+N=H+W).
[0129] Generally, this problem is exacerbated by the fact that in microprocessor-based systems (or other digital processing systems), multiplication is generally a power-consuming operation. As you can imagine, performing a large number of multiplications on a very large number of samples in a large number of blocks will result in a waste of computing power, which is generally undesirable.
[0130] Therefore, it is preferable to reduce the number of multiplications Q*P required to predict the M×N block 18.
[0131] It has been appreciated that by intelligently selecting an alternative multiplication and an operation that is easier to process, the computational power required for each intra prediction of each block 18 to be predicted can be reduced in some way.
[0132] Specifically, refer to Figures 7.1 to 7.4 It is understood that an encoder or a decoder can predict a predetermined block (e.g., 18) of a picture using a plurality of adjacent samples (e.g., 17a, 17c) by:
[0133] reducing (e.g., at step 811) (e.g., by averaging or downsampling) a plurality of adjacent samples (e.g., 17a, 17c) to obtain a reduced set of sample values, the reduced set of sample values having a smaller number of samples than the plurality of adjacent samples,
[0134] The reduced set of sample values is subjected to (eg, at step 812) a linear or affine linear transformation to obtain predicted values for predetermined samples of the predetermined block.
[0135] In some cases, the decoder or encoder may also derive predicted values of other samples of a predetermined block based on predicted values of predetermined samples and a plurality of adjacent samples, for example, by interpolation. Thus, an upsampling strategy can be obtained.
[0136] In an example, some averaging can be performed on the samples of boundary 17 (e.g., at step 811) to obtain a reduced set 102 of samples with a reduced number of samples (at least one of the samples of the reduced number of samples 102 can be the average of two samples of the original boundary samples or selected from the original boundary samples) Figures 7.1 to 7.4 ). For example, if the original boundary has P = M + N samples, the reduced set of samples can have P red = M red + N red samples, where M red < M and N red < N, such that at least one of them is satisfied, making P red < P. Thus, the boundary vector 17P actually used for prediction (e.g., at step 812b) will not have P × 1 entries, but will have P red × 1 entries, where P red < P. Similarly, the ALWIP prediction matrix 17M selected for prediction will not have a Q × P size, but the number of elements of the matrix is reduced to Q × P red (or Q red × P red , see below), at least because P red < P (by means of M red < M and N red < N, at least one of them).
[0137] In some examples (e.g., Figure 7.2 、 Figure 7.3 ), if the block obtained by ALWIP (at step 812) is a reduced block of size M′ red × N′ red , where M′ red [[ID=asc6]] < M and / or N′ red < N (i.e., the samples directly predicted by ALWIP are fewer in number than the samples of the actual block 18 to be predicted), the number of multiplications can be further reduced. Thus, set Q red = M′ red * N′ red , which will use Q red * P red multiplications instead of Q * P red multiplications (where, Q red * P red red <Q*P) to obtain the ALWIP prediction. This multiplication will perform predictions on the reduced block of size M ′red ×N′ red Although it will be possible to perform (e.g., in subsequent step 813) upsampling from the reduced M′ red ×N′ ref prediction block to the final M×N prediction block (e.g., obtained by interpolation).
[0138] Although matrix multiplication involves a reduction in the number of multiplications (Q red *P red or Q*P red ), both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or even avoiding) multiplications, and these techniques can be advantageous. For example, downsampling, averaging, and / or interpolation can be performed (e.g., in steps 811 and / or 813) by employing binary operations such as addition and shift that have no power requirements for computation.
[0139] In addition, addition is a very easy operation that can be easily performed without a large amount of computational work.
[0140] This shift operation can be used, for example, to average two boundary samples and / or to interpolate two samples (support values) of the reduced predicted block (or taken from the boundary) to obtain the final prediction block. (For interpolation, two sample values are required. Inside the block, we always have two predetermined values, but to interpolate samples along the left and upper borders of the block, we only have one predetermined value, as Figure 7.2 shown, so we use the boundary samples as support values for interpolation.)
[0141] A two-step process can be used, for example:
[0142] First, sum the values of the two samples;
[0143] Then, halve the sum value (e.g., by right shift).
[0144] Alternatively, it is possible to:
[0145] First, halve each of the samples (e.g., by left shift);
[0146] Then, sum the values of the two halved samples.
[0147] When performing downsampling (e.g., in step 811), since only one sample needs to be selected from a set of samples (e.g., samples adjacent to each other), easier operations can be performed.
[0148] Thus, it is now possible to define techniques for reducing the number of multiplications to be performed. Some of these techniques can be based, in particular, on at least one of the following principles:
[0149] Even if the actual size of the block 18 to be predicted is M×N, the block can be reduced (in at least one of the two dimensions) and an ALWIP matrix with a reduced size of Q red ×P red can be applied (where Q red = M′ red * N′ red , P red = N red + M red , where M red < M and / or N′ red < N and / or M red < M and / or N red < N). Thus, the boundary vector 17P will have a size of P red ×1, meaning that there are only P red < P multiplications (where P red = M red + N red and P = M + N).
[0150] P red ×1 boundary vector 17P can be easily obtained from the original boundary 17, for example:
[0151] By downsampling (e.g., by only selecting some samples of the boundary); and / or
[0152] By averaging multiple samples of the boundary (which can be easily obtained by addition and shifting, without multiplication).
[0153] Additionally or alternatively, instead of predicting all Q = M*N values of the block 18 to be predicted by multiplication, it is possible to predict only the reduced block with the reduced size (e.g., Q red = M′ red * N′ red , where M′ <000In the example shown, to predict a 4×4 block 18 (M = 4, N = 4, Q = M*N = 16) and the neighborhoods 17 of samples 17a (a vertical column with 4 predicted samples) and 17c (a horizontal row with 4 predicted samples) have been predicted in previous iterations (neighborhoods 17a and 17c can be jointly indicated by 17). A priori, by using Figure 6 the equation shown, the prediction matrix 17M should be a Q×P = 16×8 matrix (by virtue of Q = M*N = 4*4 and P = M+N = 4+4 = 8), and the size of the boundary vector 17P should be 8×1 (by virtue of P = 8). However, this would result in 8 multiplications being performed for each of the 16 samples of the 4×4 block 18 to be predicted, resulting in a total of 16*8 = 128 multiplications being required. (Note that the average number of multiplications per sample is a good assessment of the computational complexity. For traditional intra prediction, each sample requires four multiplications, which increases the computational workload involved. Therefore, this can potentially be used as an upper bound for ALWIP to ensure that the complexity is reasonable and does not exceed that of traditional intra prediction.)
[0155] Nevertheless, it has been understood that by using the present technique, the number of samples 17a and 17c adjacent to the block 18 to be predicted can be reduced from P to P red <P at step 811. Specifically, it has been understood that the boundary samples (17a, 17c) adjacent to each other can be averaged (e.g., at 100 in Figure 7.1 ) to obtain a reduced boundary 102 with two horizontal rows and two vertical columns, and thus the block 18 is operated as a 2×2 block (the reduced boundary is formed by the averaged values). Alternatively, downsampling can be performed, so two samples are selected for the row 17c and two samples are selected for the column 17a. Thus, the horizontal row 17c, instead of having four original samples, is processed as having two samples (e.g., the averaged samples), and the vertical column 17a, which originally had four samples, is processed as having two samples (e.g., the averaged samples). It can also be understood that after subdividing the row 17c and the column 17a into groups 110 each having two samples, a single sample (e.g., the average of the samples in the group 110 or a simple selection between the samples in the group 110) is maintained. Thus, by virtue of a set 102 having only four samples (M red = 2, N red = 2, P red = M red + N red =4, where P red <P), a so-called reduced set 102 of sample values is obtained.
[0156] It will be appreciated that operations (e.g. averaging or downsampling 100) can be performed without performing too many multiplications at the processor level: the averaging or downsampling 100 performed in step 811 can simply be obtained by direct and computationally non-power-consuming operations such as addition and shifting.
[0157] It is understood that at this point it is possible (for example, using a Figure 6 The reduced set of sample values 102 is subjected to a linear or affine linear (ALWIP) transform 19 (e.g., a prediction matrix such as the matrix 17M of FIG. 17 ). In this case, the ALWIP transform 19 directly maps the four samples 102 onto the sample values 104 of the block 18. No interpolation is required in the present case.
[0158] In this case, the size of the ALWIP matrix 17M is Q×P red =16×4: This follows from the fact that all Q=16 samples of block 18 to be predicted are directly obtained by ALWIP multiplication (no interpolation is required).
[0159] Therefore, in step 812a, the size is selected to be Q×P red The selection may be based at least in part on signaling from, for example, the data stream 12. The selected ALWIP matrix 17M may also be selected using A k Indication, where k can be understood as an index, which can be signaled in the data stream 12 (in some cases, the matrix is also indicated as See below.) The selection may be performed according to the following scheme: for each size (e.g. a height / width pair of the block 18 to be predicted), an ALWIP matrix 17M is selected from, for example, one of three sets S0, S1, S2 of matrices (each of the three sets S0, S1, S2 may group multiple ALWIP matrices 17M of the same size, and the ALWIP matrix selected for the prediction will be one of them).
[0160] In step 812b, the selected Q×P red ALWIP matrix 17M (also indicated as A k ) and P red ×1 multiplication between boundary vectors 17P.
[0161] In step 812c, the offset value (eg, b k ) is added to all obtained values 104 of the vector 18Q obtained, for example, by ALWIP. The value of the offset (b k Or in some cases instructions, see below) can be combined with a specific selected ALWIP matrix (A k) and can be based on an index (for example, the index can be signaled in the data stream 12).
[0162] So here's a comparison between using this technique and not using it again:
[0163] Without using this technology:
[0164] The size of the block 18 to be predicted is M=4, N=4;
[0165] Q = M*N = 4*4 = 16 values to be predicted;
[0166] P = M + N = 4 + 4 = 8 boundary samples
[0167] For each of the Q=16 values to be predicted, P=8 multiplications,
[0168] A total of P*Q=8*16=128 multiplications;
[0169] By using this technology, we have:
[0170] The size of the block 18 to be predicted is M=4, N=4;
[0171] Finally, we need to predict Q = M * N = 4 * 4 = 16 values;
[0172] Reduced size of the boundary vector: P red =M red +N red =2+2=4;
[0173] For each of the Q=16 values to be predicted by ALWIP,
[0174] P red = 4 multiplications,
[0175] Total P red * Q = 4 * 16 = 64 multiplications (half of 128!)
[0176] The ratio between the number of multiplications and the number of final values to be obtained is P red *Q / Q = 4, i.e. half of P = 8 multiplications for each sample to be predicted!
[0177] It will be appreciated that appropriate values can be obtained at step 812 by relying on straightforward and computationally inexpensive operations such as averaging (and, where appropriate, adding and / or shifting and / or downsampling).
[0178] refer to Figure 7.2, the block 18 to be predicted is here an 8×8 block of 64 samples (M=8, N=8). Here, a priori, the size of the prediction matrix 17M should be Q×P=64×16 (Q=64, by virtue of Q=M*N=8*8=64, M=8 and N=8, and by virtue of P=M+N=8+8=16). Therefore, a priori, for each of the Q=64 samples of the 8×8 block 18 to be predicted, P=16 multiplications will be required, resulting in 64*16=1024 multiplications for the entire 8×8 block 18!
[0179] However, if Figure 7.2 As can be seen, a method 820 can be provided according to which, instead of using all 16 samples of the boundary, only 8 values are used (e.g., 4 values in the horizontal boundary row 17c and 4 values in the vertical boundary column 17a between the original samples of the boundary). From the boundary row 17c, 4 samples can be used instead of 8 samples (e.g., they can be the average of a pair of samples and / or one sample selected from two samples). Therefore, the boundary vector is not a P×1=16×1 vector, but only P red ×1=8×1 vector (P red =M red +N red =4+4). It is understood that the samples of the horizontal rows 17c and the samples of the vertical columns 17a can be selected or averaged (eg, one pair at a time) to have only P red = 8 boundary values, instead of having the original P = 16 samples, thereby forming a reduced set 102 of sample values. This reduced set 102 will allow obtaining a reduced version of the block 18, which has Q red =M red *N red = 4*4 = 16 samples (instead of Q = M*N = 8*8 = 64). ALWIP matrix can be applied to predict the size of M red ×N red = 4×4 blocks. The reduced version of block 18 is included in Figure 7.2 The samples indicated in grey in the scheme 106: The samples indicated in grey squares (including the sample 118' and the sample 118") form a 4x4 reduced block having the Q obtained in the step of performing 812. red = 16 values. The 4x4 reduced block has been obtained by applying the linear transformation 19 in step 812. After obtaining the values of the 4x4 reduced block, the values of the remaining samples (samples indicated by white samples in scheme 106) can be obtained, for example, by interpolation.
[0180] Relative to Figure 7.1The method 810 may further include step 813: deriving the residual QQ of the M×N=8×8 block 18 to be predicted, for example, by interpolation. red =64-16=predicted values of 48 samples (white squares). The remaining QQ red = 64 - 16 = 48 samples can be obtained directly from Q by interpolation (eg, interpolation can also use boundary samples). red = 16 samples to obtain. Figure 7.2 As can be seen in FIG, although samples 118′ and 118″ have been obtained in step 812 (as indicated by the gray squares), sample 108′ (which is intermediate between sample 118′ and sample 118″ and indicated by the white square) is obtained in step 813 by interpolation between samples 118′ and 118″. It is understood that interpolation can also be obtained by operations similar to those used for averaging, such as shifting and adding. Therefore, in Figure 7.2 In the example embodiment, value 108' may generally be determined as a midpoint (which may be an average) between the value of sample 118' and the value of sample 118".
[0181] By performing interpolation, at step 813 , it is also possible to obtain a final version of the M×N=8×8 block 18 based on the plurality of sample values indicated in 104 .
[0182] Therefore, the comparison between using this technology and not using this technology is:
[0183] Without using this technology:
[0184] The size of the block 18 to be predicted is M=8, N=8, and there are Q=M*N=8*8=64 samples in the block 18 to be predicted;
[0185] P = M + N = 8 + 8 = 16 samples in boundary 17;
[0186] For each of the Q=64 values to be predicted, P=16 multiplications,
[0187] A total of P*Q=16*64=1028 multiplications
[0188] The ratio between the number of multiplications and the number of final values to be obtained is P*Q / Q=16
[0189] When using this technology:
[0190] The size of the block 18 to be predicted is M=8, N=8;
[0191] Finally, we need to predict Q = M * N = 8 * 8 = 64 values;
[0192] But to use Q red ×Pred An ALWIP matrix, where P red = M red + N red ,
[0193] Q red = M red * N red M red = 4, N red = 4
[0194] P in the boundary red = M red + N red = 4 + 4 = 8 samples, where P red < P
[0195] For each of the 16 values of Q for the 4×4 reduced block (formed by the gray squares in Scenario 106) to be predicted, P red = 8 multiplications, red In total, P [[ID=3�]] red
[0196] red * Q red = 8 * 16 = 128 multiplications (much smaller than 1024!) The ratio between the number of multiplications and the number of final values to be obtained is P red * Q red / Q = 128 / 64 = 2 (much smaller than 16 obtained without using this technology!)
[0197] Thus, the power requirement of the technology presented here is 1 / 8 of the prior art.
[0198] Figure 7.3 Another example (which can be based on Method 820) is shown, where the block eighteen to be predicted is a rectangular 4×8 block (M = 8, N = 4), with Q = 4 * 8 = 32 samples to be predicted. The boundary seventeen is formed by a horizontal row 17c with N = 8 samples and a vertical column 17a with M = 4 samples. Thus, a priori, the size of the boundary vector 17P will be P×1 = 12×1, and the predicted ALWIP matrix should be a Q×P = 32×12 matrix, so Q * P = 32 * 12 = 384 multiplications are required.
[0199] However, it is possible to, for example, average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row with only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as it is (e.g., not averaged). Overall, the size of the reduced boundary will be P red = 8, where P red < P. Thus, the boundary vector 17P will have size P red×1=8×1. The ALWIP prediction matrix 17M will be of size M*N red *P red = 4*4*8 = 64. The 4×4 reduced block (formed by the gray columns in solution 107) directly obtained when performing step 812 will have a size of Q red =M*N red = 4*4 = 16 samples (instead of Q = 4*8 = 32 of the original 4x8 block 18 to be predicted). Once the downscaled 4x4 block is obtained by ALWIP, the offset value b can be added k (Step 812c) and interpolation is performed in step 813. Figure 7.3 As can be seen from step 813 in FIG. 1 , the reduced 4×4 block is expanded to a 4×8 block 18 , wherein the value 108 ′ not obtained in step 812 is obtained in step 813 by interpolating the values 118 ′ and 118 ″ (grey square) obtained in step 812 .
[0200] Therefore, the comparison between using this technology and not using this technology is:
[0201] Without using this technology:
[0202] The size of the block 18 to be predicted is M=4, N=8
[0203] Q = M * N = 4 * 8 = 32 values to be predicted;
[0204] P = M + N = 4 + 8 = 12 samples in the boundary;
[0205] For each of the Q=32 values to be predicted, P=12 multiplications,
[0206] A total of P*Q=12*32=384 multiplications
[0207] The ratio between the number of multiplications and the number of final values to be obtained is P*Q / Q=12
[0208] When using this technology:
[0209] The size of the block 18 to be predicted is M=4, N=8
[0210] Finally, we need to predict Q = M * N = 4 * 8 = 32 values;
[0211] But you can use Q red ×P red =16×8 ALSIP matrix, where M=4,
[0212] N red =4,Q red =M*N red=16, P red =M+N red =4+4=8
[0213] Boundary P red =M+N red =4+4=8 samples, where P red <P
[0214] For the Q of the reduced block to be predicted red = Each of the 16 values,
[0215] P red = 8 multiplications,
[0216] Total Q red *P red =16*8=16*8=128 multiplications (less than 384!)
[0217] The ratio between the number of multiplications and the number of final values to be obtained is P red *Q red / Q = 128 / 32 = 4 (much smaller than the 12 obtained without using this technique!).
[0218] Therefore, using this technique, the computational effort is reduced to one third.
[0219] Figure 7.4 The case where the block 18 to be predicted has a size of M×N=16×16 and finally has Q=M*N=16*16=256 values to be predicted is shown, and the block has P=M+N=16+16=32 boundary samples. This will result in a prediction matrix of size Q×P=256×32, which means 256*32=8192 multiplications!
[0220] However, by applying the method 820, the number of boundary samples can be reduced (e.g., by averaging or downsampling) in step 811, for example, from 32 to 8: for example, for each group 120 of four consecutive samples in row 17a, a single sample is retained (e.g., selected from the four samples, or the average of the samples). Furthermore, for each group 120 of four consecutive samples in column 17c, a single sample is retained (e.g., selected from the four samples, or the average of the samples).
[0221] Here, the ALWIP matrix 17M is Q red ×P red =64×8 matrix: This is because: P has been selected red =8 (by using 8 averaged samples or selected samples of the 32 samples from the boundary), and the reduced block to be predicted in step 812 is an 8×8 block (in scheme 109, the grey squares are 64).
[0222] Therefore, once the 64 samples of the reduced 8×8 block are obtained at step 812, the residual QQ of the block to be predicted 18 can be derived at step 813. red =256-64=192 values 104.
[0223] In this case, for performing the interpolation, a choice has been made to use all samples of the boundary columns 17a and only the alternative samples in the boundary rows 17c. Other choices may be made.
[0224] In the case of this method, the ratio between the number of multiplications and the number of values finally obtained is Q red *P red / Q = 8*64 / 256 = 2, which is much less than 32 multiplications per value without using this technique!
[0225] The comparison between using this technique and not using this technique is:
[0226] Without using this technology:
[0227] The size of the block 18 to be predicted is M=16, N=16
[0228] Q = M * N = 16 * 16 = 256 values to be predicted;
[0229] P = M + N = 16 * 16 = 32 samples in the boundary;
[0230] For each of the Q=256 values to be predicted, P=32 multiplications,
[0231] A total of P*Q=32*256=8192 multiplications;
[0232] The ratio between the number of multiplications and the number of final values to be obtained is P*Q / Q=32
[0233] When using this technology:
[0234] The size of the block 18 to be predicted is M=16, N=16
[0235] Finally, we need to predict Q = M * N = 16 * 16 = 256 values;
[0236] But to use Q red ×P red =64×8ALWIP matrix, where M red =4,
[0237] N red =4, to predict Q through ALWIP red =8*8=64 samples,
[0238] P red =M red +N red =4+4=8
[0239] Boundary P red =M red +N red =4+4=8 samples, where P red <P
[0240] For the Q of the reduced block to be predicted red = each of the 64 values,
[0241] P red = 8 multiplications,
[0242] Total Q red *P red = 64 * 4 = 16 * 8 = 256 multiplications (less than 8192!) The ratio between the number of multiplications and the number of final values to be obtained is P red *Q red / Q=8*64 / 256=2 (much smaller than the 32! obtained without using this technique).
[0243] Therefore, the computing power required by this technology is 1 / 16 of that of traditional technology!
[0244] Therefore, a predetermined block (18) of a picture can be predicted using a plurality of neighboring samples (17) by the following operation:
[0245] reducing (100, 813) the plurality of adjacent samples (17) to obtain a reduced set (102) of sample values that is smaller in number of samples than the plurality of adjacent samples (17),
[0246] A linear or affine linear transformation (19, 17M) is performed (812) on the reduced set of sample values to obtain predicted values for predetermined samples (104, 118', 188") of the predetermined block (18).
[0247] Specifically, the reduction (100, 813) can be performed by downsampling the plurality of adjacent samples to obtain a reduced set of sample values (102) that is smaller in number of samples than the plurality of adjacent samples (17).
[0248] Alternatively, the reduction (100, 813) can be performed by averaging a plurality of adjacent samples to obtain a reduced set of sample values (102) that is smaller in number of samples than the plurality of adjacent samples (17).
[0249] Furthermore, predicted values for other samples (108, 108') of the predetermined block (18) can be derived (813) by interpolation based on the predicted value of the predetermined sample (104, 118', 118") and a plurality of neighboring samples (17).
[0250] The plurality of adjacent samples (17a, 17c) may be arranged one-dimensionally (e.g., in Figures 7.1 to 7.4 The predetermined samples (e.g., the predetermined samples obtained by ALWIP in step 812) may also be arranged in rows and columns and along at least one of the rows and columns, and the predetermined samples may be located at every n-th position starting from the sample (112) adjacent to both sides of the predetermined block 18 among the predetermined samples 112.
[0251] Based on the plurality of adjacent samples (17), for each of at least one of the rows and columns, a support value (118) of one adjacent position (118) of the plurality of adjacent positions can be determined, the support value (118) being aligned with a corresponding one of the at least one of the rows and columns. Predicted values 118 of other samples (108, 108') of the predetermined block (18) can also be derived by interpolation based on the predicted value of the predetermined sample (104, 118', 118") and the support values of the adjacent samples (118) aligned with at least one of the rows and columns.
[0252] The predetermined sample (104) may be located at every n-th position along the row starting from the sample (112) adjacent to both sides of the predetermined block 18, and the predetermined sample may be located at every m-th position along the column starting from the sample (112) adjacent to both sides of the predetermined block (18), where n, m>1. In some cases, n=m (e.g., in Figure 7.2 and Figure 7.3 , where samples 104 , 118 ′, 118 ″ obtained directly by ALWIP at 812 and indicated by grey squares alternate along rows and columns with samples 108 , 108 ′ obtained subsequently at step 813 ).
[0253] Along at least one of the rows (17c) and columns (17a), the determination of the support value may be performed for each support value, for example, by downsampling or averaging (122) a set (120) of adjacent samples within the plurality of adjacent samples, the set of adjacent samples including the adjacent sample (118) for which the corresponding support value is determined. Figure 7.4 In step 813 , the value of the sample 119 can be obtained by using the predetermined sample 118 ′″ (previously obtained in step 812 ) and the values of the neighboring samples 118 as support values.
[0254] The plurality of adjacent samples may extend one-dimensionally along both sides of the predetermined block (18). The reduction (811) can be performed by grouping the plurality of adjacent samples (17) into one or more groups (110) of consecutive adjacent samples and downsampling or averaging each adjacent sample in the one or more groups (110) of adjacent samples, the one or more groups (110) of adjacent samples having two or more adjacent samples.
[0255] In an example, a linear or affine linear transformation may include P red *Q red or P red *Q weighting factors, where P red is the number of sample values (102) in the reduced set of sample values, Q red Or Q is the number of predetermined samples within a predetermined block (18). At least 1 / 4P red *Q red or 1 / 4P red *Q weighting factors are non-zero weight values. For Q or Q red For each of the predetermined samples, P red *Q red or P red The Q weighting factors may include a series of P values associated with the corresponding predetermined samples. red weighting factors, wherein the series, when arranged one after another according to a raster scan order in predetermined samples of a predetermined block (18), forms an omnidirectional nonlinear envelope. red *Q or P red *Q red The weighting factors may be uncorrelated with each other via any conventional mapping rule. The mean of the maximum values of the cross-correlations between the first series of weighting factors associated with the corresponding predetermined sample, the second series of weighting factors associated with predetermined samples other than the corresponding predetermined sample, or the inverse version of the latter series (whichever results in a higher maximum value) is less than a predetermined threshold. The predetermined threshold may be 0.3 [or 0.2 or 0.1 in some cases]. red adjacent samples (17) can be located along a one-dimensional path extending along both sides of a predetermined block (18), and for Q or Q red A series of P related to the corresponding predetermined samples red The weighting factors are sorted in a manner to traverse the one-dimensional path along a predetermined direction.
[0256] 6.1 Description of methods and apparatus
[0257] To predict samples of a rectangular block of width W (also denoted by N) and height H (also denoted by M), affine linear weighted intra prediction (ALWIP) can take as input a strip of H reconstructed neighboring boundary samples to the left of the block and a strip of W reconstructed neighboring boundary samples above the block. If reconstructed samples are not available, they can be generated as in traditional intra prediction.
[0258] The generation of the prediction signal (e.g., the value for the complete block 18) may be based on at least some of the following three steps:
[0259] 1. Among the boundary samples 17, samples 102 (eg, four samples in the case of W=H=4 and / or eight samples in other cases) may be extracted by averaging or downsampling (eg, step 811).
[0260] 2. A matrix-vector multiplication may be performed on the averaged samples (or the remaining samples after downsampling) as input, followed by addition of an offset. The result may be a reduced prediction signal over a subsampled set of samples in the original block (e.g., step 812).
[0261] 3. The prediction signals at the remaining positions may be generated, for example, by upsampling (eg, by linear interpolation) based on the prediction signals on the subsampled set (eg, step 813).
[0262] Due to step 1 (811) and / or step 3 (813), the total number of multiplications required to calculate the matrix-vector product can be made so that it is always less than or equal to 4*W*H. In addition, the averaging operation on the boundary and the linear interpolation of the reduced prediction signal are performed only by using additions and shifts. In other words, in the example, at most four multiplications are required per sample in the ALWIP mode.
[0263] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b) required to generate the prediction signal are k ) can be taken from a set of matrices that can be stored in the storage units of the decoder and encoder (for example, three sets), such as S0, S1, S2.
[0264] In some examples, the set S0 may include (eg, contain) n0 (eg, n0=16 or n0=18 or another number) matrices i∈{0,…,n0-1}, each of which can have 16 rows and 4 columns and 18 offset vectors i∈{0,...,n0-1} (the size of each offset vector is 16) to perform the Figure 7.1 The set of matrices and offset vectors is used for blocks of size 4×4 18. Once the boundary vector has been reduced to P red=4 vector (such as Figure 7.1 Step 811) can be used to reduce the number of samples in the set 102 red =4 samples are directly mapped to Q = 16 samples of the 4×4 block 18 to be predicted.
[0265] In some examples, set S1 may include (eg, contain) n1 (eg, n1=8 or n1=18 or another number) matrices i∈{0,…,n1-1}, each of which can have 16 rows and 8 columns and 18 offset vectors i∈{0,…,n1-1} (the size of each offset vector is 16) to perform the Figure 7.2 or Figure 7.3 The matrix and offset vectors of the set S1 can be used for blocks of size 4×8, 4×16, 4×32, 4×64, 16×4, 32×4, 64×4, 8×4 and 8×8. In addition, it can also be used for blocks of size WxH, where max(W,H)>4 and min(W,H)=4, i.e. for blocks of size 4×16 or 16×4, 4×32 or 32×4 and 4×64 or 64×4. The 16×8 matrix refers to a reduced version of block 18, which is a 4×4 block, as in Figure 7.2 and Figure 7.3 Obtained in.
[0266] Additionally or alternatively, the set S2 may include (eg, contain) n2 (eg, n2=6 or n2=18 or another number) matrices i∈{0,...,n2-1}, each of which can have 64 rows and 8 columns and 18 offset vectors i∈{0,…,n2-1} (the size of the offset vector is 64). The 64×8 matrix refers to a reduced version of block 18, which is an 8×8 block, e.g., as in Figure 7.4 The set of matrices and offset vectors can be used for blocks of size 8×16, 8×32, 8×64, 16×8, 16×16, 16×32, 16×64, 32×8, 32×16, 32×32, 32×64, 64×8, 64×16, 64×32, 64×64.
[0267] This set of matrices and offset vectors or a portion of these matrices and offset vectors can be used for all other block shapes.
[0268] 6.2 Averaging or Downsampling of Boundaries
[0269] Here, features regarding step 811 are provided.
[0270] As explained above, the boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0271] In a first step, the input boundaries bdry top (e.g., 17c) and bdry left (e.g., 17a) can be reduced to smaller boundaries and to obtain a reduced set 102. Here, in the case of a 4×4 block, and both consist of 2 samples, and in other cases, both consist of 4 samples.
[0272] In the case of a 4×4 block, it is possible to define:
[0273]
[0274] And similarly define Thus, is the average value obtained, for example, using a shift operation.
[0275] In all other cases (e.g., for a block with a width or height not equal to 4), if the block width W is W = 4 * 2 k , then for 0 ≤ i < 4, it is defined:
[0276]
[0277] And similarly define
[0278] In some other cases, the boundary can be downsampled (e.g., by selecting a specific boundary sample from a set of boundary samples) to obtain a reduced number of samples. For example, can be selected from bdry top [0] and bdry top [1] And can be selected from bdry top [2] and bdry top [3] It is also possible to similarly define
[0279] Two reduced boundaries and can be cascaded to a reduced boundary vector bdry red (associated with the reduced set 102), also denoted by 17P. The reduced boundary vector bdry red For a shape of 4×4( Figure 7.1For example) the block size can therefore be 4 (P red =4), and for all other shapes ( Figures 7.2 to 7.4 For example, the block size of P red =8).
[0280] Here, if mode < 18 (or the number of matrices in the set of matrices), then it is possible to define:
[0281]
[0282] If mode ≥ 18, which corresponds to the transposed mode of mode-17, it is possible to define:
[0283]
[0284] Therefore, according to a specific state (one state: mode<18; another state: mode≥18), it is possible to scan along different scan orders (for example, one scan order: Another scanning order: ) assigns the predicted value of the output vector.
[0285] Other strategies can be implemented. In other examples, the mode index "mode" does not have to be in the range of 0 to 35 (other ranges can be defined). In addition, each of the three sets S0, S1, S2 does not have to have 18 matrices (so instead of an expression like mode≥18, mode≥n0,n 1, n2, which is the number of matrices in each set of matrices S0, S1, S2 respectively). In addition, each set can each have a different number of matrices (for example, it can be: S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0286] The mode and transposition information do not have to be stored and / or transmitted as one combined mode index "mode": in some examples, the transposition flag and matrix index (0-15 for S0, 0-7 for S1, and 0-5 for S2) can be signaled explicitly.
[0287] In some cases, the combination of the transposition flag and the matrix index can be interpreted as a set index. For example, there may be one bit used as a transposition flag and some bits indicating the matrix index, collectively referred to as a "set index."
[0288] 5.4 Generating Reduced Prediction Signals via Matrix-Vector Multiplication
[0289] Here, features regarding step 812 are provided.
[0290] In the reduced input vector bdry red(boundary vector 17P) can generate a reduced prediction signal pred red The latter signal can be of width W red , height is H red Here, W red and H red It can be defined as:
[0291] If max(W,H)≤8, then W red =4,H red =4,
[0292] Otherwise, W red =min(W,8),H red =min(H,8).
[0293] The reduced prediction signal pred can be calculated by calculating the matrix-vector product and adding the offset red :
[0294] pred red =A·bdry red +b.
[0295] Here, A is a matrix (e.g., prediction matrix 17M), and if W=H=4, the matrix may have W red *H red rows and 4 columns, and in all other cases has 8 columns, b is of size W red *H red vector.
[0296] If W=H=4, then A may have 4 columns and 16 rows, and therefore, each sample may require 4 multiplications to compute pred in this case. red In all other cases, A can have 8 columns, and it can be verified that in these cases it has 8*W red *H red ≤4*W*H, that is, in these cases, each sample also requires at most 4 multiplications to calculate pred red .
[0297] The matrix A and vector b can be taken from one of the following sets S0, S1, S2. The index idx = idx(W,H) is defined by setting idx(W,H): if W = H = 4, then set idx(W,H) = 0; if max(W,H) = 8, then set idx(W,H) = 1, and in all other cases set idx(W,H) = 2. In addition, if mode < 18, then m = mode, otherwise m = mode - 17. Then, if idx ≤ 1 or idx = 2 and min(W,H) > 4, then and When idx = 2 and min(W, H) = 4, let A be the matrix generated by removing every other row from . When W = 4, it corresponds to the odd x - coordinates in the down - sampling block, or when H = 4, it corresponds to the odd y - coordinates in the down - sampling block. If mode ≥ 18, replace the reduced prediction signal with its transposed signal. In an alternative example, a different strategy can be executed. For example, instead of reducing the size of a larger matrix (“removing”), use a smaller matrix S1 (idx = 1) with W red = 4 and H red = 4. That is, such a block is now assigned to S1 instead of S2.
[0298] Other strategies can be executed. In other examples, the mode index “mode” does not have to be in the range of 0 to 35 (other ranges can be defined). Additionally, each of the three sets S0, S1, S2 does not have to have 18 matrices (thus, instead of an expression like mode < 18, one can use mode < n0, n1, n2, which are the numbers of matrices in each set of matrices S0, S1, S2 respectively). Additionally, each set can have a different number of matrices (for example, it can be: S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0299] 6.4 Linear Interpolation to Generate the Final Prediction Signal
[0300] Here, the features regarding step 812 are provided.
[0301] Interpolation of the subsampled prediction signal may require a second version of the averaged boundary on large blocks. That is, if min(W, H)>8 and W ≥ H, write W = 8 * 2 l , and for 0 ≤ i < 8, define:
[0302]
[0303] If min(W, H)>8 and H > W, define similarly
[0304] Additionally or alternatively, it may be “difficult to subsample”, where is equal to: <00red Perform linear interpolation to produce (for example, Figures 7.2 to 7.4 In some examples, if W=H=4 (e.g., Figure 7.1 ), then this linear interpolation may be unnecessary.
[0308] Linear interpolation can be given as follows (although other examples are possible). Assume W ≥ H. Then, if H > H red , you can execute pred red In this case, pred red A row can be extended to the top as follows. If W = 8, then pred red Can have width W red = 4 and can be passed through the averaged boundary signal Extend to the top, e.g., as defined above. If W>8, then pred red The width is W red = 8 and it passes the averaged boundary signal Extend to the top, e.g. as defined above. For pred red The first line of pred red [x][-1]. Then, the width is W red And the height is 2*H red The signal on the block It can be given as follows:
[0309]
[0310] where 0≤x <W red and 0≤y <H red The following process can be repeated k times until 2 k *H red =H. Therefore, if H=8 or H=16, this can be done at most once. If H=32, this can be done twice. If H=64, this can be done three times. Next, the horizontal upsampling operation can be applied to the result of the vertical upsampling. The subsequent upsampling operation can use the entire boundary on the left side of the prediction signal. Finally, if H>W, a similar process can be performed by first upsampling in the horizontal direction (if necessary) and then in the vertical direction.
[0311] This is an example of interpolation using reduced boundary samples for the first interpolation (horizontally or vertically) and the original boundary samples for the second interpolation (vertically or horizontally). Depending on the block size, only the second interpolation is required or no interpolation is required. If both horizontal and vertical interpolation are required, the order depends on the width and height of the block.
[0312] However, different techniques may be implemented: for example, the original boundary samples may be used for both the first and second interpolation, and the order may be fixed, such as first horizontal then vertical (in other cases, first vertical then horizontal).
[0313] Therefore, the interpolation order (horizontal / vertical) and the use of reduced boundary samples / original boundary samples can be changed.
[0314] 6.5 Example of the entire ALWIP process
[0315] against Figures 7.1 to 7.4 The different shapes in ,show the whole process of averaging, matrix-vector multiplication and linear interpolation.,Note that the remaining shapes will be treated as one of the,depicted cases.
[0316] 1. Given a 4×4 block, ALWIP can be implemented by using Figure 7.1 The technique takes two averages along each axis of the boundary. The resulting four input samples enter the matrix-vector multiplication. The matrix is taken from the set S0. After adding the offset, this can produce 16 final prediction samples. No linear interpolation is required to generate the prediction signal. Therefore, a total of (4*16) / (4*4)=4 multiplications are performed for each sample. See for example Figure 7.1 .
[0317] 2. Given an 8×8 block, ALWIP can take four averages along each axis of the boundary. By using Figure 7.2 The technique generates eight input samples that enter the matrix-vector multiplication. The matrix is taken from the set S1. This results in 16 samples at odd positions in the prediction block. Therefore, a total of (8*16) / (8*8)=2 multiplications are performed for each sample. After adding the offset, these samples can be interpolated, for example, by vertically interpolating using the top boundary and, for example, by horizontally interpolating using the left boundary. See for example Figure 7.2 .
[0318] 3. Given an 8×4 block, ALWIP can be implemented by using Figure 7.3 The technique takes four average values along the horizontal axis of the boundary and four original boundary values on the left boundary. The resulting eight input samples enter the matrix-vector multiplication. The matrix is taken from set S1. This produces 16 samples at odd horizontal positions and every vertical position of the prediction block. Therefore, a total of (8*16) / (8*4)=4 multiplications are performed for each sample. For example, after adding the offset, these samples will be horizontally interpolated using the left boundary. See, for example Figure 7.3 .
[0319] Handle the transposed case accordingly.
[0320] 4. Given a 16×16 block, ALWIP can take four averages along each axis of the boundary. By using Figure 7.2 The eight input samples generated by the technique are fed into a matrix-vector multiplication. The matrix is taken from set S2. This generates 64 samples at odd positions in the prediction block. Therefore, a total of (8*64) / (16*16)=2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically using the top boundary and horizontally using the left boundary. See for example Figure 7.2 . See for example Figure 7.4 .
[0321] For larger shapes the process is probably essentially the same, and it's easy to check that the number of multiplications per sample is less than two.
[0322] For Wx8 blocks, only horizontal interpolation is required since samples are given at odd horizontal positions and every vertical position. Thus, in these cases, at most (8*64) / (16*8)=4 multiplications are performed per sample.
[0323] Finally, for W×4 blocks, where W>8, let A k is a matrix generated by removing each row corresponding to an odd-numbered entry along the horizontal axis of the downsampled block. Thus, the output size can be 32, and again, only horizontal interpolation remains to be performed. A maximum of (8*32) / (16*4)=4 multiplications can be performed per sample.
[0324] The transposed case can be handled accordingly.
[0325] 6.6 Required Parameter Number and Complexity Evaluation
[0326] The parameters required for all possible proposed intra prediction modes can be comprised of matrices and offset vectors belonging to the sets S0, S1, and S2. All matrix coefficients and offset vectors can be stored as 10-bit values. Therefore, based on the above description, the proposed method may require a total of 14,400 parameters, each with 10 bits of precision. This corresponds to 0.018 megabytes of memory. It should be noted that, currently, a 128×128 CTU in standard 4:2:0 chroma subsampling consists of 24,576 values, each 10 bits. Therefore, the memory requirements of the proposed intra prediction tool do not exceed those of the current picture reference tool adopted at the previous meeting. Furthermore, it should be noted that due to the PDPC tool or the 4-tap interpolation filter used for angular prediction modes with fractional angular positions, conventional intra prediction modes require four multiplications per sample. Therefore, in terms of operational complexity, the proposed method does not exceed conventional intra prediction modes.
[0327] 6.7 Signaling of the Proposed Intra Prediction Mode
[0328] For example, for luma blocks, 35 ALWIP modes are proposed (another number of modes can be used). For each coding unit (CU) in intra mode, a flag indicating whether the ALWIP mode is applied to the corresponding prediction unit (PU) is sent in the bitstream. The signaling of the latter index can be coordinated with the MRL in the same way as the first CE test. If the ALWIP mode is to be applied, an MPM list with 3 MPMSs can be used to signal the index of the ALWIP mode, predmode.
[0329] Here, the derivation of MPM can be performed using the intra modes of the upper PU and the left PU as follows. There may be tables, such as three fixed tables map_angular_to_alwip idx , idx∈{0,1,2}, which can be added to each traditional intra prediction mode predmode Angular Assign ALWIP mode:
[0330] predmode ALWIP =map_angular_to_alwip idx [predmode Angular ].
[0331] For each PU with width W and height H, define and index
[0332] idx(PU)=idx(W,H)∈{0,1,2}
[0333] It indicates from which of the three sets the ALWIP parameters will be obtained, as described in Section 4 above. above Available, belongs to the same CTU as the current PU and is in intra mode, then if idx(PU)=idx(PU above ) and if in ALWIP mode Apply ALWIP to PU above On, then:
[0334]
[0335] If the upper PU is available, belongs to the same CTU as the current PU and is in intra mode, and if the traditional intra prediction mode is set Applied to the upper PU, it makes:
[0336]
[0337] In all other cases, such that:
[0338]
[0339] This means that this mode is not available. In the same way but without the restriction that the left PU needs to belong to the same CTU as the current PU, the mode is derived:
[0340]
[0341] Finally, three fixed default lists are provided idx , idx∈{0,1,2}, each of which contains three different ALWIP modes. In the default list list idx(PU) and patterns and In , three different MPMs are constructed by replacing -1 with default values and eliminating duplicates.
[0342] The embodiments described herein are not limited by the proposed above-mentioned signaling of intra prediction modes.According to an alternative embodiment, no MPM and / or mapping table is used for MIP (ALWIP).
[0343] 6.8 MPM List Export for Traditional Luma Intra Prediction Mode and Chroma Intra Prediction Mode
[0344] The proposed ALWIP mode can be coordinated with the MPM-based coding of the traditional intra prediction mode as follows: The luma MPM list and chroma MPM list derivation process of the traditional intra prediction mode can use the fixed table map_alwip_to_angular idx , idx∈{0,1,2}, sets the ALWIP mode predmode on the given PU ALWIP Maps to one of the traditional intra prediction modes:
[0345] predmode Angular =
[0346] map_alwip_to_angular idx(PU) [predmode ALWIP ].
[0347] For Luma MPM list export, whenever ALWIP mode predmode is encountered ALWIP When the adjacent luminance block is 1, the block can be processed as if it is using the traditional intra prediction mode predmode Angular For chroma MPM list derivation, whenever the current luma block uses LWIP mode, the same mapping can be used to convert ALWIP mode to legacy intra prediction mode.
[0348] Obviously, ALWIP mode can also be coordinated with traditional intra prediction mode without using MPM and / or mapping table. For example, for chroma blocks, ALWIP mode can be mapped to planar intra prediction mode whenever the current luma block uses ALWIP mode.
[0349] 7. Implement efficient implementation examples
[0350] Let us briefly summarize the above examples, as they may form the basis for further extended embodiments described below.
[0351] For prediction of a predetermined block 18 of a picture 10, a number of neighboring samples 17a, 17c are used.
[0352] A reduction 100 has been performed by averaging a plurality of adjacent samples to obtain a reduced set 102 of sample values that is smaller in number than the plurality of adjacent samples. This reduction is optional in the embodiments herein and produces a so-called sample value vector described below. A linear or affine linear transformation 19 is performed on the reduced set of sample values to obtain predicted values for predetermined samples 104 of a predetermined block. This transformation, denoted later using a matrix A and an offset vector b, has been obtained by machine learning (ML) and should be an effectively pre-formed implementation.
[0353] By interpolation, the predicted values of the other samples 108 of the predetermined block are derived based on the predicted value of the predetermined sample and a plurality of neighboring samples. It should be noted that, in theory, the results of the affine / linear transformation can be associated with non-full-pixel sample positions of the block 18, so that all samples of the block 18 can be obtained by interpolation according to an alternative embodiment. It is also possible that no interpolation is required at all.
[0354] The plurality of adjacent samples may extend one-dimensionally along both sides of the predetermined block, the predetermined samples being arranged in rows and columns and along at least one of the rows and columns, wherein the predetermined sample may be located at every n-th position starting from a sample (112) of the predetermined samples that is adjacent to both sides of the predetermined block. Based on the plurality of adjacent samples, for each of at least one of the rows and columns, a support value of one adjacent position (118) of the plurality of adjacent positions may be determined, the support value being aligned with a corresponding one of the at least one of the rows and columns, and by interpolation, a predicted value of other samples 108 of the predetermined block may be derived based on the predicted value of the predetermined sample and the support value of the adjacent sample aligned with at least one of the rows and columns. The predetermined sample may be located at every n-th position along the row starting from a sample 112 of the predetermined sample that is adjacent to both sides of the predetermined block 18, and the predetermined sample may be located at every m-th position along the column starting from a sample 112 of the predetermined sample that is adjacent to both sides of the predetermined block, where n, m>1. It is possible that n=m. Along at least one of the rows and the columns, for each support value, a support value may be determined by averaging 122 a group of neighboring samples 120 within the plurality of neighboring samples, the group of neighboring samples 120 including the neighboring sample 118 for which the corresponding support value is determined. The plurality of neighboring samples may extend one-dimensionally along both sides of the predetermined block, and the reduction may be accomplished by grouping the plurality of neighboring samples into groups 110 of one or more consecutive neighboring samples and performing averaging on each neighboring sample in the group of one or more neighboring samples, the group of one or more neighboring samples having more than two neighboring samples.
[0355] For a predetermined block, a prediction residual can be transmitted in the data stream. It can be derived from the data stream at the decoder, and the predetermined block can be reconstructed using the prediction residual and the predicted values of the predetermined samples. At the encoder, the prediction residual is encoded into the data stream at the encoder.
[0356] The picture can be subdivided into a plurality of blocks of different block sizes, the plurality of blocks including the predetermined block. A linear or affine linear transform for the block 18 can then be selected depending on the width W and height H of the block, such that the linear or affine linear transform selected for the predetermined block is selected from: a first set of linear or affine linear transforms, as long as the width W and height H of the predetermined block are within a first set of width / height pairs; a second set of linear or affine linear transforms, as long as the width W and height H of the predetermined block are within a second set of width / height pairs, the second set of width / height pairs not intersecting the first set of width / height pairs. Again, as will become clear later, the affine / linear transform is represented by other parameters, namely the weight C and optionally the offset and scaling parameters.
[0357] The decoder and the encoder may be configured to subdivide the picture into a plurality of blocks of different block sizes, including a predetermined block, and select a linear or affine linear transform depending on a width W and a height H of the predetermined block, such that the linear or affine linear transform selected for the predetermined block is selected from the following:
[0358] a first set of linear or affine linear transforms, as long as the width W and height H of the predetermined block are within the first set of width / height pairs,
[0359] a second set of linear or affine linear transforms, as long as the width W and height H of the predetermined block are within the second set of width / height pairs that are disjoint from the first set of width / height pairs, and
[0360] A third set of linear transformations or affine linear transformations, as long as the width W and height H of the predetermined block are within the third set of one or more width / height pairs, and the third set of width / height pairs is disjoint from the first set of width / height pairs and the second set of width / height pairs.
[0361] The third set of one or more width / height pairs includes only one width / height pair W', H', and each linear or affine linear transform within the first set of linear or affine linear transforms is used to transform the N' sample values into W'*H' predicted values of a W'xH' array of sample positions.
[0362] Each of the first and second sets of width / height pairs may include a first width / height pair W p 、H p , where W p Not equal to H p , and includes a second width / height pair W q 、H q , where H q =W p And W q =H p .
[0363] Each of the first and second sets of width / height pairs may additionally include a third width / height pair W p 、H p , where W p Equal to H p And H p >H q .
[0364] For a predetermined block, a set index may be transmitted in the data stream, which index indicates which linear or affine linear transform is to be selected for the block 18 from a predetermined set of linear or affine linear transforms.
[0365] The plurality of adjacent samples may extend one-dimensionally along both sides of the predetermined block, and the reduction may be accomplished by: for a first subset of the plurality of adjacent samples adjacent to the first side of the predetermined block, grouping the first subset into a first group 110 of one or more consecutive adjacent samples; and for a second subset of the plurality of adjacent samples adjacent to the second side of the predetermined block, grouping the second subset into a second group 110 of one or more consecutive adjacent samples; and performing averaging on each of the first group of one or more adjacent samples and the second group of one or more adjacent samples, each group having more than two adjacent samples, so as to obtain a first sample value from the first group and obtain a second sample value from the second group. Then, a linear or affine linear transform can be selected depending on a set index in a predetermined set of linear or affine linear transforms, such that two different states of the set index result in selection of one of the linear or affine linear transforms in the predetermined set of linear or affine linear transforms, and in the case where the set index assumes a first state of the two different states in the form of a first vector, the predetermined linear or affine linear transform can be performed on the reduced set of sample values to produce an output vector of predicted values, and the predicted values of the output vector are distributed to predetermined samples of the predetermined block along a first scan order, and in the case where a second state of the two different states is assumed in the form of a second vector, the first vector and the second vector are different, such that a component filled by one of the first sample values in the first vector is filled by one of the second sample values in the second vector, and a component filled by one of the second sample values in the first vector is filled by one of the first sample values in the second vector, so as to produce an output vector of predicted values, and the predicted values of the output vector are distributed to predetermined samples of the predetermined block transposed relative to the first scan order along a second scan order.
[0366] For the w1×h1 array of sample positions, each linear or affine linear transform in the first set of linear or affine linear transforms can be used to transform N1 sample values into w1*h1 predicted values, and for the w2×h2 array of sample positions, each linear or affine linear transform in the second set of linear or affine linear transforms is used to transform N2 sample values into w2*h2 predicted values, wherein, for a first predetermined width / height pair in the first set of width / height pairs, w1 can exceed the width of the first predetermined width / height pair, or h1 can exceed the height of the first predetermined width / height pair, and for a second predetermined width in the first set of width / height pairs, w1 cannot exceed the width of the second predetermined width / height pair, and h1 cannot exceed the height of the second predetermined width / height pair. The reduced set of sample values (102) may then be obtained by averaging a plurality of adjacent samples (100) such that: if the predetermined block is a predetermined block of a first predetermined width / height pair and if the predetermined block is a predetermined block of a second predetermined width / height pair, the reduced set of sample values 102 has N1 sample values, and the selected linear or affine linear transformation may be performed on the reduced set of sample values using only a first sub-portion of the selected linear or affine linear transformation: if w1 exceeds the width of one width / height pair, the first sub-portion relates to a subsampling of the w1×h1 array of sample positions along the width dimension, or if h1 exceeds the height of one width / height pair, the first sub-portion relates to a subsampling of the w1×h1 array of sample positions along the height dimension, and if the predetermined block is of the first predetermined width / height pair, the reduced set of sample values is completely subjected to the selected linear or affine linear transformation.
[0367] For a w1×h1 array of sample positions (w1=h1), each linear or affine linear transform within a first set of linear or affine linear transforms can be used to transform N1 sample values into w1*h1 predicted values, and for a w2×h2 array of sample positions (w2=h2), each linear or affine linear transform within a second set of linear or affine linear transforms is used to transform N2 sample values into w2*h2 predicted values.
[0368] All of the above embodiments are illustrative only, as they may form the basis of the embodiments described below. That is, the above concepts and details should be used to understand the following embodiments and should serve as a repository for possible extensions and modifications of the embodiments described below. In particular, many of the above details are optional, such as the averaging of adjacent samples, the fact that adjacent samples are used as reference samples, etc.
[0369] More generally, the embodiments described herein assume that the prediction signal for a rectangular block is generated from reconstructed samples, for example, the intra prediction signal for a rectangular block is generated from the adjacent reconstructed samples to the left and above the block. The generation of the prediction signal is based on the following steps.
[0370] 1. Samples can be extracted by averaging the reference samples, now referred to as boundary samples, while not excluding the possibility of transferring the description to reference samples located elsewhere. Here, boundary samples are averaged for both the left and top sides of the block, or for only one of the two sides. If averaging is not performed on one side, the samples on that side remain unchanged.
[0371] 2. Perform a matrix-vector multiplication, optionally followed by an offset, where the input vector to the matrix-vector multiplication is the concatenation of the averaged boundary samples on the left side of the block with the original boundary samples above the block if averaging is applied only on the left side, or the input vector to the matrix-vector multiplication is the concatenation of the original boundary samples on the left side of the block with the averaged boundary samples above the block if averaging is applied only on the top side, or the input vector to the matrix-vector multiplication is the concatenation of the averaged boundary samples on the left side of the block with the averaged boundary samples above the block if averaging is applied on both sides of the block. Again, there will be alternatives, such as an alternative that does not use averaging at all.
[0372] 3. The result of the matrix-vector multiplication and optional offset addition may optionally be a downscaled prediction signal over a subsampled set of samples in the original block. The prediction signals at the remaining positions may be generated from the prediction signals over the subsampled set by linear interpolation.
[0373] The calculation of the matrix-vector product in step 2 is preferably performed in integer arithmetic. Thus, if x = (x1, ..., x n ) represents the input to the matrix-vector product, i.e., x represents the concatenation of the (averaged) boundary samples to the left and above the block, then in x, the (reduced) prediction signal calculated in step 2 should be calculated using only shifts, additions of offset vectors, and multiplications with integers. Ideally, the prediction signal in step 2 would be given as Ax+b, where b is an offset vector that may be zero, and where A is derived by some machine learning-based training algorithm. However, such training algorithms typically only produce matrices A=A given in floating-point precision. float . We are thus faced with specifying integer operations in the aforementioned sense so that the expression A is well approximated using these integer operations. float Here, it is important to mention that these integer operations do not have to be chosen so that they approximate the expression A under the assumption that the vector x is uniformly distributed float x, but usually consider the input expression A floatThe vector x to be approximated is the (averaged) boundary samples from a natural video signal, where the components x of x can be expected to be i Some correlations between them.
[0374] Figure 8 Improved ALWIP prediction is shown. Samples for a predetermined block can be predicted based on a first matrix-vector product between a matrix A 1100, derived through some machine learning-based training algorithm, and a sample value vector x 400. Optionally, an offset b 1110 can be added. To achieve an integer or fixed-point approximation of this first matrix-vector product, a reversible linear transformation 403 can be performed on the sample value vector to determine another vector 402. A second matrix-vector product between another matrix B 1200 and the another vector 402 can be equal to the result of the first matrix-vector product.
[0375] Due to the characteristics of the other vector 402, the second matrix-vector product can be an integer approximated by the matrix-vector product 404 between the predetermined prediction matrix C 405 and the other vector 402 plus another offset 408. The other vector 402 and the other offset 408 can be composed of integers or fixed-point values. All components of the other offset can be identical, for example. The predetermined prediction matrix 405 can be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the predetermined prediction matrix 405 and the other vector 402 can be understood as a prediction vector 406.
[0376] In the following, more details about this integer approximation are provided.
[0377] Possible solution based on Example 1: Subtracting and adding the mean
[0378] The expression A that can be used in the above situation is float One possible incorporation of the integer approximation of x is to replace the i0th component of x (ie, the sample value vector 400) with the mean of the components of x, mean(x) (ie, the predetermined value 1400). (i.e., the predetermined component 1500), and subtract the mean from all other components. In other words, if Figure 9a As shown, the reversible linear transformation 403 is defined such that the predetermined component 1500 of the further vector 402 becomes a, and each of the components of the further vector 402 other than the predetermined component 1500 is equal to the corresponding component of the sample value vector 400 minus a, where a is a predetermined value 1400, which is, for example, the average value, such as the arithmetic mean or weighted average value, of the components of the sample value vector 400. This operation on the input is given by the reversible transformation T403, which has an obvious integer implementation, especially if the dimension n of x is a power of 2.
[0379] Because A float =(AfloatT -1 )T, if such a transformation is performed on the input x, then an integer approximation to the matrix-vector product By must be found, where B = (A float T -1 ) and y=Tx. Since the matrix-vector product A float x represents a prediction for a rectangular block (ie, a predetermined block), and since x400 is included by the (ie, averaged) boundary samples of the block, it should be expected that when all sample values of x are equal, i.e., for all i, x i =mean(x), the predicted signal A float Each sample value in x should be close to mean(x) or exactly equal to mean(x). This means that the i0th column (i.e., the column corresponding to the predetermined component of B) should be expected to be very close to or equal to a column consisting of only ones. Therefore, if M(i0), i.e., the integer matrix 1300, is a matrix whose i0th row consists of ones and whose all other rows are zeros, written as By=Cy+M(i0), where C=BM(i0), then the i0th row 412 of C (i.e., the predetermined prediction matrix 405) should be expected to have small entries or be zero, as shown in FIG. Figure 9b Furthermore, since the components of x are correlated, we can expect that for each i≠i0, the i-th component y of y i =x i The -mean(x) component usually has an absolute value much smaller than the i-th component of x. Since the matrix M(i0) 1300 is an integer matrix, if an integer approximation of Cy is given, an integer approximation of By is achieved, and from the above discussion, it can be expected that the quantization error generated by quantizing each entry of C405 in a suitable manner corresponds to A float x and only slightly affects the error in the quantization result of By.
[0380] The predetermined value 1400 does not have to be the mean (x). The following alternative definition of the predetermined value 1400 can also be used to implement the expression A float The integer approximation of x described in this paper is:
[0381] In expression A float In another possible combination of integer approximations of x, the i0th component of x is remains unchanged and the same value is subtracted from all other components That is, for each i≠i0, and In other words, the preset value 1400 may be a component in the sample value vector 400 corresponding to the preset component 1500 .
[0382] Alternatively, the predetermined value 1400 is a default value or a value signaled in the data stream into which the picture is encoded.
[0383] The predetermined value 1400 is equal to 2, for example bitdepth-1 In this case, another vector 402 can be represented by y0=2 bitdepth-1 and y i =x i -x0 (for i>0) is defined.
[0384] Alternatively, the predetermined component 1500 becomes a constant minus the predetermined value 1400. The constant is, for example, equal to 2 bitdepth-1 According to an embodiment, the predetermined component of another vector y402 1500 is equal to 2 bitdepth-1 Subtract the component corresponding to the predetermined component 1500 in the sample value vector 400 And all other components of the further vector 402 are equal to the corresponding components of the sample value vector 400 minus the components of the sample value vector 400 corresponding to the predetermined components 1500 .
[0385] For example, it may be advantageous if the predetermined value 1400 has a small deviation from the predicted value of the samples of the predetermined block.
[0386] According to an embodiment, the apparatus 1000 is configured to include a plurality of reversible linear transformations 403, wherein each reversible linear transformation is associated with a component of another vector 402. In addition, the apparatus is configured, for example, to select a predetermined component 1500 from the components of the sample value vector 400 and use the reversible linear transformation 403 associated with the predetermined component 1500 from the plurality of reversible linear transformations as the predetermined reversible linear transformation. This is, for example, because the different positions of the i0-th row (i.e., the row of the reversible linear transformation 403 corresponding to the predetermined component) depend on the position of the predetermined component in the other vector. If, for example, the first component of the other vector 402 (i.e., y1) is the predetermined component, the i0-th row will replace the first row of the reversible linear transformation.
[0387] like Figure 9b As shown, the matrix component 414 of the predetermined prediction matrix C 405 within the column 412 (i.e., the i0-th column) of the predetermined prediction matrix 405 is, for example, all zero, and the matrix component 414 corresponds to the predetermined component 1500 of the other vector 402. In this case, the apparatus is configured to perform multiplication to calculate the matrix-vector product 404 by, for example, calculating the matrix-vector product 407 between the reduced prediction matrix C′405 resulting from omitting the column 412 from the predetermined prediction matrix C 405 and the other vector 410 resulting from omitting the predetermined component 1500 from the other vector 412, as shown in FIG. Figure 9b Therefore, the prediction vector 406 can be calculated with fewer multiplications.
[0388] like Figure 8 and Figure 9b As shown, the apparatus 1000 may be configured to: when predicting samples of a predetermined block based on the prediction vector 406, calculate the sum of the corresponding component and a (i.e., the predetermined value 1400) for each component of the prediction vector 406. The sum may be represented by the sum of the prediction vector 406 and the vector 409, where all components of the vector 409 are equal to the predetermined value 1400, as shown in FIG. Figure 8 and Figure 9b Alternatively, the sum can be represented by the sum of the prediction vector 406 and the matrix-vector product 1310 between the integer matrix M 1300 and the other vector 402, as shown in Figure 9b As shown, the matrix components in the integer matrix 1300 corresponding to the predetermined components 1500 of the other vector 402 are 1 in one column (ie, the i0-th column) of the integer matrix 1300, while all other components are, for example, zero.
[0389] The result of the sum of the predetermined prediction matrix C 405 and the integer matrix 1300 is equal to or approximately equal to, for example Figure 8 Another matrix B 1200 is shown in .
[0390] In other words, another matrix B 1200 generated by summing each matrix component of the predetermined prediction matrix C 405 corresponding to the predetermined component 1500 of the other vector 402 within the column 412 (i.e., the i0-th column) of the predetermined prediction matrix 405 with 1 (i.e., the matrix B), and multiplying by the reversible linear transformation 403 corresponds to, for example, a quantized version of the machine learning prediction matrix A 1100, as shown in FIG. Figure 8 、 Figure 9a and Figure 9b The sum of each matrix component in the i0-th column 412 of the predetermined prediction matrix C 405 and 1 may correspond to the sum of the predetermined prediction matrix 405 and the integer matrix 1300, as shown in FIG. Figure 9b As shown. Figure 8 As shown, the machine learning prediction matrix A 1100 can be equal to the result of multiplying another matrix 1200 by the reversible linear transformation 403. This is because A·x=BT·yT -1 The predetermined prediction matrix 405 is, for example, a quantized matrix, an integer matrix and / or a fixed-point matrix, thereby realizing a quantized version of the machine learning prediction matrix A 1100 .
[0391] Matrix multiplication using only integer operations
[0392] For low complexity implementations (in terms of the complexity of adding and multiplying scalar values, and in terms of the storage required for the entries of the matrices involved), it is desirable to perform the matrix multiplication 404 using only integer arithmetic.
[0393] To calculate the approximation z = Cy, that is:
[0394]
[0395] According to an embodiment, only integer arithmetic is used, and the real value C i,j Must be mapped to an integer value This can be done, for example, by uniform scalar quantization or by considering the value y i The integer value represents, for example, a fixed-point number, each of which can be stored in a fixed number of bits n_bits, for example, n_bits=8.
[0396] The matrix-vector product 404 with a matrix of size m×n (i.e., the predetermined prediction matrix 405) can then be performed as shown in this pseudocode, where <<, >> are arithmetic binary left and right shift operations, and +, -, and * only operate on integer values. (1)
[0398]
[0399] Here, array C, the predetermined prediction matrix 405, stores fixed-point numbers as, for example, integers. The final addition of final_offset and the right shift operation of right_shift_result reduce the precision by rounding to obtain the required fixed-point format at the output.
[0400] To allow an increased range of real values to be represented by integers in C, two additional matrices offset can be used. i,j and scale i,j ,like Figure 10 and Figure 11 As shown in the example, the following matrix vector product y j Each coefficient b i,j :
[0401]
[0402] Given by the following formula
[0403]
[0404] offset i,j and scale i,j The values of are themselves integer values. For example, these integers may represent fixed-point numbers, each of which may be represented by a fixed number of bits (e.g., 8 bits) or by a fixed bit array such as that used to store the value. The same number of bits as n_bits is used to store the data.
[0405] In other words, the apparatus 1000 is configured to use prediction parameters (eg, integer values) and the value offset i,j and scale i,j ) is used to represent a predetermined prediction matrix 405, and a matrix-vector product 404 is calculated by performing multiplication and summation on the components of the other vector 402 and the prediction parameters and the intermediate results generated thereby, wherein the absolute value of the prediction parameter can be represented by n-bit fixed point representation, wherein n is equal to or less than 14, or alternatively equal to or less than 10, or alternatively equal to or less than 8. For example, the components of the other vector 402 are multiplied by the prediction parameters to generate products as intermediate results, and then the intermediate results are summed or form addends of the summation.
[0406] According to an embodiment, the prediction parameters include weights, each weight being associated with a corresponding matrix component of the prediction matrix. In other words, the predetermined prediction matrix is, for example, replaced or represented by the prediction parameters. The weights are, for example, integer values and / or fixed-point values.
[0407] According to an embodiment, the prediction parameters also include one or more scaling factors, such as the value scale i,j , each scaling factor is associated with one or more corresponding matrix components of the predetermined prediction matrix 405 to scale the weights associated with the one or more corresponding matrix components of the predetermined prediction matrix 405, such as integer values Additionally or alternatively, the prediction parameters include one or more offsets, such as the value offset i,j , each offset is associated with one or more corresponding matrix components of the predetermined prediction matrix 405 to offset the weights associated with the one or more corresponding matrix components of the predetermined prediction matrix 405, such as integer values
[0408] To reduce the offset i,j and scale i,j To reduce the amount of storage required, their values can be chosen to be constant for a particular set of indices i, j. For example, their entries can be constant for each column, or they can be constant for each row, or they can be constant for all i, j, as in Figure 10 shown.
[0409] For example, in a preferred embodiment, offset i,j and scale i,j For a prediction mode all values of the matrix are constant, such as Figure 11 Thus, when there are K prediction modes, where k = 0..K-1, only a single value o is required. k and a single value s k To calculate the prediction for mode k.
[0410] According to an embodiment, for all matrix-based intra prediction modes, offset i,j and / or scale i,j is constant, i.e. the same. Additionally or alternatively, for all block sizes, offset i,j and / or scale i,j May be constant, i.e. the same.
[0411] In o k Indicates offset and s k In the case of scaling, the calculation in (1) can be modified as follows: (2)
[0413]
[0414]
[0415] Extended embodiments resulting from this solution
[0416] The above solutions refer to the following embodiments:
[0417] 1. A prediction method as in Part I, wherein in step 2 of Part I, the integer approximation of the matrix-vector product involved is performed as follows: n ), for a fixed i0 (where 1≤i0≤n), calculate the vector y=(y1,…,y n ), where y i =x i -mean(x) (for i≠i0) and where and where mean(x) denotes the mean of x. The vector y is then used as input to (the integer implementation of) the matrix-vector product Cy, so that the (downsampled) prediction signal pred from step 2 of part I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) denotes the signal at each sample position in the domain of the (downsampled) prediction signal, which is equal to mean(x). (See e.g. Figure 9b )
[0418] 2. A prediction method as in Part I, wherein in step 2 of Part I, the integer approximation of the matrix-vector product involved is performed as follows: n ), for a fixed i0 (where 1≤i0≤n), calculate the vector y=(y1,…,y n-1 ), where y i= x i - mean(x) (for i < i0), and where y i = x i+1 - mean(x) (for i ≥ i0), and where mean(x) represents the mean of x. The vector y is then used as the input to an integer implementation of the matrix - vector product Cy such that the (downsampled) prediction signal pred from step 2 of Part I is given by pred = Cy+meanpred(x). In this equation, meanpred(x) represents the signal at each sample position in the domain of the (downsampled) prediction signal, which is equal to mean(x). (See, for example Figure 9b )
[0419] 3. A prediction method as in Part I, wherein the integer implementation of the matrix - vector product Cy is given by using the coefficients in the matrix - vector product z i = ∑ j b i,j *y j (See, for example ) Figure 10
[0420] 4. A prediction method as in Part I, wherein step 2 uses one of K matrices such that multiple prediction patterns can be computed, each prediction pattern using a different matrix where k = 0…K - 1, and where the integer implementation of the matrix - vector product Cy is given by using the coefficients in (the matrix - vector product z k = ∑ i = ∑ h b i,j *y j ). (See, for example ) Figure 11
[0421] That is, according to an embodiment of the present application, the encoder and decoder operate as follows to predict a predetermined block 18 of picture 10, see Figure 8 . For prediction, multiple reference samples are used. As outlined above, embodiments of the present application will not be limited to intra - coding, and thus, the reference samples will not be limited to neighboring samples, i.e., samples in picture 10 adjacent to block 18. Specifically, the reference samples will not be limited to samples arranged along the outer edge of block 18, such as samples adjacent to the outer edge of the block. However, this case is of course an embodiment of the present application
[0422] To perform prediction, a sample value vector 400 is formed from reference samples, such as reference sample 17a and reference sample 17c. Possible formations have been described above. Formation may involve averaging, thereby reducing the number of samples 102 or the number of components of vector 400 compared to reference sample 17 that contributed to the formation. As described above, the formation may also depend in some way on the size or dimensions of block 18, such as its width and height.
[0423] This vector 400 should be affinely or linearly transformed in order to obtain the prediction for block 18. Different nomenclatures have been used above. Using the more recent nomenclature, the purpose is to perform a prediction by applying vector 400 to matrix A while performing a matrix-vector product with an offset vector b. The offset vector b is optional. The affine or linear transformation determined by A, or A and b, can be determined by the encoder and decoder, or more precisely, is determined by the encoder and decoder for the purpose of prediction based on the size and dimensions of block 18 as already described above.
[0424] However, to achieve the computational efficiency improvements outlined above or to make prediction more efficient in implementation, the affine or linear transformation has been quantized, and the encoder and decoder or their predictor uses the aforementioned C and T to represent and perform the linear or affine transformation, where C and T applied in the above manner represent quantized versions of the affine transformation. Specifically, rather than applying vector 400 directly to matrix A, the predictor in the encoder and decoder applies vector 402 generated from sample value vector 400 to matrix A by mapping it via a predetermined reversible linear transformation T. The transformation T used here can be the same as long as vector 400 has the same size (i.e., independent of the block size, i.e., width and height) or is at least the same for different affine / linear transformations. In the above, vector 402 is denoted as y. To perform the affine / linear transformation as determined by machine learning, the exact matrix would be B. However, the prediction in the encoder and decoder does not perform B exactly, but rather through an approximate or quantized version thereof. Specifically, the representation is done by appropriately representing C in the manner outlined above, where C+M represents a quantized version of B.
[0425] Therefore, prediction in the encoder and decoder is further performed by computing a matrix-vector product 404 between vector 402 and a predetermined prediction matrix C, appropriately represented and stored at the encoder and decoder in the manner described above. The vector 406 generated from this matrix-vector product is then used to predict the samples 104 of block 18. As described above, for prediction, each component of vector 406 can be summed with parameter a, as indicated at 408, to compensate for the corresponding definition of C. Derivation of the prediction for block 18 based on vector 406 may also involve the optional summation of vector 406 with an offset vector b. As described above, each component of vector 406, and accordingly each component of the sum of vector 406, the vector of all a's indicated at 408, and the optional vector b, can directly correspond to a sample 104 of block 18 and thus indicate a predicted value for the sample. It is also possible that only a subset of the samples 104 of the block are predicted in this manner, with the remaining samples of block 18, such as 108, being derived by interpolation.
[0426] As mentioned above, there are different embodiments for setting a. For example, it could be the arithmetic mean of the components of vector 400. For this case, see Figure 9a The reversible linear transformation T 403 can be expressed as Figure 9a As indicated above. i0 is a predetermined component of the sample value vector and the vector 402, respectively, which is replaced by a. However, as indicated above, there are other possibilities. However, as far as the representation of C is concerned, it has also been indicated above that C can be embodied in different ways. For example, the matrix-vector product 404 can have a smaller matrix-vector product of lower dimension in its actual calculation. In particular, as indicated above, due to the definition of C, the entire i0-th column 412 of C can become 0, so that the actual calculation of the product 404 can be performed by a reduced version of the vector 402, which is reduced by omitting the component (ie, multiplying the reduced vector 410 with the reduced matrix C′ generated from C by ignoring the i0-th column 412 ) produces the vector 402 .
[0427] The weights of C or the weights of C' (i.e., the components of the matrix) can be represented and stored in a fixed-point representation. However, as described above, these weights 414 can also be stored in a manner that is related to different scaling and / or offsets. The scaling and offset can be defined for the entire matrix C, that is, all weights 414 for the matrix C or the matrix C' are equal, or can be defined in a manner such that all weights 414 for the same column or all weights 414 for the same row of the matrix C and the matrix C', respectively, are constant or equal. In this regard, Figure 10 It is shown that the calculation of the matrix-vector product (ie the result of the product) may actually be performed slightly differently, ie for example by moving the multiplication with the scaling towards vector 402 or vector 410 , thereby reducing the number of further multiplications that have to be performed. Figure 11The case where one scale and one offset are used for all weights 414 for C or C' is shown, as performed in calculation (2) above.
[0428] According to an embodiment, the apparatus described herein for predicting a predetermined block of a picture may be configured to use matrix-based intra sample prediction, which includes the following features:
[0429] The apparatus is configured to form a sample value vector pTemp[x] 400 from a plurality of reference samples 17. Assuming that pTemp[x] is 2*boundarySize, pTemp[x] may be filled, for example by direct copying or by subsampling or pooling, with the following samples: a neighboring sample redT[x] (where x=0..boundarySize-1) located at the top of the predetermined block, followed by a neighboring sample redL[x] (where x=0..boundarySize-1) located to the left of the predetermined block (for example in the case of isTransposed=0), or vice versa in the case of a transposed process (for example in the case of isTransposed=1).
[0430] Deriving input values p[x], where x=0..inSize-1, i.e., the apparatus is configured to derive another vector p[x] from the sample value vector pTemp[x], mapping the sample value vector pTemp[x] to a vector p[x] by a predetermined reversible linear transformation, or more specifically a predetermined reversible affine linear transformation, as follows:
[0431] – If mipSizeId is equal to 2, the following applies:
[0432] p[x]=pTemp[x+1]-pTemp[0]
[0433] – Otherwise (mipSizeId is less than 2), the following applies:
[0434] p[0]=(1<<(BitDepth-1))-pTemp[0]
[0435] p[x]=pTemp[x]-pTemp[0]for x=1..inSize-1
[0436] Here, the variable mipSizeId indicates the size of the predetermined block. That is, according to this embodiment, the reversible transformation used to derive another vector from the sample value vector depends on the size of the predetermined block. The dependency can be given as follows:
[0437] mipSizeId boundarySize predSize 0 2 4 1 4 4 2 4 8
[0438] Where predSize indicates the number of predicted samples within a predetermined block, and according to inSize = (2*boundarySize) - (mipSizeId == 2)?1:0, 2*boundarySize indicates the size of the sample value vector and is related to inSize (i.e., the size of the other vector). More precisely, inSize indicates the number of components of the other vector that actually participate in the calculation. For smaller block sizes, inSize is the same as the size of the sample value vector; for larger block sizes, inSize is one component smaller. In the former case, one component, namely the component corresponding to the predetermined component of the other vector, can be ignored. For example, in the matrix-vector product to be calculated later, the contribution of the corresponding vector component will produce zero regardless, and therefore, no actual calculation is required. In the case of an alternative embodiment, the dependency on block size can be ignored. In this alternative embodiment, only one of the two options is used, namely, independent of block size (the option corresponding to mipSizeId less than 2, or the option corresponding to mipSizeId equal to 2).
[0439] In other words, the predetermined reversible linear transformation is defined, for example, so that a predetermined component of the other vector p becomes a, while all other components correspond to the components of the sample value vector minus a, where, for example, a=pTemp[0]. This is easy to see in the case where the first option corresponds to mipSizeId being equal to 2, and further considering the components of the other vector that are only formed in a different way. That is, in the case of the first option, the other vector is actually {p[0...inSize]; pTemp[0]}, where pTemp[0] is a, and the actual calculation part of the matrix-vector multiplication that produces the matrix-vector product (i.e., the result of the multiplication) is limited to the inSize components of the other vector and the corresponding columns of the matrix, because the matrix has zero columns that do not need to be calculated. In other cases corresponding to mipSizeId being less than 2, a=pTemp[0] is selected as all components of the other vector except p[0] (i.e., each of the other components p[x] (where x=1..inSize-1) of the other vector p except the predetermined component p[0]) being equal to the corresponding component of the sample value vector pTemp[x] minus a, but p[0] is selected as a constant minus a. The matrix-vector product is then calculated. The constant is the mean of the representable values, i.e., 2 x-1(i.e. 1<<(BitDepth-1)), where x denotes the bit depth of the computational representation used. It should be noted that if p[0] is chosen to be pTemp[0], the computed product will simply deviate from the product computed using p[0] as indicated above (p[0]=(1<<(BitDepth-1))-pTemp[0]) by a constant vector, which constant vector, i.e. the prediction vector, may be taken into account when predicting an internal block based on this product. The value a is therefore a predetermined value, e.g. pTemp[0]. The predetermined value pTemp[0] is in this case, for example, the component of the sample value vector pTemp that corresponds to the predetermined component p[0]. It may be the adjacent sample at the top of the predetermined block or to the left of the predetermined block, closest to the top left corner of the predetermined block.
[0440] For the intra-frame sample prediction process according to predModeIntra (eg, specifying an intra-frame prediction mode), the apparatus is configured to apply the following steps, for example, perform at least the first step:
[0441] 1. The matrix-based intra prediction samples predMip[x][y] (where x=0..predSize–1, y=0..predSize-1) are derived as follows:
[0442] – Set the variable modeId equal to predModeIntra.
[0443] – Derive the weight matrix mWeight[x][y] (where x=0..inSize-1, y=0..predSize*predSize-1) by calling the MIP weight matrix derivation process with mipSizeId and modeId as input.
[0444] – The matrix-based intra prediction samples predMip[x][y] (where x=0..predSize-1, y=0..predSize-1) are derived as follows:
[0445]
[0446] In other words, the device is configured to calculate the matrix-vector product between another vector p[i] or, in the case of mipSizeId equal to 2, {p[i]; pTemp[0]} and a predetermined prediction matrix mWeight or, in the case of mipSizeId less than 2, the prediction matrix mWeight with additional zero-weight lines corresponding to the omitted components of p, so as to obtain the prediction vector, which has here been assigned to the array of block positions {x, y} distributed in the interior of the predetermined block to produce the array predMip[x][y]. The prediction vector will correspond to the concatenation of the rows of predMip[x][y] or the columns of predMip[x][y], respectively.
[0447] According to the embodiment, or according to a different interpretation, only the component is understood as a prediction vector, and the device is configured to: when predicting samples of a predetermined block based on the prediction vector, calculate the sum of the corresponding component and a (for example, pTemp[0]) for each component of the prediction vector.
[0448] The apparatus may optionally be configured to: To predict the samples of a predetermined block, the following steps are additionally performed.
[0449] 2. The matrix-based intra prediction samples predMip[x][y] (where x=0..predSize-1, y=0..predSize-1) are cropped as follows:
[0450] predMip[x][y]=Clip1(predMip[x][y])
[0451] 3. When isTransposed is equal to true, the predSize×predSize array predMip[x][y] (where x=0..predSize-1, y=0..predSize-1) is transposed, for example, as follows:
[0452] predTemp[y][x]=predMip[x][y]
[0453] predMip=predTemp
[0454] 4. The predicted samples predSamples[x][y] (where x=0..nTbW-1, y=0..nTbH-1) are derived, for example, as follows:
[0455] If nTbW, which specifies the transform block width, is greater than predSize or nTbH, which specifies the transform block height, is greater than predSize, then the MIP prediction upsampling process is called with as input the input block size predSize, the matrix-based intra prediction samples predMip[x][y] (where x=0..predSize-1, y=0..predSize-1), the transform block width nTbW, the transform block height nTbH, the top reference samples refT[x] (where x=0..nTbW-1), and the left reference samples refL[y] (where y=0..nTbH-1), and the output is the array of predicted samples predSamples.
[0456] – Otherwise, set predSamples[x][y] (where x=0..nTbW-1, y=0..nTbH-1) equal to predMip[x][y].
[0457] In other words, the apparatus is configured to predict samples predSamples of a predetermined block based on the prediction vector predMip.
[0458] 8. Use block / matrix-based intra prediction modes and other intra prediction modes
[0459] The following description again presents a possibility for combining block / matrix based prediction with other intra prediction modes. It represents another presentation of possibilities based on which the embodiments described in the subsequent sections can be implemented.
[0460] Please note that in the following, the term block-based intra prediction is used to represent the intra prediction mode that can be embodied by or is equivalent to the above ALWIP.
[0461] Therefore, the following Figure 12The described embodiments relate to a decoder and an encoder supporting intra-frame prediction for decoding / encoding a predetermined block 18, wherein different intra-frame prediction modes are supported. An angular intra-frame prediction mode 500 exists, according to which reference samples 17 adjacent to the predetermined block 18 are used to fill the predetermined block 18 in order to obtain an intra-frame prediction signal for the predetermined block 18. Specifically, the reference samples 17 arranged along the boundaries of the predetermined block 18 (e.g., along the top and left edges of the predetermined block 18) represent picture content that is extrapolated or copied to the interior of the predetermined block 18 along a predetermined direction 502. Prior to the extrapolation or copying, the picture content represented by the adjacent samples 17 may be interpolated or filtered, or in other words, the picture content represented by the adjacent samples 17 may be derived from the adjacent samples 17 by means of interpolation filtering. The angular intra-frame prediction modes 500 differ from each other in the intra-frame prediction direction 502. Each angular intra prediction mode 500 may have an associated index, wherein the association of the index with the angular intra prediction mode 500 may be such that, when the angular intra prediction modes 500 are ordered according to the associated mode index, the direction 500 rotates monotonically clockwise or counterclockwise.
[0462] There can also be non-angular intra prediction modes. For example, Figure 12 504 shows a planar intra prediction mode optionally included in the set 508, according to which a two-dimensional linear function defined by a horizontal slope, a vertical slope and an offset is derived based on the neighboring samples 17, by which the predicted sample values of the predetermined block 18 are defined. The horizontal slope, the vertical slope and the offset are derived based on the neighboring samples 17. According to an embodiment, the first set of intra prediction modes 508 includes the planar intra prediction mode 504.
[0463] A specific non-angular intra prediction mode, DC mode, included in set 508 is shown at 506. Here, one value, a quasi-DC value, is derived based on neighboring samples 17, and this one DC value is attributed to all samples of a predetermined block 18 in order to obtain an intra prediction signal. Although two examples of non-intra prediction modes are shown, there may be only one or more than two examples.
[0464] The intra-prediction modes 500, 504, and 506 form a set 508 of intra-prediction modes supported by the encoder and decoder, which compete with the block-based intra-prediction modes (the examples discussed above use the abbreviation ALWIP) generally indicated with reference numeral 510 in terms of rate / distortion optimization. As described above, according to these block-based intra-prediction modes 510, a matrix-vector product 520 is performed between a vector 514 derived from neighboring samples 17 on the one hand and a predetermined prediction matrix 516 on the other hand. The result of the multiplication 520 is a prediction vector 518 for predicting the samples of the predetermined block 18. The block-based intra-prediction modes 510 differ from each other in the prediction matrix 516 associated with the respective modes.
[0465] Therefore, in short, the encoder and decoder according to the embodiments described herein include a set 508 of intra-frame prediction modes (i.e., a first set of intra-frame prediction modes) and a set 520 of block-based intra-frame prediction modes (i.e., a second set of matrix-based intra-frame prediction modes), and these two sets compete with each other.
[0466] According to an embodiment of the present application, intra prediction is used to encode / decode a predetermined block 18 in the following manner. Specifically, first, a set selection syntax element 522 selects whether to use any mode in the set 508 of intra prediction modes or any mode in the set 520 of block-based intra prediction modes for predicting the predetermined block 18. If the set selection syntax element indicates that any mode in the set 508 (i.e., the first set of intra prediction modes) is to be used to predict the predetermined block 18, a list 528 of the most likely candidates in the set 508 is constructed / formed at the decoder and encoder based on the intra prediction modes used by neighboring blocks (exemplarily indicated at 524 and 526) that have already been predicted for block 18. Neighboring blocks 524 and 526 can be determined relative to the predetermined block 18 in a predetermined manner, such as by determining those neighboring blocks that cover certain neighboring samples of block 18 (e.g., the sample on top of the upper left corner sample of block 18), as well as block 526 containing samples to the left of the corner sample just mentioned. Of course, this is merely an example. The same applies to the number of neighboring blocks used for mode prediction, which is not limited to two for all embodiments. More than two or only one can be used. If either of these blocks 524 and 526 is lost, the default intra-frame prediction mode can be used as a substitute for the intra-frame prediction mode of the lost neighboring block. This also applies if either of blocks 524 and 526 has been encoded / decoded using an inter-frame prediction mode (e.g., through motion compensated prediction).
[0467] The list of modes in set 508 (i.e., list 528 of most probable intra prediction modes) is constructed as follows. The list length of list 528, i.e., the number of most probable modes therein, can be fixed by default. The length can be as follows: Figure 12 4 as shown, or it may be different, such as 5 or 6. The latter case applies to the specific example described below. An index in the data stream, which will be described later, may indicate a mode in list 528 to be used for a predetermined block 18. Indexing is performed along a list order or sort 530, where the list index is, for example, variable-length coded so that the length of the index increases monotonically along the order 530. Therefore, it is worthwhile to initially populate list 528 with only the most probable modes from set 508, and to place modes that are more likely to be suitable for block 18 upstream along the order 530 relative to modes with lower probabilities. The modes in list 528 are derived based on the modes used for blocks 524 and 526 (i.e., neighboring blocks adjacent to predetermined block 18). If either block 524 or 526 has already been intra-predicted using a block-based mode 510 from set 520, the aforementioned mapping from such "ALWIP" or block-based mode 510 to modes within set 508 (which we call non-ALWIP modes) is used. The latter mapping may, for example, map most (ie, more than half) of the block-based pattern 510 onto the DC pattern 506 (or either the DC 506 or the planar pattern 504).
[0468] According to an embodiment, the list 528 of most probable intra prediction modes is populated with the plane intra prediction mode 504 in a manner independent of the intra prediction mode used when predicting the neighboring blocks. Thus, for example, depending on the intra prediction mode used to predict the neighboring blocks 524 and 526, only the DC intra prediction mode 506 and the angular intra prediction mode 500 are populated in the list 528. For example, independent of the intra prediction mode used when predicting the neighboring blocks 524 and 526, the plane intra prediction mode 504 is located at the first position in the list 528 of most probable intra prediction modes.
[0469] In a manner exemplarily shown in more detail below, the list construction of the most probable intra prediction mode list 528 is performed in such a way that if the neighboring blocks 524 and 526 have been previously predicted by the any-angle intra prediction mode 500, then there is no DC intra prediction mode 506 in the list 528. If one neighboring block 524 or 526 is predicted by the any-angle intra prediction mode 500 and / or if both the neighboring block 524 and the neighboring block 526 are predicted by the any-angle intra prediction mode 500, then the DC intra prediction mode 506 is not in the list 528 of the most probable intra prediction modes. According to the embodiments described herein below, for example, the list 528 is populated with the DC mode 506 only if the following conditions are true for all neighboring blocks 524 and 526: the neighboring blocks 524 and 526 have been encoded using any of the non-angular intra prediction modes 504 and 506, or the neighboring blocks 524 and 526 have been intra predicted using any block-based intra prediction mode 510 that is mapped to any of the non-angular intra prediction modes 504 and 506 by the aforementioned mapping from the block-based intra prediction modes 510 to the modes within the set 508. Only in this case is the DC intra prediction mode 506 located in the list 528. In this case, it can be positioned before any angular intra prediction mode 500 in the order 530, as can be seen from the subsequent examples.
[0470] In other words, the list 528 of most probable intra prediction modes is populated with the DC intra prediction mode 506 only if, for each of the neighboring blocks 524 and 526, the corresponding neighboring block is predicted using either of the at least one non-angular intra prediction modes 504 and 506 within the first set 508 (which includes the DC intra prediction mode 506) or is predicted using any of the block-based intra prediction modes 510 that is mapped to any of the at least one non-angular intra prediction modes 500 by mapping from the second set 520 of block-based intra prediction modes 510 to intra prediction modes within the first set 508 (which is used to form the list 528 of most probable intra prediction modes). The DC intra prediction mode 506, for example, precedes any angular intra prediction mode 500 in the list 528 of most probable intra prediction modes.
[0471] Thus, continuing with the description of how the predetermined block 18 is encoded into the data stream 12, if the set selection syntax element 522 indicates that the predetermined block 18 is to be encoded using any mode in the first set 508, then the data stream 12 optionally includes an MPM syntax element 532 indicating whether the intra-prediction mode to be used for the predetermined block 18 is within the list 528. If so, the data stream 12 includes an MPM list index 534 pointing to the list 528. The MPM list index 534 indicates the mode to be used for the predetermined block 18 in the list 528 by indexing the list 528 along the order 530. However, if a mode in the set 508 is not within the list 528, as indicated by the MPM syntax element 532, then the data stream 12 includes another syntax element 536 for the block 18, indicating which mode in the set 508 is to be used for the block 18, i.e., the predetermined intra-prediction mode. The other syntax element 536 may indicate the mode in a manner that distinguishes only between those modes in the set 508 that are not included in the list 528.
[0472] In other words, the apparatus for decoding the predetermined block 18 is configured, for example, to derive an MPM syntax element 532 from the data stream, and if the set selection syntax element 522 indicates that one of the first set of intra prediction modes 508 is to be used to predict the predetermined block 18, the MPM syntax element 532 indicates whether the predetermined intra prediction mode in the first set of intra prediction modes 508 is within the list of most probable intra prediction modes 528. If the MPM syntax element 532 indicates that the predetermined intra prediction mode in the first set of intra prediction modes 508 is within the list of most probable intra prediction modes 528, the apparatus is configured, for example, to form the list of most probable intra prediction modes 528 based on the intra prediction modes used when predicting neighboring blocks 524, 526 adjacent to the predetermined block 100, and to derive an MPM list index 534 from the data stream 12, the MPM list index 534 pointing to the predetermined intra prediction mode in the list of most probable intra prediction modes 528. If the MPM syntax element 532 from the data stream 12 indicates that the predetermined intra-prediction mode in the first set 508 of intra-prediction modes is not in the list 528 of most probable intra-prediction modes, the apparatus is configured to derive from the data stream another list index 536, the another list index 536 indicating the predetermined intra-prediction mode in the first set of intra-prediction modes. Thus, based on the MPM syntax element 532, the data stream 12 includes either the MPM list index 534 or the another list index 536 for prediction of the predetermined block 18.
[0473] By removing the case where the list 528 includes the DC intra-prediction mode 506, the following advantages are achieved. Specifically, the inventors of the present application discovered that, because the DC intra-prediction mode 506 in the set 508 competes with the block-based intra-prediction mode 510 anyway, using the DC intra-prediction mode 506 in the set 508 to "consume" valuable list positions in the list 528 for encoding / decoding the predetermined block 18 would negatively impact coding efficiency, and the predetermined block 18 should be encoded / decoded using any of the intra-prediction modes in the set 508 indicated by the syntax element 522 (i.e., the set selection syntax element). Therefore, "consuming" a list position in the list 528 with such a DC intra-prediction mode 506 in the set 508 would increase the likelihood that the intra-prediction mode ultimately to be used for the predetermined block 18 (i.e., the predetermined intra-prediction mode) is not in the list 528, necessitating the transmission of the syntax element 536 (i.e., another list index) in the data stream 12.
[0474] Specifically, since the syntax element 522 already indicates for block 18 whether any mode within the set 508 or any mode in the block-based mode 510 in the set 520 should be used to predict the block 18, it appears that if the syntax element 522 indicates that a mode within the set 508 is preferred for block 18, and therefore the block-based mode 510 is not used for block 18, then the DC prediction mode 506 in the set 508 is less likely to be applicable to block 18, so that its appearance in the list 528 should be limited to a very limited set of cascades of modes used for adjacent blocks 524 and 526, i.e., the cascade set forth above.
[0475] In other cases, that is, when the set selection syntax element 522 indicates that any one of the block-based intra prediction modes 510 is to be used to predict the predetermined block 18, encoding and decoding of the block 18 into the data stream 12 can be performed in the manner described above. To this end, an index can be used to index a selected one of the block-based intra prediction modes 510 in the set 520 (i.e., the second set of block-based intra prediction modes) to be used, or to indicate which block-based intra prediction mode in the set 520 is to be used. Another MPM syntax element 538 may indicate whether the indexing is done by an index 540 (i.e., by another MPM list index) indicating the block-based intra prediction mode 510 to be used for block 18 in a list 542 of most probable block-based intra prediction modes 510, i.e., by indexing along a list order 544, or whether the block-based intra prediction mode 510 to be used for block 18 is indicated by another syntax element 546 (i.e., by yet another list index), which may, for example, distinguish only between those modes 510 within the set 520 that are not already included in the list 542. List construction of the list 542 may be based on the modes used to predict blocks 524 and 526. If either block 524 or 526 is unavailable due to being outside a picture or due to being inter-predicted, a default intra prediction mode, such as one of the intra prediction modes in the set 508, may be used instead. For each block 524 and 526, where intra-frame prediction has been performed using a mode in set 508 instead of set 520, the aforementioned mapping from modes in set 508 to modes in set 520 is used to obtain the intra-frame prediction mode 510 for the corresponding block (i.e., the predetermined block 18), i.e., the predetermined block-based intra-frame prediction mode, and list 542 is constructed based on the block-based intra-frame prediction modes obtained for blocks 524 and 526.
[0476] According to an embodiment, the device for decoding the predetermined block 18 is configured to derive another MPM syntax element 532 from the data stream 12, and the another MPM syntax element 532 indicates: if the set selection syntax element 522 indicates that the predetermined block 18 will not be predicted using a mode in the first set 508 of intra-frame prediction modes, whether a predetermined block-based intra-frame prediction mode in the second set 520 of block-based intra-frame prediction modes 510 is in the list 542 of most probable block-based intra-frame prediction modes. If the further MPM syntax element 538 indicates that the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes 510 is in the list 542 of most probable block-based intra prediction modes, the apparatus is configured, for example, to form the list 542 of most probable block-based intra prediction modes based on the intra prediction modes used when predicting neighboring blocks 524, 526 adjacent to the predetermined block 18, and derive from the data stream 12 another MPM list index 540 pointing to the predetermined block-based intra prediction mode in the list 542 of most probable block-based intra prediction modes. If the further MPM syntax element 538 indicates that the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes is not in the list 542 of most probable block-based intra prediction modes, the apparatus is configured to derive from the data stream 12 yet another list index 546 indicating the predetermined block-based intra prediction mode in the second set 520 of block-based intra prediction modes. Therefore, based on the further MPM syntax element 538 , the data stream 12 comprises either a further MPM list index 540 or a further list index 546 for prediction of the predetermined block 18 .
[0477] Although another MPM syntax element 538, another MPM list index 540, and yet another list index 546 are in Figure 12 536, but it is apparent that data stream 12 includes another MPM syntax element 538 and an index associated with another MPM syntax element 538 (e.g., another MPM list index 540 or another list index 546), or includes MPM syntax element 532 and an index associated with an MPM syntax element (e.g., MPM list index 534 or another list index 536). Which of the syntax elements and indices data stream 12 includes depends, for example, on the set selection syntax element 522.
[0478] An example of writing the syntax element portion of the data stream 12 as pseudo code may be as follows: Figures 19a to 19d As shown, where the reference numerals indicate which syntax element corresponds to the syntax element discussed above.
[0479] The list structure of list 528 can be defined as follows: wherein candIntraPredModeA / B indicates the intra prediction mode used by either block 524 or block 526 for prediction, e.g., A for block 524 and B for block 526, or, if the corresponding block 524 or 526 has been intra predicted using any of the block-based intra prediction modes 510, indicates which mode in set 508 is mapped to the intra prediction mode. INTRA_DC is used to indicate mode 506, and angular mode 500 is indicated by INTRA_ANGULAR#, where the number (#) sorts the angular modes as exemplarily indicated above (i.e., such that the angular direction 502 is monotonically decreasing or monotonically increasing with increasing number). The ordering between modes within set 508 can be defined as in the following table, where the INTRA_PLANAR table indicates mode 504.
[0480] Note that in the above example, the index 534 is actually distributed over syntax elements 534′ and 534″: the former 534′ is specific to the first position in the order 530 in the list 528, where the INTRA_PLANAR mode 504 is inevitably located according to the example. The latter 534″ points to any of the subsequent positions in the list 528, where, as described, the DC mode 506 is only included in the specific case described.
[0481] Furthermore, in the above example, where syntax element 522 indicates the use of any of the modes in set 508, further syntax elements are included in the data stream that parameterize in some manner the intra prediction modes in set 500. For example, syntax element 600 parameterizes or varies the region in which reference sample 17 is located, based on which the modes in set 508 intra-predict the interior of block 18, for example, in terms of distance from the periphery of block 18. Additionally or alternatively, syntax element 602 parameterizes or varies whether the modes in set 508 use reference sample 17 for intra-prediction of the interior of block 18 globally or block-wide, or whether intra-prediction is performed per slice or portion into which block 18 is subdivided, and the slices or portions are intra-predicted sequentially, such that prediction residuals encoded into the data stream for one portion can be used to supplement new reference samples for intra-prediction of a subsequent portion. The subsequent coded portion controlled by syntax element 600 is only available (and a corresponding syntax element may be present in the data stream) if syntax element 600 has a predetermined state, e.g., corresponding to the region where reference sample 17 is located being adjacent to block 18. This portion may be defined by subdividing the block along a predetermined direction, e.g., horizontally, resulting in a portion as tall as block 18, or vertically, resulting in a portion as wide as block 18. If signaled partitioning is activated, syntax element 604 may be present in the data stream, controlling which partitioning direction is used. It can be seen that the position reserved for INTRA_PLANAR mode in list 528 may only be available if the mode is parameterized in some way via the parameterized syntax elements just mentioned, e.g., if syntax element 600 has a predetermined state, e.g., corresponding to the region where reference sample 17 is located being adjacent to block 18, and / or if the portion-wise intra prediction mode is not activated, as signaled by syntax element 602.
[0482] All syntax elements shown in the table and not specifically mentioned above are optional and are not discussed further herein.
[0483] – If candIntraPredModeB is equal to candIntraPredModeA and candIntraPredModeA is greater than INTRA_DC, then candModeList[x], x=0..4 is derived as follows:
[0484] candModeList[0]=candIntraPredModeA
[0485] candModeList[1]=2+((candIntraPredModeA+61)%64)
[0486] candModeList[2]=2+((candIntraPredModeA-1)%64)
[0487] candModeList[3]=2+((candIntraPredModeA+60)%64)
[0488] candModeList[4]=2+(candIntraPredModeA%64)
[0489] – Otherwise, if candIntraPredModeB is not equal to candIntraPredModeA and candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC, then the following applies:
[0490] – The variables minAB and maxAB are derived as follows:
[0491] minAB=Min(candIntraPredModeA,candIntraPredModeB)
[0492] maxAB=Max(candIntraPredModeA,candIntraPredModeB)
[0493] – If both candIntraPredModeA and candIntraPredModeB are greater than INTRA_DC, then candModeList[x], x=0..4 is derived as follows:
[0494] candModeList[0]=candIntraPredModeA
[0495] candModeList[1]=candIntraPredModeB
[0496] – If maxAB–minAB is equal to 1, the following applies:
[0497] candModeList[2]=2+((minAB+61)%64)
[0498] candModeList[3]=2+((maxAB-1)%64)
[0499] candModeList[4]=2+((minAB+60)%64)
[0500] – Otherwise, if maxAB–minAB is greater than or equal to 62, the following applies:
[0501] candModeList[2]=2+((minAB-1)%64)
[0502] candModeList[3]=2+((maxAB+61)%64)
[0503] candModeList[4]=2+(minAB%64)
[0504] – Otherwise, if maxAB–minAB is equal to 2, the following applies:
[0505] candModeList[2]=2+((minAB-1)%64)
[0506] candModeList[3]=2+((minAB+61)%64)
[0507] candModeList[4]=2+((maxAB-1)%64)
[0508] – Otherwise, the following applies:
[0509] candModeList[2]=2+((minAB+61)%64)
[0510] candModeList[3] = 2 + ( ( minAB - 1 ) % 64 ) (8-36)
[0511] candModeList[4]=2+(((maxAB+61))%64)
[0512] – Otherwise (candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC), then candModeList[x], x=0..4 is derived as follows:
[0513] candModeList[0]=maxAB
[0514] candModeList[1]=2+((maxAB+61)%64)
[0515] candModeList[2] = 2 + ( ( maxAB - 1 ) % 64 ) (8-41)
[0516] candModeList[3]=2+((maxAB+60)%64)
[0517] candModeList[4]=2+(maxAB%64)
[0518] – Otherwise, the following applies:
[0519] candModeList[0]=INTRA_DC
[0520] candModeList[1]=INTRA_ANGULAR50(
[0521] candModeList[2]=INTRA_ANGULAR18
[0522] candModeList[3]=INTRA_ANGULAR46
[0523] candModeList[4]=INTRA_ANGULAR54
[0524] Intra prediction mode Association Name 0 INTRA_PLANAR 1 INTRA_DC 2..66 INTRA_ANGULAR2..INTRA_ANGULAR66
[0525] 9. Embodiments using block / matrix-based intra prediction modes and other intra prediction modes and using secondary transforms
[0526] The following description presents embodiments for combining block / matrix based prediction with other intra prediction modes and encoding the prediction residual using a secondary transform. The above presentation of the possibility of matrix based intra prediction (ALWIP) and its combination with other intra prediction modes should be used as an example to implement the embodiments described below. For example, in Figure 12 All details regarding the construction of the MPM list including the DC mode restrictions are optional.
[0527] As described above, matrix-based intra prediction (MIP), also referred to herein as block-based intra prediction and ALWIP, generates an intra prediction signal on a rectangular block by performing matrix-vector multiplication, wherein the output of the matrix-vector multiplication can be regarded as a prediction signal for a downsampled block, and wherein the downsampled boundary samples may comprise the input of the matrix-vector multiplication. If the output is regarded as a prediction signal for a downsampled block, the prediction signal needs to undergo an upsampling (or linear interpolation) stage before obtaining the final prediction signal.
[0528] On the other hand, for conventional intra prediction modes like planar mode 504, DC mode 506 and angular mode 500 (also denoted as modes in set 508 in the above description), a non-separable secondary transform (LFNST) is a tool for transforming the prediction residuals corresponding to these intra prediction modes. Here, a set S of transform sets is given, such that each conventional intra prediction mode is associated with one of these transform sets. Then, at the decoder, it can be extracted from the bitstream whether LFNST is to be applied to a given block. If this is the case, one of the transform sets in the set S is given, depending on the intra prediction mode used on the current block, and if the transform set consists of more than one transform, it can be extracted from the bitstream which transform T of the set is to be used. Then, at the decoder, the transform T is applied as a secondary transform Ts, which means that it is applied to a subset 622 of the residual transform coefficients 620 of the separable primary transform Tp, such as Figure 14 shown.
[0529] The problem is that the above-mentioned secondary transform Ts is defined a priori only for conventional intra prediction modes. Providing a specific secondary transform Ts for each MIP mode 510 may be too costly in terms of memory requirements to additionally store the extra transforms.
[0530] Figure 13 A decoder is shown that solves this problem. The decoder decodes a predetermined block 18 of a picture using intra prediction. According to an embodiment, the encoder comprises features and / or functionalities in parallel with the decoder.
[0531] The decoder / encoder is configured to select 602 a predetermined intra prediction mode 604 from a plurality of intra prediction modes 600, the plurality of intra prediction modes 600 comprising a first set 508 of intra prediction modes and a second set 520 of matrix-based intra prediction modes 510. The intra mode selection 602 is performed by the decoder based on the data stream 12, wherein the encoder is configured to signal the predetermined intra prediction mode 604 in the data stream 12. Figure 12 Intra-mode selection 602 is performed as described.
[0532] The first set 508 of intra prediction modes includes a DC intra prediction mode 506 and an angular prediction mode 500 and an optional planar intra prediction mode 504. In case the predetermined intra prediction mode 604 is a matrix-based intra prediction mode 510 from the second set 520, the decoder / encoder is configured to use a matrix-vector product 512 between: a vector 514 derived from reference samples 17 in a neighborhood of the predetermined block 18 and a prediction matrix 516 associated with the corresponding matrix-based intra prediction mode 510 to obtain a prediction vector 518 based on which samples of the predetermined block 18 are predicted. Prediction of the predetermined block 18 using the matrix-based intra prediction mode 510 as the predetermined intra prediction mode 604 may be performed according to Figures 6 to 11 The decoder / encoder of an embodiment of the present invention is performed as follows. The decoder / encoder is configured to derive a prediction signal 606 for a predetermined block 18 using a predetermined intra prediction mode 604 .
[0533] The decoder / encoder is configured to transform Ts (eg, Ts ) from the secondary transform in a manner that depends on the predetermined intra prediction mode 604. (1) –Ts (N) ) in a set 612 and selecting 608 one or more secondary transformations T s (i1) -T s (in) (where i1 is in the range of 1 to N, i n in the range of i1 to N), such that the subset 610 is non-empty if the predetermined intra-prediction mode 604 is included in the first set 508 of intra-prediction modes and if the predetermined intra-prediction mode 604 is included in the second set 520 of matrix-based intra-prediction modes 510.
[0534] According to an embodiment, the decoder / encoder is configured to select 608 a subset 610 such that each secondary transform Ts in the set 612 of secondary transforms Ts is included in the subset 610 of one or more secondary transforms selected for at least one intra-prediction mode in the first set 508 and the second set 520. Thus, the subset 610 may be equal to the set 612 of secondary transforms. Such a subset may be selected for one or more matrix-based intra-prediction modes 510. For one or more matrix-based intra-prediction modes 510, all secondary transforms Ts in the set 612 of secondary transforms may be selected, whereby the subset 610 of the one or more matrix-based intra-prediction modes 510 includes the secondary transforms Ts that may be selected for the intra-prediction modes in the first set 508 of intra-prediction modes. Thus, for at least one of the matrix-based intra-prediction modes 510, no specific additional secondary transform is required in the set 612 of secondary transforms.
[0535] According to an embodiment, the decoder / encoder is configured to select 608 the subsets 610 in such a way that for each subset 610 of the secondary transform Ts selected for any matrix-based intra prediction mode 510 matrix Each secondary transform Ts in the first set 508 comprises a subset of secondary transforms Ts selected for at least one intra prediction mode in the first set 508 that does not belong to the angular prediction mode 500, for example 610 DC and / or 610 planar middle. Figure 15 Indicates the different possible subsets of secondary transforms Ts selected for any matrix-based intra prediction mode 510. The decoder / encoder may be configured to: for each matrix-based intra prediction mode 510, select a first union 611 of the subsets 610 of secondary transforms for the matrix-based intra prediction mode 510 matrix A subset 610 of .
[0536] like Figure 15 As shown, the subset 610 selectable for one or more matrix-based intra prediction modes 510 may be equal to the subset selectable for the DC intra prediction mode 506, for example, 610 DC1 or 610 DC2 , or may be equal to a subset selectable for planar intra prediction mode 504, such as 610 planar1 or 610p lanar2 The subset 610 selectable for one or more matrix-based intra prediction modes 510 may include only a subset of the secondary transforms Ts selectable for an intra prediction mode in the first set 508 that is not an angular prediction mode 500 (e.g., 610 DC ) in one or more secondary transforms (e.g., Ts (ax) To Ts (ay) ), such as subset 610 matrix4 As indicated.
[0537] The subset 610 selectable for the matrix-based intra prediction mode 510 may include two or more subsets selectable for the DC intra prediction mode 506 (eg, 610 DC1 and 610 DC2 ) in one or more secondary transforms, such as subset 610 matrix1 indicated, or may include two or more subsets (e.g., 610) selectable for the planar intra prediction mode 504. planar1 and 610 planar2 ) in one or more secondary transforms, such as subset 610 matrix3 As indicated.
[0538] Another possible subset 610 that may be selected for the matrix-based intra prediction mode 510 may include one or more subsets that may be selected for the DC intra prediction mode 506 (eg, 610 DC2 ), and one or more subsets (e.g., 610 ) selectable for the planar intra prediction mode 504 planar1 ) in one or more secondary transforms, such as subset 610 matrix2 As indicated.
[0539] According to an embodiment, the decoder / encoder is configured to select 608 the subset 610 in such a way that a first union 611 of the subsets 610 of the secondary transforms selected for the matrix-based intra prediction mode 510 matrix The subset 610 of the secondary transform selected for all angular intra prediction modes angular The second union of 611 angular The intersection between them is empty. The third union 611 DC Includes all subsets 610 selected for DC intra prediction mode 506 DC , the fourth set 611 planar Includes all subsets 610 selected for planar intra prediction mode 504 planar For example, since downsampling is optionally applied to the reduced prediction signal 606 , the MIP mode 510 is non-directional like the planar mode 504 and the DC mode 506 , and thus their prediction residuals 618 have more statistical similarity to the DC mode 506 and the planar mode 504 than to the angular mode 500 .
[0540] In addition, if Figure 13 As shown, the decoder is configured to derive 614 from the data stream a transformed version 616 of the prediction residual of the predetermined block 18 encoded into the data stream by the encoder, the transformed version 616 of the prediction residual of the predetermined block 18 being related to a spatial domain version of the prediction residual of the predetermined block 18 via a transform T defined by a concatenation of a primary transform Tp and a predetermined secondary transform Ts from the subset 610 of secondary transforms. Figure 14 As shown, in the case where the predetermined intra prediction mode is included in the first set 508 of intra prediction modes and in the case where the predetermined intra prediction mode is included in the second set 520 of matrix-based intra prediction modes 510, the encoder may be configured to apply the transform T to a subset 622 of coefficients 620 of the primary transform Tp. The decoder may be configured to use the inverse T of the transform T -1 To obtain a spatial domain version 618 of the prediction residual of the predetermined block 18. The primary transform is, for example, a separable 2D transform, and the secondary transform is, for example, a non-separable 2D transform.
[0541] The decoder is configured to reconstruct 624 the predetermined block 18 using the prediction signal 606 and the prediction residual 618 of the predetermined block 18 .
[0542] If the subset 610 of one or more secondary transforms Ts contains more than one secondary transform Ts, the decoder can be configured to select a predetermined secondary transform Ts from the subset 610 of one or more secondary transforms depending on a secondary transform indication syntax element transmitted in the data stream 12 for a predetermined block. The secondary transform indication syntax element can be an index pointing to the selected subset 610 of one or more secondary transforms. In this case, the encoder can be configured to transmit the secondary transform indication syntax element in the data stream 12.
[0543] According to an embodiment, the decoder is configured to infer that the transform T by which the transformed version 616 of the prediction residual for the predetermined block 18 is related to the spatial domain version 618 of the prediction residual for the predetermined block is a primary transform Tp if the size of the predetermined block 18 meets a predetermined criterion. Otherwise, transform T may be a concatenation of the primary transform Tp and a predetermined secondary transform Ts. The decoder makes the second set 520 of matrix-based intra prediction modes 510 available for selection 602 of a predetermined intra prediction mode 604, regardless of whether the size of the predetermined block 18 meets the predetermined criterion. The predetermined criterion for the size of the predetermined block 18 may be relevant only to subset selection 608 and not to intra mode selection 602. While LFNST (i.e., subset selection 608) is not available, there are block sizes available, such as for MIP / ALWIP 510. For blocks 18 of such sizes, a secondary transform indication syntax element does not need to be transmitted and read in the data stream. For example, if the size is below a predetermined threshold, the predetermined criterion is met. For small predetermined blocks 18 this may not be beneficial, which means that for the latter shapes the additional signal cost of signaling whether LFNST is to be applied to the block in MIP mode is on average higher than the gain obtained by allowing the LFNST transform for MIP mode.
[0544] According to an embodiment, the decoder is configured to read a non-zero region indication transmitted in the data stream 12 for a corresponding predetermined block 18, the non-zero region indication being used to indicate a non-zero transform domain region 623 within the transformed version 616 of the prediction residual of the predetermined block 18, such as Figure 16As shown. All non-zero coefficients are located only in the non-zero transform domain region 623. For example, the non-zero region indication is transmitted by the encoder in the data stream 12. The decoder / encoder is configured to decode / encode the coefficients within the non-zero transform domain region 623 from the data stream 12. LP syntax elements can be used as non-zero region indications. The last non-zero coefficient position along the scan path from the DC coefficient position to the opposite (or highest frequency) coefficient position is indicated by the LP syntax element herein. The LP syntax element can be used as a measure of the expected count of non-zero coefficients within the non-zero transform domain region 623.
[0545] According to an embodiment, the decoder is configured to: infer that the transformed version 616 of the prediction residual of the predetermined block 18 is a primary transform Tp via its transformation T related to the spatial domain version 618 of the prediction residual of the predetermined block 18, depending on whether the extension and / or the position of the non-zero transform domain area 623 meets a first predetermined criterion and / or the number of non-zero coefficients within the non-zero transform domain area 623 meets a second predetermined criterion.
[0546] For example, the first predetermined criterion is such that the first predetermined criterion is also met if the following condition is met: the non-zero transform domain area 623 does not only cover the subset 622 of the coefficients of the primary transform Tp applied by cascade. This is based on the idea that the secondary transform Ts should cover all non-zero coefficients of the transform coefficients of the primary transform Tp. In case the non-zero coefficients are outside the subset 622 of the coefficients of the primary transform Tp, it may not be advantageous to apply the predetermined secondary transform, so the decoder infers that the transform T is the primary transform Tp. The decoder can perform a subset selection 608 and apply a transform T defined by the primary transform Tp and the cascade of the predetermined secondary transform Ts applied to the subset 622 of the coefficients of the primary transform Tp in the subset 610 of the secondary transform to obtain a spatial domain version 618 of the prediction residual of the predetermined block 18 in case the non-zero transform domain area 623 is completely within the subset 622 of the coefficients of the primary transform Tp applied by cascade, see for example Figure 16 .
[0547] For example, if the number of non-zero coefficients within the non-zero transform domain region 623 is below a predetermined threshold, the second predetermined criterion is satisfied. This is based on the idea that if the number of non-zero coefficients within the non-zero transform domain region 623 is below the predetermined threshold, there is no need to further reduce the number of non-zero coefficients via the secondary transform Ts. In cases where the primary transform Tp has a small number of non-zero coefficients, applying the predetermined secondary transform may not be advantageous, so the decoder infers that transform T is the primary transform Tp. In cases where the number of non-zero coefficients is below the predetermined threshold, the additional signaling cost of the predetermined secondary transform would outweigh the increased coding efficiency achieved by the predetermined secondary transform. The decoder may perform subset selection 608 and apply a transform T defined by the concatenation of the primary transform Tp and the predetermined secondary transform Ts applied to the subset 622 of coefficients of the primary transform Tp within the subset 610 of secondary transforms, to obtain a spatial domain version 618 of the prediction residual for the predetermined block 18 if the number of non-zero coefficients within the non-zero transform domain region 623 equals or exceeds the predetermined threshold.
[0548] According to an embodiment, the decoder / encoder is configured to derive / encode from the data stream 12 a set selection syntax element 522 into the data stream 12, the set selection syntax element 522 indicating whether one of the first set 508 of intra prediction modes is to be used for predicting the predetermined block 18, e.g. Figure 12 If the set selection syntax element 522 indicates that one of the first set 508 of intra prediction modes is to be used to predict the predetermined block 18, the decoder / encoder is configured to form a list 542 of most probable intra prediction modes based on the intra prediction modes used to predict neighboring blocks 524, 526 adjacent to the predetermined block 18, and derive an MPM list index 540 pointing to the predetermined intra prediction mode in the list 542 of most probable intra prediction modes from the data stream 12, or signal the MPM list index 540 pointing to the predetermined intra prediction mode in the list 542 of most probable intra prediction modes to the data stream 12. If the set selection syntax element 522 indicates that one of the intra-prediction modes in the first set 508 of intra-prediction modes is not to be used to predict the predetermined block 18, the decoder is configured to derive another index 540 and / or 546 from the data stream, or the encoder is configured to encode another index 540 and / or 546 into the data stream, the another index 540 and / or 546 indicating a predetermined intra-prediction mode 604 in the second set 520 of matrix-based intra-prediction modes 510.
[0549] about Figure 13 The decoder and encoder described may include information about Figure 12 Other features and / or functionality described.
[0550] In the case that the predetermined intra prediction mode 604 of the predetermined block 18 is the matrix-based intra prediction mode 510 in the second set 520, the means for decoding the predetermined block 18 (i.e., according to Figure 13 decoder) and / or means for encoding the predetermined block 18 (i.e., according to Figure 13 encoder) may include one or more of the following features.
[0551] According to an embodiment, the apparatus is configured to form a sample value vector (e.g., with respect to Figures 6 to 9b The sample value vector 400 described in one embodiment of the present invention is used as the sample value vector 400, and the vector 514 is derived from the sample value vector so that the sample value vector is mapped to the vector 514 through a predetermined reversible linear transformation. In this case, the vector 514 can be understood as another vector. For example, as described with respect to Figures 8 to 11 One embodiment of the present invention determines and / or defines vector 514 as described for another vector 402 .
[0552] According to an embodiment, the apparatus is configured to form a sample value vector based on a plurality of reference samples 17 by, for each component of the sample value vector, employing one of a plurality of reference samples as the corresponding component of the reference sample, and / or by averaging two or more components of the sample value vector to obtain the corresponding component of the sample value vector.
[0553] A plurality of reference samples 17 are arranged within the picture, for example, along the outer edge of a predetermined block 18 .
[0554] The reversible linear transformation is defined, for example, as follows: a predetermined component of vector 514 (e.g., another vector) becomes a, and each component of vector 514 other than the predetermined component is equal to the corresponding component of the sample value vector minus a. The value a is, for example, a predetermined value of 1400.
[0555] According to an embodiment, the predetermined value 1400 is one of: an average value (e.g., an arithmetic mean or a weighted average value) of the components of the sample value vector, a default value, a value signaled in the data stream into which the picture is encoded, or a component in the sample value vector corresponding to a preset component.
[0556] The reversible linear transformation is defined, for example, such that a predetermined component of vector 514 (e.g., of another vector) becomes a, and each component of vector 514 other than the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the arithmetic mean of the components of the sample value vector.
[0557] The reversible linear transformation is defined, for example, as follows: a predetermined component of vector 514 (e.g., another vector) becomes a, and each component of vector 514 other than the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the component of the sample value vector corresponding to the predetermined component. The apparatus is configured, for example, to include a plurality of reversible linear transformations, each of the plurality of reversible linear transformations being associated with a component of vector 514, select the predetermined component from the components of the sample value vector, and use the reversible linear transformation associated with the predetermined component from the plurality of reversible linear transformations as the predetermined reversible linear transformation.
[0558] According to an embodiment, the matrix components of the prediction matrix 516 within the columns of the prediction matrix 516 corresponding to the predetermined components of the vector 514 (e.g., another vector) are all zero. The apparatus is configured to calculate the matrix-vector product 512 by performing a multiplication by calculating a matrix-vector product 512 between: a reduced prediction matrix generated from the prediction matrix 516 by removing columns, and another vector 410 generated from the vector 514 by removing the predetermined components, such as Figure 9b shown.
[0559] According to an embodiment, the apparatus is configured to, when predicting samples of the predetermined block 18 based on the prediction vector 518 , calculate, for each component of the prediction vector 518 , a sum of the corresponding component and a.
[0560] The prediction matrix 516 is generated by summing each matrix element of the prediction matrix 516 corresponding to a predetermined component of the vector 514 (e.g., another vector) within the column of the prediction matrix 516 with 1 (i.e., the sum of the matrix C 405 and the matrix M 1300, as shown in FIG. Figure 9a As shown, leading to Figure 8 ), which is multiplied by a reversible linear transformation such as the machine learning prediction matrix (i.e., Figure 8 1100) corresponds to a quantized version of the prediction matrix A shown in FIG.
[0561] According to an embodiment, the apparatus is configured to compute the matrix-vector product 512 using fixed-point arithmetic operations.
[0562] According to an embodiment, the apparatus is configured to compute the matrix-vector product 512 without floating point arithmetic operations.
[0563] According to an embodiment, the apparatus is configured to store a fixed-point representation of the prediction matrix 516 .
[0564] According to an embodiment, the apparatus is configured to represent a prediction matrix 516 using prediction parameters, and to calculate the matrix-vector product 512 by performing multiplication and addition on the components of a vector 514 (e.g., another vector) and the prediction parameters and intermediate results generated thereby, wherein the absolute value of the prediction parameter can be represented by an n-bit fixed-point number representation, wherein n is equal to or less than 14, or alternatively equal to or less than 10, or alternatively equal to or less than 8. This can be combined with Figure 10 or Figure 11 similarly or as described in Figure 10 or Figure 11 Execute as described in .
[0565] The prediction parameters include, for example, weights, each weight being associated with a corresponding matrix component of the prediction matrix 516 .
[0566] The prediction parameters also include, for example: one or more scaling factors, each of the one or more scaling factors being associated with one or more corresponding matrix components in the prediction matrix 516, for scaling the weights associated with the one or more corresponding matrix components of the prediction matrix 516; and / or one or more offsets, each of the one or more offsets being associated with one or more corresponding matrix components of the prediction matrix 516, for offsetting the weights associated with the one or more corresponding matrix components of the prediction matrix 516.
[0567] According to an embodiment, the apparatus is configured to, when predicting samples of the predetermined block 18 based on the prediction vector 518 , use interpolation to calculate at least one sample position of the predetermined block 18 based on the prediction vector 518 , each component of the prediction vector being associated with a corresponding position within the predetermined block 18 .
[0568] about Figure 13 The decoder and encoder described may include information about Figures 6 to 11 Other features and / or functionality described in one or more embodiments.
[0569] Therefore, the solution provided by the present invention is to associate a specific transform set 160 of transforms in a given transform set S612, which was originally defined for conventional intra prediction modes (i.e., intra prediction modes within the first set 508), with each MIP mode 510. One specific way of doing this is that all MIP modes 510 use the LFNST transform, a secondary transform originally designed for planar mode 504 and DC mode 506. For example, due to the optional application of downsampling to the downscaled prediction signal, MIP modes 510 are non-directional, just like planar mode 504 and DC mode 506, and therefore their prediction residuals have more statistical similarity to DC mode 506 and planar mode 504 than to angular mode 500.
[0570] It may turn out that allowing LFNST, i.e., allowing the use of a secondary transform for MIP 510, may be beneficial in terms of coding efficiency for some block shapes, but not for other block shapes, meaning that for the latter shapes, the additional signaling cost of signaling whether or not LFNST is to be applied to block 18 in MIP mode 510 is, on average, higher than the gain obtained by allowing the LFNST transform for MIP mode 510. Therefore, embodiments of the present invention allow the aforementioned combination of LFNST and MIP 510 only for a subset of all those block shapes for which such a combination is in principle possible.
[0571] Figure 17A method 6000 is shown for decoding a predetermined block (18) of a picture using intra prediction, comprising: selecting (602) a predetermined intra prediction mode (604) from a plurality (600) of intra prediction modes based on a data stream, the plurality (600) of intra prediction modes comprising a first set (508) of intra prediction modes and a second set (520) of matrix-based intra prediction modes (510), the first set (508) of intra prediction modes comprising a DC intra prediction mode (506) and an angular prediction mode (500). and an optional planar intra-frame prediction mode (504), for each matrix-based intra-frame prediction mode (510) in a second set (520) of matrix-based intra-frame prediction modes (510), obtaining a prediction vector (518) using a matrix-vector product (512) between: a vector (514) derived from reference samples (17) in the forest of the predetermined block and a prediction matrix (516) associated with the corresponding matrix-based intra-frame prediction mode, predicting samples of the predetermined block based on the prediction vector (518). A prediction signal (606) for a predetermined block is derived 6100 using a predetermined intra prediction mode, and a subset (610) of one or more secondary transforms is selected (608) from a set (612) of secondary transforms in a manner that depends on the predetermined intra prediction mode such that the subset (610) is non-empty if the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510). The method 6000 includes deriving (614) a transformed version (616) of a prediction residual for a predetermined block (18) from a data stream, if the predetermined intra prediction mode is included in a first set (508) of intra prediction modes and if the predetermined intra prediction mode is included in a second set (520) of matrix-based intra prediction modes (510), the transformed version (616) of the prediction residual for the predetermined block (18) being related to a spatial domain version (618) of the prediction residual for the predetermined block (18) via a transform (T) defined by a concatenation of a primary transform (Tp) and a predetermined secondary transform (Ts) applied to a subset (622) of coefficients (620) of the primary transform in a subset (610) of the secondary transform. Furthermore, the method 6000 includes reconstructing (624) the predetermined block using the prediction signal and the prediction residual for the predetermined block (18).
[0572] Figure 18A method 7000 is shown for encoding a predetermined block (18) of a picture using intra prediction, comprising: selecting (602) a predetermined intra prediction mode (604) from a plurality (600) of intra prediction modes, the plurality (600) of intra prediction modes comprising a first set (508) of intra prediction modes and a second set (520) of matrix-based intra prediction modes (510), the first set (508) of intra prediction modes comprising a DC intra prediction mode (506) and an angular prediction mode (500). , for each matrix-based intra prediction mode (510) in a second set (520) of matrix-based intra prediction modes (510), obtaining a prediction vector (518) using a matrix-vector product (512) between: a vector (514) derived from reference samples (17) in a neighborhood of the predetermined block (18) and a prediction matrix (516) associated with the corresponding matrix-based intra prediction mode (510), and predicting samples of the predetermined block (18) based on the prediction vector (518). Additionally. The method 7000 comprises signaling 7100 a predetermined intra-frame prediction mode (604) in a data stream, deriving 7200 a prediction signal (606) for a predetermined block (18) using the predetermined intra-frame prediction mode, and selecting (608) a subset (610) of one or more secondary transforms from a set (612) of secondary transforms in a manner that depends on the predetermined intra-frame prediction mode such that the subset (610) is non-empty if the predetermined intra-frame prediction mode is included in the first set (508) of intra-frame prediction modes and the predetermined intra-frame prediction mode is included in the second set (520) of matrix-based intra-frame prediction modes (510). The method comprises encoding (614) a transformed version (616) of a prediction residual of a predetermined block (18) into a data stream in case the predetermined intra prediction mode is included in a first set (508) of intra prediction modes and in case the predetermined intra prediction mode is included in a second set (520) of matrix-based intra prediction modes (510), the transformed version (616) of the prediction residual of the predetermined block (18) being related to a spatial domain version (618) of the prediction residual of the predetermined block (18) via a transform (T), the transform (T) being defined by a concatenation of a primary transform (Tp) and a predetermined secondary transform (Ts) applied to a subset (622) of coefficients (620) of the primary transform in a subset (610) of the secondary transform, wherein the predetermined block (18) can be reconstructed (624) using the prediction signal (606) and the prediction residual of the predetermined block (18).
[0573] References
[0574] [1] P. Helle et al., “Non-linear weighted intra prediction,” JVET-L0199, Macau, China, October 2018.
[0575] [2] F. Bossen, J. Boyce, K. Suehring, X. Li, V. Seregin, “JVET common test conditions and software reference configurations for SDR video”, JVET-K1010, Ljubljana, Slovenia, July 2018.
[0576] Other embodiments and examples
[0577] Generally, the examples can be implemented as a computer program product having program instructions, the program instructions being operable to perform one of the methods when the computer program product runs on a computer. The program instructions can be stored, for example, on a machine-readable medium.
[0578] Other examples comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0579] In other words, a method example is, therefore, a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
[0580] Therefore, another example of a method is a data carrier medium (or digital storage medium or computer-readable medium) on which is recorded a computer program for performing one of the methods described herein. A data carrier medium, digital storage medium or recorded medium is tangible and / or non-transitory, as opposed to an intangible and transitory signal.
[0581] A further example of a method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.The data stream or the sequence of signals may be transmitted, for example, via a data communication connection, for example via the Internet.
[0582] Another example comprises a processing device, such as a computer or a programmable logic device, which performs one of the methods described herein.
[0583] A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0584] Another example includes an apparatus or system for transmitting a computer program for performing one of the methods described herein to a receiver (e.g., electronically or optically). The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0585] In some examples, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods can be performed by any suitable hardware device.
[0586] The above examples are merely illustrative of the principles disclosed above. It should be understood that modifications and variations of the arrangements and details described herein will be apparent. Accordingly, it is intended that the scope of the appended claims be limiting rather than the specific details provided by way of description and explanation of the examples herein.
[0587] In the following description, the same or equivalent elements or elements having the same or equivalent functions are denoted by the same or equivalent reference numerals (even if they are shown in different drawings).
Claims
1. A method for decoding a picture from a data stream, the method comprising: selecting, for a block of the picture, a planar intra prediction mode or a matrix-based intra prediction mode based at least in part on an indication included in the data stream; deriving a prediction signal for the block using the selected intra prediction mode; Selecting a subset of secondary transforms to transform the prediction residual, the selected subset of secondary transforms including at least one low frequency non-separable secondary transform (LFNST), wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block from the data stream; transforming the prediction residuals using LFNST from the selected subset; and The block is reconstructed using the prediction signal and the transformed prediction residual for the block.
2. The method according to claim 1, wherein When the matrix-based intra prediction mode is selected, deriving the prediction signal comprises: deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the selected intra prediction mode, and The prediction vector is upsampled to obtain the prediction signal for the block.
3. The method according to claim 1, wherein A subset of the secondary transforms is selected based on: The selected intra prediction mode, and The size of the block.
4. The method according to claim 1, further comprising: decoding an indication indicating a LFNST from the selected subset of secondary transforms; as well as Based on an indication specifying one of two LFNSTs, the LFNST is selected from the subset to transform the prediction residual.
5. The method according to claim 1, wherein Transforming the prediction residual includes: Applying the LFNST (Ts) to a subset of the coefficients of the primary transform (Tp) to obtain a transform; and transforming the prediction residual using the transform, Herein, the primary transform (Tp) is a separable 2D transform.
6. The method according to claim 1, wherein Selecting an intra prediction mode for the block of the picture includes: decoding a set of syntax elements from the data stream, the set of syntax elements indicating whether to predict the block using one of a set of intra prediction modes, the set of intra prediction modes comprising the planar intra prediction mode, at least one angular prediction mode, and a DC intra prediction mode; In response to determining that the set of syntax elements indicates that one of the intra prediction modes in the set of intra prediction modes is to be used to predict the block, generating a list of most probable intra prediction modes (MPMs) based on intra prediction modes used by blocks neighboring the block, and selecting the intra prediction mode from the list; and In response to determining that the set of syntax elements indicates that the block is not predicted using any intra prediction mode in the set of intra prediction modes, the matrix-based intra prediction mode is selected from a set of matrix-based intra prediction modes.
7. The method according to claim 1, wherein Transforming the prediction residual includes: A transform defined by the concatenation of a primary transform (Tp) and a secondary transform (Ts) selected from a subset of said secondary transforms is used, said secondary transform (Ts) corresponding to LFNST.
8. An apparatus for decoding a picture from a data stream, the apparatus comprising at least one processor, the at least one processor being configured to: selecting, for a block of the picture, a planar intra prediction mode or a matrix-based intra prediction mode based at least in part on an indication included in the data stream; deriving a prediction signal for the block using the selected intra prediction mode; Selecting a subset of secondary transforms to transform the prediction residual, the selected subset of secondary transforms including at least one low frequency non-separable secondary transform (LFNST), wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block from the data stream; transforming the prediction residuals using LFNST from the selected subset; and The block is reconstructed using the prediction signal and the transformed prediction residual for the block.
9. The device according to claim 8, wherein When the matrix-based intra prediction mode is selected, to derive the prediction signal, the at least one processor is configured to: deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the selected intra prediction mode, and The prediction vector is upsampled to obtain the prediction signal for the block.
10. The device according to claim 8, wherein The at least one processor is further configured to select a subset of the secondary transforms based on: The selected intra prediction mode, and The size of the block.
11. The device according to claim 8, wherein The at least one processor is further configured to: decoding an indication indicating a LFNST from the selected subset of secondary transforms; and Based on an indication specifying one of two LFNSTs, the LFNST is selected from the subset to transform the prediction residual.
12. The device according to claim 8, wherein In order to transform the prediction residual, the at least one processor is configured to: Applying the LFNST (Ts) to a subset of the coefficients of the primary transform (Tp) to obtain a transform; and transforming the prediction residual using the transform, Herein, the primary transform (Tp) is a separable 2D transform.
13. The device according to claim 8, wherein To select an intra prediction mode for the block of the picture, the at least one processor is configured to: decoding a set of syntax elements from the data stream, the set of syntax elements indicating whether to predict the block using one of a set of intra prediction modes, the set of intra prediction modes comprising the planar intra prediction mode, at least one angular prediction mode, and a DC intra prediction mode; In response to determining that the set of syntax elements indicates that one of the intra prediction modes is to be used to predict the block, generating a list of most probable intra prediction modes (MPMs) based on intra prediction modes used by blocks neighboring the block, and selecting the intra prediction mode from the list; as well as In response to determining that the set of syntax elements indicates that the block is not predicted using any intra prediction mode in the set of intra prediction modes, the matrix-based intra prediction mode is selected from a set of matrix-based intra prediction modes.
14. The device according to claim 8, wherein In order to transform the prediction residual, the at least one processor is configured to: A transform defined by the concatenation of a primary transform (Tp) and a secondary transform (Ts) selected from a subset of said secondary transforms is used, said secondary transform (Ts) corresponding to LFNST.
15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of an electronic device, cause the electronic device to: selecting, for a block of the picture, a planar intra prediction mode or a matrix-based intra prediction mode based at least in part on an indication included in the data stream; deriving a prediction signal for the block using the selected intra prediction mode; Selecting a subset of secondary transforms to transform the prediction residual, the selected subset of secondary transforms including at least one low frequency non-separable secondary transform (LFNST), wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block from the data stream; transforming the prediction residuals using LFNST from the selected subset; and The block is reconstructed using the prediction signal and the transformed prediction residual for the block.
16. The non-transitory computer-readable medium of claim 15, wherein: The instructions that, when executed, cause the at least one processor to derive the prediction signal when the matrix-based intra prediction mode is selected include instructions that, when executed, cause the at least one processor to: deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the selected intra prediction mode, and The prediction vector is upsampled to obtain the prediction signal for the block.
17. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed, cause the at least one processor to select a subset of the secondary transforms based on: The selected intra prediction mode, and The size of the block.
18. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed, cause the at least one processor to: decoding an indication indicating a LFNST from the selected subset of secondary transforms; and Based on an indication specifying one of two LFNSTs, the LFNST is selected from the subset to transform the prediction residual.
19. The non-transitory computer-readable medium of claim 15, wherein: The instructions that, when executed, cause the at least one processor to transform the prediction residual include instructions that, when executed, cause the at least one processor to perform the following operations: Applying the LFNST (Ts) to a subset of the coefficients of the primary transform (Tp) to obtain a transform; and transforming the prediction residual using the transform, Herein, the primary transform (Tp) is a separable 2D transform.
20. The non-transitory computer-readable medium of claim 15, wherein: The instructions that, when executed, cause the at least one processor to select an intra prediction mode for the block of the picture include instructions that, when executed, cause the at least one processor to: decoding a set of syntax elements from the data stream, the set of syntax elements indicating whether to predict the block using one of a set of intra prediction modes, the set of intra prediction modes comprising the planar intra prediction mode, at least one angular prediction mode, and a DC intra prediction mode; In response to determining that the set of syntax elements indicates that one of the intra prediction modes in the set of intra prediction modes is to be used to predict the block, generating a list of most probable intra prediction modes (MPMs) based on intra prediction modes used by blocks neighboring the block, and selecting the intra prediction mode from the list; and In response to determining that the set of syntax elements indicates that the block is not predicted using any intra prediction mode in the set of intra prediction modes, the matrix-based intra prediction mode is selected from a set of matrix-based intra prediction modes.
21. The non-transitory computer-readable medium of claim 15, wherein: The instructions that, when executed, cause the at least one processor to transform the prediction residual include instructions that, when executed, cause the at least one processor to perform the following operations: A transform defined by the concatenation of a primary transform (Tp) and a secondary transform (Ts) selected from a subset of said secondary transforms is used, said secondary transform (Ts) corresponding to LFNST.
22. A method for encoding a picture into a data stream, the method comprising: generating, for a block of the picture, an indication of a selected intra prediction mode, the selected intra prediction mode being a planar intra prediction mode or a matrix-based intra prediction mode; encoding a prediction signal for the block using the selected intra prediction mode; selecting a subset of secondary transforms, wherein the selected subset of secondary transforms includes at least one low frequency non-separable secondary transform (LFNST), and wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block; transforming the prediction residuals using LFNST from the selected subset; and A transformed prediction residual for the block is encoded into the data stream.
23. The method according to claim 22, further comprising: responsive to the matrix-based intra prediction mode being the selected intra prediction mode, deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the matrix-based intra prediction mode, and The prediction signal is generated based on the prediction vector.
24. The method of claim 22, further comprising: A subset of the secondary transforms is selected based on: The selected intra prediction mode, and The size of the block.
25. An apparatus for encoding a picture into a data stream, the apparatus comprising at least one processor, the at least one processor being configured to: generating, for a block of the picture, an indication of a selected intra prediction mode, the selected intra prediction mode being a planar intra prediction mode or a matrix-based intra prediction mode; encoding a prediction signal for the block using the selected intra prediction mode; selecting a subset of secondary transforms, wherein the selected subset of secondary transforms includes at least one low frequency non-separable secondary transform (LFNST), and wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block; transforming the prediction residuals using LFNST from the selected subset; and A transformed prediction residual for the block is encoded into the data stream.
26. The method according to claim 25, wherein In response to the matrix-based intra prediction mode being the selected intra prediction mode, the at least one processor is configured to: deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the matrix-based intra prediction mode, and The prediction signal is generated based on the prediction vector.
27. The method according to claim 25, wherein The at least one processor is configured to select a subset of the secondary transforms based on: The selected intra prediction mode, and The size of the block.
28. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of an electronic device, cause the electronic device to: generating, for a block of a picture, an indication of a selected intra prediction mode, the selected intra prediction mode being a planar intra prediction mode or a matrix-based intra prediction mode; encoding a prediction signal for the block using the selected intra prediction mode; selecting a subset of secondary transforms, wherein the selected subset of secondary transforms includes at least one low frequency non-separable secondary transform (LFNST), and wherein the selected subset of secondary transforms is the same for the planar intra prediction mode and the matrix-based intra prediction mode; deriving a prediction residual for the block; transforming the prediction residuals using LFNST from the selected subset; and The transformed prediction residual for the block is encoded into a data stream.
29. The non-transitory computer-readable medium of claim 28, wherein: In response to the matrix-based intra prediction mode being the selected intra prediction mode, the instructions, when executed, cause the at least one processor to: deriving a prediction vector generated from a matrix-vector product between: a vector derived from reference samples in a neighborhood of the block and a prediction matrix associated with the matrix-based intra prediction mode, and The prediction signal is generated based on the prediction vector.
30. The non-transitory computer readable medium of claim 28, wherein: The instructions, when executed, cause the at least one processor to select a subset of the secondary transforms based on: The selected intra prediction mode, and The size of the block.