Coding Using Row-Column Based Intra Prediction and Secondary Transformation
By selecting subsets of secondary transforms that include both matrix-based and non-matrix-based intra prediction modes, the video coding technology addresses memory and efficiency challenges, achieving improved coding efficiency with managed signaling costs.
Patent Information
- Application Number
- JP2024067724
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-25
- Filing Date
- 2024-04-18
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2040-06-23
AI Technical Summary
Existing video coding technologies face challenges in efficiently supporting matrix-based intra prediction modes due to high memory requirements for storing specific secondary transforms for each mode.
The approach involves selecting subsets of secondary transforms from a set that includes transforms for both matrix-based and non-matrix-based intra prediction modes, reducing memory requirements and allowing for coding efficiency improvements.
This solution enhances coding efficiency while managing the increased signaling cost associated with indicating the use of secondary transforms for matrix-based intra prediction modes.
Smart Images

Figure 0007695440000089 
Figure 0007695440000090 
Figure 0007695440000091
Abstract
Description
Technical Field
[0001] This application relates to the field of matrix-based intra prediction and secondary transformation.
Background Art
[0002] In conventional intra prediction modes such as Planar mode, DC mode, and Angular mode, non-separable linear fractional non-separable transform (LFNST) is a tool used to transform the prediction residuals corresponding to these intra prediction modes. Here, a set S of a plurality of transform sets is given such that each conventional intra prediction mode is associated with one of these transform sets. Then, in the decoder, it can be extracted from the bitstream whether LFNST should be applied on a given block. If it should be applied, depending on the intra prediction mode used on the current block, one transform set from set S is given, and if this transform set consists of more than one transform, it can be extracted from the bitstream which transform T in this set should be used. Then, in the decoder, transform T is applied as a secondary transform, which means that it is applied to a subset of the residual transform coefficients of a separable primary transform. However, the aforementioned secondary transform is only defined a priori for conventional intra prediction modes.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] Therefore, it is desirable to provide a concept for rendering more efficient picture coding and / or video coding to support matrix-based intra prediction (MIP), i.e., a secondary transformation for block-based intra prediction.
[0005] This is achieved by the subject matter of the independent claims of this application.
[0006] Further embodiments according to the invention are defined by the subject matter of the dependent claims of this application.
Means for Solving the Problems
[0007] According to a first aspect of the present invention, the inventors of the present application recognized that one problem encountered when attempting to associate secondary transforms with matrix-based intra prediction modes is that providing a specific secondary transform for each MIP mode can be prohibitively expensive in terms of the memory requirements for storing additional transforms. According to a first aspect of the present application, this drawback is overcome by selecting one or more subsets of secondary transforms from a set of secondary transforms that includes transforms associated with matrix-based intra prediction modes and transforms associated with non-matrix-based intra prediction modes. Those secondary transforms in the set of secondary transforms may be defined for one or more prediction modes, which reduces the memory capacity required for the set of secondary transforms. The transforms defined for the Planar intra prediction mode and / or the transforms defined for the DC intra prediction mode may also be selectable for the matrix-based intra prediction mode. This specific selection of a subset of secondary transforms for the matrix-based intra prediction mode can increase coding efficiency, but the additional syntax elements required for blocks associated with the matrix-based intra prediction mode to indicate the use of the secondary transform can increase the bitstream and thus the signaling cost.
[0008] Accordingly, according to a first aspect of the present application, an apparatus for decoding a predetermined block of a picture using intra prediction, i.e., a decoder, is configured to select a predetermined intra prediction mode from a plurality of intra prediction modes including a first set of intra prediction modes and a second set of matrix-based intra prediction modes based on a data stream. The first set of intra prediction modes includes a DC intra prediction mode, an Angular prediction mode, and optionally a Planar intra prediction mode. When a certain matrix-based intra prediction mode in the second set is selected as the predetermined intra prediction mode, the decoder is configured to obtain a prediction vector using a matrix vector product of a vector derived from reference samples in the vicinity of the predetermined block and a prediction matrix associated with each matrix-based intra prediction mode, and based on this, the decoder is configured to predict samples of the predetermined block. The decoder derives a prediction signal for the predetermined block using the predetermined intra prediction mode, and selects one or a plurality of subsets of secondary transformations from a set of secondary transformations in a manner dependent on the predetermined intra prediction mode such that the subset is non-empty when the predetermined intra prediction mode is included in the first set of intra prediction modes and when the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. The first set and the second set define intra prediction modes for which secondary transformations are available. Accordingly, for a predetermined intra prediction mode selected from the first set or the second set, the decoder is configured to select one or a plurality of subsets of secondary transformations from a set of secondary transformations particularly associated with the selected predetermined intra prediction mode.In addition, the decoder is configured to derive a transformed version of the prediction residual for a given block related to a spatial region version of the prediction residual of the given block via a transformation defined by concatenating a primary transformation and a given secondary transformation applied to a subset of the coefficients of the primary transformation from the data stream when the given intra prediction mode is included in a first set of intra prediction modes and when the given intra prediction mode is included in a second set of matrix-based intra prediction modes. The primary transformation is, for example, set by default, and the given secondary transformation is selected, for example, from a subset of secondary transformations by the decoder. The decoder may be configured to select the given secondary transformation from the subset of secondary transformations by deriving a syntax element indicating the secondary transformation from the data stream. The decoder is configured to reconstruct the given block using the prediction signal and the prediction residual for the given block.
[0009] According to a first aspect of the present application, similar to a decoder, an apparatus for encoding a predetermined block of a picture using intra prediction, i.e., an encoder, is configured to select a predetermined intra prediction mode from a plurality of intra prediction modes including a first set of intra prediction modes including a DC intra prediction mode, an Angular prediction mode, and optionally a Planar intra prediction mode, and a second set of matrix-based intra prediction modes. According to each of the matrix-based intra prediction modes, the matrix-vector product of a vector derived from reference samples in the vicinity of a predetermined block and a prediction matrix associated with each matrix-based intra prediction mode is used to obtain a prediction vector, and based on this, the samples of the predetermined block are predicted. The encoder is configured to signal a predetermined intra prediction mode in a data stream and derive a prediction signal for a predetermined block using the predetermined intra prediction mode. In addition, the encoder is configured to select one or a plurality of subsets of secondary transformations from a set of secondary transformations in a manner dependent on the predetermined intra prediction mode such that the subset is non-empty when the predetermined intra prediction mode is included in the first set of intra prediction modes and when the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. The encoder is configured to encode a transformed version of the prediction residual for a predetermined block related to the spatial region version of the prediction residual of the predetermined block into the data stream via a transformation defined by the concatenation of a primary transformation and a predetermined secondary transformation from a subset of the secondary transformation applied to a subset of the coefficients of the primary transformation when the predetermined intra prediction mode is included in the first set of intra prediction modes and when the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. The predetermined block is reconstructable using the prediction signal and the prediction residual for the predetermined block.
[0010] According to one embodiment, the decoder / encoder is configured to select one or more subsets of the set of secondary transforms in a manner dependent on a given intra prediction mode such that each secondary transform of the set of secondary transforms is included in one or more subsets of secondary transforms selected for at least one of the intra prediction modes in the first and second sets. For one or more intra prediction modes within the first set or the second set, each selected subset may include all of the secondary transforms of the set of secondary transforms.
[0011] According to an embodiment, the decoder / encoder is configured to select one or more subsets of the set of secondary transforms in a manner dependent on a given intra prediction mode such that each secondary transform of each subset of secondary transforms selected for any matrix-based intra prediction mode is included in a subset of secondary transforms selected for at least one intra prediction mode within a first set that does not belong to an Angular prediction mode. The subset selected for a matrix-based intra prediction mode may include secondary transforms associated with one or more non-Angular intra prediction modes within the first set, such as a DC intra prediction mode and / or a Planar intra prediction mode. Thus, those secondary transforms from the set of secondary transforms may be part of more than one subset of one or more secondary transforms. Additional secondary transforms that are only available for blocks where the given intra prediction mode is a matrix-based intra prediction mode may not be included in the set of secondary transforms. For a given block where the given intra prediction mode is a matrix-based intra prediction mode, the decoder / encoder is configured to select the same secondary transforms from the set of secondary transforms for a given intra prediction mode that is a non-Angular prediction mode from the first set. The subset selected for a matrix-based intra prediction mode may be equal to the subset of secondary transforms selected for an intra prediction mode within the first set that does not belong to an Angular prediction mode, or may include some of the secondary transforms of the subset of secondary transforms selected for an intra prediction mode within the first set that does not belong to an Angular prediction mode, or may include some or all of the secondary transforms of two or more subsets of secondary transforms selected for an intra prediction mode within the first set that does not belong to an Angular prediction mode.
[0012] A method for encoding or decoding is based on the same considerations as the above-described apparatus for encoding or decoding. Incidentally, the method may be completed using all the features and functions described also with respect to the apparatus for encoding or decoding.
[0013] The drawings are not necessarily to scale; instead, it is generally emphasized that they show the principles of the present invention. In the following description, various embodiments of the present invention will be described with reference to the following drawings.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7.1
Figure 7.2
Figure 7.3
Figure 7.4
Figure 8
Figure 9a
Figure 9b
Figure 9c
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19a
Figure 19b
Figure 19c
Figure 19d
[0015] Equal or equivalent elements, or elements having equal or equivalent functions, are denoted in the following description by equal or equivalent reference numerals, even if they are present in different figures.
[0016] In the following description, numerous specific details are set forth in order to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to one skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present invention. Additionally, features of different embodiments described later in this specification may be combined with each other unless otherwise specifically stated.
[0017] 1 Introduction In the following, different inventive examples, embodiments, and aspects are described. At least some of these examples, embodiments, and aspects refer, inter alia, to methods and / or apparatuses for video coding and / or for performing intra prediction using, for example, a linear or affine transform with reduction of neighboring samples and / or for optimizing video delivery (such as broadcast, streaming, file playback, etc.) for, for example, video applications and / or virtual reality applications.
[0018] Furthermore, examples, embodiments, and aspects may refer to High Efficiency Video Coding (HEVC) or successors. Or, further embodiments, examples, and aspects are defined by the appended claims.
[0019] Note that any embodiment, example, and aspect as defined by the claims may be supplemented by any of the details (features and functions) described in the following chapters.
[0020] Also, the embodiments, examples, and aspects described in the following chapters may be used individually, or may be supplemented by any of the features in other chapters or any of the features included in the claims.
[0021] Also, note that the individual examples, embodiments, and aspects described herein may be used individually or in combination. Thus, details may be added to each of the individual aspects without adding details to another one of the examples, embodiments, and aspects.
[0022] It should also be noted that the present disclosure explicitly or implicitly describes features of encoding and / or decoding systems and / or methods.
[0023] Moreover, the features and functions disclosed herein with respect to the method may also be used in an apparatus. Further, any of the features and functions disclosed herein with respect to the apparatus may be used in the corresponding method. In other words, the methods disclosed herein may be supplemented by any of the features and functions described with respect to the apparatus.
[0024] Also, any of the features and functions described herein may be implemented in hardware or software, or using a combination of hardware and software, as described in the "Alternative Implementations" section.
[0025] Moreover, any of the features described in parentheses ("(...) " or "[...]") may be considered optional in some examples, embodiments, or aspects.
[0026] 2 Encoder, Decoder Examples that may help achieve more effective compression when using block-based prediction are described below. Some examples achieve high compression efficiency by using a set of intra prediction modes. The latter may be added to other heuristically designed intra prediction modes, for example, or provided exclusively. And other examples utilize both of the characteristics discussed immediately above. As a variation of these embodiments, however, intra prediction may be changed to inter prediction by instead using reference samples in another picture.
[0027] To facilitate understanding of the following examples of the present application, the description begins with the presentation of an encoder and a decoder that may conform to them, and the examples outlined later in the present application may be built upon them. FIG. 1 shows an apparatus for encoding picture 10 into data stream 12 block by block. The apparatus is shown using reference numeral 14 and may be a still picture encoder or a video encoder. In other words, picture 10 may be the current picture from video 16 when encoder 14 is configured to encode video 16 including picture 10 into data stream 12, or when encoder 14 may encode picture 10 into data stream 12 exclusively.
[0028] As mentioned, encoder 14 performs encoding in a block-by-block fashion, or on a block basis. To this end, encoder 14 subdivides picture 10 into blocks, and on a block-by-block basis, encoder 14 encodes picture 10 into data stream 12. Examples of possible subdivisions of picture 10 into blocks 18 are described in more detail below. In general, the subdivision involves a hierarchical quadtree subdivision, such as starting with a quadtree subdivision from the overall picture area of picture 10, or from a pre-partitioning of picture 10 into an array of tree blocks, resulting in blocks 18 of a certain size, such as an array of blocks arranged in rows and columns, or blocks 18 of different block sizes, and these examples should not be treated as excluding other possible ways of subdividing picture 10 into blocks 18.
[0029] Furthermore, encoder 14 is a predictive encoder configured to predictively encode picture 10 into data stream 12. For a given block 18, this means that encoder 14 determines a prediction signal for block 18 and encodes the prediction residual, i.e., the prediction error by which the prediction signal deviates from the actual picture content within block 18, into data stream 12.
[0030] Encoder 14 may support different prediction modes to derive a prediction signal for a certain block 18. The prediction mode that is important in the following example is the intra prediction mode in which the inside of block 18 is spatially predicted from already encoded samples in the vicinity of picture 10 accordingly. The encoding of picture 10 into data stream 12, and thus the corresponding decoding procedure, may be based on a coding order 20 defined among blocks 18. For example, the coding order 20 may scan blocks 18 in a raster scan order, such as row by row from top to bottom while scanning each row from left to right. In the case of hierarchical quadtree-based subdivision, a raster scan order may be applied within each hierarchical level, where a depth-first scan order may also be applied, i.e., a leaf node within a block of a certain hierarchical level may be before a block of the same hierarchical level having the same parent block according to coding order 20. Depending on coding order 20, the already encoded samples in the vicinity of block 18 may typically be located on one or more sides of block 18. In the case of the example presented herein, for example, the already encoded samples in the vicinity of block 18 are located above and to the left of block 18.
[0031] The intra prediction mode may not be the only mode supported by encoder 14. If encoder 14 is a video encoder, for example, encoder 14 may also support an inter prediction mode in which block 18 is temporally predicted from a previously encoded picture of video 16 accordingly. Such an inter prediction mode may be a motion-compensated prediction mode, according to which a motion vector indicating a relative spatial offset of the portion from which the prediction signal of block 18 should be derived as a copy is signaled for such a block 18. Additionally or alternatively, other non-intra prediction modes may be available, such as an inter prediction mode when encoder 14 is a multi-view encoder, or a non-prediction mode in which the inside of block 18 is encoded as is, i.e., without any prediction.
[0032] Before beginning to focus the description of the present application into the intra prediction mode, a more specific example for a possible block-based encoder, i.e., a possible implementation form of encoder 14, is described with reference to FIG. 2, and then two corresponding examples of decoders that conform to FIGS. 1 and 2 are presented respectively.
[0033] FIG. 2 shows a possible implementation of the encoder 14 of FIG. 1, i.e., an implementation in which the encoder is configured to use transform coding to encode the prediction residual, which is mostly an example, and the present application is not limited to that kind of prediction residual coding. According to FIG. 2, the encoder 14 includes a subtractor 22 configured to subtract the current block 18 and the corresponding prediction signal 24 from the incoming signal, i.e., the picture 10, or block by block, in order to obtain a prediction residual signal 26 that will later be encoded into the data stream 12 by the prediction residual encoder 28. The prediction residual encoder 28 consists of an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible stage 28a includes a quantizer 30 that receives the prediction residual signal 26 and quantizes the samples of the prediction residual signal 26. As already mentioned above, this example uses transform coding of the prediction residual signal 26, and accordingly, the irreversible coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to use the quantization of the quantizer 30 for the transformed coefficients while presenting the residual signal 26 to transform such spectrally decomposed prediction residual 26. The transform can be a DCT, DST, FFT, Hadamard transform, etc. The transformed and quantized prediction residual signal 34 then undergoes reversible coding by the reversible coding stage 28b, which is an entropy encoder that entropy encodes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal in a manner available also at the decoder, i.e., taking into account the coding loss. For this purpose, the prediction residual reconstruction stage 36 includes an inverse quantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs an inverse transformation with respect to the transformation performed by a transformer 32 such as an inverse of the spectral decomposition of any of the specific transform examples mentioned above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transformer 40 and the prediction signal 24 to output the reconstructed signal, i.e., the reconstructed samples.This output is supplied to the predictor 44 of the encoder 14, and the encoder 14 then determines the prediction signal 24 based thereon. The predictor 44 supports all the prediction modes already discussed above with respect to FIG. 1. FIG. 2 also shows that when the encoder 14 is a video encoder, the encoder 14 also includes an in-loop filter with a filter-fully reconstructed picture, and after the picture is filtered, it forms a reference picture for the predictor 44 with respect to the inter-predicted blocks.
[0034] As described above, encoder 14 operates on a block basis. In the following description, the basis of the block of interest is such that picture 10 is re - divided into blocks for which an intra - prediction mode is selected from a set of intra - prediction modes or a plurality of intra - prediction modes respectively supported by predictor 44 or encoder 14, and the selected intra - prediction modes are executed individually. However, there may also be other types of blocks into which picture 10 is re - divided. For example, the above - mentioned determination of whether picture 10 is inter - coded or intra - coded can be made at a granularity or block unit that deviates from block 18. For example, the inter / intra - mode determination may be made at the level of coding blocks into which picture 10 is re - divided, and each coding block is re - divided into prediction blocks. Prediction blocks associated with coding blocks for which intra - prediction is determined to be used are each re - divided into a determination of an intra - prediction mode. For each of these prediction blocks, it is determined which supported intra - prediction mode should be used for each prediction block. These prediction blocks form block 18, which is of interest here. Prediction blocks within coding blocks associated with inter - prediction are treated differently by predictor 44. They are inter - predicted from a reference picture by determining a motion vector and copying a prediction signal for this block from a position in the reference picture indicated by the motion vector. Another block re - division is related to re - division into transform blocks, and the transformation by transformer 32 and inverse transformer 40 is performed in units of transform blocks. The transformed blocks can be, for example, the result of further re - division coding blocks. Of course, the examples described herein should be treated as non - limiting, and there are other examples. For the sake of completeness, the re - division into coding blocks may use, for example, quadtree re - division, and prediction blocks and / or transform blocks can also be obtained by further re - dividing coding blocks using quadtree re - division.
[0035] A decoder 54 or apparatus for per-block decoding adapted to the encoder 14 of FIG. 1 is illustrated in FIG. 3. This decoder 54 does the opposite of the encoder 14, i.e., it decodes the picture 10 from the data stream 12 per block and, for this purpose, supports a plurality of intra prediction modes. The decoder 54 may include, for example, a residual provider 156. All other possibilities discussed above with respect to FIG. 1 are also valid for the decoder 54. For this purpose, the decoder 54 may be a still picture decoder or a video decoder, and all prediction modes and prediction possibilities are also supported by the decoder 54. The difference between the encoder 14 and the decoder 54 is mainly in the fact that the encoder 14 makes or selects coding decisions according to some optimization, for example, to minimize some cost function that may depend on the coding rate and / or coding distortion. One of these coding options or coding parameters may involve the selection of the intra prediction mode to be used for the current block 18 among the available or supported intra prediction modes. The selected intra prediction mode may then be signaled by the encoder 14 for the current block 18 in the data stream 12, and the decoder 54 uses this signaling in the data stream 12 for the block 18 to make the selection again. Similarly, the subdivision of the picture 10 into the block 18 may be optimized within the encoder 14, and the corresponding subdivision information may be conveyed in the data stream 12, and the decoder 54 restores the subdivision of the picture 10 into the block 18 based on the subdivision information. Summarizing the above, the decoder 54 may be a prediction decoder operating per block, and except for the intra prediction mode, the decoder 54 may support other prediction modes, such as the inter prediction mode when the decoder 54 is a video decoder.In decoding, decoder 54 may also use the coding order 20 discussed with respect to FIG. 1, and since both encoder 14 and decoder 54 follow this coding order 20, the same neighboring samples are available for the current block 18 in both encoder 14 and decoder 54. Thus, to avoid unnecessary repetition, the description of the operation mode of encoder 14 shall also apply to decoder 54 as long as the subdivision of picture 10 into blocks is involved, for example, as long as prediction is involved, and as long as the coding of prediction residuals is involved. The difference is that encoder 14, by optimization, selects some coding option or coding parameter, signals or inserts the coding parameter within data stream 12, and the coding parameter is then derived from data stream 12 by decoder 54 for performing prediction, subdivision, etc. again.
[0036] Figure 4 shows a possible implementation of the decoder 54 of FIG. 3, i.e., an implementation form that conforms to the implementation form of the encoder 14 of FIG. 1 as shown in FIG. 2. Since many elements of the encoder 54 in FIG. 4 are the same as those in the corresponding encoder in FIG. 2, the same reference symbols given with apostrophes are used in FIG. 4 to indicate these elements. Specifically, the adder 42', the optional in-loop filter 46', and the predictor 44' are connected to the prediction loop in the same manner as in the encoder of FIG. 2. The reconstructed, i.e., inverse quantized and inverse transformed, prediction residual signal applied to the adder 42' is derived by the sequence of the entropy decoder 56, and the entropy decoder 56 performs the inverse of the entropy coding of the entropy encoder 28b in the same way as in the encoding side, followed by a residual signal reconstruction stage 36' consisting of an inverse quantizer 38' and an inverse transformer 40'. The output of the decoder is the reconstruction of the picture 10. The reconstruction of the picture 10 may be directly available at the output of the adder 42' or, alternatively, at the output of the in-loop filter 46'. Some post-filters may be arranged at the output of the decoder to apply some post-filtering to the reconstruction of the picture 10 to improve the picture quality, but this option is not shown in FIG. 4.
[0037] Regarding FIG. 4 again, it is assumed that the description presented above with respect to FIG. 2 is also valid for FIG. 4, except that only the encoder performs the decisions associated with the optimization tasks and coding options. However, all the descriptions regarding block re-division, prediction, inverse quantization, and retransmission are also valid for the decoder 54 in FIG. 4.
[0038] 3 ALWIP (Affine Linear Weighted Intra Predictor) ALWIP is not always necessary to implement the techniques discussed here, but some non-limiting examples regarding ALWIP are discussed here.
[0039] That is, the present application relates to the concept of an improved block-based prediction mode for picture coding on a per-block basis, such as those that can be used in a video codec such as HEVC or any successor of HEVC. The prediction mode can be an intra prediction mode, but theoretically, the concepts described herein can also be applied to an inter prediction mode where the reference samples are part of another picture.
[0040] There is a need for the concept of block-based prediction that enables an efficient implementation, such as a hardware-friendly implementation.
[0041] This object is achieved by the subject matter of the independent claims of the present application.
[0042] The intra prediction mode is widely used in the coding of pictures and videos. In video coding, the intra prediction mode competes with other prediction modes such as the inter prediction mode, such as the motion-compensated prediction mode. In the intra prediction mode, the current block is predicted based on neighboring samples, that is, samples that have already been encoded as far as the encoder side is concerned and samples that have already been decoded as far as the decoder side is concerned. To form a prediction signal for the current block, the neighboring sample values are extrapolated to the current block, and the prediction residual is transmitted in the data stream for the current block. The better the prediction signal, the smaller the prediction residual, and thus the fewer the number of bits required to code the prediction residual.
[0043] In order to be effective, several aspects should be considered to form an effective framework for intra prediction in the context of per-block picture coding. For example, the larger the number of intra prediction modes supported by the codec, the greater the consumption of the side information rate for signaling the decoder's choice. On the other hand, the set of supported intra prediction modes should be able to provide a good prediction signal, that is, a prediction signal with a low prediction residual.
[0044] In the following, as a comparative embodiment or a basic example, an apparatus (encoder or decoder) for decoding pictures block by block from a data stream is disclosed, the apparatus supporting at least one intra prediction mode in which an intra prediction signal for a block of a predetermined size of a picture is determined accordingly by applying a first template of samples adjacent to the current block to an affine linear predictor, the affine linear predictor thus being referred to as an affine linear weighted intra predictor (ALWIP).
[0045] The apparatus may have at least one of the following properties (the same may apply to a method or another technique implemented in a non-transitory storage unit storing instructions that, when executed by a processor, cause the processor to implement the method and / or operate as the apparatus).
[0046] 3.1 The predictor may supplement other predictors The intra prediction modes that may form the subject of improvement of the implementation further described below may supplement other intra prediction modes of the codec. Thus, they may supplement the DC prediction mode, the Planar prediction mode, or the Angular prediction mode defined in the HEVC codec and the JEM reference software, respectively. The latter three types of intra prediction modes are hereinafter referred to as conventional intra prediction modes. Thus, for a given block in the intra mode, a flag indicating whether one of the intra prediction modes supported by the apparatus should be used needs to be parsed by the decoder.
[0047] 3.2 More than one proposed prediction mode The device may include more than one ALWIP mode. Thus, if the decoder knows that one of the ALWIP modes supported by the device should be used, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the device should be used.
[0048] The signaling of the supported modes may have the property that the coding of some ALWIP modes may require fewer bins than other ALWIP modes. Which of these modes require fewer bins and which modes require more bins may depend on information that can be extracted from the already decoded bitstream or may be predefined.
[0049] 4 Some aspects FIG. 2 shows a decoder 54 for decoding a picture from a data stream 12. The decoder 54 may be configured to decode a predetermined block 18 of the picture. Specifically, the predictor 44 may be configured to map a set of P neighboring samples in the vicinity of the predetermined block 18 to a set of Q predicted values for the samples of the predetermined block using a linear transformation or an affine linear transformation [e.g., ALWIP].
[0050] As shown in FIG. 5, a predetermined block 18 includes Q values to be predicted (which will become "predicted values" at the end of the operation). If block 18 has M rows and N columns, then Q = M·N. The Q values of block 18 may be in the spatial domain (e.g., pixels) or in the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of block 18 can generally be predicted based on P values taken from neighboring blocks 17a - 17c adjacent to block 18. The P values of neighboring blocks 17a - 17c can be in the positions closest to block 18 (e.g., adjacent). The P values of neighboring blocks 17a - 17c have already been processed and predicted. The P values are shown as values in portions 17'a - 17'c to distinguish the portions 17'a - 17'c from the blocks of which they are a part (in some examples, 17'b is not used).
[0051] As shown in FIG. 6, in order to perform prediction, it is possible to operate with a first vector 17P with P entries (each entry is associated with a specific position in neighboring portions 17'a - 17'c), a second vector 18Q with Q entries (each entry is associated with a specific position in block 18), and a mapping matrix 17M (each row is associated with a specific position in block 18 and each column is associated with a specific position in neighboring portions 17'a - 17'c). Thus, the mapping matrix 17M performs the prediction of the P values of neighboring portions 17'a - 17'c to the values of block 18 according to a predetermined mode. Thus, the entries in the mapping matrix 17M can be understood as weighting factors. In the following section, the symbols 17a - 17c are used instead of 17'a - 17'c to refer to the portions near the boundary.
[0052] In this technical field, several conventional modes are known, such as the DC mode, the Planar mode, and 65 directional prediction modes. For example, there can be 67 known modes.
[0053] However, it should be noted that it is also possible here to utilize different modes, referred to as linear transformation or affine linear transformation. The linear transformation or affine linear transformation includes P·Q weighting coefficients, at least (1 / 4)P·Q of which are non-zero weighting values, and these include, for each of the Q predicted values, a series of P weighting coefficients related to that respective predicted value. When this series is arranged one below the other according to the raster scan order among the samples of a given block, it forms an envelope that is non-linear in all directions.
[0054] It is possible to map the P positions of the neighboring values 17'a to 17'c (templates), map the Q positions of the neighboring samples 17'a to 17'c, and map at the values of the P*Q weighting coefficients of the matrix 17M. The plane is an example of a series of envelopes for DC transformation (this is the plane for DC transformation). Since the envelope is clearly planar, it is excluded from the definition of linear transformation or affine linear transformation (ALWIP). Another example is a matrix that results in the emulation of the Angular mode. The envelope is excluded from the definition of ALWIP and, frankly speaking, looks like a hill that extends diagonally from top to bottom along a direction in the P / Q plane. The Planar mode and 65 directional prediction modes have different envelopes, however, these are linear in at least one direction, i.e., in all directions for the exemplified DC, and in the direction of the hill for the Angular mode for example.
[0055] In contrast, the envelope of the linear transformation or affine transformation is not linear in all directions. It is understood that in some situations such types of transformation may be optimal for performing the prediction for block 18. It should be noted that it is preferable that at least one quarter of the weighting coefficients are different from 0 (i.e., at least 25% of the P*Q weighting coefficients are different from 0).
[0056] The weighting factors may not be related to each other according to any normal mapping rules. Therefore, matrix 17M can be such that the values of its entries do not have an obvious recognizable relationship. For example, the weighting factors cannot be described by any analytical or differential functions.
[0057] In an example, the ALWIP transform can be such that the average of the larger of the maximum cross-correlations between a first series of weighting factors for each predicted value and a second series of weighting factors for predicted values other than each predicted value, or the maximum cross-correlation with the reversed version of the latter series, is less than a predetermined threshold (for example, a threshold in the range between 0.2 or 0.3 or 0.35 or 0.1, for example between 0.05 and 0.035). For example, for each pair (i1,i2) of rows of the ALWIP matrix 17M, the cross-correlation can be calculated by multiplying the P value of the i1-th row by the P value of the i2-th row. For each obtained cross-correlation, a maximum value can be obtained. Therefore, for the entire matrix 17M, an average (mean) can be obtained (that is, the maximum cross-correlations in all combinations are averaged). Then, the threshold can be a threshold in the range between 0.2 or 0.3 or 0.35 or 0.1, for example between 0.05 and 0.035.
[0058] The P neighboring samples of blocks 17a to 17c can be located along a one-dimensional path extending along the boundary of a predetermined block 18 (for example, 18c, 18a). For each of the Q predicted values of the predetermined block 18, a series of P weighting factors for each predicted value can be arranged in a manner that scans the one-dimensional path in a predetermined direction (for example, from left to right, from top to bottom, etc.).
[0059] In an example, the ALWIP matrix 17M can be non-diagonal or non-block diagonal.
[0060] An example of the ALWIP matrix 17M for predicting the 4×4 block 18 from four already predicted neighboring samples could be as follows. { {37, 59, 77, 28}, {32, 92, 85, 25}, {31, 69, 100, 24}, {33, 36, 106, 29}, {24, 49, 104, 48}, {24, 21, 94, 59}, {29, 0, 80, 72}, {35, 2, 66, 84}, {32, 13, 35, 99}, {39, 11, 34, 103}, {45, 21, 34, 106}, {51, 24, 40, 105}, {50, 28, 43, 101}, {56, 32, 49, 101}, {61, 31, 53, 102}, {61, 32, 54, 100} }
[0061] (Here, {37, 59, 77, 28} is the first row, {32, 92, 85, 25} is the second row, and {61, 32, 54, 100} is the 16th row of the matrix 17M). The matrix 17M has dimension 16×4 and contains 64 weighting coefficients (as a result of 16*4 = 64). This is because the matrix 17M has dimension Q×P, where Q = M*N, and these are the number of samples of the block 18 to be predicted (the block 18 is a 4×4 block), and P is the number of samples of the already predicted samples. Here, M = 4, N = 4, Q = 16 (as a result of M*N = 4*4 = 16), and P = 4. The matrix is non - diagonal and non - block - diagonal and is not described by a specific rule.
[0062] As can be understood, less than one quarter of the weighting coefficients are zero (in the case of the matrix shown above, one of the 64 weighting coefficients is zero). The envelopes formed by these values form a non-linear envelope in all directions when arranged one below the other in raster scan order.
[0063] The above description has mainly been discussed with reference to a decoder (e.g., decoder 54), but the same can be performed in an encoder (e.g., encoder 14).
[0064] In some examples, for each block size (in a set of block sizes), the ALWIP transformation of the intra prediction modes within the second set of intra prediction modes for each block size is different from each other. Additionally, or alternatively, the density of the second set of intra prediction modes for the block sizes in the set of block sizes may be the same, but the associated linear transformation or affine linear transformation of the intra prediction modes within the second set of intra prediction modes for different block sizes may not be convertible to each other by scaling.
[0065] In some examples, the ALWIP transformation can be defined as having "nothing to share" with conventional transformations (e.g., even if the ALWIP transformation is mapped through one of the above mappings, there may be "nothing" to share with the corresponding conventional transformation).
[0066] In some examples, the ALWIP mode is used for both the luma and chroma components, but in other examples, the ALWIP mode is used for the luma component but not for the chroma component.
[0067] 5 Affine linear weighted intra prediction mode with encoder speedup (e.g., Test CE3-1.2.1) 5.1 Description of the method or apparatus The affine linear weighted intra prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as that proposed in JVET-L0199 under test CE3-2.2.2, except for the following changes. · Multiple reference line (MRL) intra prediction, especially in harmony with encoder estimation and signaling. That is, MRL is not combined with ALWIP, and the transmission of MRL indices is restricted to non-ALWIP blocks. · Subsampling, which is now mandatory for all blocks W×H≧32×32 (previously optional for 32×32). Accordingly, additional tests in the encoder and the transmission of subsampling flags are eliminated. · ALWIP for 64×N and N×64 blocks (where N≦32) is added by downsampling to 32×N and N×32 respectively and applying the corresponding ALWIP mode.
[0068] Moreover, test CE3-1.2.1 includes the following encoder optimizations for ALWIP. · Integrated mode estimation: The conventional mode and the ALWIP mode use a shared Hadamard candidate list for full RD estimation. That is, ALWIP mode candidates are added to the same list as conventional (and MRL) mode candidates based on Hadamard cost. · EMT intra fast and PB intra fast are supported for the integrated mode list, with further optimizations to reduce the number of full RD checks. · Only the MPM of the available left and upper blocks is added to the list for full RD estimation for ALWIP, following the same approach as the conventional mode.
[0069] 5.2 Complexity evaluation In Test CE3-1.2.1, excluding the calculations that call the discrete cosine transform, at most 12 multiplications per sample were required to generate the prediction signal. Moreover, a total of 136,492 parameters, each 16 bits, were required. This corresponds to 0.273 megabytes of memory.
[0070] 5.3 Experimental Results The evaluation of the tests was performed according to the common test conditions JVET-J1010[2] for the intra-only (AI) configuration and the random access (RA) configuration using the VTM software version 3.0.1. The corresponding simulations were performed on an Intel Xeon cluster (E5-2697A v4, AVX2 on, turbo boost off) using the Linux® OS and the GCC 7.2.1 compiler.
[0071] [Table 1]
[0072] [Table 2]
[0073] 5.4 Affine Linear Weighted Intra Prediction with Reduced Complexity (e.g., Test CE3-1.2.2) The technique tested in CE2 relates to the "affine linear intra prediction" described in JVET-L0199[1], but simplifies it in terms of memory requirements and computational complexity. · There can only be three different sets of prediction matrices (e.g., see S0, S1, S2 below) and bias vectors (e.g., for providing offset values) that cover all block shapes. As a result, the number of parameters is reduced to 14,400 10-bit values, which is less memory than what is stored in a 128×128 CTU. · The input size and output size of the predictor are further reduced. Moreover, instead of transforming the boundary via DCT, averaging or downsampling may be performed on the boundary samples, and linear interpolation may be used to generate the prediction signal instead of the inverse DCT. As a result, a maximum of four multiplications per sample may be required to generate the prediction signal.
[0074] 6. Example Here, discuss how to perform several predictions (e.g., as shown in FIG. 6) using ALWIP prediction.
[0075] In principle, referring to FIG. 6, to obtain the Q = M*N values of the M×N block 18 to be predicted, the multiplication of the Q×P samples of the ALWIP prediction matrix 17M by the P samples of the P×1 neighborhood vector 17P should be performed. Thus, generally, at least P = M + N multiplications are required to obtain each of the Q = M*N values of the M×N block 18 to be predicted.
[0076] These multiplications have highly undesirable effects. The dimension P of the boundary vector 17P generally depends on the number M + N of boundary samples (bins or pixels) 17a, 17c in the neighborhood (e.g., adjacent to) of the M×N block 18 to be predicted. This is because when the size of the block 18 to be predicted is large, the number M + N of boundary pixels (17a, 17c) also increases accordingly, increasing the dimension P = M + N of the P×1 boundary vector 17P and the length of each row of the Q×P ALWIP prediction matrix 17M, and thus increasing the number of multiplications required (generally, Q = M*N = W*H, where W (width) is another symbol for N and H (height) is another symbol for M. When the boundary vector is formed by only one row and / or one column of samples, P = P = M + N = H + W).
[0077] This problem is generally exacerbated in microprocessor-based systems (or other digital processing systems) by the fact that multiplication is generally an operation that consumes processing power. One can imagine that the numerous multiplications performed on a large number of samples for a large number of blocks cause a waste of processing power, which is generally not desirable.
[0078] Therefore, it is preferable to reduce the number Q*P of multiplications required to predict the M×N block 18.
[0079] It is understood that by intelligently selecting a more easily processed operation instead of multiplication, it is possible to somewhat reduce the processing power required for each intra-prediction of each block 18 to be predicted.
[0080] Specifically, referring to FIGS. 7.1 to 7.4, an encoder or decoder uses a plurality of neighboring samples (e.g., 17a, 17c) to reduce a plurality of neighboring samples (e.g., by averaging or downsampling) (e.g., in step 811) to obtain a reduced set of sample values with fewer samples compared to the plurality of neighboring samples (e.g., 17a, 17c), and apply a linear transformation or an affine linear transformation (e.g., in step 812) to the reduced set of sample values to obtain a predicted value for a predetermined sample of a predetermined block and it is understood that prediction can be achieved thereby.
[0081] In some cases, the decoder or encoder can also derive, for example, by interpolation, a predicted value for a further sample of a predetermined block based on, for example, the predicted values for a predetermined sample and a plurality of neighboring samples. Therefore, an upsampling strategy can be obtained.
[0082] In an example, in order to reach a reduced set 102 of samples with fewer samples (e.g., FIGS. 7.1 to 7.4), it is possible to perform some average (e.g., at step 811) on the samples of the boundary 17 (at least one of the samples of the fewer number of samples 102 can be the average of two samples of the original boundary samples or a selection of the original boundary samples). For example, if the original boundary has P = M + N samples, the reduced set of samples may have P red = M red + N red where at least one of M red < M and N red < N holds, so P red < P. Thus, the boundary vector 17P actually used for prediction (e.g., at step 812b) has P red × 1 entries instead of P red < P, and P red (or Q red × P red , see below), and since at least one of M red < M and N red < N holds, at least P red < P, so the number of elements in the matrix is fewer.
[0083] In some examples (e.g., FIGS. 7.2, 7.3), the block obtained by ALWIP (at step 812) is a reduced block having size M' red × N' red .
[0084]
Number
[0085] and / or
[0086]
Number
[0087] That is, when the number of samples directly predicted by ALWIP is less than the number of samples of block 18 to be actually predicted, it is even possible to further reduce the number of multiplications. Therefore,
[0088]
Number
[0089] Setting it to this results in obtaining the ALWIP prediction by using Q * P red multiplications instead of Q red * P red multiplications (Q red * P red < Q * P red < Q * P). This multiplication predicts a reduced block of dimension
[0090]
Number
[0091] Despite this, it would be possible to perform (for example, by interpolation) upsampling from the predicted block of the reduced
[0092]
Number
[0093] to the final predicted block of M × N (for example, in subsequent step 813).
[0094] These techniques can be advantageous in that there are fewer matrix multiplications (Q red * P red or Q * P redOn the other hand, while involving multiplication, both the first reduction (e.g., averaging or downsampling) and the last transformation (e.g., interpolation) can be performed by reducing (or even avoiding) multiplication. For example, downsampling, averaging, and / or interpolation can be performed (e.g., in steps 811 and / or 813) by adopting binary operations that do not require processing power, such as addition and shift.
[0095] Also, addition is a very simple operation that can be easily performed without much computational effort.
[0096] This shift operation can be used, for example, to average two boundary samples and / or to interpolate two samples (support values) of the downsampled predicted block (or taken from the boundary) in order to obtain the final predicted block. (Two sample values are required for interpolation. As shown in Figure 7.2, within the block, there are always two predetermined values, but for interpolating samples along the left and upper boundaries of the block, there is only one predetermined value, so boundary samples are used as support values for interpolation.)
[0097] A two-step procedure may be used, i.e., First, add the values of two samples. Then divide the total value by half (e.g., by a right shift). and so on.
[0098] Alternatively, the following is possible. First divide each sample by half (e.g., by a left shift). Then add the values of the two halved samples.
[0099] When downsampling (e.g., in step 811), even simpler operations can be performed because it only requires selecting one sample from a group of samples (e.g., samples adjacent to each other).
[0100] Therefore, it is possible here to define techniques for reducing the number of multiplications to be performed. That is, some of these techniques can be based on at least one of the following techniques. Even when the block 18 to be actually predicted has a size of M×N, the block can be reduced (in at least one of the two dimensions), and the reduced size is Q red ×P red with an ALWIP matrix (
[0101]
Number
[0102] where P red =N red +M red and
[0103]
Number
[0104] and / or
[0105]
Number
[0106] and / or M red <M and / or N red <N) can be applied. Therefore, the boundary vector 17P has a size of P red ×1 and suggests that there are only P red <P multiplications (P red =M red +N redand P = M + N). P red The boundary vector 17P of ×1 can be easily obtained from the original boundary 17, for example, by downsampling (for example, by selecting only samples of a part of the boundary) and / or by averaging a plurality of samples of the boundary (which can be easily obtained by addition and shift without multiplication). It can be easily obtained from the original boundary 17. Additionally or alternatively, instead of predicting all Q = M * N values of the block 18 to be predicted by multiplication, it is possible to predict only the reduced block with reduced dimensions (for example,
[0107] [Number]
[0108] and
[0109] [Number]
[0110] and / or
[0111] [Number]
[0112] and). The remaining samples of the block 18 to be predicted are obtained by interpolation, for example, using Q red samples as support values for the remaining Q - Q red values to be predicted.
[0113] According to the example shown in FIG. 7.1, a 4×4 block 18 (M = 4, N = 4, Q = M*N = 16) is to be predicted, and the vicinity 17 of the sample, namely 17a (a vertical column with four already predicted samples) and 17c (a horizontal row with four already predicted samples), have already been predicted in previous iterations (the vicinities 17a and 17c can be collectively denoted by 17). A priori, by using the formula shown in FIG. 6, the prediction matrix 17M should be a matrix of Q×P = 16×8 (where Q = M*N = 4*4 and P = M+N = 4+4 = 8), and the boundary vector 17P should have a dimension of 8×1 (where P = 8). However, this results in the need to perform 8 multiplications for each of the 16 samples of the 4×4 block 18 to be predicted, so a total of 16*8 = 128 multiplications need to be performed. (Note that the average number of multiplications per sample is a good indicator of the computational complexity. In conventional intra prediction, 4 multiplications per sample are required, which increases the computational amount involved. Therefore, by using this as an upper limit for ALWIP, it is possible to ensure that the complexity is reasonable and does not exceed that of conventional intra prediction.)
[0114] Nevertheless, by using this technique, in step 811, the number of samples 17a and 17c in the vicinity of the block 18 to be predicted is changed from P to P redIt is understood that it is possible to reduce to . Specifically, in order to obtain a reduced boundary 102 with two horizontal rows and two vertical columns, adjacent boundary samples (17a, 17c) are averaged (e.g., at 100 in FIG. 7.1), and thus it is understood that it is possible to operate as if block 18 were a 2×2 block (the reduced boundary is formed by the average value). Alternatively, it is possible to perform downsampling, and thus select two samples for row 17c and two samples for column 17a. Thus, the horizontal row 17c is processed as having two samples (e.g., averaged samples) instead of four original samples, while the vertical column 17a, which originally had four samples, is processed as having two samples (e.g., averaged samples). After re - dividing row 17c and column 17a in each group 110 of two samples each, it is also possible to understand that one single sample is maintained (e.g., the average of the samples in group 110 or a simple selection from the samples in group 110). Thus, a so - called reduced set 102 of sample values is obtained by the fact that the set 102 has only four samples (M red = 2, N red = 2, P red = M red + N red = 4, provided that P red < P).
[0115] It is understood that it is possible to perform operations (such as averaging or downsampling 100) without performing too many multiplications at the processor level. The averaging or downsampling 100 performed in step 811 can be easily obtained by simple operations such as addition and shift that do not consume processing power.
[0116] At this point, it is understood that it is possible to apply a linear or affine linear (ALWIP) transform 19 to the reduced set 102 of sample values (e.g., using a prediction matrix such as the matrix 17M of FIG. 6). In this case, the ALWIP transform 19 directly maps four samples 102 to the sample values 104 of block 18. In this case, no interpolation is necessary.
[0117] In this case, the ALWIP matrix 17M has dimensions Q×P red = 16×4. This follows from the fact that all Q = 16 samples of block 18 to be predicted are obtained directly by ALWIP multiplication (no interpolation is required).
[0118] Thus, in step 812a, an appropriate ALWIP matrix 17M with dimensions Q×P red is selected. This selection may be based at least in part on signaling from the data stream 12. The selected ALWIP matrix 17M may also be denoted using A k where k may be understood as an index and this may be signaled in the data stream 12 (in some cases, the matrix is
[0119]
Number
[0120] also shown as. See below). This selection may be performed according to the following manner. For each dimension (e.g., the height / width pair of block 18 to be predicted), the ALWIP matrix 17M is selected from one of, for example, three sets of matrices S0, S1, S2 (each of the three sets S0, S1, S2 may group multiple ALWIP matrices 17M of the same dimension and the ALWIP matrix to be selected for prediction is one of them).
[0121] In step 812b, the selected Q×P red ALWIP matrix 17M (Ak as also shown) and P red Multiplication with the boundary vector 17P of ×1 is performed.
[0122] In step 812c, an offset value (e.g., b k ) can be added to all the obtained values 104 of the vector 18Q obtained, for example, by ALWIP. The value of the offset (b k , or in some cases
[0123]
Number
[0124] is also shown using. See below) may be associated with a particular selected ALWIP matrix (A k ) and may be based on an index (which may be signaled, for example, in the data stream 12).
[0125] Therefore, the comparison between using this technique and not using this technique is resumed here. When not using this technique: Block 18 will be predicted, and this block has dimensions M = 4, N = 4 Q = M * N = 4 * 4 = 16 values will be predicted P = M + N = 4 + 4 = 8 boundary samples P = 8 multiplications for each of the Q = 16 values to be predicted A total of P * Q = 8 * 16 = 128 multiplications When using this technique: Block 18 will be predicted, and this block has dimensions M = 4, N = 4 Q = M * N = 4 * 4 = 16 values will finally be predicted Reduced dimension of the boundary vector: P red = M red + N re d = 2 + 2 = 4 For each of the Q = 16 values to be predicted by ALWIP, P red = 4 multiplications In total, P red * Q = 4 * 16 = 64 multiplications (half of 128) The ratio of the number of multiplications to the number of final values to be obtained is P red * Q / Q = 4, that is, half of P = 8 multiplications for each sample to be predicted
[0126] As can be understood, by using operations that are simple and do not require processing power, such as averaging (and in some cases, addition and / or shift and / or downsampling), it is possible to obtain appropriate values in step 812.
[0127] Referring to FIG. 7.2, the block 18 to be predicted here is an 8 × 8 block of 64 samples (M = 8, N = 8). Here, a priori, the prediction matrix 17M should have a size Q × P = 64 × 16 (since Q = M * N = 8 * 8 = 64, M = 8, and N = 8, and since P = M + N = 8 + 8 = 16, Q = 64). Thus, a priori, for each of the Q = 64 samples of the 8 × 8 block 18 to be predicted, P = 16 multiplications are required to reach 64 × 16 = 1024 multiplications for the entire 8 × 8 block 18.
[0128] However, as can be seen in FIG. 7.2, a method 820 can be provided such that instead of using all 16 samples at the boundary, only 8 values (for example, 4 values in the horizontal boundary row 17c between the original samples at the boundary and 4 values in the vertical boundary column 17a) are used. From the boundary row 17c, 4 instead of 8 samples can be used (for example, they can be 2 - to - 2 averaging and / or selection of 1 sample from 2 samples). Thus, the boundary vector is not a P × 1 = 16 × 1 vector, but P red × 1 = 8 × 1 vector only (P red = Mred +N red =4 + 4). Instead of the original P = 16 samples, P red = 8 boundary values only, it is understood that it is possible to select or average the samples in the horizontal row 17c and the samples in the vertical column 17a (e.g., 2 vs. 2) to form a reduced set 102 of sample values. This reduced set 102 enables obtaining a reduced version of block 18, and the reduced version has Q (instead of Q = M * N = 8 * 8 = 64) red = M red * N red = 4 * 4 = 16 samples. It is possible to apply an ALWIP matrix for predicting a block with size M red × N red = 4 × 4. The reduced version of block 18 includes the samples shown in gray in the manner 106 of FIG. 7.2. The samples shown by the gray squares (including samples 118' and 118'') form a 4 × 4 reduced block with Q red = 16 values obtained in step 812 to be applied. The 4 × 4 reduced block is obtained by applying a linear transformation 19 in step 812 to be applied. After obtaining the values of the 4 × 4 reduced block, it is possible to obtain the values of the remaining samples (the samples shown by the white samples in manner 106), for example, by interpolation.
[0129] Regarding the method 810 of FIG. 7.1, this method 820 may additionally include step 813 of deriving predicted values for the remaining Q - Q red = 64 - 16 = 48 samples (white squares) of the M × N = 8 × 8 block 18 to be predicted, for example, by interpolation. The remaining Q - Q red = 64 - 16 = 48 samples are interpolated to Q redcan be obtained from 16 directly obtained samples (for example, interpolation can also use the values of the boundary samples). As can be seen in FIG. 7.2, samples 118' and 118'' (as indicated by the gray squares) are obtained at step 812, while sample 108' (in the middle of samples 118' and 118'', indicated by the white square) is obtained by interpolation between samples 118' and 118'' at step 813. It is understood that the interpolation can also be obtained by operations similar to those for averaging such as shifting and adding. Thus, in FIG. 7.2, the value 108' can generally be determined (can be an average) as a value intermediate between the value of sample 118' and the value of sample 118''.
[0130] By performing interpolation, it is also possible to reach the final version of block 18 of M×N = 8×8 based on the plurality of sample values shown at 104 at step 813.
[0131] Thus, the comparison between using this technique and not using this technique is as follows. When not using this technique: Block 18 will be predicted, and this block has dimensions M = 8, N = 8 Q = M*N = 8*8 = 64 samples in block 18 will be predicted P = M + N = 8 + 8 = 16 samples in boundary 17 P = 16 multiplications for each of the Q = 64 values to be predicted A total of P*Q = 16*64 = 1028 multiplications The ratio of the number of multiplications to the number of final values to be obtained is P*Q / Q = 16 When using this technique: Block 18 will be predicted, and this block has dimensions M = 8, N = 8 Finally, Q = M*N = 8*8 = 64 values will be predicted However, Q red ×P red of the ALWIP matrix will be used, Pred =M red +N red 、Q red =M red *N red 、M red =4、N red =4 is P within the boundary red =M red +N red =4 + 4 = 8 samples, and P red <P is (Formed by the gray square in Method 106) Q of the predicted 4×4 reduced block red = P for each of the 16 values red = 8 multiplications In total P red *Q red = 8 * 16 = 128 multiplications (much less than 1024) The ratio of the number of multiplications to the number of final values to be obtained is P red *Q red / Q = 128 / 64 = 2 (much smaller than 16 obtained without this technique)
[0132] Therefore, the technique presented in this specification requires only one-eighth of the processing power compared to previous techniques.
[0133] Figure 7.3 shows another example (obtained based on Method 820) where the block 18 to be predicted is a rectangular 4×8 block (M = 8, N = 4), and Q = 4 * 8 = 32 samples are to be predicted. The boundary 17 is formed by the horizontal row 17c with N = 8 samples and the vertical column 17a with M = 4 samples. Therefore, a priori, the boundary vector 17P has dimension P×1 = 12×1, but the predicted ALWIP matrix should be a matrix of Q×P = 32×12, and thus Q * P = 32 * 12 = 384 multiplications are required.
[0134] However, for example, in order to obtain a reduced horizontal row with only 4 samples (e.g., averaged samples), it is possible to average or downsample at least 8 samples of horizontal row 17c. In some examples, vertical column 17a remains as is (e.g., without averaging). Overall, the reduced boundary has dimension P red = 8, and P red < P. Thus, boundary vector 17P has dimension P red × 1 = 8 × 1. The ALWIP prediction matrix 17M is a matrix with dimension M * N red * P red = 4 * 4 * 8 = 64. The directly obtained 4 × 4 reduced block (formed by the gray columns in illustration 107) in the application step 812 has Q red = M * N red = 4 * 4 = 16 samples instead of Q k = 4 * 8 = 32 of the original 4 × 8 block 18 to be predicted. When the reduced 4 × 4 block is obtained by ALWIP, it is possible to add the offset value b
[0135] Thus, the comparison between using this technique and not using this technique is as follows. When not using this technique: Block 18 will be predicted, and this block has dimensions M = 4, N = 8 Q = M * N = 4 * 8 = 32 values will be predicted There are P = M + N = 4 + 8 = 12 samples in the boundary P = 12 multiplications for each of the Q = 32 values to be predicted A total of P * Q = 12 * 32 = 384 multiplications The ratio of the number of multiplications to the number of final values to be obtained is P * Q / Q = 12. When using this technique: Block 18 will be predicted, and this block has dimensions M = 4, N = 8. Q = M * N = 4 * 8 = 32 values will be finally predicted. However, Q red × P red = 16 × 8 ALWIP matrix can be used, where M = 4, N red = 4, Q red = M * N red = 16, P red = M + N red = 4 + 4 = 8. There are P red = M + N red = 4 + 4 = 8 samples within the boundary, and P red < P. Q of the reduced block to be predicted red = P for each of the 16 values red = 8 multiplications In total, Q red * P red = 16 * 8 = 128 multiplications (less than 384). The ratio of the number of multiplications to the number of final values to be obtained is P red * Q red / Q = 128 / 32 = 4 (much smaller than 12 obtained without this technique).
[0136] Therefore, using this technique reduces the computational complexity by a factor of three.
[0137] Figure 7.4 shows an example of block 18 to be predicted, with dimensions M × N = 16 × 16, and finally having Q = M * N = 16 * 16 = 256 values to be predicted, along with P = M + N = 16 + 16 = 32 boundary samples. This results in a prediction matrix with dimensions Q × P = 256 × 32, which implies 256 × 32 = 8192 multiplications.
[0138] However, by applying method 820, in step 811, it is possible to reduce the number of boundary samples, for example, from 32 to 8 (e.g., by averaging or downsampling). For example, for each group 120 of four consecutive samples in row 17a, one single sample (e.g., selected from the four samples or the average of the samples) remains. Also, for each group of four consecutive samples in column 17c, one single sample (e.g., selected from the four samples or the average of the samples) remains.
[0139] Here, the ALWIP matrix 17M is Q red ×P red = a 64×8 matrix. This is due to the fact that P red = 8 was chosen (by using eight averaged or selected samples from the 32 boundary samples), and the fact that the reduced block to be predicted in step 812 is an 8×8 block (in method 109, the gray square is 64).
[0140] Thus, when 64 samples of the reduced 8×8 block are obtained in step 812, in step 813, it is possible to derive the remaining Q - Q red = 256 - 64 = 192 values 104 of block 18 to be predicted.
[0141] In this case, to perform interpolation, it has been chosen to use only all the samples of boundary column 17a and alternative samples in boundary row 17c. Other selections may be made.
[0142] Using this method, the ratio of the number of multiplications to the number of finally obtained values is Q red *P red / Q = 8 * 64 / 256 = 2, which is much less than 32 multiplications for each value when not using this technique.
[0143] The comparison between using this technique and not using this technique is as follows. When not using this technique: Block 18 will be predicted, and this block has dimensions M = 16, N = 16 Q = M * N = 16 * 16 = 256 values will be predicted There are P = M + N = 16 + 16 = 32 samples within the boundary P = 32 multiplications for each of the Q = 256 values to be predicted In total, P * Q = 32 * 256 = 8192 multiplications The ratio of the number of multiplications to the number of final values to be obtained is P * Q / Q = 32 When using this technique: Block 18 will be predicted, and this block has dimensions M = 16, N = 16 Q = M * N = 16 * 16 = 256 values will be finally predicted However, Q red ×P red = The ALWIP matrix of 64×8 will be used, and Q red = 8 * 8 = 64 samples will be predicted by M red = 4, N red = 4, by ALWIP, and P red = M red +N red = 4 + 4 = 8 There are P red = M red +N red = 4 + 4 = 8 samples within the boundary, and P red < P For the reduced block Q to be predicted red = P for each of the 64 values red = 8 multiplications In total, Q red *P red = 64 * 4 = 256 multiplications (much less than 8192) The ratio of the number of multiplications to the number of final values to be obtained is P red *Q red / Q = 8 * 64 / 256 = 2 (much smaller than 32 obtained without this technique)
[0144] Therefore, the processing power required by this technique is 16 times less than that of the conventional technique.
[0145] Therefore, using a plurality of neighboring samples (17), a predetermined block (18) of the picture To obtain a reduced set (102) of sample values with fewer samples compared to a plurality of neighboring samples (17), reduce the plurality of neighboring samples (100, 813), To obtain a predicted value for a predetermined sample (104, 118', 188'') of a predetermined block (18), apply a linear transformation or an affine linear transformation (19, 17M) to the reduced set (102) of sample values (812) It is possible to predict by doing so.
[0146] Specifically, in order to obtain a reduced set (102) of sample values with fewer samples compared to a plurality of neighboring samples (17), it is possible to perform the reduction (100, 813) by downsampling the plurality of neighboring samples.
[0147] Alternatively, in order to obtain a reduced set (102) of sample values with fewer samples compared to a plurality of neighboring samples (17), it is possible to perform the reduction (100, 813) by averaging the plurality of neighboring samples.
[0148] Furthermore, it is possible to derive predicted values for additional samples (108, 108') of a predetermined block (18) based on the predicted values (104, 118', 118'') for a predetermined sample and a plurality of neighboring samples (17) by interpolation (813).
[0149] A plurality of neighboring samples (17a, 17c) can extend one-dimensionally along two sides of a predetermined block (18) (e.g., toward the right and downward in FIGS. 7.1 to 7.4). A predetermined sample (e.g., the one obtained by ALWIP in step 812) may also be arranged in rows and columns, and along at least one of the rows and columns, the predetermined sample may be arranged at every nth position from the samples (112) of the predetermined sample 112 adjacent to two sides of the predetermined block 18.
[0150] Based on the plurality of neighboring samples (17), it is possible to determine a support value (118) for one of the plurality of neighboring positions (118) for each of at least one of the rows and columns, and these are aligned with one each of at least one of the rows and columns. By interpolation, based on the predicted values for the predetermined samples (104, 118', 118'') and the support values for the neighboring samples (118) aligned with at least one of the rows and columns, it is also possible to derive predicted values 118 for further samples (108, 108') of the predetermined block (18).
[0151] The predetermined sample (104) may be arranged at every nth position from the samples (112) adjacent to two sides of the predetermined block 18 along the row, and the predetermined sample is arranged at every mth position from the samples (112) adjacent to two sides of the predetermined block (18) along the column, where n, m > 1. In some cases, n = m (e.g., in FIGS. 7.2 and 7.3, the samples 104, 118', 118'' directly obtained by ALWIP in 812 and shown using gray squares are staggered with the samples 108, 108' obtained later in step 813 along the rows and columns).
[0152] It may be possible to perform determining support values along at least one of a row (17c) and a column (17a), for example, by downsampling and averaging (122), and for each support value, a group (120) of neighboring samples within a plurality of neighboring samples includes the neighboring sample (118) for which the respective support value is determined. Thus, in FIG. 7.4, at step 813, it is possible to obtain the value of sample 119 by using the value of a predetermined sample 118''' (previously obtained at step 812) and the neighboring sample 118 as support values.
[0153] The plurality of neighboring samples may extend one-dimensionally along two sides of a predetermined block (18). Reduction (811) may be possible by grouping the plurality of neighboring samples (17) into one or more groups (110) of consecutive neighboring samples and performing downsampling or averaging on each of the one or more groups (110) of neighboring samples having two or more neighboring samples.
[0154] In an example, the linear transformation or the affine linear transformation may include P red *Q red weights or P red *Q weights, where P red is the number of sample values (102) in a reduced set of sample values, and Q red or Q is the number of predetermined samples within a predetermined block (18). At least (1 / 4)P red *Q red weights or (1 / 4)P red *Q weights are non-zero weighted values. P red *Q red weights or P red *Q weights are for each of Q or Q red predetermined samples, and P redIt may include a series of weighting coefficients, and when the series are arranged one below the other in raster scan order among the predetermined samples of a predetermined block (18), they form a non-linear envelope in all directions. P red *Q or P red *Q red The weighting coefficients may not be related to each other via any normal mapping rule. The average of the larger of the maximum cross-correlation between the first series of weighting coefficients for each predetermined sample and the second series of weighting coefficients for predetermined samples other than each predetermined sample, or the maximum cross-correlation between the latter series and its reversed version, is lower than a predetermined threshold. The predetermined threshold may be 0.3 [or in some cases 0.2 or 0.1]. P red The Q neighboring samples (17) may be located along a one-dimensional path extending along two sides of the predetermined block (18), Q or Q red For each of the Q predetermined samples, P for each predetermined sample red The series of P weighting coefficients are ordered in a manner that scans a one-dimensional path in a predetermined direction.
[0155] 6.1 Description of the method and apparatus To predict the samples of a rectangular block of width W (also denoted as N) and height H (also denoted as M), affine linear weighted intra prediction (ALWIP) may take as input one row of the H reconstructed neighboring boundary samples to the left of the block and one row of the W reconstructed neighboring boundary samples above the block. If the reconstructed samples are not available, they may be generated as done in conventional intra prediction.
[0156] The generation of the prediction signal (e.g., the values for the complete block 18) may be based on at least some of the following three steps.
[0157] 1. From among the boundary samples 17, samples 102 (e.g., 4 samples when W = H = 4 and / or 8 samples in other cases) can be extracted by averaging or downsampling (e.g., step 811).
[0158] 2. Matrix-vector multiplication, followed by addition of an offset, can be performed using the averaged samples (or the samples remaining from downsampling) as an input. The result can be a reduced prediction signal for the subsampled set of samples in the original block (e.g., step 812).
[0159] 3. The prediction signal at the remaining positions can be generated from the prediction signal for the subsampled set, for example, by upsampling, for example, by linear interpolation (e.g., step 813).
[0160] By step 1 (811) and / or 3 (813), the total number of multiplications required for the calculation of the matrix-vector product can be such that it is always 4 * W * H or less. Moreover, the averaging operation for the boundary and the linear interpolation of the reduced prediction signal are performed by using only addition and bit shift. In other words, in the example, a maximum of 4 multiplications per sample are required for the ALWIP mode.
[0161] In some examples, the matrices (e.g., 17M) and offset vectors (e.g., b k ) required to generate the prediction signal may be taken from a set of matrices (e.g., 3 sets), such as S0, S1, S2, which can be stored, for example, in the memory units of the decoder and encoder.
[0162] In some examples, set S0 has n0 matrices (e.g., n0 = 16 or n0 = 18 or another number)
[0163]
Number
[0164] may include (e.g., may consist of), and for performing the technique according to FIG. 7.1, each of these has 16 rows, 4 columns, and 18 offset vectors each of size 16
[0165]
Number
[0166] and may have. This set of matrices and offset vectors is used for 18 blocks of size 4×4. When the boundary vector is reduced to P red = 4 vectors (as in step 811 of FIG. 7.1), the reduced set of P of the samples 102 is directly mapped to the Q = 16 samples of the 4×4 block 18 to be predicted. red = 4 samples.
[0167] In some examples, the set S1 may include n1 (e.g., n1 = 8 or n1 = 18 or another number) matrices
[0168]
Number
[0169] may include (e.g., may consist of), and for performing the technique according to FIG. 7.2 or FIG. 7.3, each of these has 16 rows, 8 columns, and 18 offset vectors each of size 16
[0170]
Number
[0171] may have. The matrices and offset vectors of this set S1 can be used for blocks of sizes 4×8, 4×16, 4×32, 4×64, 16×4, 32×4, 64×4, 8×4, and 8×8. Additionally, it can also be used for blocks of size W×H where max(W,H)>4 and min(W,H)=4, i.e., blocks of size 4×16 or 16×4, 4×32 or 32×4, and 4×64 or 64×4. The 16×8 matrix refers to a reduced version of block 18, which is a 4×4 block as obtained in FIGS. 7.2 and 7.3.
[0172] Additionally or alternatively, set S2 may include (e.g., may consist of) n2 (e.g., n2 = 6 or n2 = 18 or another number) matrices
[0173] [Number]
[0174] each having 64 rows, 8 columns, and 18 offset vectors each of size 64
[0175] [Number]
[0176] may have. The 64×8 matrix refers to a reduced version of block 18, which is an 8×8 block as obtained, for example, in FIG. 7.4. The matrices and offset vectors of this set can be used for blocks of sizes 8×16, 8×32, 8×64, 16×8, 16×16, 16×32, 16×64, 32×8, 32×16, 32×32, 32×64, 64×8, 64×16, 64×32, 64×64.
[0177] The matrices and offset vectors of that set or a portion of these matrices and offset vectors can be used for all other block shapes.
[0178] 6.2 Boundary averaging or downsampling Here, features regarding step 811 are provided.
[0179] As described above, boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0180] In a first step, input boundaries bdry top (e.g., 17c) and bdry left (e.g., 17a) reach a reduced set 102 by smaller boundaries
[0181]
Number
[0182] and
[0183]
Number
[0184] can be reduced to. Here,[[]]
[0185]
Number
[0186] and
[0187]
Number
[0188] both consist of 2 samples in the case of a 4×4 block and both consist of 4 samples in other cases.
[0189] In the case of a 4×4 block,
[0190]
Number
[0191] is defined, and
[0192]
Number
[0193] it is also possible to define them in the same way. Therefore,
[0194]
Number
[0195] ,
[0196]
Number
[0197] ,
[0198]
Number
[0199] ,
[0200]
Number
[0201] is the average value obtained, for example, using bit shift operations.
[0202] In all other cases (for example, for blocks where either the width or the height is different from 4), the block width W is given by W = 4 * 2 k and for 0 ≤ i < 4,
[0203]
Number
[0204] is defined, and similarly
[0205]
Number
[0206] is defined.
[0207] In other cases, it is possible to downsample the boundary (for example, by selecting one specific boundary sample from a group of boundary samples) in order to reach a smaller number of samples. For example,
[0208]
Number
[0209] can be selected from bdry top [0] and bdry top [1], and
[0210]
Number
[0211] can be selected from bdry top [2] and bdry top [3]. Similarly
[0212]
Number
[0213] it is also possible to define.
[0214] Two reduced boundaries
[0215]
Number
[0216] and
[0217]
Number
[0218] is the reduced boundary vector bdry, also denoted as 17P red (associated with the reduced set 102) can be connected. Thus, the reduced boundary vector bdry red is of size 4 (P red = 4) for blocks of shape 4×4 (example in Figure 7.1), and can be of size 8 (P red = 8) for blocks of all other shapes (examples in Figures 7.2 to 7.4).
[0219] Here, if mode < 18 (or the number of matrices in the set of matrices),
[0220]
Number
[0221] it is possible to define.
[0222] When mode ≥ 18, corresponding to the transposed mode of mode - 17,
[0223]
Number
[0224] it is possible to define.
[0225] Therefore, according to a specific state (one state: mode < 18, one other state: mode ≥ 18), it is possible to distribute the predicted values of the output vector along different scanning orders (for example, one scanning order:
[0226]
Number
[0227] , one other scanning order:
[0228]
Number
[0229] ).
[0230] Other strategies may be executed. In other examples, the mode index "mode" is not necessarily in the range from 0 to 35 (other ranges may be defined). Further, it is not necessary for each of the three sets S0, S1, S2 to have 18 matrices (therefore, instead of expressions such as mode ≥ 18, it is possible that mode ≥ n0, n1, n2, where these are the numbers of matrices for each set S0, S1, S2 of matrices respectively). Further, the sets may each have a different number of matrices (for example, it may be the case that S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0231] The mode and the transposed information are not necessarily stored and / or transmitted as one integrated mode index "mode". In some examples, it may be explicitly signaled as the transposed flag and matrix index (0 to 15 for S0, 0 to 7 for S1, and 0 to 5 for S2).
[0232] In some cases, a combination of a transposed flag and a matrix index can be interpreted as a set index. For example, there may be one bit operating as a transposed flag and some bits indicating a matrix index, which are collectively shown as a "set index".
[0233] 5.4 Generation of Reduced Prediction Signal by Matrix-Vector Multiplication Here, features regarding step 812 are provided.
[0234] Reduced input vector bdry red (Boundary vector 17P), a reduced prediction signal pred red can be generated. The latter signal can be a signal for a downsampled block with W red and height H red . Here, W red and H red can be defined as follows. If max(W, H) ≤ 8, then W red = 4, H red = 4 Otherwise, W red = min(W, 8), H red = min(H, 8)
[0235] The reduced prediction signal pred red can be calculated by computing a matrix-vector product and adding an offset. pred red = A · bdry red + b
[0236] Here, A is a matrix that can have W red * H red rows and, when W = H = 4, 4 columns, and in all other cases, 8 columns (e.g., prediction matrix 17M), and b is a vector that can be of size W red * H red .
[0237] When W = H = 4, A can have 4 columns and 16 rows, in which case, for pred red To calculate, 4 multiplications per sample may be required. In all other cases, A may have 8 columns, and in these cases, 8 * W red * H red ≦ 4 * W * H, that is, in these cases, it can also be confirmed that at most 4 multiplications per sample are required to calculate pred red .
[0238] Matrix A and vector b may be taken from one of the sets S0, S1, S2 as follows. Define the index idx = idx(W, H) by setting idx(W, H) = 0 when W = H = 4, idx(W, H) = 1 when max(W, H) = 8, and idx(W, H) = 2 in all other cases. Moreover, when mode < 18, m = mode, and in other cases, m = mode - 17. Then, when idx ≤ 1 or idx = 2 and min(W, H) > 4
[0239]
Number
[0240] and
[0241]
Number
[0242] it can be set as. When idx = 2 and min(W, H) = 4, A is
[0243]
Number
[0244] as the matrix resulting from excluding each row thereof, and those rows correspond to the odd x - coordinates in the down - sampled block when W = 4, or to the odd y - coordinates in the down - sampled block when H = 4. When mode≧18, replace the down - scaled prediction signal with the transposed signal. In an alternative example, different strategies may be implemented. For example, instead of reducing ( "excluding") the size of the larger matrix, W red = 4 and H red = 4, the smaller matrix (idx = 1) of S1 is used. That is, such a block is now assigned to S1 instead of S2.
[0245] Other strategies may be implemented. In other examples, the mode index "mode" is not necessarily in the range from 0 to 35 (other ranges may be defined). Further, it is not necessary for each of the three sets S0, S1, S2 to have 18 matrices (thus, instead of expressions such as mode < 18, it is possible that mode < n0,n1,n2, where these are the number of matrices for each set S0, S1, S2 of matrices respectively). Further, the sets may have different numbers of matrices (for example, it may be that S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).
[0246] 6.4 Linear interpolation for generating the final prediction signal Here, features regarding step 812 are provided.
[0247] For interpolation of the subsampled prediction signal, on the larger block, a second version of the averaged boundary may be required. That is, when min(W,H)>8 and W≧H, W = 8*2 l is written, and for 0≦i<8,
[0248]
Number
[0249] Define it.
[0250] If min(W, H) > 8 and H > W, similarly
[0251] [Number]
[0252] Define it.
[0253] Additionally or alternatively, it is possible to have "hard downsampling", and in hard downsampling,
[0254] [Number]
[0255] is
[0256] [Number]
[0257] is equal to.
[0258] Also,
[0259] [Number]
[0260] can be similarly defined.
[0261] pred red At the sample positions excluded in the generation of, the final prediction signal can result from linear interpolation from pred red (e.g., step 813 in the examples from Figure 7.2 to Figure 7.4). In some examples, when W = H = 4 (e.g., the example in Figure 7.1), this linear interpolation may not be necessary.
[0262] Linear interpolation can be given as follows (although other examples are possible). Assume that W ≥ H. And, H > H red If so, pred red vertical upsampling can be performed. In that case, pred red can be extended upward by only one row as follows. If W = 8, pred red may have a width W red = 4, and for example, as defined above, the averaged boundary signal
[0263]
Number
[0264] may be extended upward by it. If W > 8, pred red has a width of W red = 8, and for example, as defined above, the averaged boundary signal
[0265]
Number
[0266] is extended upward by it. For the first row of pred red , pred red [x][-1] can be written. And for the block with width W red and height 2*H red the signal on
[0267]
Number
[0268] is
[0269]
Number
[0270] may be given as, 0 ≦ x < W red where 0 ≦ y < H red is true. The latter process can be performed k times until 2k * H red = H. Therefore, if H = 8 or H = 16, it can be performed at most once. If H = 32, it can be performed twice. If H = 64, it can be performed three times. Next, the horizontal upsampling operation can be applied to the result of the vertical upsampling. The latter upsampling operation can use the complete boundary on the left of the prediction signal. Finally, if H > W, it can proceed in the same way by first upsampling horizontally (if necessary) and then vertically upsampling.
[0271] This is an example of interpolation using the reduced boundary samples (horizontal or vertical) for the first interpolation and the original boundary samples (vertical or horizontal) for the second interpolation. Depending on the block size, only the second interpolation may be required, or no interpolation may be required. If both horizontal and vertical interpolations are required, the order depends on the width and height of the block.
[0272] However, different techniques may be implemented. For example, the original boundary samples may be used for both the first and second interpolations, and the order may be fixed. For example, it may be first horizontal and then vertical (in other cases, first vertical and then horizontal).
[0273] Therefore, the interpolation order (horizontal / vertical) and the use of reduced / original boundary samples can vary.
[0274] 6.5 Description of an Example of the Entire ALWIP Process The entire process of averaging matrix-vector multiplication and linear interpolation for different shapes in FIGS. 7.1 to 7.4 is shown. Note that the remaining shapes are treated like one of the illustrated cases.
[0275] 1. Assuming a 4×4 block, ALWIP can take two averages along each axis of the boundary by using the technique of FIG. 7.1. The four resulting input samples enter a matrix-vector multiplication. The matrix is taken from set S0. After adding an offset, this can produce 16 final predicted samples. Linear interpolation is not necessary to generate the predicted signal. Thus, a total of (4*16) / (4*4) = 4 multiplications are performed per sample. See, for example, FIG. 7.1.
[0276] 2. Assuming an 8×8 block, ALWIP can take four averages along each axis of the boundary. The eight resulting input samples enter a matrix-vector multiplication by using the technique of FIG. 7.2. The matrix is taken from set S1. This produces 16 samples at the odd positions of the prediction block. Thus, a total of (8*16) / (8*8) = 2 multiplications are performed per sample. After adding an offset, these samples can be interpolated vertically, for example by using the upper boundary, and horizontally, for example by using the left boundary. See, for example, FIG. 7.2.
[0277] 3. Assuming an 8×4 block, ALWIP can take four averages along the four original boundary values on the horizontal axis of the boundary and the left boundary by using the technique of FIG. 7.3. The eight resulting input samples enter a matrix-vector multiplication. The matrix is taken from set S1. This produces 16 samples at the odd horizontal positions of the prediction block and at each vertical position. Thus, a total of (8*16) / (8*4) = 4 multiplications are performed per sample. After adding an offset, these samples can be interpolated horizontally, for example by using the left boundary. See, for example, FIG. 7.3.
[0278] Accordingly, the case of transposition is handled.
[0279] 4. Assuming a 16×16 block, ALWIP can take four averages along each axis of the boundary. By using the technique of FIG. 7.2, the eight resulting input samples enter a matrix-vector multiplication. The matrix is taken from set S2. This produces 64 samples at the odd positions of the prediction block. Thus, a total of (8 * 64) / (16 * 16) = 2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically, for example, by using the upper boundary, and horizontally, for example, by using the left boundary. See, for example, FIG. 7.2. See, for example, FIG. 7.4.
[0280] For larger shapes, the procedure may basically be the same, and it is easy to verify that the number of multiplications per sample is less than 2.
[0281] For a W×8 block, samples are given at the odd horizontal positions and each vertical position, so only horizontal interpolation is required. Thus, in these cases, a maximum of (8 * 64) / (16 * 8) = 4 multiplications are performed per sample.
[0282] Finally, for a W×4 block where W > 8, let A k be the matrix resulting from excluding each row corresponding to the odd entries along the horizontal axis of the downsampled block. Thus, the output size may be 32, and again, only horizontal interpolation continues to be performed. A maximum of (8 * 32) / (16 * 4) = 4 multiplications can be performed per sample.
[0283] Accordingly, the transposed case can be handled.
[0284] 6.6 Evaluation of the Number of Required Parameters and Complexity The parameters required for all possible proposed intra prediction modes can be included in matrices and offset vectors belonging to sets S0, S1, S2. All matrix coefficients and offset vectors can be stored as 10-bit values. Thus, according to the above description, a total of 14,400 parameters, each with 10-bit precision, may be required for the proposed method. This corresponds to 0.018 megabytes of memory. Currently, a CTU of size 128×128 in standard 4:2:0 chroma subsampling consists of 24,576 values of 10 bits each. Thus, the memory requirements of the proposed intra prediction tool do not exceed those of the current picture reference tool adopted in the last meeting. Also, it is pointed out that conventional intra prediction modes require four multiplications per sample by a PDPC tool or a 4-tap interpolation filter for the Angular prediction mode with fractional angular positions. Thus, in terms of the complexity of operation, the proposed method does not exceed conventional intra prediction modes.
[0285] 6.7 Signaling of the Proposed Intra Prediction Mode For a luma block, for example, 35 ALWIP modes are proposed (other numbers of modes may be used). For each coding unit (CU) of the intra mode, a flag indicating whether the ALWIP mode should be applied on the corresponding prediction unit (PU) is sent in the bitstream. The signaling of the latter index can be coordinated with the MRL in the same way as in the first CE test. If the ALWIP mode should be applied, the index predmode of the ALWIP mode can be signaled using an MPM list with three MPMs.
[0286] Here, the derivation of the MPM can be performed using the intra modes of the upper and left PUs as follows. For each conventional intra prediction mode predmode Angular a table that can assign an ALWIP mode to it, for example, three fixed tables map_angular_to_alwip idx , idx ∈ {0, 1, 2} can exist. predmode ALWIP =map_angular_to_alwip idx [predmode Angular
[0287] For each PU of width W and height H, an index indicating from which of the three sets the ALWIP parameters should be taken, as in section 4 above idx(PU)=idx(W,H)∈{0,1,2} is defined. The above prediction unit PU above is available, belongs to the same CTU as the current PU, and is in the intra mode. If idx(PU)=idx(PU above ), and the ALWIP mode
[0288]
Number
[0289] is used and ALWIP is applied to PU above then
[0290]
Number
[0291] it is.
[0292] If the above PU is available, belongs to the same CTU as the current PU, is in the intra mode, and the conventional intra prediction mode
[0293]
Number
[0294] is applied to the above PU,
[0295]
Number
[0296] is.
[0297] In all other cases,
[0298]
Number
[0299] is, which means that this mode is not available. Similarly, but without the constraint that the left PU must belong to the same CTU as the current PU, mode
[0300]
Number
[0301] is derived.
[0302] Finally, three fixed default lists list idx , idx ∈ {0, 1, 2} are given, each of which contains three separate ALWIP modes. Default list list idx(PU) and mode
[0303]
Number
[0304] and
[0305]
Number
[0306] From, replace -1 with the default value and eliminate the iteration to construct three separate MPMs.
[0307] The embodiments described herein are not limited to the above signaling of the proposed intra prediction mode. According to an alternative embodiment, for MIP (ALWIP), the MPM and / or the mapping table are not used.
[0308] 6.8 Derivation of an Adapted MPM List for Conventional Luma and Chroma Intra Prediction Modes The MPM-based coding of the conventional intra prediction mode and the proposed ALWIP mode can be reconciled as follows. The luma and chroma MPM list derivation processes for the conventional intra prediction mode may use a fixed table map_alwip_to_angular idx , idx ∈ {0, 1, 2}, which maps the ALWIP mode predmode ALWIP on a given PU to one of the conventional intra prediction modes. predmode Angular = map_alwip_to_angular idx(PU) [predmode ALWIP
[0309] For the derivation of the luma MPM list, whenever a neighboring luma block using the ALWIP mode predmode ALWIP is encountered, this block can be treated as if it were using the conventional intra prediction mode predmode Angular . For the derivation of the chroma MPM list, whenever the current luma block uses the ALWIP mode, the same mapping can be used to convert the ALWIP mode to the conventional intra prediction mode.
[0310] It is clear that the ALWIP mode can be reconciled with the conventional intra prediction mode without using the MPM and / or the mapping table. For example, for a chroma block, whenever the current luma block uses the ALWIP mode, it is possible to map the ALWIP mode to the Planar intra prediction mode.
[0311] 7. Embodiments with Efficient Implementation The above example can serve as a basis for further expanding the embodiments described below in this specification, so let's briefly summarize it.
[0312] To predict a given block 18 of picture 10, a plurality of neighboring samples 17a, 17c are used.
[0313] Reduction 100 by averaging of the plurality of neighboring samples is performed to obtain a reduced set 102 of sample values with fewer samples compared to the plurality of neighboring samples. This reduction is optional in the embodiments of this specification and gives rise to what is hereinafter referred to as the so-called sample values. To obtain a predicted value for a given sample 104 of a given block, the reduced set of sample values undergoes a linear transformation or an affine linear transformation 19. This transformation, which is to be shown hereinafter using matrix A and offset vector b, is expected to be obtained and executed efficiently by machine learning (ML).
[0314] By interpolation, predicted values for further samples 108 of a given block are derived based on the predicted value for the given sample and the plurality of neighboring samples. Theoretically, it should be said that the result of the affine / linear transformation can be associated with the non-full-sample positions of block 18 such that all samples of block 18 can be obtained by interpolation according to an alternative embodiment. It may also be the case that no interpolation is necessary at all.
[0315] A plurality of neighboring samples may extend one-dimensionally along two sides of a predetermined block, and predetermined samples may be arranged in rows and columns, and along at least one of the rows and columns, the predetermined samples may be arranged at every nth position from a sample (112) that is a predetermined sample adjacent to two sides of the predetermined block. Based on the plurality of neighboring samples, for each of at least one of the rows and columns, a support value for one of the plurality of neighboring positions (118) may be determined, which is aligned to one for each of at least one of the rows and columns, and by interpolation, a predicted value for a further sample 108 of the predetermined block may be derived based on the predicted value for the predetermined sample and the support values for the neighboring samples aligned to at least one of the rows and columns. The predetermined samples may be arranged at every nth position from a sample 112 that is a predetermined sample adjacent to two sides of the predetermined block along a row, and the predetermined samples may be arranged at every mth position from a sample 112 that is a predetermined sample adjacent to two sides of the predetermined block along a column, where n, m > 1. It is also possible that n = m. Along at least one of the rows and columns, the determination of the support value may be performed by averaging (122) a group 120 of neighboring samples within the plurality of neighboring samples that includes the neighboring sample 118 for which each support value is determined, for each support value. The plurality of neighboring samples may extend one-dimensionally along two sides of the predetermined block, and reduction may be performed by grouping the plurality of neighboring samples into one or more groups 110 of consecutive neighboring samples and by performing averaging for each of the one or more groups of neighboring samples having more than two neighboring samples.
[0316] For a predetermined block, a prediction residual may be transmitted in a data stream. It may be derived therefrom in a decoder, and the predetermined block may be reconstructed using the prediction residual and the predicted value for the predetermined sample. In an encoder, the prediction residual is encoded into the data stream in the encoder.
[0317] The picture may be subdivided into a plurality of blocks of different block sizes, and these plurality of blocks include a predetermined block. And the linear transformation or affine linear transformation selected for the predetermined block is selected from a first set of linear transformations or affine linear transformations as long as the width W and height H of the predetermined block are within a first set of width / height pairs, and is selected from a second set of linear transformations or affine linear transformations as long as the width W and height H of the predetermined block are within a second set of width / height pairs separated from the first set of width / height pairs. It is possible that the linear transformation or affine linear transformation for block 18 is selected according to the width W and height H of the predetermined block. Again, later, it becomes clear that the affine / linear transformation is represented by other parameters, namely the weight of C, and optionally the offset and scale parameters.
[0318] The decoder and encoder may be configured to subdivide the picture into a plurality of blocks of different block sizes, including a predetermined block, and select a linear transformation or affine linear transformation according to the width W and height H of the predetermined block, whereby the linear transformation or affine linear transformation selected for the predetermined block is a first set of linear transformations or affine linear transformations as long as the width W and height H of the predetermined block are within a first set of width / height pairs, a second set of linear transformations or affine linear transformations as long as the width W and height H of the predetermined block are within a second set of width / height pairs separated from the first set of width / height pairs, and a third set of linear transformations or affine linear transformations as long as the width W and height H of the predetermined block are within one or more third sets of width / height pairs separated from the first and second sets of width / height pairs, is selected from.
[0319] A third set of one or more width / height pairs simply includes a single width / height pair W', H', and each linear or affine linear transformation within a first set of linear or affine linear transformations is for transforming N' sample values into W'*H' predicted values for a W'×H' array of sample positions.
[0320] Each of the first and second sets of width / height pairs may include a first width / height pair W p , H p and may not be equal to a second width / height pair W p where H p = W q and H q = W q = W p and W q = H p = H.
[0321] Each of the first and second sets of width / height pairs may additionally include a third width / height pair W p , H p where W p = H p and H p > H q = H.
[0322] For a given block, a set index may be transmitted in the data stream, which indicates the linear or affine linear transformation to be selected for block 18 from a given set of linear or affine linear transformations.
[0323] The plurality of neighboring samples may extend one-dimensionally along two sides of a predetermined block, and to obtain a first sample value from a first group and a second sample value for a second group, for a first subset of a plurality of adjacent samples adjacent to a first side of the predetermined block, the first subset is grouped into a first group 110 of one or more consecutive adjacent samples, and for a second subset of a plurality of neighboring samples adjacent to a second side of the predetermined block, the second subset is grouped into a second group 110 of one or more consecutive neighboring samples, and reduction may be performed by performing averaging on each of the first group and the second group of one or more neighboring samples having more than two neighboring samples. Next, a linear transformation or an affine linear transformation may be selected from a predetermined set of linear transformations or affine linear transformations according to a set index, such that two different states of the set index result in a selection of one of the linear transformations or affine linear transformations of the predetermined set. When the set index takes the first state of two different states in the form of a first vector, an output vector of predicted values is generated, and to spread the predicted values of the output vector along a first scanning order to predetermined samples of a predetermined block, and when the set index takes the second state of two different states in the form of a second vector, and a component filled by one of the first sample values in the first vector is filled by one of the second sample values in the second vector, and a component filled by one of the second sample values in the first vector is filled by one of the first sample values in the second vector, and the first vector and the second vector are different, an output vector of predicted values is generated, and to spread the predicted values of the output vector along a second scanning order to predetermined samples of a predetermined block transposed with respect to the first scanning order, the reduced set of sample values may be subjected to a predetermined linear transformation or affine linear transformation.
[0324] Each linear transformation or affine linear transformation within the first set of linear transformations or affine linear transformations may be for transforming N1 sample values into w1*h1 predicted values for a w1×h1 array of sample positions, and each linear transformation or affine linear transformation within the second set of linear transformations or affine linear transformations is for transforming N2 sample values into w2*h2 predicted values for a w2×h2 array of sample positions, and for a first predetermined pair of the first set of width / height pairs, w1 may exceed the width of the first predetermined width / height pair, or h1 may exceed the height of the first predetermined width / height pair, and for a second predetermined pair of the first set of width / height pairs, w1 may not exceed the width of the second predetermined width / height pair, nor may h1 exceed the height of the second predetermined width / height pair. When the predetermined block is of the first predetermined width / height pair, and when the predetermined block is of the second predetermined width / height pair, reduction (100) by averaging a plurality of neighboring samples to obtain a reduced set (102) of sample values may then be performed such that the reduced set 102 of sample values has N1 sample values, and applying a selected linear transformation or affine linear transformation to the reduced set of sample values is performed by using only a first subpart of the selected linear transformation or affine linear transformation regarding subsampling of the w1×h1 array of sample positions along the width dimension if w1 exceeds the width of a width / height pair, or along the height dimension if h1 exceeds the height of a width / height pair and the predetermined block is of the first predetermined width / height pair, and may be performed by using the selected linear transformation or affine linear transformation completely when the predetermined block is of the second predetermined width / height pair.
[0325] Each linear transformation or affine linear transformation within the first set of linear transformations or affine linear transformations may be for transforming N1 sample values into w1*h1 predicted values for a w1×h1 array of sample positions where w1 = h1, and each linear transformation or affine linear transformation within the second set of linear transformations or affine linear transformations is for transforming N2 sample values into w2*h2 predicted values for a w2×h2 array of sample positions where w2 = h2.
[0326] All of the above-described embodiments are merely exemplary in that all of them can form the basis of the embodiments described hereinafter. That is, the above concepts and details will help in understanding the following embodiments and will serve as an accumulation of possible extensions and modifications of the embodiments described hereinafter. Specifically, many of the above details, such as the averaging of neighboring samples and the fact that neighboring samples are used as reference samples, are optional.
[0327] More generally, the embodiments described herein assume that the predicted signal on a rectangular block is generated from samples that have already been reconstructed, for example, that the intra prediction signal on a rectangular block is generated from already reconstructed samples in the vicinity to the left and above of that block. The generation of the predicted signal is based on the following steps.
[0328] 1. Samples can be extracted by averaging from among the reference samples, herein called boundary samples, while precluding the possibility of diverting the description to reference samples located elsewhere. Here, the averaging is performed either on the boundary samples on both the left and the top of the block or on only the boundary samples on one of the two sides. If the averaging is not performed on a certain side, the samples on that side remain unchanged.
[0329] 2. A matrix-vector multiplication is performed, optionally followed by an addition of an offset. The input vector for the matrix-vector multiplication is either the concatenation of the averaged boundary samples on the left of the block and the original boundary samples above the block when averaging is applied only to the left side, or the concatenation of the original boundary samples on the left of the block and the averaged boundary samples above the block when averaging is applied only to the upper side, or the concatenation of the averaged boundary samples on the left of the block and the averaged boundary samples above the block when averaging is applied to both sides of the block. Again, there are alternative forms, such as a form where no averaging is used at all.
[0330] 3. Optionally, the result of the matrix-vector multiplication and the optional offset addition can be a reduced prediction signal on a subsampled set of samples within the original block. The prediction signal at the remaining positions can be generated from the prediction signal on the subsampled set by linear interpolation.
[0331] The calculation of the matrix-vector product in step 2 should preferably be performed with integer calculations. Thus, if x = (x1, … x n ) represents the input to the matrix-vector product, i.e., if x represents the concatenation of the (averaged) boundary samples on the left and above the block, the (reduced) prediction signal calculated in step 2 should be calculated from x using only bit shifts, addition of offset vectors, and multiplications with integers. Ideally, the prediction signal in step 2 is given as Ax + b, where b is an offset vector that can be 0 and A is derived by some machine learning-based training algorithm. However, such a training algorithm usually only results in a matrix A = A float given in floating-point precision. Thus, we face the problem of defining integer operations in the above-mentioned sense such that the expression A float x is well approximated using these integer operations. Here, these integer operations are not necessarily chosen to approximate A float x assuming a uniform distribution of the vector x, and the expression Afloat The input vector x, for which x is the object to be approximated, is an (averaged) boundary sample from a natural video signal, and there may usually be some correlation expected between the components of x i It is important to note that there may usually be some correlation expected between the components of x.
[0332] FIG. 8 shows an improved ALWIP prediction. Samples of a given block can be predicted based on a first matrix-vector product of a matrix A1100 derived by some machine learning-based training algorithm and a sample value vector x400. Optionally, an offset b1110 can be added. To achieve an integer approximation or a fixed-point approximation of this first matrix-vector product, the sample value vector can undergo a regular linear transformation 403 to determine a further vector 402. A second matrix-vector product between a further matrix B1200 and the further vector 402 may be equal to the result of the first matrix-vector product.
[0333] Due to the characteristics of the further vector 402, the second matrix-vector product can be an integer approximated by a matrix-vector product 404 of a given prediction matrix C405 and the further vector 402 with a further offset 408 added thereto. The further vector 402 and the further offset 408 can consist of integer values or fixed-point values. All components of the further offset are, for example, the same. The given prediction matrix 405 can be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the given prediction matrix 405 and the further vector 402 can be understood as a prediction vector 406.
[0334] Further details regarding this integer approximation are given below.
[0335] Possible solution according to Example I: Subtraction and addition of the mean value Expression A that can be used in the above scenario float One possible way to incorporate the integer approximation of x, i.e., the i0-th component of the sample value vector 400
[0336]
Number
[0337] That is, a predetermined component 1500 is replaced by the average value mean(x) of the x components, i.e., a predetermined value 1400, and this average value is subtracted from all other components. In other words, the regular linear transformation 403 as shown in FIG. 9a is defined such that a predetermined component 1500 of a further vector 402 becomes a, and each of the other components of the further vector 402 excluding the predetermined component 1500 is equal to the corresponding component of the sample value vector 400 minus a, where a is the predetermined value 1400, which is an average, such as the arithmetic mean or weighted mean of the components of the sample value vector 400. This operation on the input is given by a regular transformation T403 that has a particularly obvious integer implementation form when the dimension n of x is a power of 2.
[0338] A float =(A float T -1 )T, so when performing such a transformation on the input x, an integer approximation of the matrix-vector product By must be found, where B=(A float T -1 ) and y = Tx. Since the matrix-vector product A float x represents a prediction for a rectangular block, i.e., a predetermined block, and x400 is included in the (e.g., averaged) boundary samples of that block, when all sample values of x are equal, i.e., x i = mean(x) for all i, the prediction signal A floatEach sample value in x should be expected to be close to, or exactly equal to, mean(x). This means that the i0-th column of B, i.e., the column corresponding to a given component, should be expected to be very close to, or equal to, a column consisting only of 1s. Therefore, if M(i0), i.e., the integer matrix 1300, is a matrix such that its i0-th column consists of 1s and all other columns are 0s, then By = Cy + M(i0)y can be written, where C = B - M(i0), and the i0-th column 412 of C, i.e., the given prediction matrix 405, should be expected to have relatively small entries, or be 0, as shown in Figure 9b. Moreover, since the components of x are correlated, for each i ≠ i0, the i-th component y i = x i - mean(x) can often be expected to have an absolute value much smaller than the i-th component of x. Since the matrix M(i0)1300 is an integer matrix, the integer approximation of By is achieved if the integer approximation of Cy is given, and by the above discussion, the quantization error resulting from quantizing each entry of C405 in an appropriate way should only slightly affect the error in the resulting quantization of each By of x. float It can be expected that it will only slightly affect the error in the resulting quantization of each By of x.
[0339] The given value 1400 is not necessarily the average value mean(x). The expression A described herein float The integer approximation of x can also be achieved using the following alternative definition of the given value 1400.
[0340] Expression A float In another possible way of incorporating the integer approximation of x, the i0-th component of x
[0341]
Number
[0342] remains unchanged and has the same value
[0343]
Number
[0344] is subtracted from all other components. That is, for each i≠i0,
[0345]
Number
[0346] and
[0347]
Number
[0348] is. In other words, the predetermined value 1400 can be a component of the sample value vector 400 corresponding to the predetermined component 1500.
[0349] Alternatively, the predetermined value 1400 is a default value, or a value signaled in the data stream in which the picture is coded.
[0350] The predetermined value 1400 is, for example, 2 bitdepth-1 is equal to. In this case, a further vector 402 may be defined by y0 = 2 bitdepth-1 and for i>0, y i = x i - x0.
[0351] Alternatively, the predetermined component 1500 is the result of subtracting the predetermined value 1400 from a constant. The constant is, for example, 2 bitdepth-1 is equal to. According to an embodiment, a predetermined component of a further vector y402
[0352]
Number
[0353] 1500 is 2 bitdepth-1equal to the component of the sample value vector 400 corresponding to a predetermined component 1500 minus, and all other components of the further vector 402 are equal to the corresponding components of the sample value vector 400 minus the component of the sample value vector 400 corresponding to the predetermined component 1500.
[0354] [Number]
[0355] For example, it is advantageous for a predetermined value 1400 to have a small deviation from the predicted value of the samples of a predetermined block.
[0356] According to an embodiment, the apparatus 1000 is configured to include a plurality of regular linear transformations 403, each of which is associated with one component of the further vector 402. Further, the apparatus is configured to select, for example, a predetermined component 1500 from the components of the sample value vector 400 and use the regular linear transformation 403 from the plurality of regular linear transformations associated with the predetermined component 1500 as the predetermined regular linear transformation. This is due to, for example, the position of the i0-th row, i.e., the position of the row of the regular linear transformation 403 corresponding to the predetermined component, being different depending on the position of the predetermined component in the further vector. For example, if the first component of the further vector 402, i.e., y1, is the predetermined component, the i0-th row replaces the first row of the regular linear transformation.
[0357]
[0358] As shown in FIG. 9b, the column 412 of the predetermined prediction matrix 405 corresponding to the predetermined component 1500 of the further vector 402, that is, the matrix component 414 of the predetermined prediction matrix C405 within the i0-th column, is, for example, all 0. In this case, the apparatus is configured to calculate the matrix-vector product 404, for example, by performing a multiplication by calculating a matrix-vector product 407 of a reduced prediction matrix C'405 obtained from the predetermined prediction matrix C405 by excluding the column 412 and a further vector 410 obtained from the further vector 402 by excluding the predetermined component 1500, as shown in FIG. 9c. Therefore, the prediction vector 406 can be calculated with fewer multiplications.
[0359] As shown in FIGS. 8, 9b, and 9c, when predicting samples of a predetermined block based on the prediction vector 406, the apparatus 1000 may be configured to calculate, for each component of the prediction vector 406, the sum of each component and a, that is, a predetermined value 1400. As shown in FIGS. 8 and 9c, this addition may be represented by an addition of the prediction vector 406 and the vector 409, where all components of the vector 409 are equal to the predetermined value 1400. Alternatively, as shown in FIG. 9b, this addition may be represented by the sum of the prediction vector 406 and the matrix-vector product 1310 of the integer matrix M1300 and the further vector 402, where the matrix components of the integer matrix 1300 are an example of an integer matrix corresponding to the predetermined component 1500 of the further vector 402, that is, 1 within the i0-th column and, for example, 0 for all other components.
[0360] The result of the addition of the predetermined prediction matrix C405 and the integer matrix 1300 is equal to, or close to, for example, a further matrix B1200 shown in FIG. 8.
[0361] In other words, as shown in FIGS. 8, 9a, and 9b, a matrix obtained by adding matrix components with a value of 1 in the i0-th column 412 of a predetermined prediction matrix 405 corresponding to a predetermined component 1500 of a further vector 402, that is, a matrix obtained by adding each matrix component of a predetermined prediction matrix C405 within the i0-th column in which the value is 1, that is, a matrix obtained by multiplying a further matrix B1200 (that is, a plurality of matrices B) by a regular linear transformation 403, corresponds to, for example, a quantized version of a machine learning prediction matrix A1100. As shown in FIG. 9b, the addition of each matrix component of the predetermined prediction matrix C405 within the i0-th column 412 with a value of 1 may correspond to the addition of the predetermined prediction matrix 405 and an integer matrix 1300. As shown in FIG. 8, the machine learning prediction matrix A1100 may be equal to the result of multiplying a further matrix 1200 by a regular linear transformation 403. This is because A·x = BT·yT -1 This is because the predetermined prediction matrix 405 is, for example, a quantized matrix, an integer matrix, and / or a fixed-point matrix, whereby a quantized version of the machine learning prediction matrix A1100 can be realized.
[0362] Matrix multiplication using only integer operations (Regarding the complexity of scalar value addition and multiplication, as well as the storage required for the entries of the relevant matrices), in an implementation with low complexity, it is desirable to perform matrix multiplication 404 using only integer calculations.
[0363] Using only operations on integers, for the approximation z = Cy, that is
[0364]
Number
[0365] To calculate, according to one embodiment, the real value C i,j is an integer value
[0366]
Number
[0367] must be mapped to. This can be done, for example, by uniform scalar quantization or by considering a specific correlation between the values y i The integer values can each be stored using a fixed number of bits n_bits, for example n_bits = 8, which represents a fixed-point number. For example, the integer values can each be stored using a fixed number of bits n_bits, for example n_bits = 8, which represents a fixed-point number.
[0368] The matrix-vector product 404 with a matrix of size m×n, i.e., a predetermined prediction matrix 405, may then be performed as shown in this pseudocode, where <<, >> are binary left shift and right shift operations, and +, -, and * act only on integer values.
[0369] (1) final_offset = 1 << (right_shift_result - 1); for i in 0..m-1 { accumulator = 0 for j in 0..n-1 { accumulator := accumulator + y[j]*C[i,j] } z[i] = (accumulator + final_offset) >> right_shift_result; }
[0370] Here, the array C, i.e., the predetermined prediction matrix 405, stores fixed-point numbers as integers. The addition of the final final_offset and the right shift operation using right_shift_result reduce the precision by rounding to obtain the fixed-point format required in the output.
[0371] To enable an expansion of the range of real values representable by integers in C, as shown in the embodiments of FIGS. 10 and 11, two additional matrices offset i,jand scale i,j can be used, so the matrix-vector product
[0372] [Number]
[0373] the y in j each coefficient b of i,j is
[0374] [Number]
[0375] given by
[0376] value offset i,j and scale i,j themselves are integer values. For example, these integers can be stored using a certain number of bits, for example 8 bits, or can each be stored using the same number of bits n_bits used to store the value
[0377] [Number]
[0378] which can represent fixed-point numbers.
[0379] In other words, the apparatus 1000 has prediction parameters, such as integer values
[0380] [Number]
[0381] as well as the value offset i,j and scale i,jis used to represent a predetermined prediction matrix 405 and is configured to calculate a matrix-vector product 404 by performing multiplication and addition on the components of a further vector 402, prediction parameters, and intermediate results obtained therefrom, and the absolute value of the prediction parameter can be represented by an n-bit fixed-point decimal representation with n being 14 or less, or alternatively 10 or less, or alternatively 8 or less. For example, the components of the further vector 402 are multiplied by the prediction parameters to produce a product as an intermediate result, and this product then undergoes addition or serves as an addend for addition.
[0382] According to an embodiment, the prediction parameters include weights, each of which is associated with a corresponding matrix component of the prediction matrix. In other words, a predetermined prediction matrix is replaced or represented by, for example, the prediction parameters. The weights are, for example, integer values and / or fixed-point values.
[0383] According to an embodiment, the prediction parameters further include one or more scaling factors, such as the value scale i,j and each of them is associated with a weight, such as an integer value, corresponding to one or more corresponding matrix components of the predetermined prediction matrix 405
[0384]
Number
[0385] for scaling one or more corresponding matrix components of the predetermined prediction matrix 405. Additionally or alternatively, the prediction parameters include one or more offsets, such as the value offset i,j and each of them is associated with a weight, such as an integer value, corresponding to one or more corresponding matrix components of the predetermined prediction matrix 405
[0386]
Number
[0387] is associated with one or more corresponding matrix components of a given prediction matrix 405 for offsetting.
[0388] offset i,j and scale i,j To reduce the amount of storage required for them, those values can be chosen to be constant for a particular set of indices i, j. For example, as shown in FIG. 10, those entries may be constant for each column, or they may be constant for each row, or they may be constant for all i, j.
[0389] For example, in one preferred embodiment, as shown in FIG. 11, offset i,j and scale i,j are constant for all values of the matrix of one prediction mode. Thus, when there are K prediction modes and k = 0... K-1, only a single value o k and a single value s k are required to calculate the prediction for mode k.
[0390] According to an embodiment, offset i,j and / or scale i,j are constant, i.e., the same, for all matrix-based intra prediction modes. Additionally or alternatively, it is possible that offset i,j and / or scale i,j are constant, i.e., the same, for all block sizes.
[0391] Since offset represents o k and scale represents s k the calculation of (1) can be modified as follows. (2) final_offset = 0; for i in 0..n-1 { final_offset := final_offset - y[i];} final_offset *= final_offset * offset * scale; final_offset += 1 << (right_shift_result - 1); for i in 0..m-1 { accumulator = 0 for j in 0..n-1 { accumulator := accumulator + y[j]*C[i,j] } z[i] = (accumulator*scale + final_offset) >> right_shift_result; }
[0392] The extended embodiments resulting from that solution The above solution suggests the following embodiments.
[0393] 1. A prediction method as in Section I, wherein in step 2 of Section I, the following is done for an integer approximation of the relevant matrix-vector product. For a certain i0 with 1 ≦ i0 ≦ n, from among the (averaged) boundary samples x = (x1,…,x n ), the vector y = (y1,…,y n ) is calculated, where for i ≠ i0, y i = x i - mean(x),
[0394]
Number
[0395] and mean(x) represents the average value of x. Then, since the vector y serves as the input for the matrix-vector product Cy (embodied by an integer), the (downsampled) predicted signal pred from step 2 of section I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents a signal equal to mean(x) for each sample position in the region of the (downsampled) predicted signal (see, for example, Fig. 9b).
[0396] 2. A prediction method as in section I, wherein in step 2 of section I, the following is done for the integer approximation of the relevant matrix-vector product. For a certain i0 with 1 ≤ i0 ≤ n, from among the (averaged) boundary samples x = (x1,..., x n ), a vector y = (y1,..., y n-1 ) is calculated, where for i < i0, y i = x i - mean(x), and for i ≥ i0, y i = x i+1 - mean(x), and mean(x) represents the average value of x. Then, since the vector y serves as the input for the matrix-vector product Cy (embodied by an integer), the (downsampled) predicted signal pred from step 2 of section I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents a signal equal to mean(x) for each sample position in the region of the (downsampled) predicted signal (see, for example, Fig. 9c).
[0397] 3. A prediction method as in section I, wherein the integer embodiment of the matrix-vector product Cy is the coefficient i = Σ j b i,j * y j in
[0398]
Number
[0399] given by using (see, for example, FIG. 10).
[0400] 4. A prediction method as in Section I, wherein step 2 uses one of K matrices, whereby a plurality of prediction modes are calculated using different matrices for k = 0…K-1
[0401]
Number
[0402] can each be calculated using, and the matrix-vector product C k the integer embodiment of y is the matrix-vector product
[0403]
Number
[0404] the coefficients in
[0405]
Number
[0406] given by using (see, for example, FIG. 11).
[0407] That is, according to an embodiment of the present application, the encoder and decoder operate as follows to predict a predetermined block 18 of picture 10. See FIG. 8. For prediction, a plurality of reference samples are used. As outlined above, since the embodiments of the present application are not limited to intra-coding, the reference samples are not limited to neighboring samples, i.e., samples of picture 10 in the vicinity of block 18. Specifically, the reference samples are not limited to samples arranged along the outer edge of block 18, such as samples that contact the outer edge of the block. However, this situation is of course one embodiment of the present application.
[0408] To perform the prediction, a sample value vector 400 is formed from reference samples such as reference samples 17a and 17c. Possible formations have been described above. This formation may involve averaging and thereby reducing the number of samples 102 or the number of components of vector 400 as compared to the reference samples contributing to the formation. This formation may also depend in some way on the dimensions or size of block 18, such as its width and height, as described above.
[0409] This vector 400 should undergo an affine or linear transformation to obtain the prediction of block 18. Various terms have been used above. Using the most recent ones, the aim is to perform the prediction by applying vector 400 to matrix A by means of a matrix-vector product when performing the addition with offset vector b. The offset vector b is optional. The affine or linear transformation determined by A or A and b can be determined by the encoder and decoder, or more precisely, for the prediction based on the size and dimensions of block 18 as already described above.
[0410] However, in order to achieve the improvement in computational efficiency outlined above, or to make predictions more effective with respect to implementation, the affine and linear transforms are quantized, and the encoders and decoders, or their predictors, use C and T to represent and execute a linear or affine transform that represents a quantized version of the affine transform applied in the manner described above. Specifically, instead of applying vector 400 directly to matrix A, the predictors in the encoders and decoders apply a mapping via a given regular linear transform T to sample value vector 400, and apply vector 402 obtained from sample value vector 400. The transform T used here can be the same as long as vector 400 has the same size, i.e., is independent of the dimensions of the block, i.e., the width and height, or is at least the same for different affine / linear transforms. Above, vector 402 is denoted as y. The exact matrix for performing the affine / linear transform as determined by machine learning should have been B. However, instead of executing B exactly, the prediction in the encoders and decoders is performed by an approximation or its quantized version. Specifically, the representation is done by appropriately representing C in the manner outlined above using C+M that represents the quantized version of B.
[0411] Accordingly, the prediction in the encoder and decoder is further performed by calculating a matrix-vector product 404 of the vector 402 and a predetermined prediction matrix C that is appropriately represented and stored in the encoder and decoder in the manner described above. The vector 406 obtained from this matrix-vector product is then used to predict the samples 104 of block 18. As described above, for prediction, each component of the vector 406 may undergo an addition with a parameter a as shown at 408 to perform padding for the corresponding definition of C. An optional addition of the vector 406 with the offset vector b may also be involved in deriving the prediction of block 18 based on the vector 406. As described above, each component of the vector 406, and thus each component of the addition of the vector 406, the vector all of whose components shown at 408 are a, and the optional vector b may directly correspond to the samples 104 of block 18, and thus may indicate the predicted value of the samples. It may also be the case that only a subset of the samples 104 of the block is predicted in that manner, and the remaining samples of block 18, such as 108, are derived by interpolation.
[0412] As described above, there are various embodiments for setting a. For example, it can be the arithmetic mean of the components of the vector 400. In that case, refer to FIG. 9a. The regular linear transformation T403 can be as shown in FIG. 9a. i0 is a predetermined component of the sample value vector and the vector 402, and these are each replaced by a. However, as also described above, there are other possibilities. However, as long as it relates to the representation of C, it has also been described that the same thing can be embodied differently. For example, the matrix-vector product 404 may, in its actual calculation, be the actual calculation of a smaller matrix-vector product with fewer dimensions. Specifically, as described above, due to the definition of C, the entire i0-th column 412 thereof becomes 0, so the actual calculation of the product 404 is the component
[0413]
Number
[0414] It may be possible to perform it by a reduced version of vector 402 obtained from vector 402 by omission, i.e., by excluding the i0-th column 412, that is, by multiplying the reduced matrix C' obtained from C by this reduced vector 410.
[0415] The weights of C or C', i.e., the components of this matrix, can be represented and stored in fixed-point decimal representation. However, these weights 414 can also be stored in a manner related to different scales and / or offsets as described above. The scale and offset may be defined for the entire matrix C, i.e., may be equal for all weights 414 of matrix C or matrix C', or may be defined in such a way that they are constant or equal for all weights 414 in the same row or the same column of matrix C and matrix C'. FIG. 10 shows, in this regard, that the calculation of the matrix-vector product, i.e., the result of the product, may be performed in such a way that it is actually slightly different, i.e., for example, by shifting the multiplication by the scale towards vector 402 or 410, thereby reducing the number of multiplications that have to be performed. FIG. 11 shows an example of using one scale and one offset for all weights 414 of C or C' as done in the above equation (2).
[0416] According to an embodiment, the apparatus described herein for predicting a predetermined block of a picture can be configured to use matrix-based intra-sample prediction including the following features.
[0417] The device is configured to form a sample value vector pTemp[x] 400 from a plurality of reference samples 17. Assuming that pTemp[x] is 2*boundarySize, then, for example, by direct copying, or by subsampling, or by pooling, pTemp[x] can be filled by redT[x] which are neighboring samples located above a predetermined block, where x = 0…boundarySize-1, and then by redL[x] which are neighboring samples located to the left of the predetermined block, where x = 0…boundarySize-1 (for example, when Transposed = 0), or vice versa in the case of a transposed process (for example, when Transposed = 1).
[0418] An input value p[x] where x = 0…inSize-1 is derived, that is, the device is configured to derive a further vector p[x] from the sample value vector pTemp[x] such that the sample value vector pTemp[x] is mapped as follows by a predetermined regular linear transformation, or more specifically by a predetermined regular affine linear transformation. - When mipSizeId is equal to 2, the following applies. p[x]=pTemp[x+1]-pTemp[0] - Otherwise (when mipSizeId is less than 2), the following applies. p[0]=(1<<(BitDepth-1))-pTemp[0] For x = 1…inSize-1, p[x]=pTemp[x]-pTemp[0]
[0419] Here, the variable mipSizeId indicates the size of a predetermined block. That is, according to this embodiment, the regular transformation by which a further vector is derived from the sample value vector using it depends on the size of the predetermined block. The dependency is
[0420]
Table 3
[0421] can be provided in accordance with.
[0422] If predSize indicates the number of predicted samples within a given block, 2 * boundarySize indicates the size of the sample value vector, and inSize is related to the size of inSize, i.e., the size of the further vector S, according to inSize = (2 * boundarySize) - (mipSizeId == 2)? 1:0. More precisely, inSize indicates the number of components of the further vector that actually participate in the calculation. inSize is the same size as the sample value vector for smaller block sizes and one component smaller than the sample value vector for larger block sizes. In the former case, one component, i.e., the component corresponding to a given component of the further vector, may be ignored, as its contribution to the corresponding vector component will be 0 in any case in the matrix-vector product to be calculated later and thus does not actually need to be calculated. The dependence on the block size may not be eliminated in an alternative embodiment, in which case only one of the two options will necessarily be used independently of the block size (the option corresponding to mipSizeId less than 2 or the option corresponding to mipSizeId equal to 2).
[0423] In other words, a given regular linear transformation is defined such that, for example, a given component of a further vector p becomes a, while all other components correspond to the components of the sample value vector minus a. For example, a = pTemp[0]. In the case of the first option corresponding to a mipSizeId equal to 2, this is readily apparent, and only the differentially formed components of the further vector are further considered. That is, in the case of the first option, the further vector is actually {p[0…inSize]; pTemp[0]}, where pTemp[0] is a, and the actually computed part of the matrix-vector product, i.e., the matrix-vector multiplication that produces the result of the multiplication, is limited to the inSize components of the further vector and the corresponding columns of the matrix, because there is a zero column in the matrix that does not require computation. In other cases corresponding to a mipSizeId smaller than 2, a = pTemp[0] is chosen, because each of the other components of the further vector except p[0], i.e., the other components p[x] of the further vector p except the given component p[0] (for x = 1…inSize - 1), is equal to the corresponding component of the sample value vector pTemp[x] minus a, while p[0] is chosen to be the constant minus a. Then, the matrix-vector product is computed. This constant is the average of the representable values, i.e., 2 x-1 (i.e., 1<<(BitDepth - 1)), where x represents the bit depth of the computational representation used. Note that if p[0] is chosen to be pTemp[0] instead, the computed product simply deviates from that computed using p[0] as described above by the amount of the constant vector that can be considered when predicting the inside of the block based on the product, i.e., the prediction vector (p[0] = (1<<(BitDepth - 1)) - pTemp[0]). Thus, the value a is a given value, for example, pTemp[0]. The given value pTemp[0] is, in this case, for example, the component of the sample value vector pTemp corresponding to the given component p[0]. It can be the nearest neighbor sample to the upper left corner of the given block that is above or to the left of the given block.
[0424] For example, for intra-sample prediction processing according to predModeIntra that specifies an intra prediction mode, the apparatus is configured to apply, for example, at least the following steps, for example, to execute at least the first step.
[0425] 1. The matrix-based intra prediction sample predMip[x][y] where x = 0...predSize-1 and y = 0...predSize-1 is derived as follows. - The variable modeId is set equal to predModeIntra. - The weight matrix mWeight[x][y] where x = 0...inSize-1 and y = 0...predSize*predSize-1 is derived by calling a MIP weight matrix derivation process with mipSizeId and modeId as inputs. - The matrix-based intra prediction sample predMip[x][y] where x = 0...predSize-1 and y = 0...predSize-1 is derived as follows.
[0426]
Number
[0427] In other words, the apparatus is configured to calculate a matrix-vector product of a further vector p[i], or {p[i]; pTemp[0]} if mipSizeId is equal to 2, and a predetermined prediction matrix mWeight, or a prediction matrix mWeight with additional zero-weight lines corresponding to the omitted components of p if mipSizeId is less than 2, to obtain a prediction vector, where the prediction vector is already assigned to an array of block positions {x, y} distributed inside a predetermined block so as to result in an array predMip[x][y]. The prediction vectors respectively correspond to the concatenation of the rows of predMip[x][y] and the columns of predMip[x][y].
[0428] According to one embodiment, or according to a different interpretation, the component
[0429]
Number
[0430] is only understood as a prediction vector, and when the device predicts samples of a predetermined block based on the prediction vector, for each component of the prediction vector, the sum of each component and a, for example pTemp[0], is calculated.
[0431] Optionally, the device may be configured to additionally execute the following steps when predicting samples of a predetermined block based on a prediction vector, such as predMip or
[0432]
Number
[0433] based on.
[0434] 2. The matrix-based intra prediction sample predMip[x][y] where x = 0...predSize-1 and y = 0...predSize-1 is clipped, for example, as follows. predMip[x][y]=Clip1(predMip[x][y])
[0435] 3. When isTransposed is equal to TRUE, the predSize×predSize array predMip[x][y] where x = 0...predSize-1 and y = 0...predSize-1 is transposed, for example, as follows. predTemp[y][x]=predMip[x][y] predMip=predTemp
[0436] 4. The predicted samples predSamples[x][y] where x = 0…nTbW-1 and y = 0…nTbH-1 are derived, for example, as follows. - If nTbW, which specifies the width of the transform block, is greater than predSize, or if nTbH, which specifies the height of the transform block, is greater than predSize, a MIP prediction upsampling process is called with as input the matrix-based intra prediction samples predMip[x][y] with input block size predSize, where x = 0…predSize-1 and y = 0…predSize-1, the width nTbW of the transform block, the height nTbH of the transform block, the upper reference samples refT[x] where x = 0…nTbW-1, and the left reference samples refL[y] where y = 0…nTbH-1, and the output is the predicted sample array predSamples. - Otherwise, predSamples[x][y] where x = 0…nTbW-1 and y = 0…nTbH-1 is set equal to predMip[x][y].
[0437] In other words, the apparatus is configured to predict the samples predSamples of a given block based on the prediction vector predMip.
[0438] 8. Using the block / matrix-based intra prediction mode together with other intra prediction modes The following description again presents the possibilities for combining other intra prediction modes with block / matrix-based prediction. It is a further presentation of the possibilities on which the embodiments described in the subsequent sections may be implemented.
[0439] Note that in the following, the term block-based intra prediction is used to denote an intra prediction mode that may be implemented in or equivalent to what is shown by the above ALWIP.
[0440] Accordingly, the embodiments described below with respect to FIG. 12 relate to a decoder and an encoder that support intra prediction for decoding / encoding a predetermined block 18, and various intra prediction modes are supported. There is an Angular intra prediction mode 500 in which reference samples 17 in the vicinity of a predetermined block 18 are used accordingly to fill the predetermined block 18 and obtain an intra prediction signal for the predetermined block 18. Specifically, the reference samples 17 arranged along the boundary of the predetermined block 18, such as along the upper and left edges of the predetermined block 18, represent picture content that is extrapolated or copied into the predetermined block 18 along a predetermined direction 502. Before extrapolation or copying, the picture content represented by the neighboring samples 17 may undergo interpolation filtering, or in other words, may be derived from the neighboring samples 17 by interpolation filtering. The Angular intra prediction modes 500 are different from each other in the intra prediction direction 502. Each Angular intra prediction mode 500 may have an associated index, and the association of the index to the Angular intra prediction mode 500 may be such that when the directions 500 are arranged in the Angular intra prediction mode 500 according to the associated mode index, they rotate monotonically clockwise or counterclockwise.
[0441] There may also be non-Angular intra prediction modes. For example, 504 shows a Planar intra prediction mode that is optionally included in the set 508 in FIG. 12. Accordingly, a two-dimensional linear function defined by a horizontal slope, a vertical slope, and an offset is derived based on the neighboring samples 17, and this linear function defines the predicted sample values of the predetermined block 18. The horizontal slope, the vertical slope, and the offset are derived based on the neighboring samples 17. According to an embodiment, the first set 508 of intra prediction modes includes the Planar intra prediction mode 504.
[0442] A particular non-Angular intra prediction mode, namely the DC mode, included in set 508 is shown at 506. Here, one value, which is a pseudo-DC value, is derived based on neighboring sample 17, and this one DC value is attributable to all samples of a given block 18 for obtaining an intra prediction signal. Two examples for non-intra prediction modes are shown, but only one, or more than two, may exist.
[0443] Intra prediction modes 500, 504, and 506 form a set 508 of intra prediction modes supported by an encoder and a decoder that compete with block-based intra prediction modes in the sense of rate / distortion optimization, which are generally shown using reference numeral 510 with examples described above using the abbreviation ALWIP. As described above, according to these block-based intra prediction modes 510, matrix-vector product 520 is performed between a vector 514 derived from neighboring samples 17 at one edge and a given prediction matrix 516 at the other edge. The result of multiplication 520 is a prediction vector 518 used to predict samples of a given block 18. The block-based intra prediction modes 510 are different from each other in the prediction matrix 516 associated with each mode.
[0444] Thus, briefly summarized, an encoder and a decoder according to embodiments described herein include a set 508 of intra prediction modes, namely a first set of intra prediction modes, and a set 520 of block-based intra prediction modes, namely a second set of matrix-based intra prediction modes, and they compete with each other.
[0445] According to an embodiment of the present application, a predetermined block 18 is coded / decoded using intra prediction in the following manner. Specifically, first, a set selection syntax element 522 as to whether the predetermined block 18 should be predicted using any of a set 508 of intra prediction modes or using any of the modes of a set 520 of block-based intra prediction modes. If the set selected syntax element indicates that the predetermined block 18 should be predicted using any of the modes of the set 508, i.e., the first set of intra prediction modes, a list 528 of the most likely candidates in the set 508 is constructed / formed in the decoder and encoder based on the intra prediction mode by which the neighboring blocks, i.e., the neighboring blocks 18 exemplified at 524 and 526, are predicted using it. The neighboring blocks 524 and 526 can be determined relative to the position of the predetermined block 18 in a predetermined manner, such as by determining neighboring blocks that overlap some of the neighboring samples of the block 18, such as samples above the samples at the upper left edge of the block 18, and a block 526 that includes samples to the left of the corner samples just mentioned. Of course, this is only an example. The same applies to the number of neighboring blocks used for mode prediction, and this number is not limited to 2 for all embodiments. More than two or only one may be used. If any of these blocks 524 and 526 are missing, a default intra prediction mode may be used as an alternative to the intra prediction mode of the missing neighboring block. The same may apply if any of the blocks 524 and 526 are coded / encoded using an inter prediction mode, such as by motion compensation prediction.
[0446] The construction of the list of modes in the set 508 of maximum likelihood intra prediction modes, i.e., list 528, is as follows. The length of list 528, i.e., the number of maximum likelihood modes in the list, may be constant by default. This length may be 4 as shown in FIG. 12, or may be different from 4, such as 5 or 6. In the latter case, it applies to the specific examples described below in this specification. The index in the data stream, which will be described later, may indicate one of the modes in list 528 to be used for a given block 18. Indexing is performed along the list order or ranking 530, and the list index is of variable length, coded such that, for example, the length of the index monotonically increases along order 530. Thus, first, fill list 528 with only the maximum likelihood modes in set 508, and it is valuable to place more likely modes upstream of less likely modes appropriate for block 18 along order 530. The modes of list 528 are derived based on the modes used for blocks 524 and 526, i.e., neighboring blocks in the vicinity of a given block 18. If either of blocks 524 and 526 is intra predicted using the block-based mode in set 520, the aforementioned mapping from such an "ALWIP" or block-based mode 510 to a mode within set 508, e.g., a non-ALWIP mode, is used. The latter mapping may map, for example, the majority (i.e., more than half) of the block-based modes 510 to the DC mode 506 (or either DC506 or Planar mode 504).
[0447] According to one embodiment, the list 528 of the most likely intra prediction modes is filled using the Planar intra prediction mode 504 in a manner independent of the intra prediction mode by which the neighboring blocks are predicted using it. Thus, for example, only the DC intra prediction mode 506 and the Angular intra prediction mode 500 are filled in the list 528 depending on the intra prediction mode used for the prediction of the neighboring blocks 524 and 526. The Planar intra prediction mode 504 is placed, for example, at the first position of the list 528 of the most likely intra prediction modes independently of the intra prediction mode by which the neighboring blocks 524 and 526 are predicted using it.
[0448] In the approach exemplified in more detail below, the list construction of the list 528 of the most likely intra prediction modes is performed in such a way that if the neighboring blocks 524 and 526 are predicted only by any of the Angular intra prediction modes 500, the DC intra prediction mode 506 is not in the list 528. If one of the neighboring blocks 524 or 526 is predicted by any of the Angular intra prediction modes 500, and / or if both of the neighboring blocks 524 and 526 are predicted by any of the Angular intra prediction modes 500, the DC intra prediction mode 506 is not in the list 528 of the most likely intra prediction modes. According to the embodiments described below in this specification, for example, the list 528 is for all neighboring blocks 524 and 526 in the following situations, namely, all neighboring blocks 524 and 526 are coded using either of the non-Angular intra prediction modes 504 and 506, or all neighboring blocks 524 and 526 are intra predicted using any block-based intra prediction mode 510, and the block-based intra prediction mode 510 is mapped to either of the non-Angular intra prediction modes 504 and 506 by the aforementioned mapping from the block-based intra prediction mode 510 to the modes within the set 508, only if any of these is true. Only in that case is the DC mode 506 used to fill. Only in that case is the DC intra prediction mode 506 placed in the list 528. In that case, as may be seen in subsequent examples, the DC intra prediction mode 506 may be placed before any of the Angular intra prediction modes 500 in the order 530.
[0449] In other words, the list 528 of the most likely intra prediction modes is used, for example, for each of the neighboring blocks 524 and 526, only when each respective neighboring block is predicted using at least one of the non-Angular intra prediction modes 504 and 506, together with the first set 508 including the DC intra prediction mode 506, or only when predicted using any of the block-based intra prediction modes 510 that are mapped to any of the at least one non-Angular intra prediction modes 500 by a mapping from the second set 520 of the block-based intra prediction modes 510 to the intra prediction modes within the first set 508 that are used for forming the list 528 of the most likely intra prediction modes, is filled with the DC intra prediction mode 506. The DC intra prediction mode 506 is placed, for example, before any of the Angular intra prediction modes 500 in the list 528 of the most likely intra prediction modes.
[0450] Resuming the description of how a given block 18 is coded into the data stream 12, the data stream 12 optionally includes an MPM syntax element 532 indicating whether an intra prediction mode to be used for the given block 18 is in the list 528 if the set selection syntax element 522 indicates that the given block 18 should be coded by any mode in the first set 508. If it is in the list 528, the data stream 12 includes an MPM list index 534 pointing within the list 528 indicating the mode to be used for the given block 18 in the list 528, i.e., the given intra prediction mode, by indexing the modes along the order 530. However, if the mode in the set 508 is not in the list 528 as indicated by the MPM syntax element 532, the data stream 12 includes a further syntax element 536 for the block 18 indicating which mode in the set 508, i.e., which given intra prediction mode, should be used for the block 18. The further syntax element 536 may indicate the mode in a way by simply differentiating the modes in the set 508 not included in the list 528.
[0451] In other words, the apparatus for decoding a given block 18 is configured to derive, from the data stream, an MPM syntax element 532 indicating whether a given intra prediction mode of the first set 508 of intra prediction modes is within a list 528 of most probable intra prediction modes, when, for example, a set selection syntax element 522 indicates that the given block 18 is to be predicted using one of the first set 508 of intra prediction modes. When the MPM syntax element 532 indicates that a given intra prediction mode of the first set 508 of intra prediction modes is within the list 528 of most probable intra prediction modes, the apparatus is configured to, for example, perform the formation of the list 528 of most probable intra prediction modes based on the intra prediction modes by which neighboring blocks 524, 526 in the neighborhood of the given block 100 are predicted using it, and to perform the derivation from the data stream 12 of an MPM list index 534 indicating the given intra prediction mode of the list 528 of most probable intra prediction modes. When the MPM syntax element 532 from the data stream 12 indicates that a given intra prediction mode of the first set 508 of intra prediction modes is not within the list 528 of most probable intra prediction modes, the apparatus is configured to derive from the data stream a further list index 536 indicating the given intra prediction mode from the first set of intra prediction modes. Thus, based on the MPM syntax element 532, the data stream 12 includes either an MPM list index 534 or a further list index 536 for the prediction of the given block 18.
[0452] By eliminating situations where list 528 includes the DC intra prediction mode 506, the following advantages are achieved. Specifically, the inventors of the present application have found that using a valuable list position of list 528 by the DC intra prediction mode 506 in set 508 for coding / decoding a predetermined block 18, such as indicated by the syntax element 522, i.e., the set selection syntax element, any of the intra prediction modes of set 508, "consumes" the list position, which has an adverse effect on coding efficiency because such a DC intra prediction mode 506 in set 508 always competes with the block-based intra prediction mode 510. Therefore, "consuming" the list position of list 528 by such a DC intra prediction mode 506 in set 508 increases the probability of a situation where the syntax element 536, i.e., an additional list index, needs to be transmitted in the data stream 12 because the intra prediction mode that should ultimately be used for the predetermined block 18, i.e., the predetermined intra prediction mode, is not in list 528.
[0453] Specifically, since the syntax element 522 already indicates whether block 18 should be predicted using any of the modes in set 508 or any of the block-based modes 510 of set 520 for block 18, if the syntax element 522 indicates that a mode in set 508 is preferred for block 18 and thus the block-based mode 510 should not be used for block 18, the probability that the DC prediction mode 506 in set 508 can be appropriate for block 18 seems to be as low as it should be limited to a very limited set of modes used for neighboring blocks 524 and 526, i.e., the set presented above, due to the presence of the DC prediction mode in list 528.
[0454] In other cases, i.e., when the set selection syntax element 522 indicates that a given block 18 should be predicted using any of the block-based intra prediction modes 510, the coding of block 18 into the data stream 12 and its subsequent decoding can be performed in the manner presented above. For this purpose, indexing may be used to index one of the selected block-based intra prediction modes 510 out of the second set of block-based intra prediction modes in set 520, or to indicate which of the block-based intra prediction modes should be used. A further MPM syntax element 538 indicates whether the indexing is done by index 540, i.e., by a further MPM list index indicating the block-based intra prediction mode 510 to be used for block 18 in the list 542 of most likely block-based intra prediction modes, i.e., by indexing along the list order 544, or whether the block-based intra prediction mode 510 to be used for block 18 is indicated by a further syntax element 546, i.e., by a further list index indicating the block-based intra prediction mode 510 in set 520, where the latter syntax element 546 may, for example, only distinguish modes 510 within set 520 that are not yet included in list 542. The construction of list 542 can be performed based on the modes for which blocks 524 and 526 are predicted using it. If either of blocks 524 and 526 is not available because it is outside the picture or is inter-predicted, a default intra prediction mode, such as one of those in set 508, may be used instead.For each of blocks 524 and 526 that is intra-predicted using a mode in set 508 instead of set 520, the aforementioned mapping from the mode of set 508 to a mode in set 520 is used to obtain an intra-prediction mode 510, i.e., a block-based intra-prediction mode for a respective block, i.e., a predetermined block 18, and list 542 is interpreted based on the obtained block-based intra-prediction modes for blocks 524 and 526.
[0455] According to an embodiment, an apparatus for decoding a given block 18 is configured to derive from the data stream 12 a further MPM syntax element 538 indicating whether a given block-based intra prediction mode of a second set 520 of block-based intra prediction modes 510 is within a list 542 of most likely block-based intra prediction modes when a set selection syntax element 522 indicates that the given block 18 should not be predicted using one of a first set 508 of intra prediction modes. When the further MPM syntax element 538 indicates that a given block-based intra prediction mode of the second set 520 of block-based intra prediction modes 510 is within the list 542 of most likely block-based intra prediction modes, the apparatus is configured to form, for example, the list 542 of most likely block-based intra prediction modes based on the intra prediction mode by which neighboring blocks 524, 526 in the neighborhood of the given block 18 are predicted, and to derive from the data stream 12 a further MPM list index 540 indicating the given block-based intra prediction mode on the list 542 of most likely block-based intra prediction modes. When the further MPM syntax element 538 indicates that a given block-based intra prediction mode of the second set 520 of block-based intra prediction modes 510 is not within the list 542 of most likely block-based intra prediction modes, the apparatus is configured to derive from the data stream 12 a further list index 546 indicating the given block-based intra prediction mode within the second set 520 of block-based intra prediction modes. Thus, based on the further MPM syntax element 538, the data stream 12 includes either a further MPM list index 540 or a further list index 546 for the prediction of the given block 18.
[0456] A further MPM syntax element 538, a further MPM list index 540, and yet another list index 546 are represented in data stream 12 of FIG. 12 as corresponding to MPM syntax element 532, MPM list index 534, and yet another list index 536, but it is clear that either the further MPM syntax element 538 or an index associated with the further MPM syntax element 538, such as the further MPM list index 540 or yet another list index 546, is included in data stream 12 or in the MPM syntax element 532 and an index associated with the MPM syntax element 532, such as the MPM list index 534 or yet another list index 536. Which of this syntax element and index is included in data stream 12 depends, for example, on the set selection syntax element 522.
[0457] An example of the syntax element portion of data stream 12 written as pseudocode may be as shown in FIGS. 19a through 19d, and the reference symbols indicate which syntax element corresponds to the syntax element discussed previously.
[0458] The list construction of list 528 may be defined as follows. candIntraPredModeA / B indicates an intra prediction mode for block 524 such as A for block 524 and B for block 526, or indicates which mode in set 508 the intra prediction mode is mapped to when the corresponding block 524 or 526 is intra predicted using any of the block-based intra prediction modes 510. INTRA_DC is used to indicate mode 506, and the Angular mode 500 is indicated by INTRA_ANGULAR#, where the number (#) arranges the Angular modes as exemplified above, i.e., in such a way that the angular direction 502 monotonically decreases or increases with an increase in the number. This order among the modes within set 508 can be such as defined in a subsequent table where INTAR_PLANAR indicates mode 504.
[0459] Note that in the above example, index 534 is actually split into syntax elements 534' and 534''. The former 534' is unique to the first position of list 528 in order 530, and according to this example, the INTRA_PLANAR mode 504 is necessarily placed. The latter 534'' points to any of the subsequent positions of list 528, and as explained, only the DC mode 506 is included in this described special situation.
[0460] Furthermore, in the above example, if the syntax element 522 indicates that any of the modes within the set 508 are used, additional syntax elements are included in the data stream, and these additional syntax elements parameterize the intra prediction modes within the set 500 in some way. For example, the syntax element 600 parameterizes, or varies, the region in which the reference sample 17 is placed based on which mode within the set 508 intra-predicts the inside of the block 18, with respect to, for example, the distance to the outer perimeter of the block 18. Additionally or alternatively, the syntax element 602 parameterizes, or varies, whether the reference sample 17 is used for globally or en-block intra-predicting the inside of the block 18 by the modes within the set 508, or whether intra-prediction is performed in the fragments or portions into which the block 18 is subdivided, and those fragments or portions are sequentially intra-predicted such that the prediction residuals coded into the data stream for one portion can help gather new reference samples for intra-predicting the next portion. The latter coding option controlled by the syntax element may be available only if the syntax element 600 has a predetermined state corresponding to, for example, the region where the reference sample 17 exists, which is adjacent to the block 18 (and the corresponding syntax element may exist only in the data stream). The portions may be defined by subdividing the block along a predetermined direction, such as horizontally so that they are the same height as the block 18, or vertically so that they are the same width as the block 18. If the segmentation is signaled as being valid, the syntax element 604 may be present in the data stream, which controls which split direction is used.As can be understood, the positions in list 528 reserved for the INTRA_PLANAR mode are available only if the syntax element 600 has a predetermined state corresponding to a region where, for example, reference sample 17 touches block 18, and / or only if the per-part intra prediction mode signaled by syntax element 602 is not valid, etc., i.e., in the case of some parameterization of the mode by parameterization of the syntax elements mentioned just above.
[0461] All syntax elements shown in the table but not specifically mentioned above are optional and will not be further discussed herein. - If candIntraPredModeB is equal to candIntraPredModeA and candIntraPredModeA is greater than INTRA_DC, candModeList[x] for x = 0…4 is derived as follows. candModeList[0]=candIntraPredModeA candModeList[1]=2+((candIntraPredModeA+61)%64) candModeList[2]=2+((candIntraPredModeA-1)%64) candModeList[3]=2+((candIntraPredModeA+60)%64) candModeList[4]=2+(candIntraPredModeA%64) - Otherwise, if candIntraPredModeB is not equal to candIntraPredModeA and either candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC, the following applies. - Variables minAB and maxAB are derived as follows. minAB=Min(candIntraPredModeA, candIntraPredModeB) maxAB = Max(candIntraPredModeA, candIntraPredModeB) - When both candIntraPredModeA and candIntraPredModeB are greater than INTRA_DC, candModeList[x] where x = 0…4 is derived as follows. candModeList[0] = candIntraPredModeA candModeList[1] = candIntraPredModeB - When maxAB - minAB is equal to 1, the following applies. candModeList[2] = 2 + ((minAB + 61) % 64) candModeList[3] = 2 + ((maxAB - 1) % 64) candModeList[4] = 2 + ((minAB + 60) % 64) - Otherwise, when maxAB - minAB is 62 or more, the following applies. candModeList[2] = 2 + ((minAB - 1) % 64) candModeList[3] = 2 + ((maxAB + 61) % 64) candModeList[4] = 2 + (minAB % 64) - Otherwise, when maxAB - minAB is equal to 2, the following applies. candModeList[2] = 2 + ((minAB - 1) % 64) candModeList[3] = 2 + ((minAB + 61) % 64) candModeList[4] = 2 + ((maxAB - 1) % 64) - In other cases, the following applies. candModeList[2] = 2 + ((minAB + 61) % 64) candModeList[3] = 2 + ((minAB - 1) % 64) (8 - 36) candModeList[4] = 2 + (((maxAB + 61)) % 64) - In other cases (when candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC), candModeList[x] where x = 0…4 is derived as follows. candModeList[0]=maxAB candModeList[1]=2+((maxAB+61)%64) candModeList[2]=2+((maxAB-1)%64) (8-41) candModeList[3]=2+((maxAB+60)%64) candModeList[4]=2+(maxAB%64) - In other cases, the following applies. candModeList[0]=INTRA_DC candModeList[1]=INTRA_ANGULAR50 candModeList[2]=INTRA_ANGULAR18 candModeList[3]=INTRA_ANGULAR46 candModeList[4]=INTRA_ANGULAR54
[0462]
Table 4
[0463] 9. Embodiments that utilize block / matrix-based intra prediction modes together with other intra prediction modes and utilize secondary transformation The following description presents embodiments for combining the use of a secondary transform to code prediction residuals with block / matrix-based prediction using other intra prediction modes. The foregoing presentation of possibilities for matrix-based intra prediction (ALWIP) and combinations thereof with other intra prediction modes shall serve as examples for implementing the embodiments described hereinafter in this specification. In FIG. 12, for example, all details regarding MPM list construction, regarding the constraint that the DC mode is included in the MPM list, are optional.
[0464] As described above, matrix-based intra prediction (MIP), also referred to herein as block-based intra prediction and ALWIP, generates an intra prediction signal on a rectangular block by performing a matrix-vector multiplication, and the output of the matrix-vector multiplication may be regarded as the prediction signal on the downsampled block, and the input to the matrix-vector multiplication may be included in the downsampled boundary samples. If the output is regarded as the prediction signal on the downsampled block, this prediction signal needs to go through an upsampling (or linear interpolation) stage before the final prediction signal is obtained.
[0465] On one hand, in the conventional intra prediction modes such as Planar mode 504, DC mode 506, and Angular mode 500, also denoted as modes of set 508 in the above description, the non-separable second-order transform (LFNST) is a tool used to transform the prediction residuals corresponding to these intra prediction modes. Here, a set S of transform sets is given such that each conventional intra prediction mode is associated with one of these sets of transforms. Then, in the decoder, it can be extracted from the bitstream whether the LFNST should be applied on a given block. If it should be applied, depending on the intra prediction mode used on the current block, one transform set from set S is given, and if this transform set consists of more than one transform, which transform T from this set should be used can be extracted from the bitstream. Then, in the decoder, transform T is applied as the second-order transform Ts, which means that, for example, as shown in FIG. 14, it is applied to a subset 622 of the residual transform coefficients 620 of the separable first-order transform Tp.
[0466] The problem is that the aforementioned second-order transform Ts is defined a priori only for the conventional intra prediction modes. Providing a specific second-order transform Ts for each MIP mode 510 can be prohibitively expensive in terms of the memory requirements for storing additional transforms.
[0467] FIG. 13 shows a decoder that solves this problem. The decoder decodes a given block 18 of a picture using intra prediction. According to one embodiment, the encoder includes similar features and / or functions as the decoder.
[0468] The decoder / encoder is configured to select (602) a predetermined intra prediction mode 604 from a plurality of intra prediction modes 600 including a first set 508 of intra prediction modes and a second set 520 of matrix-based intra prediction modes 510. This intra mode selection 602 is performed by the decoder based on the data stream 12, and the encoder is configured to signal the predetermined intra prediction mode 604 in the data stream 12. The intra mode selection 602 may be performed as described with respect to FIG. 12.
[0469] The first set 508 of intra prediction modes includes a DC intra prediction mode 506, an Angular prediction mode 500, and optionally a Planar intra prediction mode 504. When the predetermined intra prediction mode 604 is a matrix-based intra prediction mode 510 in the second set 520, the decoder / encoder is configured to use a matrix-vector product 512 of a vector 514 derived from reference samples 17 in the vicinity of a predetermined block 18 and a prediction matrix 516 associated with each matrix-based intra prediction mode 510 to obtain a prediction vector 518, and based on the prediction vector 518, samples of the predetermined block 18 are predicted. The prediction of the predetermined block 18 using the matrix-based intra prediction mode 510 as the predetermined intra prediction mode 604 may be performed by the decoder / encoder according to the embodiments of FIGS. 6 to 11. The decoder / encoder is configured to derive a prediction signal 606 for the predetermined block 18 using the predetermined intra prediction mode 604.
[0470] The decoder / encoder depends on the predetermined intra prediction mode 604 such that the subset 610 is non-empty when the predetermined intra prediction mode 604 is included in the first set 508 of intra prediction modes and when the predetermined intra prediction mode 604 is included in the second set 520 of matrix-based intra prediction modes 510, a set 612 of secondary transforms Ts, e.g., Ts (1) -Ts (N) from one or more secondary transforms Ts(i1) -Ts (in) configured to select (608) a subset 610 of the secondary transforms Ts, where i1 is in the range from 1 to N, and i n is in the range from i1 to N.
[0471] According to an embodiment, the decoder / encoder is configured to select (608) a subset 610 such that each secondary transform Ts of the set 612 of secondary transforms Ts is included in one or more subsets 610 of secondary transforms Ts selected for at least one of the intra prediction modes in the first set 508 and the second set 520. Thus, the subset 610 may be equal to the set 612 of secondary transforms. Such a subset may be selected for one or more matrix-based intra prediction modes 510. In one or more matrix-based intra prediction modes 510, it may be possible to select all secondary transforms Ts of the set 612 of secondary transforms, and the subset 610 for this one or more matrix-based intra prediction modes 510 includes secondary transforms Ts that are selectable for the intra prediction modes in the first set 508 of intra prediction modes. Thus, for at least one of the matrix-based intra prediction modes 510, no specific additional secondary transforms are required in the set 612 of secondary transforms.
[0472] According to an embodiment, the decoder / encoder for each subset 610 of secondary transforms Ts selected for any matrix-based intra prediction mode 510 matrix each secondary transform Ts of which is a subset of secondary transforms Ts selected for at least one intra prediction mode in the first set 508 that does not belong to the Angular prediction mode 500, e.g., 610 DC and / or 610 planarconfigured to select (608) a subset 610 in a manner such as that included in. FIG. 15 shows various possible subsets of the secondary transform Ts selected for any matrix-based intra prediction mode 510. The decoder / encoder, for each matrix-based intra prediction mode 510, a first union 611 of the subset 610 of the secondary transform for the matrix-based intra prediction mode 510 matrix may be configured to select a subset 610 of.
[0473] As shown in FIG. 15, the subset 610 selectable for one or more matrix-based intra prediction modes 510 may be the subset selectable for the DC intra prediction mode 506, such as 610 DC1 or 610 DC2 or may be equal to the subset selectable for the Planar intra prediction mode 504, such as 610 planar1 or 610 planar2 or may be equal to. The subset 610 selectable for one or more matrix-based intra prediction modes 510 may be the subset 610 matrix4 As shown by, a subset of the secondary transform Ts selected for one intra prediction mode within a first set 508 that does not belong to the Angular prediction mode 500, such as 610 DC one or more secondary transforms of, such as Ts (ax) from Ts (aY) and may only include Ts.
[0474] The subset 610 selectable for the matrix-based intra prediction mode 510 may include one or more secondary transforms among two or more subsets selectable for the DC intra prediction mode 506, such as 610 matrix1 and 610 DC1 as shown by, or may include one or more secondary transforms among two or more subsets selectable for the Planar intra prediction mode 504, such as 610 DC2 and 610 matrix3 as shown by, planar1 and 610planar2 The transformation may include one or more secondary transformations in
[0475] Further possible subsets 610 that can be selected for the matrix-based intra-prediction mode 510 include the subsets 610 matrix2 One or more subsets selectable for DC intra-prediction mode 506, such as those shown by DC2 , and one or more subsets selectable for the Planar intra-prediction mode 504, e.g., 610 planar1 It may include one or more secondary transformations from
[0476] According to an embodiment, the decoder / encoder performs a first union 611 of the subset 610 of secondary transforms selected for the matrix-based intra-prediction mode 510. matrix and a subset of the secondary transforms selected for all Angular intra prediction modes 610 angular The Second Union 611 angular The third union 611 is configured to select (608) a subset 610 in such a way that the intersection of DC is all the subsets 610 selected for the DC intra prediction mode 506 DC Including the fourth coalition 611 planar is all the subsets 610 selected for the Planar intra prediction mode 504. planar For example, because MIP modes 510 are non-directional like planar mode 504 and DC mode 506 due to the downsampling optionally applied to the reduced prediction signal 606, their prediction residuals 618 have more statistical similarity to the prediction residuals of DC mode 506 and planar mode 504 than to the prediction residuals of angular mode 500.
[0477] Furthermore, as shown in FIG. 13, the decoder is configured to derive (614) from the data stream a transformed version 616 of the prediction residual for a given block 18 encoded by the encoder into the data stream, this transformed version being related to a spatial domain version 618 of the prediction residual of the given block 18 via a transformation T defined by the concatenation of a primary transformation Tp and a given secondary transformation Ts from a subset 610 of secondary transformations. As shown in FIG. 14, the encoder may be configured to apply the transformation T to a subset 622 of the coefficients 620 of the primary transformation Tp when a given intra prediction mode is included in a first set 508 of intra prediction modes and when a given intra prediction mode is included in a second set 520 of matrix-based intra prediction modes 510. The decoder may be configured to use the inverse function T -1 of the transformation T to obtain the spatial domain version 618 of the prediction residual of the given block 18. The primary transformation is, for example, a separable 2D transformation and the secondary transformation is, for example, a non-separable 2D transformation.
[0478] The decoder is configured to reconstruct (624) the given block 18 using the prediction signal 606 and the prediction residual 618 for the given block 18.
[0479] If a subset 610 of one or more secondary transformations Ts includes more than one secondary transformation Ts, the decoder may be configured to select a given secondary transformation Ts from the subset 610 of one or more secondary transformations according to a secondary transformation indication syntax element transmitted in the data stream 12 for the given block. The secondary transformation indication syntax element may be an index indicating into a selected subset 610 of one or more secondary transformations. In this case, the encoder may be configured to transmit the secondary transformation indication syntax element in the data stream 12.
[0480] According to one embodiment, when the dimensions of a given block 18 meet a given criterion, the decoder is configured to infer that the transformed version 616 of the prediction residual for the given block 18 and the transformation T associated with the spatial region version 618 of the prediction residual of the given block 18 is the primary transformation Tp. Otherwise, the transformation T can be a concatenation of the primary transformation Tp and a given secondary transformation Ts. Independently of whether the dimensions of the given block 18 meet the given criterion, the decoder makes available a second set 520 of matrix-based intra prediction modes 510 for the selection 602 of a given intra prediction mode 604. The given criterion for the dimensions of the given block 18 may be related only to the selection 608 of the subset and may not be related to the selection 602 of the intra mode. For example, there are block dimensions for which MIP / ALWIP 510 is available but LFNST, i.e., the selection 608 of the subset, is not available. For such dimensioned blocks 18, the secondary transformation indication syntax element need not be sent in the data stream and need not be read from the data stream. The given criterion is met, for example, when the dimensions are below a given threshold. For small given blocks 18, this may not be beneficial, meaning that the additional signaling cost for signaling whether LFNST should be applied to the block in the MIP mode for the latter shape is, on average, higher than the benefit obtained by allowing the LFNST transformation for the MIP mode.
[0481] According to an embodiment, as shown in FIG. 16, the decoder is configured to read a non-zero zone indication transmitted in the data stream 12 for each predetermined block 18, which indicates a non-zero transformed region area 623 within the transformed version 616 of the prediction residual for the predetermined block 18. All non-zero coefficients are located only in the non-zero transformed region area 623. The non-zero zone indication is transmitted, for example, in the data stream 12 by the encoder. The decoder / encoder is configured to decode / encode the coefficients within the non-zero transformed region area 623 from / to the data stream 12. The LP syntax element can operate as a non-zero zone indication. The last non-zero coefficient position along the scan path from the DC coefficient position to the coefficient position of the highest frequency, or the opposite side of the DC coefficient position, is indicated by the LP syntax element herein. The LP syntax element can be a measure of the expected count of non-zero coefficients within the non-zero transformed region area 623, quasi.
[0482] According to an embodiment, the decoder is configured to infer that the transform T, in which the transformed version 616 of the prediction residual for the predetermined block 18 is associated with the spatial region version 618 of the prediction residual of the predetermined block 18, is a primary transform Tp according to the expansion and / or position of the non-zero transformed region area 623 that meets the first predetermined criterion and / or the number of non-zero coefficients within the non-zero transformed region area 623 that meets the second predetermined criterion.
[0483] The first predetermined criterion is, for example, satisfied when the non-zero transform area 623 does not exclusively encompass the subset 622 of the coefficients of the primary transform Tp to which the secondary transform Tp is applied by concatenation. This is based on the idea that the secondary transform Ts should encompass all non-zero coefficients of the transform coefficients of the primary transform Tp. When non-zero coefficients are outside the subset 622 of the coefficients of the primary transform Tp, it may not be advantageous to apply a predetermined secondary transform, and for that reason, the decoder presumes that the transform T is the primary transform Tp. The decoder performs the subset selection 608, and when the non-zero transform area 623 is completely located inside the subset 622 of the coefficients of the primary transform Tp to which the secondary transform Tp is applied by the following concatenation, the decoder applies the transform T defined by the concatenation of the primary transform Tp, the subset 610 of the secondary transforms applied to the subset 622 of the coefficients of the primary transform Tp, and the predetermined secondary transform Ts, to obtain the spatial region version 618 of the prediction residual of the predetermined block 18. See, for example, FIG. 16.
[0484] The second predetermined criterion is, for example, satisfied when the number of non-zero coefficients within the non-zero transform area 623 is below a predetermined threshold. This is based on the idea that when the number of non-zero coefficients within the non-zero transform area 623 is below a predetermined threshold, there is no need to further reduce it by the secondary transform Ts. When the number of non-zero coefficients of the primary transform Tp is small, it may not be advantageous to additionally apply a predetermined secondary transform, and for that reason, the decoder presumes that the transform T is the primary transform Tp. When the number of non-zero coefficients is below a predetermined threshold, the additional signaling cost for a predetermined secondary transform is greater than the coding efficiency improvement achieved by the predetermined secondary transform. The decoder performs the subset selection 608, and when the number of non-zero coefficients within the non-zero transform area 623 is equal to or greater than a predetermined threshold, the decoder applies the transform T defined by the concatenation of the primary transform Tp, the subset 610 of the secondary transforms applied to the subset 622 of the coefficients of the primary transform Tp, and the predetermined secondary transform Ts, to obtain the spatial region version 618 of the prediction residual of the predetermined block 18.
[0485] According to an embodiment, as described with respect to FIG. 12 for example, the decoder / encoder is configured to derive from the data stream 12 / encode into the data stream 12 a set selection syntax element 522 indicating whether a given block 18 is to be predicted using one of a first set 508 of intra prediction modes. If the set selection syntax element 522 indicates that the given block 18 is to be predicted using one of the first set 508 of intra prediction modes, the decoder / encoder forms a list 528 of most likely intra prediction modes based on the neighboring blocks 524 and 526 in the neighborhood of the given block 18 and the intra prediction mode by which they are predicted, and is configured to derive from the data stream 12 / signal into the data stream 12 an MPM list index 534 that indicates a given intra prediction mode 604 on the list 528 of most likely intra prediction modes. If the set selection syntax element 522 indicates that the given block 18 is to be predicted using one of the first set 508 of intra prediction modes, the decoder / encoder is configured to derive from the data stream / encode into the data stream additional indexes 540 and / or 546 that indicate a given intra prediction mode 604 within a second set 520 of matrix-based intra prediction modes 510.
[0486] The decoder and encoder, described with respect to FIG. 13, may include additional features and / or functions as described with respect to FIG. 12.
[0487] If the given intra prediction mode 604 for a given block 18 is a matrix-based intra prediction mode 510 within the second set 520, the apparatus for decoding the given block 18, i.e., the decoder according to FIG. 13, and / or the apparatus for encoding the given block 18, i.e., the encoder according to FIG. 13, may include one or more of the following features.
[0488] According to one embodiment, the apparatus is configured to form a sample value vector, for example, the sample value vector 400 as described with respect to one of the embodiments of FIGS. 6 to 9, from among a plurality of reference samples 17, and to derive a vector 514 from the sample value vector such that the sample value vector is mapped by a predetermined regular linear transformation to the vector 514. In this case, the vector 514 can be understood as a further vector. The vector 514 is determined and / or defined, for example, as described for a further vector 402 with respect to one of the embodiments of FIGS. 8 to 11.
[0489] According to one embodiment, the apparatus is configured to form a sample value vector from a plurality of reference samples 17 by, for each component of the sample value vector, employing one of the plurality of reference samples as each respective component of the sample value vector, and / or by averaging two or more components of the sample value vector to obtain each respective component of the sample value vector.
[0490] The plurality of reference samples 17 are arranged, for example, in a picture along the outer edge of a predetermined block 18.
[0491] The regular linear transformation is defined such that, for example, a predetermined component of the vector 514, for example of a further vector, becomes a, and each of the other components of the vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a. The value a is, for example, a predetermined value 1400.
[0492] According to one embodiment, the predetermined value 1400 is an average, such as an arithmetic mean or a weighted mean, of the components of the sample value vector, a default value, a value signaled in a data stream in which the picture is coded, and one of the components of the sample value vector corresponding to the predetermined component.
[0493] The regular linear transformation is defined such that, for example, a predetermined component of vector 514, for example, a further vector, becomes a, and each of the other components of vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the arithmetic mean of the components of the sample value vector.
[0494] The regular linear transformation is defined such that, for example, a predetermined component of vector 514, for example, a further vector, becomes a, and each of the other components of vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is a component of the sample value vector corresponding to the predetermined component. The apparatus includes, for example, a plurality of regular linear transformations each associated with one component of vector 514, selects a predetermined component from the components of the sample value vector, and is configured to use a regular linear transformation from the plurality of regular linear transformations associated with the predetermined component as the predetermined regular linear transformation.
[0495] According to an embodiment, the matrix components of the prediction matrix 516 within the column corresponding to the predetermined component of vector 514, for example, a further vector, are all 0. The apparatus is configured to perform multiplication by calculating a matrix-vector product 512 of a reduced prediction matrix obtained from the prediction matrix 516 by excluding that column and a further vector 410 obtained from vector 514 by excluding the predetermined component, as shown in FIG. 9c.
[0496] According to an embodiment, when predicting samples of a predetermined block 18 based on the prediction vector 518, the apparatus is configured to calculate the sum of each component of the prediction vector 518 and a for each component.
[0497] Multiplying the matrix obtained by adding each matrix component of the prediction matrix 516 within the column of the prediction matrix 516 with a value of 1, corresponding to a predetermined component of the vector 514 (i.e., adding the matrix C405 shown in FIG. 9a and the matrix M1300, resulting in the matrix B in FIG. 8), by a regular linear transformation corresponds to, for example, a quantized version of the machine learning prediction matrix (i.e., the prediction matrix A1100 shown in FIG. 8).
[0498] According to an embodiment, the apparatus is configured to calculate the matrix-vector product 512 using fixed-point arithmetic operations.
[0499] According to an embodiment, the apparatus is configured to calculate the matrix-vector product 512 without using floating-point arithmetic operations.
[0500] According to an embodiment, the apparatus is configured to store a fixed-point number representation of the prediction matrix 516.
[0501] According to an embodiment, the apparatus is configured to represent the prediction matrix 516 using prediction parameters and calculate the matrix-vector product 512 by performing multiplications and additions on the components of the vector 514, such as components of a further vector, as well as the prediction parameters and intermediate results obtained therefrom. The absolute value of the prediction parameters can be represented by an n-bit fixed-point number representation, where n is 14 or less, or alternatively 10 or less, or alternatively 8 or less. This can be performed in a similar manner or as described in FIG. 10 or FIG. 11.
[0502] The prediction parameters include, for example, weights each associated with a corresponding matrix component of the prediction matrix 516.
[0503] The prediction parameters further include, for example, one or more scaling coefficients each associated with one or more corresponding matrix components of the prediction matrix 516 for scaling the weights associated with one or more corresponding matrix components of the prediction matrix 516, and / or one or more offsets each associated with one or more corresponding matrix components of the prediction matrix 516 for offsetting the weights associated with one or more corresponding matrix components of the prediction matrix 516.
[0504] According to an embodiment, when predicting samples of a given block 18 based on the prediction vector 518, the apparatus is configured to use interpolation to calculate at least one sample position of the given block 18 based on the prediction vector 518, where each component thereof is associated with a corresponding position within the given block 18.
[0505] The decoder and encoder described with respect to FIG. 13 may include additional features and / or functions as described with respect to one or more embodiments of FIGS. 6 to 11.
[0506] Therefore, the solution provided by the present invention is to associate each MIP mode 510 with a specific set 160 of one of the transforms in a given set of transforms S612, the latter originally being defined for the conventional intra prediction modes, i.e., the intra prediction modes within the first set 508. One specific way to do this is to use the LFNST transform, i.e., the secondary transform, originally designed for the Planar mode 504 and the DC mode 506 for all MIP modes 510. For example, due to the downsampling optionally applied to the downscaled prediction signal, the MIP modes 510 are non-directional like the Planar mode 504 and the DC mode 506, so their prediction residuals have a greater statistical similarity with the prediction residuals of the DC mode 506 and the Planar mode 504 than with the prediction residuals of the Angular mode 500.
[0507] Allowing LFNST for MIP510 in the foregoing manner, i.e., allowing the use of the secondary transformation, may be beneficial in terms of coding efficiency for some block shapes and may not be beneficial for other block shapes. It has been found that, for the latter shapes, the cost of additional signaling for signaling whether LFNST should be applied for block 18 in MIP mode 510 is, on average, greater than the benefit obtained by allowing the LFNST transformation for MIP mode 510. Accordingly, in some embodiments of the present invention, the foregoing combination of LFNST and MIP510 is allowed only in a subset of all block shapes for which such a combination is possible in principle.
[0508] FIG. 17 shows a method 6000 for decoding a predetermined block (18) of a picture using intra prediction, which includes selecting (602) a predetermined intra prediction mode (604) from a plurality of intra prediction modes (600) including a first set (508) of intra prediction modes including a DC intra prediction mode (506), an Angular prediction mode (500), and optionally a Planar intra prediction mode (504), and a second set (520) of matrix-based intra prediction modes (510). According to each of the matrix-based intra prediction modes, a matrix vector product (512) of a vector (514) derived from reference samples (17) in the neighborhood of the predetermined block and a prediction matrix (516) associated with the respective matrix-based intra prediction mode is used to obtain a prediction vector (518), and based on the prediction vector, samples of the predetermined block are predicted. A prediction signal (606) for the predetermined block is derived (6100) using the predetermined intra prediction mode, and one or more subsets (610) of a set of secondary transforms (612) are selected (608) in a manner dependent on the predetermined intra prediction mode such that the subset (610) is non-empty when the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and when the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510). The method 6000 includes deriving (614) a transformed version (616) of the prediction residual for a predetermined block (18) related to a spatial domain version (618) of the prediction residual of the predetermined block via a transform (T) defined by a concatenation of a primary transform (Tp) and a predetermined secondary transform (Ts) from a subset (610) of secondary transforms applied to a subset (622) of coefficients (620) of the primary transform from the data stream when the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and when the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510).Additionally, method 6000 includes a step (624) of reconstructing a given block using a prediction signal and a prediction residual for the given block (18).
[0509] FIG. 18 shows a method 7000 for encoding a predetermined block (18) of a picture using intra prediction, which includes a step (602) of selecting a predetermined intra prediction mode (604) from a plurality of intra prediction modes (600) including a first set (508) of intra prediction modes including a DC intra prediction mode (506) and an Angular intra prediction mode (500), and a second set (520) of matrix-based intra prediction modes (510). According to each of those matrix-based intra prediction modes, a matrix-vector product (512) of a vector (514) derived from reference samples (17) in the neighborhood of the predetermined block and a prediction matrix (516) associated with each matrix-based intra prediction mode is used to obtain a prediction vector (518), based on which samples of the predetermined block are predicted. In addition, the method 7000 includes a step (7100) of signaling the predetermined intra prediction mode (604) in a data stream, a step (7200) of deriving a prediction signal (606) for the predetermined block using the predetermined intra prediction mode, and a step (608) of selecting one or a subset (610) of a plurality of secondary transforms from a set of secondary transforms (612) in a manner dependent on the predetermined intra prediction mode such that the subset (610) is non-empty when the predetermined intra prediction mode is included in the first set (508) of intra prediction modes and the predetermined intra prediction mode is included in the second set (520) of matrix-based intra prediction modes (510).The method includes encoding into a data stream a transformed version (616) of a prediction residual for a given block (18) where the transformed version is related to a spatial domain version (618) of the prediction residual of the given block, via a transform (T) defined by concatenating a primary transform (Tp) and a given secondary transform (Ts) from a subset (610) of secondary transforms applied to a subset (622) of coefficients (620) of the primary transform, when the given intra prediction mode is included in a first set (508) of intra prediction modes and when the given intra prediction mode is included in a second set (520) of matrix-based intra prediction modes (510), and the given block is reconstructable (624) using a prediction signal and a prediction residual for the given block (18).
[0510] (References)
[0511] Further embodiments and examples In general, an example may be implemented as a computer program product with program instructions that, when the computer program product is executed on a computer, are operative to execute one of the methods. The program instructions may be stored, for example, on a machine-readable medium.
[0512] Another example includes a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0513] In other words, an example of a method is thus a computer program having program instructions for executing one of the methods described herein when the computer program is executed on a computer.
[0514] A further example of a method is thus a data carrier medium (or digital storage medium, or computer-readable medium) on which a computer program for performing one of the methods described herein is recorded. The data carrier medium, digital storage medium, or recorded medium is not an intangible and transient signal, but is tangible and / or non-transient.
[0515] A further example of a method is thus a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals can be transmitted, for example, via a data communication connection, for example via the Internet.
[0516] A further example includes processing means, for example a computer, or a programmable logic device for performing one of the methods described herein.
[0517] A further example includes a computer on which a computer program for performing one of the methods described herein is installed.
[0518] A further example includes an apparatus or system for transferring (for example, electrically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transferring the computer program to the receiver.
[0519] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods may be performed by any suitable hardware device.
[0520] The above examples are merely illustrative of the principles discussed above. It is understood that modifications and variations of the configurations and details described herein will be apparent. Accordingly, it is intended to be limited by the following claims rather than by the specific details presented for purposes of illustration and the descriptions of the examples herein.
[0521] One or more elements having equal or equivalent functions, even if they are present in different drawings, are denoted in the following description by equal or equivalent reference numerals.
Description of Reference Numerals
[0522] 10 Picture 12 Data Stream 14 Encoder 16 Video 17 Neighborhood Block, Neighborhood Sample 18 Block, Predetermined Block 19 ALWIP Conversion 20 Coding Order 22 Subtractor 24 Prediction Signal 26 Prediction Residual Signal 28 Prediction Residual Encoder 30 Quantizer 32 Conversion Stage 34 Quantized Prediction Residual Signal 36 Prediction Residual Reconstruction Stage 38 Inverse Quantizer 40 Inverse Transformer 42 Adder 44 Predictor 46 In-loop filter 54 Decoder 56 Entropy decoder 102 Reduced set 104 Sample value, sample 108 Sample 110 Group 112 Sample 118 Sample 119 Sample 120 Group 150 Predetermined component 156 Residual provider 400 Sample value vector 402 Further vector 403 Orthonormal linear transform 404 Matrix-vector product 405 Predetermined prediction matrix 406 Prediction vector 407 Matrix-vector product 408 Further offset 409 Vector 410 Yet further vector 412 Column 414 Weight, matrix component 500 Angular intra prediction mode 502 Angular direction 504 INTRA_PLANAR mode, Planar intra prediction mode 506 DC mode, DC intra prediction mode 508 First set 510 Matrix-based intra prediction mode 512 Matrix-vector product 514 Vector 516 Prediction matrix 518 Prediction vector 520 Second set, multiplication 522 Set selection syntax element 524 Neighboring block 526 Neighboring block 528 List 530 Order 532 MPM Syntax Element 534 MPM List Index 536 Further Syntax Element 538 Further MPM Syntax Element 540 Index, MPM List Index 542 List 544 List Order 546 Further Syntax Element 600 Syntax Element 602 Syntax Element 604 Predetermined Intra Prediction Mode 606 Prediction Signal 610 Subset 611 First Association, Second Association 612 Set 616 Converted Version 618 Spatial Region Version 620 Coefficient 622 Subset 623 Non-Zero Transformed Region Area 1100 Machine Learning Prediction Matrix 1110 Offset B 1200 Further Matrix B 1300 Integer Matrix, Matrix M(i0) 1310 Matrix-Vector Product 1400 Predetermined Value 1500 Predetermined Component
Claims
1. 1. A method for decoding a picture from a data stream, comprising: selecting, for a block of the picture, an intra-prediction mode from a first set of intra-prediction modes or a second set of intra-prediction modes based on an indication included in a data stream, the first set of intra-prediction modes comprising at least one angular, planar, and DC intra-prediction mode, and the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; deriving a prediction signal for the block using the selected intra-prediction mode; and selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the selected subset of secondary transforms comprising at least one of the plurality of LFNSTs; deriving a prediction residual for said block from said data stream; transforming the prediction residual using a LFNST from the selected subset; reconstructing the block using the prediction signal and the transformed prediction residual for the block; Equipped with A method according to claim 1, wherein the same subset of secondary transforms is selected for both a Planar intra prediction mode and the at least one matrix-based intra prediction mode.
2. Deriving the prediction signal when the intra-prediction mode is selected from the second set of intra-prediction modes comprises: deriving a prediction vector generated from a matrix-vector product of a vector derived from neighboring reference samples of the block and a prediction matrix associated with a selected intra-prediction mode; upsampling the prediction vector to obtain a prediction signal for the block; The method of claim 1 , comprising:
3. Selecting the subset of secondary transformations comprises: The selected intra prediction mode; and The size of the block; and The method of claim 1 , based on
4. decoding an indication of a LFNST from a selected subset of the secondary transforms; selecting, based on the indication specifying one of two LFNSTs, the LFNST from the subset for transforming the prediction residual; The method of claim 1 further comprising:
5. Transforming the prediction residuals includes: applying the LFNST (Ts) to a subset of the coefficients of a linear transform (Tp) to obtain a transform; transforming the prediction residual using the transform; The method of claim 1 , comprising:
6. The method of claim 5 , wherein the linear transformation (Tp) is a separable 2D transformation.
7. Selecting the intra prediction mode for the block of the picture comprises: decoding from the data stream a set of syntax elements indicating whether the block should be predicted using one of the first set of intra-prediction modes; In response to determining that the set of syntax elements indicates that the block is predicted using one of the first set of intra-prediction modes, generating a list of maximum likelihood intra-prediction modes (MPMs) based on intra-prediction modes used by blocks neighboring the block, and selecting the intra-prediction mode from the list; selecting one of the at least one matrix-based intra-prediction mode from the second set of intra-prediction modes in response to determining that the set of syntax elements indicates that the block is to be predicted using one of the second set of intra-prediction modes; and The method of claim 1 , comprising:
8. Transforming the prediction residuals includes:
2. The method of claim 1, comprising using a transformation defined by the concatenation of a primary transformation (Tp) and a secondary transformation (Ts) selected from a subset of secondary transformations, the secondary transformation (Ts) corresponding to the LFNST.
9. after selecting the intra-prediction mode from the second set of intra-prediction modes, selecting a matrix-based intra-prediction mode from the at least one matrix-based intra-prediction mode to derive the prediction signal. The method of claim 1.
10. deriving the prediction signal using the matrix-based intra-prediction mode, generating a prediction vector from a matrix-vector product of a vector derived from neighboring reference samples of the block and a prediction matrix associated with the selected matrix-based intra-prediction mode from which samples of the block are predicted.
10. The method of claim 9.
11. 1. An apparatus for decoding a picture from a data stream, comprising: selecting an intra-prediction mode from a first set of intra-prediction modes or a second set of intra-prediction modes based on an indication included in a data stream for a block of the picture, the first set of intra-prediction modes comprising at least one angular prediction mode, a planar intra-prediction mode, and a DC intra-prediction mode, and the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; deriving a prediction signal for the block using the selected intra-prediction mode; selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the selected subset of secondary transforms comprising at least one of the plurality of LFNSTs; Derive a prediction residual for the block from the data stream; Transforming the prediction residual using a LFNST from the selected subset; reconstructing the block using the prediction signal and the transformed prediction residual for the block; a processor configured to: The apparatus, wherein the same subset of secondary transforms is selected for both a Planar intra-prediction mode and the at least one matrix-based intra-prediction mode.
12. To derive the prediction signal when the intra-prediction mode is selected from the second set of intra-prediction modes, the processor: deriving a prediction vector generated from a matrix-vector product of a vector derived from neighboring reference samples of the block and a prediction matrix associated with a selected intra-prediction mode; upsampling the prediction vector to obtain a prediction signal for the block; The apparatus of claim 11 configured to:
13. The processor, The selected intra prediction mode; and The size of the block; and The apparatus of claim 11 , further configured to select the subset of secondary transformations based on:
14. The processor, Decoding an indication of a LFNST from the selected subset of the secondary transforms; selecting the LFNST from the subset for transforming the prediction residual based on the indication specifying one of two LFNSTs; The apparatus of claim 11 further configured to:
15. To transform the prediction residual, the processor applying the LFNST (Ts) to a subset of the coefficients of a linear transform (Tp) to obtain a transform; transforming the prediction residuals using the transform; The apparatus of claim 11 configured to:
16. The apparatus of claim 15, wherein the linear transform (Tp) is a separable 2D transform.
17. To select the intra prediction mode for the block of the picture, the processor: decoding from the data stream a set of syntax elements indicating whether the block should be predicted using one of the first set of intra-prediction modes; responsive to determining that the set of syntax elements indicates that the block is predicted using one of the first set of intra-prediction modes, generating a list of maximum likelihood intra-prediction modes (MPMs) based on intra-prediction modes used by blocks neighboring the block, and selecting the intra-prediction mode from the list; selecting one of the at least one matrix-based intra-prediction mode from the second set of intra-prediction modes in response to determining that the set of syntax elements indicates that the block is to be predicted using one of the second set of intra-prediction modes. The apparatus of claim 11 configured to:
18. To transform the prediction residual, the processor 12. The apparatus of claim 11, configured to use a transformation defined by the concatenation of a primary transformation (Tp) and a secondary transformation (Ts) selected from a subset of secondary transformations, the secondary transformation (Ts) corresponding to the LFNST.
19. After selecting the intra prediction mode from the second set of intra prediction modes, the processor is configured to select a matrix-based intra prediction mode from the at least one matrix-based intra prediction mode for deriving the prediction signal.
12. The apparatus of claim 11.
20. To derive the prediction signal using the matrix based intra prediction mode, the processor: and generating a prediction vector from a matrix-vector product of a vector derived from reference samples of a neighborhood of the block and a prediction matrix associated with the selected matrix-based intra-prediction mode for which samples of the block are predicted.
20. The apparatus of claim 19.
21. When executed by a processor of an electronic device, the electronic device selecting, for blocks of a picture, an intra-prediction mode from a first set of intra-prediction modes or a second set of intra-prediction modes based on an instruction included in a data stream, the first set of intra-prediction modes comprising at least one angular prediction mode, a planar intra-prediction mode, and a DC intra-prediction mode, and the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; deriving a prediction signal for the block using the selected intra-prediction mode; selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the selected subset of secondary transforms comprising at least one of the plurality of LFNSTs; deriving a prediction residual for said block from said data stream; transforming the prediction residual using an LFNST from the selected subset; reconstructing the block using the prediction signal and the transformed prediction residual for the block; With orders, 11. A non-transitory computer-readable medium, wherein the same subset of secondary transforms is selected for both a Planar intra-prediction mode and the at least one matrix-based intra-prediction mode.
22. The instructions, when executed, cause at least one of the processors to derive the prediction signal when the intra-prediction mode is selected from the second set of intra-prediction modes, the instructions, when executed, cause at least one of the processors to: deriving a prediction vector generated from a matrix-vector product of a vector derived from neighboring reference samples of the block and a prediction matrix associated with a selected intra-prediction mode; upsampling the prediction vector to obtain a prediction signal for the block; 22. The non-transitory computer-readable medium of claim 21 comprising instructions.
23. The instructions, which when executed cause at least one of the processors to select a subset of the secondary transformations, include: The selected intra prediction mode; and The size of the block; and The non-transitory computer-readable medium of claim 21 ,
24. When executed, the method causes at least one of the processors to: Decoding an indication of a LFNST from the selected subset of the secondary transforms; selecting the LFNST from the subset for transforming the prediction residual based on the indication specifying one of two LFNSTs; 22. The non-transitory computer readable medium of claim 21 further comprising instructions.
25. The instructions, when executed, cause at least one of the processors to transform the prediction residual, when executed, cause the at least one of the processors to: applying the LFNST (Ts) to a subset of the coefficients of a linear transform (Tp) to obtain a transform; transforming the prediction residuals using the transform; 22. The non-transitory computer-readable medium of claim 21 comprising instructions.
26. The linear transform (Tp) is a separable 2D transform.
26. The non-transitory computer readable medium of claim 25.
27. The instructions, which when executed cause at least one of the processors to select the intra prediction mode for the blocks of the picture, may also cause at least one of the processors, when executed, to: decode from the data stream a set of syntax elements indicating whether the block should be predicted using one of the first set of intra-prediction modes; in response to determining that the set of syntax elements indicates that the block is predicted using one of the first set of intra prediction modes, generating a list of maximum likelihood intra prediction modes (MPMs) based on intra prediction modes used in blocks neighboring the block, and selecting the intra prediction mode from the list; selecting one of the at least one matrix-based intra-prediction mode from the second set of intra-prediction modes in response to determining that the set of syntax elements indicates that the block is to be predicted using one of the second set of intra-prediction modes.
22. The non-transitory computer-readable medium of claim 21 comprising instructions.
28. The instructions, when executed, cause at least one of the processors to transform the prediction residual, when executed, cause the at least one of the processors to:
22. The non-transitory computer-readable medium of claim 21, comprising instructions to use a transformation defined by the concatenation of a primary transformation (Tp) and a secondary transformation (Ts) selected from a subset of secondary transformations, the secondary transformation (Ts) corresponding to the LFNST.
29. When executed, the method causes at least one of the processors to: and instructions for selecting a matrix-based intra-prediction mode from the at least one matrix-based intra-prediction mode for deriving the prediction signal after selecting the intra-prediction mode from the second set of intra-prediction modes.
22. The non-transitory computer-readable medium of claim 21.
30. The instructions, when executed, cause at least one of the processors to derive the prediction signal using the matrix-based intra-prediction mode, the instructions, when executed, cause at least one of the processors to: generating a prediction vector from a matrix-vector product of a vector derived from neighboring reference samples of the block and a prediction matrix associated with the selected matrix-based intra-prediction mode for which samples of the block are predicted; 30. The non-transitory computer readable medium of claim 29 comprising instructions.
31. A method for encoding a picture into a data stream, comprising the steps of: generating, for a block of the picture, an indication of an intra-prediction mode selected from a first set of intra-prediction modes or a second set of intra-prediction modes, the first set of intra-prediction modes including at least one angular, planar, and DC intra-prediction mode, and the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; encoding a prediction signal for the block using the selected intra-prediction mode; and selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the subset of selected secondary transforms comprising at least one of the plurality of LFNSTs; deriving a prediction residual for said block; transforming a prediction residual using a LFNST from the selected subset; encoding the transformed prediction residual for the block into the data stream; Equipped with the same subset of secondary transforms is selected for both a Planar intra prediction mode and at least one matrix-based intra prediction mode. method.
32. deriving a prediction vector generated from a matrix-vector product between a vector derived from neighboring reference samples of the block and a prediction matrix associated with the selected intra-prediction mode in response to the intra-prediction mode being selected from the second set of intra-prediction modes; generating the predicted signal based on the predicted vector; 32. The method of claim 31 further comprising:
33. Selecting the subset of secondary transformations comprises: The selected intra prediction mode; and The size of the block; and The method of claim 31 , which is based on 34. An apparatus for encoding a picture into a data stream, comprising: generating, for a block of the picture, an indication of an intra-prediction mode selected from a first set of intra-prediction modes or a second set of intra-prediction modes, the first set of intra-prediction modes comprising at least one angular prediction mode, a planar intra-prediction mode, and a DC intra-prediction mode, the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; encoding a prediction signal for the block using the selected intra-prediction mode; selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the subset of selected secondary transforms comprising at least one of the plurality of LFNSTs; Deriving a prediction residual for the block; Transforming the prediction residual using a LFNST from the selected subset; encoding the transformed prediction residual for the block into the data stream; a processor configured to: The apparatus, wherein the same subset of secondary transforms is selected for both a Planar intra-prediction mode and at least one matrix-based intra-prediction mode.
35. In response to the intra prediction mode being selected from the second set of intra prediction modes, the processor: deriving a prediction vector generated from a matrix-vector product between a vector derived from neighboring reference samples of the block and a prediction matrix associated with the selected intra-prediction mode; generating the predicted signal based on the predicted vector; 35. The apparatus of claim 34 configured to:
36. The processor, The selected intra prediction mode; and The size of the block; and 35. The apparatus of claim 34 configured to select the subset of secondary transformations based on:
37. When executed by a processor of an electronic device, the electronic device generating, for a block of a picture, an indication of an intra-prediction mode selected from a first set of intra-prediction modes or a second set of intra-prediction modes, the first set of intra-prediction modes including at least one angular prediction mode, a planar intra-prediction mode, and a DC intra-prediction mode, and the second set of intra-prediction modes comprising at least one matrix-based intra-prediction mode; encoding a prediction signal for the block using the selected intra-prediction mode; selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, the subset of selected secondary transforms comprising at least one of the plurality of LFNSTs; deriving a prediction residual for said block; Transforming the prediction residual using an LFNST from the selected subset; encoding the transformed prediction residual for the block into a data stream; With orders, 11. A non-transitory computer-readable medium, wherein the same subset of secondary transforms is selected for both a Planar intra-prediction mode and at least one matrix-based intra-prediction mode.
38. In response to the intra-prediction mode being selected from the second set of intra-prediction modes, the instructions, when executed, cause at least one of the processors to: deriving a prediction vector generated from a matrix-vector product between a vector derived from neighboring reference samples of the block and a prediction matrix associated with the selected intra-prediction mode; generating the predicted signal based on the predicted vector; 38. The non-transitory computer readable medium of claim 37.
Citation Information
Patent Citations
Encoder, decoder, and corresponding method for reconciling matrix-based intra prediction and quadratic transform core selection
JP2022529030A
Image encoding / decoding method and device for performing MIP and LFNST, and bitstream transmission method
JP2022532114A
An encoder, a decoder and corresponding methods harmonzting matrix-based intra prediction and secoundary transform core selection
WO2020211765A1
Image encoding / decoding method and device for performing MIP and lfnst, and method for transmitting bitstream
WO2020226424A1