Coding using matrix-based intra prediction and secondary conversion

By selecting a subset of secondary transforms for matrix-based intra prediction, the inefficiencies in memory and bitstream signaling are addressed, improving coding efficiency in video coding systems.

JP2025124881APending Publication Date: 2025-08-26FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025094800
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-25
Filing Date
2025-06-06
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Conventional intra prediction modes in video coding require excessive memory and bitstream signaling for secondary transforms, making them inefficient for matrix-based intra prediction.

Method used

Select a subset of secondary transforms from a set that includes both matrix-based and non-matrix-based intra-prediction modes, reducing memory requirements and optimizing coding efficiency by selectively applying these transforms based on the prediction mode.

Benefits of technology

This approach reduces memory usage and bitstream signaling costs while enhancing coding efficiency for matrix-based intra prediction by strategically selecting secondary transforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025124881000001_ABST
    Figure 2025124881000001_ABST
Patent Text Reader

Abstract

To provide a conception for supporting a secondary conversion for a matrix-based intra prediction.SOLUTION: A decoder selects a predetermined intra prediction mode 604 from a plurality of intra prediction modes 600 comprising a first set 508 of each intra prediction mode and a second set 520 of a matrix-based intra prediction mode 510 based on a data stream 12, derives a prediction signal 606 for a predetermined block 18 using a predetermined intra prediction mode, and selects a subset 610 of one or more of the secondary conversions of a set 612 of the secondary conversion in a manner dependent on the predetermined intra prediction mode such that the subset is non-empty if the predetermined intra prediction mode is included in the first set of each intra prediction mode and if the predetermined intra prediction mode is included in the second set of the matrix-based intra prediction mode.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to the field of matrix-based intra prediction and quadratic transforms. [Background technology]

[0002] In conventional intra prediction modes such as Planar mode, DC mode, and Angular mode, a non-separable secondary transform (LFNST) is a tool used to transform prediction residuals corresponding to these intra prediction modes. Here, a set S of multiple transform sets is provided such that each conventional intra prediction mode is associated with one of these transform sets. Then, in a decoder, whether LFNST should be applied to a given block can be extracted from the bitstream. If it should be applied, one transform set from set S is provided according to the intra prediction mode used on the current block. If this transform set consists of more than one transform, which transform T in this set should be used can be extracted from the bitstream. Then, in a decoder, transform T is applied as a secondary transform, which means that it is applied to a subset of residual transform coefficients of the separable primary transform. However, the aforementioned secondary transform is defined a priori only for conventional intra prediction modes. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] P. Helle et al., “Non-linear weighted intra prediction”, JVET-L0199, Macao, China, October 2018 [Non-patent document 2] F.Bossen, J.Boyce, K.Suehring, X.Li, V.Seregin, "JVET common test conditions and software reference configurations for SDR video", JVET-K1010, Ljubljana, Sl, July 2018. Summary of the Invention [Problem to be solved by the invention]

[0004] It is therefore desirable to provide concepts for rendering picture coding and / or video coding more efficient to support secondary transforms for matrix-based intra prediction (MIP), i.e., block-based intra prediction.

[0005] This is achieved by the subject matter of the independent claims of the present application.

[0006] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application. [Means for solving the problem]

[0007] According to a first aspect of the present invention, the inventors of the present application have recognized that one problem encountered when attempting to associate a secondary transform with a matrix-based intra-prediction mode stems from the fact that providing a specific secondary transform for each MIP mode may be too expensive in terms of the memory requirements for additionally storing the additional transforms. According to a first aspect of the present application, this difficulty is overcome by selecting a subset of one or more secondary transforms from a set of secondary transforms including transforms associated with matrix-based intra-prediction modes and transforms associated with non-matrix-based intra-prediction modes. The secondary transforms in the set of secondary transforms may be defined for one or more prediction modes, which reduces the memory capacity required for the set of secondary transforms. The transforms defined for planar intra-prediction modes and / or the transforms defined for DC intra-prediction modes may also be selectable for matrix-based intra-prediction modes. While this specific selection of a subset of secondary transforms for a matrix-based intra-prediction mode may increase coding efficiency, the additional syntax elements required for blocks associated with matrix-based intra-prediction modes to indicate the use of a secondary transform may increase bitstream and therefore signaling costs.

[0008] According to a first aspect of the present application, an apparatus for decoding a predetermined block of a picture using intra prediction, i.e., a decoder, is configured to select, based on a data stream, a predetermined intra prediction mode from a plurality of intra prediction modes, including a first set of intra prediction modes and a second set of matrix-based intra prediction modes. The first set of intra prediction modes includes a DC intra prediction mode, an angular prediction mode, and optionally a planar intra prediction mode. When a matrix-based intra prediction mode from the second set is selected as the predetermined intra prediction mode, the decoder is configured to obtain a prediction vector using a matrix-vector product of a vector derived from a reference sample in a neighborhood of the predetermined block and a prediction matrix associated with each matrix-based intra prediction mode, and to predict samples of the predetermined block based on the prediction vector. The decoder is configured to derive a prediction signal for a given block using a given intra-prediction mode and to select a subset of one or more secondary transforms from the set of secondary transforms in a manner dependent on the given intra-prediction mode such that the subset is non-empty when the given intra-prediction mode is included in the first set of intra-prediction modes and when the given intra-prediction mode is included in the second set of matrix-based intra-prediction modes. The first and second sets define the intra-prediction modes for which the secondary transforms are available. Thus, for a given intra-prediction mode selected from the first or second set, the decoder is configured to select a subset of one or more secondary transforms from the set of secondary transforms specifically associated with the selected given intra-prediction mode.Additionally, the decoder is configured to derive, from the data stream, a transformed version of the prediction residual for a given block associated with a spatial domain version of the prediction residual for the given block via a transform defined by concatenation of a primary transform and a predetermined secondary transform from a subset of secondary transforms applied to a subset of coefficients of the primary transform when the given intra prediction mode is included in the first set of intra prediction modes and when the given intra prediction mode is included in the second set of matrix-based intra prediction modes. The primary transform is set, for example, by default, and the predetermined secondary transform is selected, for example, by the decoder, from the subset of secondary transforms. The decoder may be configured to select the predetermined secondary transform from the subset of secondary transforms by deriving a syntax element indicating the secondary transform from the data stream. The decoder is configured to reconstruct the given block using the prediction signal and the prediction residual for the given block.

[0009] According to a first aspect of the present application, an apparatus for encoding a predetermined block of a picture using intra prediction, i.e., an encoder, similar to a decoder, is configured to select the predetermined intra prediction mode from a plurality of intra prediction modes, including a first set of intra prediction modes including a DC intra prediction mode, an angular prediction mode, and optionally a planar intra prediction mode, and a second set of matrix-based intra prediction modes, wherein, according to each matrix-based intra prediction mode, a matrix-vector product of a vector derived from a reference sample in a neighborhood of the predetermined block and a prediction matrix associated with the respective matrix-based intra prediction mode is used to obtain a prediction vector, on the basis of which samples of the predetermined block are predicted. The encoder is configured to signal the predetermined intra prediction mode in a data stream and to derive a prediction signal for the predetermined block using the predetermined intra prediction mode. Additionally, the encoder is configured to select a subset of one or more secondary transforms from the set of secondary transforms in a manner dependent on a predetermined intra prediction mode, such that the subset is non-empty when the predetermined intra prediction mode is included in the first set of intra prediction modes and when the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes. When the predetermined intra prediction mode is included in the first set of intra prediction modes and when the predetermined intra prediction mode is included in the second set of matrix-based intra prediction modes, the encoder is configured to encode a transformed version of the prediction residual for a predetermined block, associated with a spatial domain version of the prediction residual for the predetermined block, into a data stream via a transform defined by a concatenation of the primary transform and a predetermined secondary transform from the subset of secondary transforms applied to a subset of coefficients of the primary transform. The predetermined block is reconstructable using the prediction signal and the prediction residual for the predetermined block.

[0010] According to an embodiment, the decoder / encoder is configured to select a subset of one or more secondary transforms from the set of secondary transforms in a manner dependent on a given intra-prediction mode, such that each secondary transform of the set of secondary transforms is included in the subset of one or more secondary transforms selected for at least one of the intra-prediction modes in the first set and the second set, and for one or more intra-prediction modes in the first set or the second set, each selected subset may include all secondary transforms of the set of secondary transforms.

[0011] According to an embodiment, the decoder / encoder is configured to select a subset of one or more secondary transforms from the set of secondary transforms in a manner dependent on a given intra prediction mode, such that each secondary transform of each subset of secondary transforms selected for any matrix-based intra prediction mode is included in the subset of secondary transforms selected for at least one intra prediction mode in the first set that does not belong to an angular prediction mode. The subset selected for a matrix-based intra prediction mode may include secondary transforms associated with one or more non-angular intra prediction modes in the first set, such as DC intra prediction mode and / or planar intra prediction mode. Thus, those secondary transforms from the set of secondary transforms may be part of more than one subset of one or more secondary transforms. Additional secondary transforms that are usable only for blocks for which the given intra prediction mode is a matrix-based intra prediction mode may not be included in the set of secondary transforms. For a given block for which the given intra prediction mode is a matrix-based intra prediction mode, the decoder / encoder is configured to select the same secondary transform from the set of secondary transforms for the given intra prediction mode that is a non-angular prediction mode from the first set. The subset selected for the matrix-based intra prediction mode may be equal to the subset selected for the intra prediction modes in the first set that do not belong to the angular prediction mode, or may include some secondary transforms of the subset selected for the intra prediction modes in the first set that do not belong to the angular prediction mode, or may include some or all secondary transforms of two or more subsets selected for the intra prediction modes in the first set that do not belong to the angular prediction mode.

[0012] The method for encoding or decoding is based on the same considerations as the above-mentioned apparatus for encoding or decoding, however, the method can be completed using all the features and functions also described with respect to the apparatus for encoding or decoding.

[0013] The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which: [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 illustrates an embodiment of encoding into a data stream. [Figure 2] FIG. 1 illustrates an embodiment of an encoder. [Figure 3] FIG. 1 illustrates an embodiment of picture reconstruction. [Figure 4] FIG. 1 illustrates an embodiment of a decoder. [Figure 5] 1 is a schematic diagram of prediction of a block for encoding and / or decoding according to an embodiment; [Figure 6] FIG. 1 illustrates a matrix operation for prediction of a block for encoding and / or decoding according to an embodiment. [Figure 7.1] FIG. 1 illustrates prediction of a block using a reduced sample value vector according to an embodiment. [Figure 7.2] FIG. 1 illustrates block prediction using sample interpolation according to an embodiment. [Figure 7.3] FIG. 10 illustrates prediction of a block using a reduced sample value vector in which only some boundary samples are averaged, according to an embodiment. [Figure 7.4] FIG. 10 illustrates prediction of a block using a reduced sample value vector in which groups of four boundary samples are averaged, according to an embodiment. [Figure 8] FIG. 1 illustrates a matrix operation performed by an apparatus according to an embodiment. [Figure 9a] FIG. 2 illustrates detailed matrix operations performed by the device, according to an embodiment. [Figure 9b] FIG. 2 illustrates detailed matrix operations performed by the device, according to an embodiment. [Figure 9c]FIG. 2 illustrates detailed matrix operations performed by the device, according to an embodiment. [Figure 10] FIG. 10 illustrates detailed matrix operations performed by the device using offset and scaling parameters, according to an embodiment. [Figure 11] 5A-5C illustrate detailed matrix operations performed by the device using offset and scaling parameters according to different embodiments. [Figure 12] FIG. 1 is a schematic diagram with details of intra-prediction of a given block using a prediction mode from a maximum likelihood mode list, according to an embodiment. [Figure 13] FIG. 2 is a schematic diagram of decoding a given block using a quadratic transform according to an embodiment. [Figure 14] FIG. 1 is a schematic diagram of the application of a linear and a secondary transformation, according to an embodiment. [Figure 15] FIG. 10 is a schematic diagram of the selection of a subset of quadratic transformations according to an embodiment. [Figure 16] FIG. 2 is a schematic diagram of a predetermined block with non-zero transform domain areas, according to an embodiment. [Figure 17] FIG. 2 is a block diagram of a method for decoding a predetermined block, according to an embodiment. [Figure 18] 1 is a block diagram of a method for encoding a predetermined block according to an embodiment. [Figure 19a] FIG. 10 is a diagram illustrating syntax element portions of a data stream. [Figure 19b] FIG. 10 is a diagram illustrating syntax element portions of a data stream. [Figure 19c] FIG. 10 is a diagram illustrating syntax element portions of a data stream. [Figure 19d] FIG. 10 is a diagram illustrating syntax element portions of a data stream. DETAILED DESCRIPTION OF THE INVENTION

[0015] Equal or equivalent elements, or elements with equal or equivalent functions, are designated in the following description by equal or equivalent reference numerals, even if they are present in different figures.

[0016] In the following description, numerous details are set forth to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, to avoid obscuring embodiments of the present invention. In addition, features of different embodiments described hereinafter may be combined with each other unless otherwise specifically stated.

[0017] 1 Introduction In the following, different inventive examples, embodiments, and aspects are described, at least some of which refer to methods and / or apparatuses for, inter alia, video coding and / or for performing intra prediction using, for example, linear or affine transforms with neighborhood reduction, and / or for optimizing video distribution (e.g., broadcast, streaming, file playback, etc.), for example, video applications and / or virtual reality applications.

[0018] Additionally, examples, embodiments, and aspects may refer to High Efficiency Video Coding (HEVC) or a successor, or further embodiments, examples, and aspects are defined by the accompanying claims.

[0019] It should be noted that any embodiment, example, and aspect as defined by the claims may be supplemented by any of the details (features and functions) described in the following sections.

[0020] Also, the embodiments, examples, and aspects described in the following sections may be used individually or may be supplemented by any of the features in another section or by any of the features contained in the claims.

[0021] It should also be noted that the individual examples, embodiments, and aspects described herein may be used individually or in combination, and thus details may be added to each of the individual aspects without adding details to another one of the examples, embodiments, and aspects.

[0022] It should also be noted that this disclosure explicitly or implicitly describes features of encoding and / or decoding systems and / or methods.

[0023] Moreover, features and functions disclosed herein with respect to a method may also be used in an apparatus. Furthermore, any feature and function disclosed herein with respect to an apparatus may also be used in the corresponding method. In other words, the method disclosed herein may be supplemented with any of the features and functions described with respect to an apparatus.

[0024] Also, any of the features and functions described herein may be implemented in hardware or software, or using a combination of hardware and software, as described in the "Alternative Implementations" section.

[0025] Moreover, any feature listed in parentheses (“(...)” or "[...]") may be considered optional in some examples, embodiments, or aspects.

[0026] 2 Encoder and decoder In the following, various examples are described that can help achieve more effective compression when using block-based prediction. Some examples achieve high compression efficiency by using a set of intra-prediction modes. The latter may be in addition to, for example, or provided exclusively with, other heuristically designed intra-prediction modes. And other examples also take advantage of both of the properties discussed immediately above. As a variation of these embodiments, however, intra-prediction can be turned into inter-prediction by instead using reference samples in another picture.

[0027] To facilitate understanding of the following examples of this application, the description begins with a presentation of possible suitable encoders and decoders, onto which the examples outlined later in this application may build. FIG. 1 shows an apparatus for block-by-block encoding of a picture 10 into a data stream 12. The apparatus is indicated using reference symbol 14 and may be a still picture encoder or a video encoder. In other words, picture 10 may be the current picture from video 16 when encoder 14 is configured to encode a video 16 containing picture 10 into data stream 12, or when encoder 14 may encode picture 10 exclusively into data stream 12.

[0028] As mentioned, the encoder 14 performs the encoding in a block-by-block manner, or on a block basis. To this end, the encoder 14 subdivides the picture 10 into blocks, and on a block-by-block basis, the encoder 14 encodes the picture 10 into the data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are described in more detail below. In general, the subdivision may result in blocks 18 of a fixed size, such as an array of blocks arranged in rows and columns, or blocks 18 of different block sizes, such as by using hierarchical multi-tree subdivision, which involves starting from the entire picture area of ​​the picture 10 or from a pre-partitioning of the picture 10 into an array of tree blocks; these examples should not be treated as exclusive of other possible ways of subdividing the picture 10 into blocks 18.

[0029] Furthermore, encoder 14 is a predictive encoder that is configured to predictively encode picture 10 into data stream 12. For a given block 18, this means that encoder 14 determines a prediction for block 18 and encodes into data stream 12 the prediction residual, i.e., the prediction error by which the prediction deviates from the actual picture content within block 18.

[0030] The encoder 14 may support different prediction modes for deriving a prediction signal for a block 18. The prediction mode of interest in the following example is intra-prediction mode, according to which the inside of the block 18 is spatially predicted from already-encoded samples of neighboring blocks of the picture 10. The encoding of the picture 10 into the data stream 12, and therefore the corresponding decoding procedure, may be based on a coding order 20 defined among the blocks 18. For example, the coding order 20 may scan the blocks 18 in a raster scan order, such as from top to bottom row by row, while scanning each row from left to right. In the case of hierarchical multi-tree-based subdivision, a raster scan order may be applied within each hierarchical level, where a depth-first scan order may be applied, i.e., a leaf node in a block at a hierarchical level may precede a block at the same hierarchical level that has the same parent block according to the coding order 20. Depending on the coding order 20, already-encoded samples of neighboring blocks of the block 18 may typically be located on one or more sides of the block 18. In the example presented here, for example, the already coded samples in the neighborhood of block 18 are located above and to the left of block 18 .

[0031] Intra-prediction modes may not be the only modes supported by encoder 14. If encoder 14 is a video encoder, for example, encoder 14 may also support inter-prediction modes according to which blocks 18 are temporally predicted from previously encoded pictures of video 16. Such inter-prediction modes may be motion-compensated prediction modes according to which motion vectors are signaled for such blocks 18, indicating the relative spatial offset of the portions from which the prediction signal of block 18 should be derived as a copy. Additionally or alternatively, other non-intra-prediction modes may also be available, such as inter-prediction modes if encoder 14 is a multiview encoder, or non-prediction modes according to which the inside of block 18 is coded as is, i.e., without any prediction.

[0032] Before beginning to focus the description of this application on intra-prediction modes, a more specific example of a possible implementation of a possible block-based encoder, i.e., encoder 14, will be described with reference to FIG. 2, followed by two corresponding examples of decoders compatible with FIGS. 1 and 2, respectively.

[0033] 2 shows a possible implementation of the encoder 14 of FIG. 1, i.e., an implementation in which the encoder is configured to use transform coding for encoding the prediction residual, but this is mostly by way of example and the application is not limited to that type of prediction residual coding. According to FIG. 2, the encoder 14 includes a subtractor 22 configured to subtract a current block 18, a corresponding prediction signal 24, from an incoming signal, i.e., picture 10, or block by block, to obtain a prediction residual signal 26, which is then encoded into a data stream 12 by a prediction residual encoder 28. The prediction residual encoder 28 consists of a lossy coding stage 28a and a lossless coding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes samples of the prediction residual signal 26. As already mentioned above, this example uses transform coding of the prediction residual signal 26, according to which the lossy encoding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such spectrally decomposed prediction residual 26, with the quantizer 30 quantizing the transformed coefficients while presenting the residual signal 26. The transform may be a DCT, DST, FFT, Hadamard transform, etc. The transformed and quantized prediction residual signal 34 is then subjected to lossless coding by the lossless encoding stage 28b, which is an entropy coder that entropy codes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner that is also usable by the decoder, i.e., taking into account coding losses. To this end, the prediction residual reconstruction stage 36 includes an inverse quantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs an inverse transformation to that performed by the transformer 32, such as the inverse of the spectral decomposition, such as the inverse of any of the specific example transformations mentioned above. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal as output by the inverse transformer 40 and the prediction signal 24 to output a reconstructed signal, i.e., reconstructed samples.This output is provided to a predictor 44 of the encoder 14, which then determines a prediction signal 24 based thereon. It is the predictor 44 that supports all of the prediction modes already discussed above with respect to Figure 1. Figure 2 also shows that, if the encoder 14 is a video encoder, the encoder 14 also includes an in-loop filter that filters fully reconstructed pictures, which, after being filtered, form reference pictures for the predictor 44 for inter-predicted blocks.

[0034] As described above, the encoder 14 operates on a block-by-block basis. In the following description, the basis for a block of interest is a subdivision of the picture 10 into blocks for which an intra-prediction mode is selected from a set of intra-prediction modes or multiple intra-prediction modes supported by the predictor 44 or the encoder 14, respectively, and the selected intra-prediction mode is executed individually. However, there may be other types of blocks into which the picture 10 is subdivided. For example, the above-mentioned decision of whether the picture 10 is inter-coded or intra-coded may be made at a granularity or block unit that deviates from the block 18. For example, the inter / intra mode decision may be performed at the coding block level into which the picture 10 is subdivided, and each coding block is subdivided into prediction blocks. Prediction blocks with coding blocks for which it is determined that intra-prediction is used are each subdivided into a determination of the intra-prediction mode. For this purpose, for each of these prediction blocks, a decision is made as to which supported intra-prediction mode should be used for the respective prediction block. These prediction blocks form the block 18 of interest here. Predictive blocks within a coding block associated with inter-prediction are treated differently by the predictor 44. They are inter-predicted from a reference picture by determining a motion vector and copying the prediction signal for this block from the location in the reference picture pointed to by the motion vector. Another block subdivision involves subdivision into transform blocks, on the basis of which the transformer 32 and the inverse transformer 40 perform transformation. The transformed blocks may be, for example, the result of further subdivision of coding blocks. Of course, the examples described herein should not be treated as limiting, and other examples exist. For the sake of completeness only, it should be noted that the subdivision into coding blocks may use, for example, multi-tree subdivision, and that prediction blocks and / or transform blocks may also be obtained by further subdividing the coding block using multi-tree subdivision.

[0035] A decoder 54 or device for block-by-block decoding compatible with the encoder 14 of FIG. 1 is illustrated in FIG. 3. This decoder 54 does the opposite of the encoder 14, i.e., it decodes the picture 10 from the data stream 12 block by block and, to this end, supports multiple intra-prediction modes. The decoder 54 may, for example, include a residual provider 156. All other possibilities discussed above with respect to FIG. 1 are also valid for the decoder 54. To this end, the decoder 54 may be a still picture decoder or a video decoder, with all prediction modes and prediction possibilities supported by the decoder 54. The difference between the encoder 14 and the decoder 54 lies primarily in the fact that the encoder 14 chooses or selects coding decisions according to some optimization, such as to minimize some cost function that may depend on the coding rate and / or coding distortion. One of these coding options or coding parameters may involve selecting the intra-prediction mode to be used for the current block 18 among the available or supported intra-prediction modes. The selected intra-prediction mode may then be signaled by encoder 14 for the current block 18 in data stream 12, and decoder 54 uses this signaling in data stream 12 for block 18 to make the selection again. Similarly, the subdivision of picture 10 into blocks 18 may undergo optimization in encoder 14, and corresponding subdivision information may be conveyed in data stream 12, and decoder 54 reconstructs the subdivision of picture 10 into blocks 18 based on the subdivision information. In summary of the above, decoder 54 may be a predictive decoder that operates on a block-by-block basis, and except for intra-prediction mode, decoder 54 may support other prediction modes, such as, for example, inter-prediction mode if decoder 54 is a video decoder.1, and since this coding order 20 is followed in both the encoder 14 and the decoder 54, the same neighboring samples are available for the current block 18 in both the encoder 14 and the decoder 54. Therefore, to avoid unnecessary repetition, the description of the modes of operation of the encoder 14 shall also apply to the decoder 54 as far as the subdivision of the picture 10 into blocks is concerned, e.g., as far as prediction is concerned and as far as coding of prediction residuals is concerned. The difference lies in the fact that the encoder 14, by optimization, chooses some coding options or coding parameters and signals or inserts the coding parameters in or into the data stream 12, which are then derived from the data stream 12 by the decoder 54 in order to perform the prediction, subdivision, etc. again.

[0036] FIG. 4 illustrates a possible implementation of the decoder 54 of FIG. 3, i.e., one that matches the implementation of the encoder 14 of FIG. 1 as shown in FIG. 2. Because many elements of the encoder 54 of FIG. 4 are the same as those present in the corresponding encoder of FIG. 2, the same reference symbols, given with an apostrophe, are used in FIG. 4 to indicate these elements. Specifically, the adder 42′, optional in-loop filter 46′, and predictor 44′ are connected to the prediction loop in the same manner as in the encoder of FIG. 2. The reconstructed, i.e., dequantized, retransformed prediction residual signal applied to the adder 42′ is derived by a sequence of entropy decoders 56, which, as on the encoding side, perform the inverse of the entropy coding of the entropy encoder 28b, followed by a residual signal reconstruction stage 36′ consisting of an inverse quantizer 38′ and an inverse transformer 40′. The output of the decoder is a reconstruction of the picture 10. The reconstruction of picture 10 may be available at the output of adder 42', or alternatively directly at the output of in-loop filter 46'. Some post-filter may be lined up at the output of the decoder to apply some post-filtering to the reconstruction of picture 10 to improve picture quality, although this option is not shown in FIG.

[0037] Referring again to Figure 4, the explanations presented above with respect to Figure 2 shall also be valid for Figure 4, except that only the encoder performs the optimization tasks and associated decisions regarding coding options. However, all explanations regarding block subdivision, prediction, inverse quantization, and retransmission are also valid for the decoder 54 of Figure 4.

[0038] 3. ALWIP (Affine Linear Weighted Intra Predictor) Although ALWIP is not always required to implement the techniques discussed herein, some non-limiting examples of ALWIP are discussed herein.

[0039] That is, this application relates to the concept of improved block-based prediction modes for block-by-block picture coding, such as those available in video codecs such as HEVC or any successor to HEVC. The prediction modes may be intra-prediction modes, but theoretically the concepts described herein may also be applied to inter-prediction modes where the reference samples are part of another picture.

[0040] There is a need for a block-based prediction concept that allows for efficient implementations, including hardware-friendly implementations.

[0041] This object is achieved by the subject matter of the independent claims of the present application.

[0042] Intra-prediction modes are widely used in picture and video coding. In video coding, intra-prediction modes compete with other prediction modes, such as inter-prediction modes, such as motion-compensated prediction modes. In intra-prediction modes, a current block is predicted based on neighboring samples, i.e., samples that have already been coded as far as the encoder side is concerned, and samples that have already been decoded as far as the decoder side is concerned. To form a prediction signal for the current block, the neighboring sample values ​​are extrapolated to the current block, and the prediction residual is transmitted in the data stream for the current block. The better the prediction signal, the smaller the prediction residual, and therefore the fewer bits required to code the prediction residual.

[0043] To be effective, several aspects should be taken into consideration to form an effective framework for intra prediction in the context of block-by-block picture coding. For example, the more intra prediction modes supported by a codec, the higher the side information rate consumption for signaling the decoder's selection. On the other hand, the set of supported intra prediction modes should be able to provide a good prediction signal, i.e., a prediction signal that results in a low prediction residual.

[0044] In the following, as an example of a comparative embodiment or basis, an apparatus (encoder or decoder) for decoding a picture from a data stream block by block is disclosed, which apparatus supports at least one intra prediction mode according to which an intra prediction signal for a block of a predetermined size of the picture is determined by applying a first template of samples neighboring the current block to an affine linear predictor, which affine linear predictor shall consequently be called an affine linear weighted intra predictor (ALWIP).

[0045] The apparatus may have at least one of the following characteristics (the same may apply to a method or another technique implemented in a non-transitory storage unit storing instructions that, when executed by a processor, cause the processor to perform the method and / or operate as an apparatus):

[0046] 3.1 Predictors can be complementary to other predictors The intra-prediction modes, which may form the subject of implementation improvements described further below, may be complementary to other intra-prediction modes of the codec. Thus, they may be complementary to the DC, planar, or angular prediction modes defined in the HEVC codec and the JEM reference software, respectively. The latter three types of intra-prediction modes will hereinafter be referred to as conventional intra-prediction modes. Therefore, a flag indicating whether one of the intra-prediction modes supported by the device should be used for a given block in an intra-mode needs to be analyzed by the decoder.

[0047] 3.2 More than one proposed prediction mode A device may contain more than one ALWIP mode, so if the decoder knows that one of the ALWIP modes supported by the device should be used, the decoder needs to parse additional information that indicates which of the ALWIP modes supported by the device should be used.

[0048] The signaling of supported modes may have the property that coding of some ALWIP modes may require fewer bins than other ALWIP modes. Which of these modes require fewer bins and which modes require more bins may either depend on information that can be extracted from an already decoded bitstream or may be predetermined.

[0049] 4. Some Aspects 2 shows a decoder 54 for decoding a picture from data stream 12. Decoder 54 may be configured to decode a given block 18 of the picture. Specifically, predictor 44 may be configured to map a set of P neighboring samples around given block 18 using a linear or affine-linear transformation [e.g., ALWIP] to a set of Q predicted values ​​for the samples of the given block.

[0050] As shown in FIG. 5, a given block 18 contains a Q value to be predicted (which becomes the "predicted value" at the end of the operation). If block 18 has M rows and N columns, then Q=M·N. The Q value of block 18 may be in the spatial domain (e.g., pixels) or in the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q value of block 18 may be predicted based on P values ​​taken from neighboring blocks 17a-17c, which are generally adjacent to block 18. The P values ​​of neighboring blocks 17a-17c may be located closest to (e.g., adjacent to) block 18. The P values ​​of neighboring blocks 17a-17c have already been processed and predicted. The P values ​​are shown as values ​​in portions 17'a-17'c to distinguish portions 17'a-17'c from the blocks of which they are a part (in some examples, 17'b is not used).

[0051] As shown in FIG. 6, to perform the prediction, it is possible to work with a first vector 17P with P entries (each entry is associated with a specific position in the neighborhood portions 17′a-17′c), a second vector 18Q with Q entries (each entry is associated with a specific position in the block 18), and a mapping matrix 17M (each row is associated with a specific position in the block 18 and each column is associated with a specific position in the neighborhood portions 17′a-17′c). The mapping matrix 17M thus performs a prediction of the P values ​​of the neighborhood portions 17′a-17′c to the values ​​of the block 18 according to a predetermined mode. The entries in the mapping matrix 17M can therefore be understood as weighting coefficients. In the following passage, the symbols 17a-17c are used instead of 17′a-17′c to refer to the neighborhood portions of the boundary.

[0052] Several conventional modes are known in the art, such as DC mode, Planar mode, and 65 directional prediction modes. For example, there may be 67 known modes.

[0053] However, it should be noted that a different mode, referred to herein as a linear or affine-linear transformation, can also be used. A linear or affine-linear transformation includes P·Q weighting coefficients, at least ¼P·Q of which are non-zero, and for each of the Q predicted values ​​includes a sequence of P weighting coefficients for each predicted value. When ordered one under the other according to the raster scan order between the samples of a given block, this sequence forms an omnidirectionally nonlinear envelope.

[0054] It is possible to map P positions of the neighboring values ​​17'a to 17'c (template), map Q positions of the neighboring samples 17'a to 17'c, and map them at P*Q weighting coefficient values ​​of the matrix 17M. The plane is an example of the envelope of the sequence for the DC transform (this is the plane for the DC transform). Because the envelope is clearly planar, it is excluded from the definition of linear or affine-linear transformation (ALWIP). Another example is the matrix that emulates the angular mode. The envelope is excluded from the definition of ALWIP and, frankly, looks like a hill extending diagonally from top to bottom along a direction in the P / Q plane. The planar mode and the 65 directional prediction modes have different envelopes; however, they are linear in at least one direction, i.e., all directions for the DC example and the hill direction for the angular mode.

[0055] In contrast, the envelope of a linear or affine transformation is not linear in all directions. It is understood that in some situations, such types of transformations may be optimal for performing prediction for block 18. Note that it is preferable that at least one-quarter of the weighting coefficients are different from 0 (i.e., at least 25% of the P*Q weighting coefficients are different from 0).

[0056] The weighting coefficients may not be related to one another according to any ordinary mapping rules. Thus, the matrix 17M may be such that the values ​​of its entries have no obvious discernible relationship. For example, the weighting coefficients cannot be described by any analytical or difference function.

[0057] In an example, the ALWIP transformation is such that the average of the maximum values ​​of the cross-correlations between a first series of weighting coefficients for each predicted value and a second series of weighting coefficients for predicted values ​​other than each predicted value, or the maximum values ​​of the cross-correlations with an inverted version of the latter series, whichever is greater, may be less than a predetermined threshold (e.g., 0.2, 0.3, 0.35, or 0.1, e.g., a threshold ranging between 0.05 and 0.035). For example, for each pair (i, i) of rows of the ALWIP matrix 17M, the cross-correlation may be calculated by multiplying the P value of the i row by the P value of the i row. For each obtained cross-correlation, the maximum value may be obtained. Thus, the mean (average) may be obtained for the entire matrix 17M (i.e., the maximum values ​​of the cross-correlations in all combinations are averaged). The threshold may then be, for example, 0.2 or 0.3 or 0.35 or 0.1, for example a threshold in the range between 0.05 and 0.035.

[0058] The P neighboring samples of blocks 17a-17c may be located along a one-dimensional path extending along the boundaries (e.g., 18c, 18a) of a given block 18. For each of the Q predicted values ​​of a given block 18, the sequence of P weighting coefficients for the respective predicted value may be ordered in a manner that traverses the one-dimensional path in a predetermined direction (e.g., from left to right, from top to bottom, etc.).

[0059] In the example, the ALWIP matrix 17M may be non-diagonal or non-block diagonal.

[0060] An example of an ALWIP matrix 17M for predicting a 4x4 block 18 from four already predicted neighboring samples may be as follows: { {37,59,77,28}, {32,92,85,25}, {31,69,100,24}, {33,36,106,29}, {24,49,104,48}, {24,21,94,59}, {29,0,80,72}, {35,2,66,84}, {32,13,35,99}, {39,11,34,103}, {45,21,34,106}, {51,24,40,105}, {50,28,43,101}, {56,32,49,101}, {61,31,53,102}, {61,32,54,100} }

[0061] (Here, {37, 59, 77, 28} is the first row, {32, 92, 85, 25} is the second row, and {61, 32, 54, 100} is the 16th row of matrix 17M. Matrix 17M has dimensions 16x4 and contains 64 weighting coefficients (as a result of 16*4=64). This is because matrix 17M has dimensions QxP, where Q=M*N, which is the number of samples in block 18 to be predicted (block 18 is a 4x4 block), and P is the number of samples in the already predicted samples. Here, M=4, N=4, Q=16 (as a result of M*N=4*4=16), and P=4. The matrix is ​​non-diagonal and non-block diagonal and is not described by any particular rule.

[0062] As can be seen, fewer than a quarter of the weighting coefficients are 0 (in the matrix shown above, 1 in 64 weighting coefficients is 0). The envelopes formed by these values, when arranged one under the other according to raster scan order, form an envelope that is non-linear in all directions.

[0063] Although the above description is primarily discussed with reference to a decoder (eg, decoder 54), the same may be performed in an encoder (eg, encoder 14).

[0064] In some examples, for each block size (in the set of block sizes), the ALWIP transforms of the intra prediction modes in the second set of intra prediction modes for the respective block sizes are different from each other. Additionally or alternatively, the cardinality of the second sets of intra prediction modes for block sizes in the set of block sizes may match, but the associated linear transforms or affine-linear transforms of the intra prediction modes in the second sets of intra prediction modes for different block sizes may not be transferable to each other by scaling.

[0065] In some examples, an ALWIP transformation may be defined to have "nothing in common" with a conventional transformation (e.g., an ALWIP transformation may have "nothing in common" with a corresponding conventional transformation, even if it is mapped via one of the mappings described above).

[0066] In some examples, the ALWIP mode is used for both the luma and chroma components, while in other examples, the ALWIP mode is used for the luma component but not for the chroma component.

[0067] 5 Affine linear weighted intra prediction mode with encoder speedup (e.g., Test CE3-1.2.1) 5.1 Description of the method or apparatus The Affine Linear Weighted Intra Prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as that proposed in JVET-L0199 under test CE3-2.2.2, except for the following changes: Multiple Reference Line (MRL) intra prediction, especially in harmony with encoder estimation and signaling, i.e. MRL is not combined with ALWIP and transmission of MRL indices is restricted to non-ALWIP blocks. Subsampling is now mandatory for all blocks WxH >= 32x32 (previously it was optional for 32x32). Therefore, the additional test and sending of the subsampling flag in the encoder has been eliminated. ALWIP for 64xN and Nx64 blocks (where N≦32) has been added by downsampling to 32xN and Nx32 respectively and applying the corresponding ALWIP mode.

[0068] Additionally, test CE3-1.2.1 includes the following encoder optimizations for ALWIP: Unified mode estimation: Conventional and ALWIP modes use a shared Hadamard candidate list for complete RD estimation, i.e., ALWIP mode candidates are added to the same list as conventional (and MRL) mode candidates based on their Hadamard cost. EMT intra fast and PB intra fast are supported for a unified mode list, with further optimizations to reduce the number of full RD checks. Only the MPMs of the available left and top blocks are added to the list for full RD estimation for ALWIP, following the same approach as in conventional mode.

[0069] 5.2 Complexity Assessment In Test CE3-1.2.1, at most 12 multiplications per sample were required to generate the predicted signal, excluding calculations involving the discrete cosine transform. Furthermore, a total of 136,492 parameters, each 16 bits long, were required, which corresponds to 0.273 MB of memory.

[0070] 5.3 Experimental results The test evaluation was performed under the common test conditions JVET-J1010 [2] for Intra-only (AI) and Random Access (RA) configurations using VTM software version 3.0.1. The corresponding simulations were performed on an Intel Xeon cluster (E5-2697A v4, AVX2 on, Turbo Boost off) using the Linux OS and GCC 7.2.1 compiler.

[0071] [Table 1]

[0072] [Table 2]

[0073] 5.4 Reduced-Complexity Affine Linear Weighted Intra Prediction (e.g., Test CE3-1.2.2) The technique tested in CE2 is related to the "Affine Linear Intra Prediction" described in JVET-L0199 [1], but simplifies it in terms of memory requirements and computational complexity. There can be only three different sets of prediction matrices (e.g., S0, S1, S2, see below) and bias vectors (e.g., to provide offset values), covering all block shapes. As a result, the number of parameters is reduced to 14400 10-bit values, which is less memory than would be stored in a 128x128 CTU. The input and output sizes of the predictor are further reduced. Furthermore, instead of transforming the boundaries via DCT, averaging or downsampling may be performed on the boundary samples, and the generation of the prediction signal may use linear interpolation instead of an inverse DCT. As a result, a maximum of four multiplications per sample may be required to generate the prediction signal.

[0074] 6. Examples We now discuss how to perform some predictions (eg, as shown in FIG. 6) using ALWIP prediction.

[0075] 6, multiplications of Q×P samples of the Q×P ALWIP prediction matrix 17M with P samples of the P×1 neighborhood vector 17P should be performed to obtain Q=M*N values ​​of the M×N block 18 to be predicted. Thus, in general, at least P=M+N value multiplications are required to obtain each of the Q=M*N values ​​of the M×N block 18 to be predicted.

[0076] These multiplications have a highly undesirable effect: the dimension P of the boundary vector 17P generally depends on the number M+N of boundary samples (bins or pixels) 17a, 17c in the vicinity (e.g., adjacent to) of the M×N block 18 to be predicted. This means that if the size of the block 18 to be predicted is large, the number M+N of boundary pixels (17a, 17c) also increases accordingly, increasing the dimension P=M+N of the P×1 boundary vector 17P and the length of each row of the Q×P ALWIP prediction matrix 17M, and therefore the number of required multiplications. (In general, Q=M*N=W*H, where W (width) is another symbol of N and H (height) is another symbol of M. If the boundary vector is formed by only one row and / or one column of samples, then P=M+N=H+W.)

[0077] This problem is generally exacerbated by the fact that in microprocessor-based systems (or other digital processing systems), multiplication is generally a processing-power-consuming operation. It can be imagined that a large number of multiplications performed on an extremely large number of samples for a large number of blocks would result in a waste of processing power, which is generally undesirable.

[0078] It is therefore desirable to reduce the number of multiplications Q*P required to predict the M×N block 18 .

[0079] It has been recognized that by intelligently choosing simpler operations to replace multiplications, it is possible to reduce somewhat the processing power required for each intra prediction of each block 18 to be predicted.

[0080] Specifically, referring to Figures 7.1 to 7.4, an encoder or decoder may use multiple neighboring samples (e.g., 17a, 17c) to encode a given block (e.g., 18) of a picture as follows: reducing (e.g., in step 811) the plurality of neighboring samples (e.g., by averaging or downsampling) to obtain a reduced set of sample values ​​having a smaller number of samples compared to the plurality of neighboring samples (e.g., 17a, 17c); Applying a linear or affine-linear transformation (e.g., in step 812) to the reduced set of sample values ​​to obtain a predicted value for a given sample of a given block. It is understood that this can be predicted by

[0081] In some cases, the decoder or encoder may also derive predicted values ​​for further samples of a given block, e.g., by interpolation, based on predicted values ​​for the given sample and multiple neighboring samples, thus resulting in an upsampling strategy.

[0082] In an example, in order to reach a reduced set 102 of samples with fewer samples (for example, FIGS. 7.1 to 7.4), it is possible to perform some averaging (for example, at step 811) on the samples of boundary 17 (at least one of the samples of the fewer number of samples 102 can be the average of two samples of the original boundary samples or a selection of the original boundary samples). For example, if the original boundary has P = M + N samples, the reduced set of samples may have P ,

[0083] , = M red + N red where at least one of M red < M and N red < N holds, so P red < P. Thus, the boundary vector 17P actually used for prediction (for example, at step 812b) has entries of P red × 1 instead of P red < P and P red (or Q red × P red , see below), and since at least one of M red < M and N red < N holds, it has at least P red < P, so the number of elements in the matrix is fewer.

[0083] In some examples (for example, FIGS. 7.2, 7.3), the block obtained by ALWIP (at step 812) is a reduced block having size M' red × N' red . <null><null><null>

Number

Number

[0087] That is, when the number of samples directly predicted by ALWIP is less than the number of samples of block 18 to be actually predicted, it is even possible to further reduce the number of multiplications. Therefore,

[0088]

Number

[0089] When set to, this results in obtaining an ALWIP prediction by using not Q * P red multiplications but Q red * P red multiplications (Q red * P red < Q * P red < Q * P). This multiplication predicts a reduced block of dimension

[0090]

Number

[0091] Even so, it would be possible to perform (for example, by interpolation) upsampling from the predicted block of the reduced

[0092]

Number

[0093] to the final predicted block of M × N (for example, in subsequent step 813).

[0094] These techniques can be advantageous in that there are fewer matrix multiplications (Q red * P red or Q * P red), while both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or even avoiding) multiplications. For example, downsampling, averaging, and / or interpolation can be performed (e.g., in steps 811 and / or 813) by employing binary operations that do not require processing power, such as additions and shifts.

[0095] Also, addition is a very simple operation that can be easily performed without much computational effort.

[0096] This shifting operation can be used, for example, to average two boundary samples to obtain the final predicted block, and / or to interpolate two samples (support values) of (or taken from) the downscaled predicted block. (Two sample values ​​are needed for the interpolation; as in Figure 7.2, within a block there are always two predetermined values, but to interpolate samples along the left and top boundaries of a block there is only one predetermined value, so we use the boundary sample as the support value for the interpolation.)

[0097] A two-step procedure may be used, namely: First, add the values ​​of the two samples. The sum value is then halved (eg, by right shifting). And so on.

[0098] Alternatively, the following is possible: Each of the samples is first halved (eg, by left shifting). The values ​​of the two halved samples are then added together.

[0099] When downsampling (e.g., in step 811), even simpler operations can be performed, because it only requires selecting one sample from a group of samples (e.g., samples adjacent to each other).

[0100] [[ID=­3]] Therefore, it is possible to define techniques here to reduce the number of multiplications to be performed. That is, some of these techniques can be based on at least one of the following techniques. Even when the actually predicted block 18 has a size of M×N, the block can be reduced (in at least one of the two dimensions), with a reduced size of Q red ×P red with an ALWIP matrix (

[0101]

Number

[0102] where P red = N red + M red and

[0103] [[ID=3­2]]<00008(17>

Number

[0104] and / or

[0105]

Number

[0106] and / or M red < M and / or N red < N applies). Therefore, the boundary vector 17P has a size of P red ×1, suggesting that there are only P red < P multiplications (P red = M red + N redand P=M+N). P red The boundary vector 17P of ×1 is, for example, by downsampling (e.g., by choosing only some samples on the boundary), and / or By averaging multiple samples of the boundary (which can be easily obtained by addition and shifting without multiplication), It can be easily obtained from the original boundary 17. Additionally or alternatively, rather than predicting all Q=M*N values ​​of the block 18 to be predicted by multiplication, it is possible to predict only the reduced block with reduced dimensions (e.g.,

[0107]

number

[0108] and

[0109]

number

[0110] and / or

[0111]

number

[0112] The remaining samples of the block 18 to be predicted can be obtained by interpolation, for example, by subtracting the remaining QQ red Q as the support value for the value red is obtained by using samples.

[0113] According to the example shown in FIG. 7.1, a 4×4 block 18 (M=4, N=4, Q=M*N=16) is to be predicted, and the sample neighborhood 17, 17a (vertical column with four already predicted samples) and 17c (horizontal row with four already predicted samples), have already been predicted in the previous iteration (neighborhoods 17a and 17c can be collectively denoted by 17). A priori, by using the formulas shown in FIG. 6, the prediction matrix 17M should be a Q×P=16×8 matrix (because Q=M*N=4*4 and P=M+N=4+4=8), and the boundary vector 17P should have dimensions 8×1 (because P=8). However, this requires performing 8 multiplications for each of the 16 samples of the 4×4 block 18 to be predicted, resulting in a total of 16*8=128 multiplications. (Note that the average number of multiplications per sample is a good indicator of computational complexity. Traditional intra prediction requires four multiplications per sample, which increases the amount of computation involved. Therefore, using this as an upper bound for ALWIP can ensure that the complexity is reasonable and does not exceed that of traditional intra prediction.)

[0114] Nevertheless, by using this technique, in step 811, the number of samples 17a and 17c in the neighborhood of the block 18 to be predicted can be reduced from P redIt is understood that it is possible to reduce to . Specifically, in order to obtain a reduced boundary 102 with two horizontal rows and two vertical columns, adjacent boundary samples (17a, 17c) are averaged (e.g., at 100 in FIG. 7.1), and thus it is understood that it is possible to operate as if block 18 were a 2×2 block (the reduced boundary is formed by the average value). Alternatively, it is possible to perform downsampling, and thus select two samples for row 17c and two samples for column 17a. Thus, the horizontal row 17c is processed as having two samples (e.g., averaged samples) rather than four original samples, while the vertical column 17a, which originally had four samples, is processed as having two samples (e.g., averaged samples). After re-dividing row 17c and column 17a within each group 110 of two samples each, it is also possible to understand that one single sample is maintained (e.g., the average of the samples in group 110 or a simple selection from the samples in group 110). Thus, a so-called reduced set 102 of sample values is obtained by the fact that the set 102 has only four samples (M red = 2, N red = 2, P red = M red + N[[ID=⑨]] red = 4, provided that P red < P).

[0115] It is understood that it is possible to perform operations (such as averaging or downsampling 100) without performing too many multiplications at the processor level. The averaging or downsampling 100 performed in step 811 can be easily obtained by simple operations such as addition and shift that do not consume processing power.

[0116] At this point, it is understood that a linear or affine-linear (ALWIP) transform 19 can be applied to the reduced set of sample values ​​102 (e.g., using a prediction matrix such as matrix 17M of FIG. 6 ). In this case, the ALWIP transform 19 directly maps the four samples 102 to the sample values ​​104 of block 18. No interpolation is required in this case.

[0117] In this case, the ALWIP matrix 17M has dimensions Q × P red = 16 × 4. This follows from the fact that all Q = 16 samples of the block 18 to be predicted are obtained directly by ALWIP multiplication (no interpolation is required).

[0118] Therefore, in step 812a, the dimension Q×P red An appropriate ALWIP matrix 17M is selected with A. This selection may be based, for example, at least in part, on signaling from data stream 12. The selected ALWIP matrix 17M also k where k may be understood as an index, which may be signaled in the data stream 12 (in some cases the matrix

[0119]

number

[0120] (Also denoted as ∇ ...

[0121] In step 812b, the selected Q×P red ALWIP matrix 17M(Ak (also shown as P red A multiplication with the x1 boundary vector 17P is performed.

[0122] In step 812c, an offset value (e.g., b k ) can be added to all the obtained values ​​104 of the vector 18Q obtained by, for example, ALWIP. k , or in some cases

[0123]

number

[0124] (see below) is a function of the particular chosen ALWIP matrix (A k ) or may be based on an index (which may be signaled in data stream 12, for example).

[0125] Therefore, the comparison between using the present technique and not using the present technique is now resumed. Without this technique: Block 18 is to be predicted, this block has dimensions M=4, N=4 Q=M*N=4*4=16 values ​​will be predicted P = M + N = 4 + 4 = 8 boundary samples P=8 multiplications for each of the Q=16 values ​​to be predicted A total of P*Q=8*16=128 multiplications When using this technique: Block 18 is to be predicted, this block has dimensions M=4, N=4 Q=M*N=4*4=16 values ​​will be predicted in the end Reduced dimension of boundary vector: P red =M red +N re d=2+2=4 For each of the Q=16 values ​​to be predicted by ALWIP, red = 4 multiplications In total, P red *Q=4*16=64 multiplications (half of 128) The ratio between the number of multiplications and the number of final values ​​to be obtained is P red *Q / Q=4, i.e., half of P=8 multiplications for each sample to be predicted

[0126] As can be appreciated, it is possible to obtain the appropriate value in step 812 by utilizing simple, low-power operations such as averaging (and possibly adding and / or shifting and / or downsampling).

[0127] Referring to Figure 7.2, the block 18 to be predicted is now an 8x8 block of 64 samples (M=8, N=8). Here, a priori, the prediction matrix 17M should have size QxP=64x16 (Q=64 since Q=M*N=8*8=64, M=8, and N=8, and P=M+N=8+8=16). Thus, a priori, P=16 multiplications are required for each of the Q=64 samples of the 8x8 block 18 to be predicted, to arrive at 64x16=1024 multiplications for the entire 8x8 block 18.

[0128] However, as can be seen in Figure 7.2, a method 820 can be provided according to which, rather than using all 16 samples of the boundary, only 8 values ​​are used (e.g., 4 values ​​in the horizontal boundary row 17c and 4 values ​​in the vertical boundary column 17a between the original samples of the boundary). From the boundary row 17c, 4 samples can be used instead of 8 (e.g., they can be a 2-by-2 average and / or a selection of 1 sample from 2 samples). Thus, the boundary vector is not a P x 1 = 16 x 1 vector, but a P red ×1=8×1 vector (P red =Mred +N red =4+4). Instead of the original P=16 samples, P red It is understood that it is possible to select or average (e.g., 2 by 2) the samples in the horizontal rows 17c and the samples in the vertical columns 17a so that there are only Q=M*N=8*8=64 boundary values ​​to form a reduced set 102 of sample values. This reduced set 102 makes it possible to obtain a reduced version of the block 18, which has Q=M*N=8*8=64 (instead of Q=M*N=8*8=64). red =M red *N red = 4 * 4 = 16 samples of size M red ×N red It is possible to apply the ALWIP matrix to predict a block with Q = 4 × 4. The reduced version of block 18 includes the samples shown in gray in scheme 106 of Figure 7.2. The samples shown in gray squares (including samples 118' and 118'') are obtained in step 812 by applying Q red 8. A 4x4 reduced block is formed with 16 values. The 4x4 reduced block is obtained by applying a linear transformation 19 in an applying step 812. After obtaining the values ​​of the 4x4 reduced block, it is possible to obtain the values ​​of the remaining samples (shown as white samples in scheme 106), for example by interpolation.

[0129] As with method 810 of FIG. 7.1, method 820 calculates the remaining QQ of the M×N=8×8 block 18 to be predicted, for example by interpolation. red 8. The method may additionally include step 813 of deriving predicted values ​​for the remaining QQ = 64 - 16 = 48 samples (white squares). red =64-16=48 samples are interpolated to Q red= 16 directly obtained samples (e.g., interpolation may also utilize the values ​​of boundary samples). As can be seen in Figure 7.2, samples 118' and 118" (as indicated by the grey squares) are obtained in step 812, while sample 108' (shown by the white square, halfway between samples 118' and 118") is obtained by interpolation between samples 118' and 118" in step 813. It is understood that interpolation may also be obtained by operations similar to those for averaging, such as shifting and adding. Thus, in Figure 7.2, value 108' may generally be determined as a value halfway between the value of sample 118' and the value of sample 118". (It may be the average.)

[0130] It is also possible to perform interpolation to arrive at the final version of the M×N=8×8 block 18 in step 813 based on the multiple sample values ​​shown at 104 .

[0131] Therefore, the comparison between using the technique and not using the technique is as follows: Without this technique: Block 18 is to be predicted, this block has dimensions M=8, N=8 Q=M*N=8*8=64 samples in block 18 will be predicted P=M+N=8+8=16 samples within boundary 17 P=16 multiplications for each of the Q=64 values ​​to be predicted A total of P*Q=16*64=1028 multiplications The ratio between the number of multiplications and the number of final values ​​to be obtained is P*Q / Q=16 When using this technique: Block 18 is to be predicted, this block has dimensions M=8, N=8 Q=M*N=8*8=64 values ​​will be predicted in the end But Q red ×P red The ALWIP matrix will be used, Pred =M red +N red , Q red =M red *N red , M red =4, N red =4 P in the boundary red =M red +N red = 4 + 4 = 8 samples, P red <Pである Q of the 4x4 scaled block to be predicted (formed by the grey square in scheme 106) red = P for each of the 16 values red = 8 multiplications In total, P red *Q red =8*16=128 multiplications (much less than 1024) The ratio between the number of multiplications and the number of final values ​​to be obtained is P red *Q red / Q=128 / 64=2 (much smaller than the 16 obtained without this technique)

[0132] Thus, the technique presented here requires eight times less processing power than previous techniques.

[0133] Figure 7.3 shows another example (which may be based on method 820) in which the block 18 to be predicted is a rectangular 4x8 block (M=8, N=4) and Q=4*8=32 samples are to be predicted. Horizontal rows 17c with N=8 samples and vertical columns 17a with M=4 samples form the boundary 17. Thus, a priori, the boundary vector 17P has dimensions Px1=12x1, but the predicted ALWIP matrix should be a QxP=32x12 matrix, thus requiring Q*P=32*12=384 multiplications.

[0134] However, for example, in order to obtain a reduced horizontal row with only 4 samples (e.g., averaged samples), it is possible to average or downsample at least 8 samples of the horizontal row 17c. In some examples, the vertical column 17a remains as it is (e.g., without averaging). Overall, the reduced boundary has dimension P red = 8, and P red < P. Thus, the boundary vector 17P has dimension P red × 1 = 8 × 1. The ALWIP prediction matrix 17M is a matrix with dimension M * N red * P red = 4 * 4 * 8 = 64. The directly obtained 4 × 4 reduced block (formed by the gray columns in Illustration 107) in the application step 812 has Q red = M * N red = 4 * 4 = 16 samples instead of the Q = 4 * 8 = 32 of the original 4 × 8 block 18 to be predicted. When the reduced 4 × 4 block is obtained by ALWIP, it is possible to add the offset value b k and perform interpolation in step 813. As can be seen in step 813 of FIG. 7.3, the reduced 4 × 4 block is expanded to a 4 × 8 block 18, and the value 108' not obtained in step 812 is obtained in step 813 by interpolating the values 118' and 118'' (gray squares) obtained in step 812.

[0135] Therefore, the comparison between using this technique and not using this technique is as follows. When not using this technique: The block 18 will be predicted, and this block has dimensions M = 4, N = 8 Q = M * N = 4 * 8 = 32 values will be predicted There are P = M + N = 4 + 8 = 12 samples in the boundary P = 12 multiplications for each of the Q = 32 values to be predicted A total of P * Q = 12 * 32 = 384 multiplications The ratio between the number of multiplications and the number of final values ​​to be obtained is P*Q / Q=12 When using this technique: Block 18 is to be predicted, this block has dimensions M=4, N=8 Q=M*N=4*8=32 values ​​will be predicted at the end But Q red ×P red = 16 × 8 ALWIP matrix can be used, M = 4, N red =4, Q red =M*N red =16, P red =M+N red =4+4=8 P in the boundary red =M+N red = 4 + 4 = 8 samples, P red <Pである Q of the reduced block to be predicted red = P for each of the 16 values red = 8 multiplications Q in total red *P red =16*8=128 multiplications (less than 384) The ratio between the number of multiplications and the number of final values ​​to be obtained is P red *Q red / Q=128 / 32=4 (much less than the 12 obtained without this technique)

[0136] Therefore, the amount of calculation is reduced by a factor of three using this technique.

[0137] Figure 7.4 shows the case of a block 18 to be predicted, with dimensions M×N=16×16, and finally with Q=M*N=16*16=256 values ​​to be predicted, and P=M+N=16+16=32 boundary samples. This results in a prediction matrix with dimensions Q×P=256×32, which implies 256×32=8192 multiplications.

[0138] However, by applying method 820, it is possible to reduce the number of boundary samples in step 811 (e.g., by averaging or downsampling), for example, from 32 to 8. For example, for each group 120 of four consecutive samples in row 17a, one single sample (e.g., selected from the four samples, or the average of the samples) remains. Also, for each group of four consecutive samples in column 17c, one single sample (e.g., selected from the four samples, or the average of the samples) remains.

[0139] Here, the ALWIP matrix 17M is Q red ×P red = 64 × 8 matrix. This means that P red This is due to the fact that =8 was chosen (by using 8 averaged or selected samples from the 32 samples of the boundary) and the fact that the reduced block to be predicted in step 812 is an 8x8 block (in scheme 109 the grey square is 64).

[0140] Thus, once the 64 samples of the reduced 8x8 block are obtained in step 812, the remaining QQ of the block 18 to be predicted are obtained in step 813. red =256-64=192 values ​​104 can be derived.

[0141] In this case, it has been chosen to use all samples in boundary column 17a and only the alternate samples in boundary row 17c to perform the interpolation, although other choices may be made.

[0142] Using this method, the ratio of the number of multiplications to the number of final values ​​is Q red *P red / Q=8*64 / 256=2, which is much less than the 32 multiplications for each value without this technique.

[0143] A comparison of using the present technique and not using the present technique is as follows. Without this technique: Block 18 is to be predicted, this block has dimensions M=16, N=16 Q=M*N=16*16=256 values ​​will be predicted There are P=M+N=16+16=32 samples within the boundary P=32 multiplications for each of the Q=256 values ​​to be predicted A total of P*Q=32*256=8192 multiplications The ratio between the number of multiplications and the number of final values ​​to be obtained is P*Q / Q=32 When using this technique: Block 18 is to be predicted, this block has dimensions M=16, N=16 Q=M*N=16*16=256 values ​​will be predicted at the end But Q red ×P red = 64 × 8 ALWIP matrix is ​​used, and Q red =8*8=64 samples are M red =4, N red = 4, predicted by ALWIP, P red =M red +N red =4+4=8 P in the boundary red =M red +N red = 4 + 4 = 8 samples, P red <Pである Q of the reduced block to be predicted red = P for each of the 64 values red = 8 multiplications Q in total red *P red =64*4=256 multiplications (much less than 8192) The ratio between the number of multiplications and the number of final values ​​to be obtained is P red *Q red / Q=8*64 / 256=2 (much smaller than the 32 obtained without this technique)

[0144] Therefore, the processing power required by the present technique is 16 times less than that of conventional techniques.

[0145] Thus, a given block (18) of a picture is decoded using a number of neighboring samples (17): reducing the number of neighboring samples (100, 813) to obtain a reduced set of sample values ​​(102) with fewer samples compared to the number of neighboring samples (17); Applying (812) a linear or affine linear transformation (19, 17M) to the reduced set (102) of sample values ​​to obtain a predicted value for a given sample (104, 118', 188'') of a given block (18). It is possible to predict this by

[0146] In particular, the reduction (100, 813) can be performed by downsampling a plurality of neighboring samples to obtain a reduced set (102) of sample values ​​having a smaller number of samples compared to the plurality of neighboring samples (17).

[0147] Alternatively, the reduction (100, 813) can be performed by averaging multiple neighboring samples to obtain a reduced set (102) of sample values ​​with a smaller number of samples compared to multiple neighboring samples (17).

[0148] Furthermore, by interpolation, it is possible to derive (813) predicted values ​​for further samples (108, 108') of a given block (18) based on the predicted values ​​(104, 118', 118'') for the given sample and multiple neighboring samples (17).

[0149] The plurality of neighboring samples (17a, 17c) may extend in one dimension along two sides of the predetermined block (18) (e.g., toward the right and downward in Figures 7.1-7.4). The predetermined samples (e.g., those obtained by ALWIP in step 812) may also be arranged in rows and columns, and along at least one of the rows and columns, the predetermined samples may be located every nth position from the predetermined samples 112 (112) adjacent to the two sides of the predetermined block 18.

[0150] Based on the plurality of neighboring samples (17), a support value (118) for one of a plurality of neighboring positions (118) can be determined for each of at least one of the rows and columns, which are aligned with a respective one of the at least one of the rows and columns. It is also possible to derive a predicted value 118 for a further sample (108, 108') of the given block (18) by interpolation based on a predicted value for the given sample (104, 118', 118'') and the support values ​​for the neighboring samples (118) aligned with at least one of the rows and columns.

[0151] A given sample (104) may be located every nth sample (112) along a row from the samples (112) adjacent to two sides of a given block 18, and a given sample may be located every mth sample (112) along a column from the samples (112) adjacent to two sides of a given block 18, where n, m > 1. In some cases, n = m (e.g., in Figures 7.2 and 7.3, samples 104, 118', 118'' obtained directly by ALWIP in 812 and shown with gray boxes are staggered along rows and columns with samples 108, 108' obtained later in step 813).

[0152] It may be possible to perform determining the support values ​​along at least one of the rows (17c) and columns (17a), for example by downsampling and averaging (122), and for each support value, a group (120) of neighboring samples in the plurality of neighboring samples includes the neighboring sample (118) for which the respective support value is determined. Thus, in Figure 7.4, in step 813, it is possible to obtain the value of sample 119 by using the value of the given sample 118''' (previously obtained in step 812) and the neighboring samples 118 as support values.

[0153] The plurality of neighboring samples may extend in one dimension along two sides of a given block 18. It may be possible to perform the reduction 811 by grouping the plurality of neighboring samples 17 into one or more groups 110 of contiguous neighboring samples and performing downsampling or averaging on each of the one or more groups 110 of neighboring samples having two or more neighboring samples.

[0154] In the example, a linear or affine linear transformation is P red *Q red Pieces or P red * may contain Q weighting coefficients, P red is the number of sample values ​​(102) in the reduced set of sample values, and Q red Or Q is the number of given samples in a given block (18). red *Q red Pieces or (1 / 4)P red *Q weighting coefficients are non-zero weights. P red *Q red Pieces or P red *Q weighting coefficients are Q or Q red For each of the given samples, P redThe weighting coefficients may include a sequence of weighting coefficients that, when arranged one below the other in raster scan order within a given sample of a given block (18), form an omnidirectionally nonlinear envelope. red *Q or P red *Q red The weighting coefficients may not be related to one another via any conventional mapping rule. The average of the greater of the maximum cross-correlation values ​​between a first series of weighting coefficients for each given sample and a second series of weighting coefficients for given samples other than the given sample, or an inverted version of the latter series, is less than a predetermined threshold. The predetermined threshold may be 0.3 [or possibly 0.2 or 0.1]. P red The neighboring samples (17) may be located along a one-dimensional path extending along two edges of a given block (18), and may be Q or Q red For each of the given samples, P red The sequence of weighting factors is ordered in a manner that traverses a one-dimensional path in a given direction.

[0155] 6.1 Method and Apparatus Description To predict samples of a rectangular block of width W (also denoted N) and height H (also denoted M), affine linear weighted intra prediction (ALWIP) may take as input one row of H reconstructed neighboring boundary samples to the left of the block and one row of W reconstructed neighboring boundary samples above the block. If reconstructed samples are not available, they may be generated as is done in conventional intra prediction.

[0156] The generation of a prediction signal (eg, a value for a complete block 18) may be based on at least some of the following three steps.

[0157] 1. From among the boundary samples 17, samples 102 (e.g., four samples when W=H=4 and / or eight samples in other cases) may be extracted by averaging or downsampling (e.g., step 811).

[0158] 2. A matrix-vector multiplication followed by the addition of an offset may be performed using the averaged samples (or samples remaining from downsampling) as input. The result may be a reduced prediction signal for the subsampled set of samples in the original block (e.g., step 812).

[0159] 3. Predicted signals at the remaining positions may be generated from the predicted signals for the subsampled set, for example by upsampling, for example by linear interpolation (eg, step 813).

[0160] By steps 1 (811) and / or 3 (813), the total number of multiplications required to calculate the matrix-vector product can always be less than or equal to 4*W*H. Furthermore, the averaging operations on the boundaries and the linear interpolation of the downsized prediction signal are performed by using only additions and bit shifts. In other words, in the example, a maximum of four multiplications per sample are required for the ALWIP mode.

[0161] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b k ) may be taken from a set (eg, a set of three) of matrices, eg, S0, S1, S2, which may be stored, for example, in storage units of the decoder and encoder.

[0162] In some examples, the set S0 is a set of n0 matrices (e.g., n0=16 or n0=18 or another number).

[0163]

number

[0164] , each of which may have 16 rows, 4 columns, and 18 offset vectors each of size 16, to implement the technique according to Figure 7.1.

[0165]

number

[0166] This set of matrices and offset vectors is used for a block 18 of size 4x4. red = 4 vectors (as in step 811 in Figure 7.1), the P of the reduced set of samples 102 is directly converted to Q = 16 samples of the 4 × 4 block 18 to be predicted. red = 4 samples can be mapped.

[0167] In some examples, the set S1 is a set of n1 matrices (e.g., n1=8 or n1=18 or another number).

[0168]

number

[0169] , each of which may have 16 rows, 8 columns, and 18 offset vectors each of size 16, to implement the technique according to Figure 7.2 or Figure 7.3.

[0170]

number

[0171] The matrices and offset vectors of this set S1 can be used for blocks of size 4x8, 4x16, 4x32, 4x64, 16x4, 32x4, 64x4, 8x4, and 8x8. In addition, it can also be used for blocks of size WxH where max(W,H)>4 and min(W,H)=4, i.e., blocks of size 4x16 or 16x4, 4x32 or 32x4, and 4x64 or 64x4. The 16x8 matrix refers to a reduced version of block 18, which is a 4x4 block, as can be seen in Figures 7.2 and 7.3.

[0172] Additionally or alternatively, the set S2 may be a set of n2 matrices (e.g., n2=6 or n2=18 or another number).

[0173]

number

[0174] each of which may have 64 rows, 8 columns, and 18 offset vectors each of size 64

[0175]

number

[0176] The 64x8 matrix refers to a reduced version of block 18, which is for example an 8x8 block as obtained in Figure 7.4. This set of matrices and offset vectors can be used for blocks of size 8x16, 8x32, 8x64, 16x8, 16x16, 16x32, 16x64, 32x8, 32x16, 32x32, 32x64, 64x8, 64x16, 64x32, 64x64.

[0177] The matrices and offset vectors of that set, or a subset of these matrices and offset vectors, can be used for all other block shapes.

[0178] 6.2 Boundary averaging or downsampling Here, features regarding step 811 are provided.

[0179] As described above, the boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to <P samples). red to <P samples).

[0180] In a first step, the input boundaries bdry top (e.g., 17c) and bdry left (e.g., 17a) reach a smaller boundary in order to reach the reduced set 102

[0181]

Number

[0182] and

[0183]

Number

[0184] can be reduced to. Here,[[]]

[0185]

Number

[0186] and

[0187]

Number

[0188] both consist of 2 samples in the case of a 4×4 block and both consist of 4 samples in other cases.

[0189] In the case of a 4×4 block,

[0190]

number

[0191] Define

[0192]

number

[0193] can be similarly defined.

[0194]

number

[0195] ,

[0196]

number

[0197] ,

[0198]

number

[0199] ,

[0200]

number

[0201] is the average value obtained using, for example, a bit shift operation.

[0202] In all other cases (e.g., for blocks with either width or height different from 4), the block width W is W=4*2 k and for 0≦i<4,

[0203]

number

[0204] is defined, and similarly

[0205]

number

[0206] is defined.

[0207] In still other cases, it is possible to downsample the boundary (e.g., by selecting one particular boundary sample from a group of boundary samples) to arrive at a smaller number of samples. For example,

[0208]

number

[0209] , bdry top [0] and bdry top [1] may be selected from

[0210]

number

[0211] , bdry top [2] and bdry top [3] may be selected from the following.

[0212]

number

[0213] It is also possible to define

[0214] Two reduced boundaries

[0215]

number

[0216] and

[0217]

number

[0218] is the reduced boundary vector bdry, also denoted as 17P. red (associated with the reduced set 102). red For a block of shape 4x4, the size is 4 (P red =4) (example in Figure 7.1), and for all other shaped blocks the size is 8 (P red = 8) (examples in Figures 7.2 to 7.4).

[0219] where mode<18 (or the number of matrices in the set of matrices),

[0220]

number

[0221] It is possible to define

[0222] If mode ≥ 18, which corresponds to the transposed mode of mode-17,

[0223]

number

[0224] It is possible to define

[0225] Therefore, according to a specific state (one state: mode<18, one other state: mode≧18), different scanning orders (for example, one scanning order:

[0226]

number

[0227] , one other traversal order:

[0228]

number

[0229] ) along which the predicted values ​​of the output vector can be distributed.

[0230] Other strategies may be implemented. In other examples, the mode index "mode" does not necessarily range from 0 to 35 (other ranges may be defined). Furthermore, it is not necessary for each of the three sets S0, S1, and S2 to have 18 matrices (thus, instead of an expression such as mode≧18, mode≧n0, n1, and n2 are possible, which are the number of matrices for each set S0, S1, and S2 of matrices, respectively). Furthermore, the sets may each have a different number of matrices (for example, S0 may have 16 matrices, S1 may have 8 matrices, and S2 may have 6 matrices).

[0231] The mode and transpose information is not necessarily stored and / or transmitted as one unified mode index "mode." In some instances, it may be signaled explicitly as a transpose flag and matrix index (0-15 for S0, 0-7 for S1, and 0-5 for S2).

[0232] In some cases, the combination of the transposed flag and the matrix index may be interpreted as a set index. For example, there may be one bit that acts as a transposed flag and several bits that indicate the matrix index, collectively referred to as a "set index."

[0233] 5.4 Generating a Reduced Prediction Signal by Matrix-Vector Multiplication Now, characteristics regarding step 812 are provided.

[0234] The reduced input vector bdry red (Boundary vector 17P) from the reduced predicted signal pred red The latter signal can be generated by red and height H red where W red and H red can be defined as follows: If max(W,H)≦8, W red =4, H red =4 Otherwise, W red =min(W,8),H red =min(H,8)

[0235] The reduced prediction signal pred red can be calculated by calculating the matrix vector product and adding an offset. pred red =A·bdry red +b

[0236] where A is W red *H red is a matrix (e.g., a prediction matrix 17M) that may have 4 rows and 4 columns if W=H=4, and 8 columns in all other cases, and b is of size W red *H red is a vector that can be

[0237] If W=H=4, A can have 4 columns and 16 rows, so in that case, pred red In all other cases, A may have 8 columns, and in these cases, 8*W red *H red ≦4*W*H, i.e., in these cases also pred red It can be seen that up to four multiplications per sample are required to calculate .times. ...

[0238] The matrix A and vector b may be taken from one of the sets S0, S1, S2 as follows: Define the index idx=idx(W,H) by setting idx(W,H)=0 if W=H=4, idx(W,H)=1 if max(W,H)=8, and idx(W,H)=2 in all other cases. Furthermore, let m=mode if mode<18, and m=mode-17 otherwise. Then, if idx≦1 or idx=2 and min(W,H)>4, then

[0239]

number

[0240] and

[0241]

number

[0242] If idx=2 and min(W,H)=4, then A can be expressed as

[0243]

number

[0244] As the matrix resulting from excluding each row thereof, those rows correspond to the odd x - coordinates in the down - sampled block when W = 4, or to the odd y - coordinates in the down - sampled block when H = 4. When mode≧18, the down - sized predicted signal is replaced by the transposed signal. In an alternative example, different strategies may be implemented. For example, instead of reducing ( "excluding") the size of the larger matrix, W red = 4 and H red = 4, the smaller matrix (idx = 1) of S1 is used. That is, such a block is now assigned to S1 instead of S2.

[0245] Other strategies may be implemented. In other examples, the mode index "mode" is not necessarily in the range from 0 to 35 (other ranges may be defined). Further, it is not necessary for each of the three sets S0, S1, S2 to have 18 matrices (thus, instead of expressions like mode<18, it is possible for mode < n0,n1,n2, where these are the number of matrices for each set S0, S1, S2 of matrices respectively). Further, the sets may have different numbers of matrices (for example, it may be that S0 has 16 matrices, S1 has 8 matrices, and S2 has 6 matrices).

[0246] 6.4 Linear interpolation for generating the final predicted signal Here, features regarding step 812 are provided.

[0247] Interpolation of the subsampled predicted signal, on the large block, a second version of the averaged boundary may be required. That is, when min(W,H)>8 and W≧H, W = 8*2 l is written, and for 0≦i<8,

[0248]

Number

[0249] Define

[0250] Similarly, if min(W,H)>8 and H>W

[0251]

number

[0252] Define

[0253] Additionally or alternatively, it is possible to have a "hard downsampling", in which:

[0254]

number

[0255] teeth

[0256]

number

[0257] is equal to.

[0258] Also,

[0259]

number

[0260] can be similarly defined.

[0261] pred red At the sample positions omitted in the generation of pred red 7.2 to 7.4 by linear interpolation (e.g., step 813 in the examples of Figures 7.2 to 7.4). In some examples, this linear interpolation may not be necessary when W=H=4 (e.g., the example of Figure 7.1).

[0262] Linear interpolation can be given as follows (although other examples are possible): It is assumed that W≧H and H>H red If pred red A vertical upsampling of pred can be performed. red can be expanded by one line upwards as follows: If W=8, then pred red is the width W red = 4, for example, as defined above, the averaged boundary signal

[0263]

number

[0264] If W>8, pred red is width W red = 8, for example, as defined above, the averaged boundary signal

[0265]

number

[0266] It is extended upwards by pred red For the first row of pred red [x][-1] can be written, and the width W red and height 2*H red Signals on the block

[0267]

number

[0268] teeth,

[0269]

number

[0270] may be given as, 0≦x <W red and 0≦y <H red The latter process is 2k*H red = H. Thus, if H=8 or H=16, it may be done at most once. If H=32, it may be done twice. If H=64, it may be done three times. Next, a horizontal upsampling operation may be applied to the result of the vertical upsampling. The latter upsampling operation may use the complete left boundary of the predicted signal. Finally, if H>W, one may proceed similarly by first upsampling horizontally (if necessary) and then upsampling vertically.

[0271] This is an example of interpolation using scaled boundary samples (horizontal or vertical) for the first interpolation and the original boundary samples (vertical or horizontal) for the second interpolation. Depending on the block size, only the second interpolation is needed, or no interpolation is needed. If both horizontal and vertical interpolation are needed, the order depends on the width and height of the block.

[0272] However, different techniques may be implemented, for example the original boundary samples may be used for both the first and second interpolations, or the order may be fixed, for example horizontally first, then vertically (or in other cases vertically first, then horizontally).

[0273] Therefore, the interpolation order (horizontal / vertical) and the use of reduced / original boundary samples can vary.

[0274] 6.5 Explaining the complete ALWIP process example The entire process of averaging matrix-vector multiplication and linear interpolation is shown for different shapes in Figures 7.1 to 7.4. Note that the remaining shapes are treated like one of the illustrated cases.

[0275] 1. Assuming a 4x4 block, ALWIP can take two averages along each axis of the boundary by using the technique in Figure 7.1. The resulting four input samples go into a matrix-vector multiplication. The matrix is ​​taken from set S0. After adding an offset, this can yield 16 final predicted samples. No linear interpolation is required to generate the predicted signal. Therefore, a total of (4*16) / (4*4)=4 multiplications are performed per sample. See Figure 7.1 for an example.

[0276] 2. Assuming an 8x8 block, ALWIP can take four averages along each axis of the boundary. The resulting eight input samples go into a matrix-vector multiplication by using the technique in Figure 7.2. The matrix is ​​taken from set S1. This produces 16 samples in the odd-numbered positions of the prediction block. Thus, a total of (8*16) / (8*8)=2 multiplications are performed per sample. After adding the offsets, these samples can be interpolated vertically, for example, by using the top boundary, and horizontally, for example, by using the left boundary. See, for example, Figure 7.2.

[0277] 3. Assuming an 8x4 block, ALWIP can take four averages along the horizontal axis of the boundary and the four original boundary values ​​on the left boundary by using the technique in Figure 7.3. The resulting eight input samples enter a matrix-vector multiplication. The matrix is ​​taken from set S1. This produces 16 samples at odd horizontal positions and each vertical position of the prediction block. Therefore, a total of (8*16) / (8*4)=4 multiplications are performed per sample. After adding the offset, these samples can be horizontally interpolated, for example, by using the left boundary. See, for example, Figure 7.3.

[0278] The transposed case is treated accordingly.

[0279] 4. Assuming a 16x16 block, ALWIP can take four averages along each axis of the boundary. By using the technique of Figure 7.2, the resulting eight input samples go into a matrix-vector multiplication. The matrix is ​​taken from set S2. This produces 64 samples in odd-numbered positions of the predicted block. Therefore, a total of (8*64) / (16*16)=2 multiplications are performed per sample. After adding the offsets, these samples are interpolated, for example, vertically by using the top boundary and horizontally by using the left boundary. See, for example, Figure 7.2. See, for example, Figure 7.4.

[0280] For larger shapes, the procedure may be essentially the same, and it is easy to ensure that the number of multiplications per sample is less than two.

[0281] For Wx8 blocks, only horizontal interpolation is required, since samples are provided at odd horizontal positions and at each vertical position, so in these cases a maximum of (8*64) / (16*8)=4 multiplications are performed per sample.

[0282] Finally, for W×4 blocks where W>8, A k Let be the matrix resulting from excluding each row corresponding to an odd-numbered entry along the horizontal axis of the downsampled block. Thus, the output size may be 32, and again, only horizontal interpolation continues to be performed. A maximum of (8*32) / (16*4)=4 multiplications may be performed per sample.

[0283] The transposed case can be handled accordingly.

[0284] 6.6 Evaluating the number and complexity of parameters required The parameters required for all possible proposed intra prediction modes may be included in matrices and offset vectors belonging to sets S0, S1, and S2. All matrix coefficients and offset vectors may be stored as 10-bit values. Therefore, according to the above description, a total of 14,400 parameters, each with 10-bit precision, may be required for the proposed method. This corresponds to 0.018 MB of memory. Currently, a CTU of size 128 x 128 in standard 4:2:0 chroma subsampling consists of 24,576 values, each with 10 bits. Therefore, the memory requirements of the proposed intra prediction tool do not exceed those of the current picture reference tool adopted in the last conference. It should also be noted that conventional intra prediction modes require four multiplications per sample by the PDPC tool or a four-tap interpolation filter for angular prediction modes with fractional angular positions. Therefore, in terms of operational complexity, the proposed method does not exceed conventional intra prediction modes.

[0285] 6.7 Proposed Intra Prediction Mode Signaling For example, 35 ALWIP modes are proposed for luma blocks (other numbers of modes may also be used). For each coding unit (CU) of intra mode, a flag is transmitted in the bitstream indicating whether the ALWIP mode should be applied on the corresponding prediction unit (PU). The signaling of the latter index may be coordinated with the MRL in the same way as the first CE test. If the ALWIP mode should be applied, the ALWIP mode index predmode may be signaled using an MPM list with three MPMs.

[0286] Here, the derivation of the MPM may be performed using the intra modes of the top and left PUs as follows: Angular Tables that can be assigned ALWIP mode, for example, three fixed tables: map_angular_to_alwip idx , idx∈{0,1,2} are possible. predmode ALWIP =map_angular_to_alwip idx [predmode Angular ]

[0287] For each PU of width W and height H, an index indicating from which of the three sets the ALWIP parameters should be taken, as in section 4 above. idx(PU)=idx(W,H)∈{0,1,2} Define the prediction unit PU above is available, belongs to the same CTU as the current PU, and is in intra mode, then idx(PU) = idx(PU above ) and ALWIP mode

[0288]

number

[0289] Using PU above If ALWIP applies to

[0290]

number

[0291] is.

[0292] If the upper PU is available, belongs to the same CTU as the current PU, and is in intra mode, and in conventional intra prediction mode

[0293]

number

[0294] If is applied in the above PU,

[0295]

number

[0296] is.

[0297] In all other cases,

[0298]

number

[0299] , which means that this mode is unavailable. Similarly, but without the constraint that the left PU must belong to the same CTU as the current PU, the mode

[0300]

number

[0301] is derived.

[0302] Finally, there are three fixed default lists: idx , idx∈{0,1,2}, each of which contains three distinct ALWIP modes. The default list, list idx(PU) and mode

[0303]

number

[0304] and

[0305]

number

[0306] From the above, we construct three separate MPMs by replacing -1 with a default value as well as eliminating the repetitions.

[0307] The embodiments described herein are not limited to the above-mentioned signaling of proposed intra-prediction modes. According to an alternative embodiment, for MIP (ALWIP), MPM and / or mapping tables are not used.

[0308] 6.8 Deriving Adapted MPM Lists for Conventional Luma and Chroma Intra Prediction Modes The proposed ALWIP mode can be reconciled with the conventional MPM-based coding of intra prediction modes as follows: The luma and chroma MPM list derivation process for conventional intra prediction modes is performed using a fixed table, map_alwip_to_angular idx , idx∈{0,1,2}, which corresponds to the ALWIP mode predmode on a given PU. ALWIP to one of the conventional intra prediction modes. predmode Angular =map_alwip_to_angular idx(PU) [predmode ALWIP ]

[0309] For the derivation of the Luma MPM list, the ALWIP mode predmode ALWIP Whenever a neighboring luma block using the conventional intra prediction mode predmode is encountered, this block is Angular For the derivation of the chroma MPM list, whenever the current luma block uses ALWIP mode, the same mapping may be used to convert the ALWIP mode to a conventional intra prediction mode.

[0310] It should be clear that the ALWIP mode can be matched with conventional intra prediction modes without the use of an MPM and / or a mapping table. For example, for a chroma block, the ALWIP mode can be mapped to the Planar intra prediction mode whenever the current luma block uses the ALWIP mode.

[0311] 7. Implementation Efficiency Let us briefly summarize the above example, as it can be the basis for further extending the embodiments described herein below.

[0312] To predict a given block 18 of a picture 10, a number of neighboring samples 17a, 17c are used.

[0313] A reduction 100 by averaging of multiple neighboring samples is performed to obtain a reduced set of sample values ​​102, which has a smaller number of samples compared to the multiple neighboring samples. This reduction is optional in embodiments herein and produces so-called sample values, which will be referred to below. To obtain a predicted value for a given sample 104 of a given block, the reduced set of sample values ​​undergoes a linear or affine-linear transformation 19. It is this transformation, shown hereafter using a matrix A and an offset vector b, that should be obtained and efficiently performed by machine learning (ML).

[0314] By interpolation, predicted values ​​for further samples 108 of a given block are derived based on the predicted value for the given sample and multiple neighboring samples. Theoretically, it should be possible to say that the results of the affine / linear transformation can be associated with non-full-pel sample positions of block 18, such that all samples of block 18 can be obtained by interpolation according to alternative embodiments. Interpolation may not be necessary at all.

[0315] A plurality of neighboring samples may extend one-dimensionally along two sides of the predetermined block, the predetermined samples arranged in rows and columns, and along at least one of the rows and columns, the predetermined sample may be located every nth position from a sample (112) that is a predetermined sample adjacent to the two sides of the predetermined block. Based on the plurality of neighboring samples, a support value for one of the plurality of neighboring positions (118) may be determined for each of at least one of the rows and columns, which is aligned with a respective one of the rows and columns. By interpolation, predicted values ​​for further samples 108 of the predetermined block may be derived based on the predicted value for the predetermined sample and the support values ​​for the neighboring samples aligned with at least one of the rows and columns. The predetermined sample may be located every nth position from the sample 112 that is a predetermined sample adjacent to the two sides of the predetermined block along the rows, and the predetermined sample may be located every mth position from the sample 112 that is a predetermined sample adjacent to the two sides of the predetermined block along the columns, where n, m > 1. It is also possible for n = m. Along at least one of the rows and columns, the determination of the support values ​​may be made by averaging 122 a group 120 of neighboring samples within a plurality of neighboring samples that includes the neighboring sample 118 for which the respective support value is determined. The neighboring samples may extend in one dimension along two sides of a given block, and reduction may be made by grouping the neighboring samples into one or more contiguous groups 110 of neighboring samples and by performing averaging for each of one or more groups of neighboring samples having more than two neighboring samples.

[0316] For a given block, a prediction residual may be transmitted in the data stream, which may be derived at the decoder, and the given block may be reconstructed using the prediction residual and predicted values ​​for the given samples. At the encoder, the prediction residual is coded into the data stream.

[0317] A picture may be subdivided into multiple blocks of different block sizes, including a given block, and a linear or affine-linear transformation for block 18 may be selected depending on the width W and height H of the given block such that the linear or affine-linear transformation selected for the given block is selected from a first set of linear or affine-linear transformations as long as the width W and height H of the given block are within a first set of width / height pairs, and is selected from a second set of linear or affine-linear transformations as long as the width W and height H of the given block are within a second set of width / height pairs separate from the first set of width / height pairs. Again, it will become clear later that the affine / linear transformation is expressed by other parameters, namely a weight C, and optionally offset and scale parameters.

[0318] The decoder and encoder may be configured to subdivide a picture into a plurality of blocks of different block sizes, including a predetermined block, and select a linear transformation or an affine-linear transformation according to a width W and a height H of the predetermined block, such that the linear transformation or the affine-linear transformation selected for the predetermined block is a first set of linear or affine-linear transformations as long as the width W and height H of a given block are within a first set of width / height pairs; a second set of linear or affine-linear transformations, as long as the width W and height H of a given block are within a second set of width / height pairs disjoint from the first set of width / height pairs; and a third set of linear or affine-linear transformations, so long as the width W and height H of a given block are within a third set of one or more width / height pairs separate from the first and second sets of width / height pairs; is selected from.

[0319] The third set of one or more width / height pairs simply includes one width / height pair W', H', and each linear or affine-linear transform in the first set of linear or affine-linear transformations is for transforming N' sample values ​​into W'*H' predicted values ​​for a W'×H' array of sample locations.

[0320] Each of the first and second sets of width / height pairs is a first width / height pair W p , H p may include W p is H p and a second width / height pair W q , H q is not equal to H q =W p and W q =H p is.

[0321] Each of the first and second sets of width / height pairs additionally includes a third width / height pair W p , H p may include W p is H p is equal to H p >H q is.

[0322] For a given block, a set index may be transmitted in the data stream, which indicates the linear or affine-linear transformation to be selected for the block 18 from a predetermined set of linear or affine-linear transformations.

[0323] The plurality of neighboring samples may extend in one dimension along two sides of the predetermined block, and reduction may be performed by grouping a first subset of the plurality of neighboring samples adjacent to the first side of the predetermined block into a first group 110 of one or more consecutive neighboring samples, and by grouping a second subset of the plurality of neighboring samples adjacent to the second side of the predetermined block into a second group 110 of one or more consecutive neighboring samples, and by performing averaging on each of the one or more first and second groups of neighboring samples having more than two neighboring samples to obtain a first sample value from the first group and a second sample value for the second group. Then, a linear or affine-linear transformation may be selected from among a predetermined set of linear or affine-linear transformations according to the set index, such that two different states of the set index result in the selection of one of the linear or affine-linear transformations of the predetermined set of linear or affine-linear transformations, when the set index takes a first state of the two different states in the form of a first vector, to produce an output vector of predicted values ​​and distribute the predicted values ​​of the output vector along the first scanning order to predetermined samples of the predetermined block, and when the set index takes a second state of the two different states in the form of a second vector, If the first vector and the second vector are different such that an element filled by one of the first sample values ​​in the first vector is filled by one of the second sample values ​​in the second vector, and such that an element filled by one of the second sample values ​​in the first vector is filled by one of the first sample values ​​in the second vector, the reduced set of sample values ​​may be subjected to a predetermined linear transformation or an affine-linear transformation to produce an output vector of predicted values ​​and to distribute the predicted values ​​of the output vector along the second scanning order to predetermined samples of a predetermined block transposed with respect to the first scanning order.

[0324] Each linear or affine-linear transform in the first set of linear or affine-linear transformations may be for transforming the N1 sample values ​​into w1*h1 predicted values ​​for a w1×h1 array of sample locations, and each linear or affine-linear transform in the second set of linear or affine-linear transformations is for transforming the N2 sample values ​​into w2*h2 predicted values ​​for a w2×h2 array of sample locations, where for a first given pair in the first set of width / height pairs, w1 may exceed the width of the first predetermined width / height pair or h1 may exceed the height of the first predetermined width / height pair, and where for a second given pair in the first set of width / height pairs, w1 may not exceed the width of the second predetermined width / height pair or h1 may not exceed the height of the second predetermined width / height pair. If the given block is of a first predetermined width / height pair, and if the given block is of a second predetermined width / height pair, then a reduction (100) by averaging multiple neighboring samples to obtain a reduced set of sample values ​​(102) may be performed, such that the reduced set of sample values ​​102 has N1 sample values, and applying the selected linear or affine-linear transformation to the reduced set of sample values ​​may be performed by using only a first subportion of the selected linear or affine-linear transformation for a subsampling of the w1 x h1 array of sample locations along the width dimension if w1 exceeds the width of the width / height pair, or along the height dimension if h1 exceeds the height of the width / height pair and the given block is of the first predetermined width / height pair, or by using the selected linear or affine-linear transformation in its entirety if the given block is of the second predetermined width / height pair.

[0325] Each linear or affine-linear transformation in the first set of linear or affine-linear transformations may be for transforming N1 sample values ​​into w1*h1 predicted values ​​for a w1×h1 array of sample locations where w1=h1, and each linear or affine-linear transformation in the second set of linear or affine-linear transformations is for transforming N2 sample values ​​into w2*h2 predicted values ​​for a w2×h2 array of sample locations where w2=h2.

[0326] All of the above-described embodiments are merely exemplary in that they may form the basis for the embodiments described hereinafter. That is, the above concepts and details will be helpful in understanding the following embodiments and will serve as a basis for possible extensions and modifications of the embodiments described hereinafter. In particular, many of the above-described details are optional, such as the averaging of neighboring samples, the fact that neighboring samples are used as reference samples, etc.

[0327] More generally, the embodiments described herein assume that a prediction signal on a rectangular block is generated from already reconstructed samples, e.g., an intra prediction signal on a rectangular block is generated from already reconstructed samples of the left and top neighbors of the block. The generation of the prediction signal is based on the following steps:

[0328] 1. From among the reference samples, referred to herein as boundary samples, but excluding the possibility of transferring the description to reference samples located elsewhere, samples can be extracted by averaging, where averaging is performed either over both the left and top boundary samples of the block, or over only the boundary samples of one of the two sides. If averaging is not performed on a side, the sample on that side remains unchanged.

[0329] 2. A matrix-vector multiplication, optionally followed by the addition of an offset, in which the input vector is either the concatenation of the averaged boundary sample to the left of the block and the original boundary sample above the block, if averaging is applied to the left side only, or the concatenation of the original boundary sample to the left of the block and the averaged boundary sample above the block, if averaging is applied to the top side only, or the concatenation of the averaged boundary sample to the left of the block and the averaged boundary sample above the block, if averaging is applied to both sides of the block. Again, alternatives exist, such as no averaging at all.

[0330] 3. The result of the matrix-vector multiplication and optional offset addition may optionally be a downsized prediction signal on a subsampled set of samples in the original block. Predictions at the remaining positions may be generated from the prediction signals on the subsampled set by linear interpolation.

[0331] The matrix-vector product calculation in step 2 should preferably be done in integer arithmetic. Thus, x = (x1,...x n ) denotes the input of a matrix-vector product, i.e., if x denotes the concatenation of the left and top (averaged) boundary samples of a block, then from x, the (downsized) predicted signal calculated in step 2 should be calculated using only bit shifts, addition of an offset vector, and multiplication with integers. Ideally, the predicted signal in step 2 is given as Ax+b, where b is an offset vector that can be 0, and A is derived by some machine learning-based training algorithm. However, such training algorithms usually use a matrix A=A given with floating-point precision. float Therefore, the expression A float We are faced with the problem of defining integer arithmetic in the sense above such that x can be well approximated using these integer arithmetic. Now, these integer arithmetic can be written as follows, assuming a uniform distribution of the vector x: float is not necessarily chosen to approximate x, but the representation Afloat Let x be the input vector x to be approximated, and x be the (averaged) boundary samples from the natural video signal. i It is important to note that it is usually considered that some correlation between

[0332] 8 illustrates an improved ALWIP prediction. Samples of a given block may be predicted based on a first matrix-vector product of a matrix A 1100 derived by some machine learning-based training algorithm and a sample value vector x 400. Optionally, an offset b 1110 may be added. To achieve an integer or fixed-point approximation of this first matrix-vector product, the sample value vector may undergo a regular linear transformation 403 to determine a further vector 402. A second matrix-vector product between a further matrix B 1200 and the further vector 402 may be equal to the result of the first matrix-vector product.

[0333] Depending on the characteristics of the further vector 402, the second matrix-vector product may be an integer approximated by a matrix-vector product 404 of the predetermined prediction matrix C 405 and the further vector 402 plus the further offset 408. The further vector 402 and the further offset 408 may consist of integer or fixed-point values. All components of the further offset are, for example, the same. The predetermined prediction matrix 405 may be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the predetermined prediction matrix 405 and the further vector 402 may be understood as a prediction vector 406.

[0334] Further details regarding this integer approximation are given below.

[0335] Possible solutions according to example I: Subtracting and adding average values In the above scenario, possible expressions are: float One possible implementation of the integer approximation of x is x, i.e., the i0th component of the sample vector 400.

[0336]

number

[0337] , i.e., replace the predetermined component 1500 with the mean value of the components of x, mean(x), i.e., the predetermined value 1400, and subtract this mean value from all other components. In other words, a regular linear transformation 403 as shown in Figure 9a is defined such that the predetermined component 1500 of the further vector 402 is a and each of the other components of the further vector 402, except for the predetermined component 1500, is equal to the corresponding component of the sampled vector 400 minus a, where a is the predetermined value 1400, which is an average, e.g., an arithmetic or weighted average, of the components of the sampled vector 400. This operation on the input is given by a regular linear transformation T403, which has an obvious integer implementation, especially when the dimension n of x is a power of two.

[0338] A float =(A float T -1 )T, so to perform such a transformation on an input x, we must find an integer approximation to the matrix-vector product By, where B=(A float T -1 ) and y=Tx. Matrix-vector product A float Since x represents a rectangular block, i.e., a prediction for a given block, and since x is included in the (e.g., averaged) boundary samples of that block, if all sample values ​​of x are equal, i.e., x for all i, i =mean(x), the predicted signal A floatWe should expect each sample value in x to be close to or exactly equal to mean(x). This means that we should expect the i0th column of B, i.e., the column corresponding to a given component, to be very close to or equal to a column consisting of only ones. Thus, if M(i0), i.e., integer matrix 1300, is a matrix whose i0th column consists of ones and all other columns are zeros, then we write By = Cy + M(i0)y, and C = BM(i0), and we should expect C, i.e., the i0th column 412 of a given prediction matrix 405, to have relatively few entries or to be zero, as shown in Figure 9b. Moreover, since the components of x are correlated, for each i≠i0, the i0th component y of y, y, i =x i It can be expected that -mean(x) will often have an absolute value much smaller than the i-th component of x. Since the matrix M(i0) 1300 is an integer matrix, the integer approximation of By can be achieved given the integer approximation of Cy, and by the above discussion, the quantization error caused by quantizing each entry of C 405 in an appropriate way is A float It can be expected that this should only slightly affect the error in the resulting quantization of each By of x.

[0339] The predetermined value 1400 is not necessarily the mean value mean(x). float Integer approximation of x can also be achieved using the following alternative definition of the predetermined value 1400:

[0340] Expression A float Another possible implementation of integer approximations of x is the i0th component of x.

[0341]

number

[0342] remains unchanged and has the same value

[0343]

number

[0344] is subtracted from all other components, i.e., for each i≠i0,

[0345]

number

[0346] and

[0347]

number

[0348] In other words, the predetermined value 1400 may be the component of the sample value vector 400 that corresponds to the predetermined component 1500.

[0349] Alternatively, the predetermined value 1400 is a default value or a value signaled in the data stream in which the picture is coded.

[0350] The given value 1400 is, for example, 2 bitdepth-1 In this case, the further vector 402 is equal to y0=2 bitdepth-1 and for i>0, y i =x i -x0.

[0351] Alternatively, the predetermined component 1500 is a constant minus a predetermined value 1400. The constant may be, for example, 2 bitdepth-1 According to one embodiment, a predetermined component of the further vector y 402 is equal to

[0352]

number

[0353] 1500 is 2 bitdepth-1The component of the sample value vector 400 corresponding to the given component 1500 from

[0354]

number

[0355] minus the component of the sampled vector 400 that corresponds to the given component 1500 , and all other components of the further vector 402 are equal to the corresponding component of the sampled vector 400 minus the component of the sampled vector 400 that corresponds to the given component 1500 .

[0356] For example, it may be advantageous for the predetermined value 1400 to have a small deviation from the expected value for the samples of a given block.

[0357] According to an embodiment, the device 1000 is configured to include a plurality of regular linear transformations 403, each of which is associated with one component of the further vector 402. Furthermore, the device is configured, for example, to select a predetermined component 1500 from the components of the sample value vector 400 and to use the regular linear transformation 403 from the plurality of regular linear transformations associated with the predetermined component 1500 as the predetermined regular linear transformation. This is because, for example, depending on the position of the predetermined component in the further vector, the position of the i0th row, i.e., the row of the regular linear transformation 403 corresponding to the predetermined component, differs. For example, if the first component of the further vector 402, i.e., y1, is the predetermined component, the i0th row replaces the first row of the regular linear transformation.

[0358] 9b, column 412 of predetermined prediction matrix 405 corresponding to predetermined component 1500 of further vector 402, i.e., matrix component 414 of predetermined prediction matrix C 405 in the i0-th column, is, for example, all zeros. In this case, the apparatus is configured to calculate matrix-vector product 404 by performing multiplication of reduced prediction matrix C′ 405 obtained from predetermined prediction matrix C 405 by excluding column 412 and yet another vector 410 obtained from further vector 402 by excluding predetermined component 1500, by calculating matrix-vector product 407, as shown in FIG. 9c. Thus, prediction vector 406 can be calculated with fewer multiplications.

[0359] 8, 9b, and 9c, the apparatus 1000 may be configured, when predicting samples of a predetermined block based on the prediction vector 406, to calculate, for each component of the prediction vector 406, the sum of the respective component and a, i.e., a predetermined value 1400. As shown in Figures 8 and 9c, this addition may be represented by the addition of the prediction vector 406 and the vector 409, where all components of the vector 409 are equal to the predetermined value 1400. Alternatively, as shown in Figure 9b, this addition may be represented by the sum of the prediction vector 406 and a matrix-vector product 1310 of an integer matrix M 1300 and the further vector 402, where the matrix components of the integer matrix 1300 are 1 in one instance, i.e., the i0th column, of the integer matrix corresponding to the predetermined component 1500 of the further vector 402, and all other components are, for example, 0.

[0360] The result of adding the predetermined prediction matrix C 405 and the integer matrix 1300 is, for example, equal to or close to the further matrix B 1200 shown in FIG.

[0361] In other words, as shown in Figures 8, 9a, and 9b, the matrix obtained by adding each matrix element with a value of 1 of the predetermined prediction matrix C 405 in the column 412, i.e., the i0th column, of the predetermined prediction matrix 405 corresponding to the predetermined element 1500 of the further vector 402, i.e., the matrix obtained by multiplying the predetermined prediction matrix C 405 in the i0th column with the regular linear transformation 403, i.e., the further matrix B 1200 (i.e., the plurality of matrices B), corresponds to, for example, a quantized version of the machine learning prediction matrix A 1100. As shown in Figure 9b, the addition of each matrix element with a value of 1 of the predetermined prediction matrix C 405 in the i0th column 412 may correspond to the addition of the predetermined prediction matrix 405 and the integer matrix 1300. As shown in Figure 8, the machine learning prediction matrix A 1100 may be equal to the result of multiplying the further matrix 1200 with the regular linear transformation 403. This is expressed as A x = B T y T -1 The predetermined prediction matrix 405 may be, for example, a quantized matrix, an integer matrix, and / or a fixed-point matrix, thereby realizing a quantized version of the machine learning prediction matrix A1100.

[0362] Matrix multiplication using only integer arithmetic For a low complexity implementation (in terms of the complexity of adding and multiplying scalar values, and in terms of the storage required for the matrix entries involved), it is desirable to perform matrix multiplication 404 using only integer arithmetic.

[0363] Using only integer arithmetic, we can approximate z=Cy, i.e.

[0364]

number

[0365] To calculate i,j is an integer value

[0366]

number

[0367] This can be done, for example, by uniform scalar quantization or by using the values ​​y i This can be done by considering the specific correlation between the integer values, which may each be stored using a fixed number of bits n_bits, for example n_bits=8, e.g., represent fixed-point numbers.

[0368] A matrix-vector product 404 with a matrix of size m×n, i.e., a predetermined prediction matrix 405, may then be performed as shown in this pseudocode, where <<, >> are binary left-shift and right-shift operations, and +, −, and * only operate on integer values.

[0369] (1) final_offset = 1 << (right_shift_result - 1); for i in 0..m-1 { accumulator = 0 for j in 0..n-1 { accumulator := accumulator + y[j]*C[i,j] } z[i] = (accumulator + final_offset) >> right_shift_result; }

[0370] Here, array C, the predetermined prediction matrix 405, stores, for example, fixed-point numbers as integers. The final addition of final_offset and right-shift operation with right_shift_result reduces precision by rounding to obtain the required fixed-point format at the output.

[0371] To allow for an extension of the range of real values ​​representable by integers in C, two additional matrices, offset i,jand scale i,j can be used, so the matrix vector product

[0372]

number

[0373] in y j Each coefficient b i,j teeth,

[0374]

number

[0375] is given by

[0376] offset i,j and scale i,j are themselves integer values. For example, these integers can be expressed using a certain number of bits, e.g., 8 bits, or by the value

[0377]

number

[0378] may represent fixed-point numbers, each of which may be stored using the same number of bits n_bits as used to store n_bits.

[0379] In other words, the apparatus 1000 calculates the predicted parameter, e.g., an integer value.

[0380]

number

[0381] and the value offset i,j and scale i,jto represent a predetermined prediction matrix 405, and is configured to calculate a matrix-vector product 404 by performing multiplications and additions on the elements of the further vector 402 and the prediction parameters and intermediate results obtained therefrom, the absolute values ​​of the prediction parameters being representable by an n-bit fixed-point representation, where n is less than or equal to 14, or alternatively less than or equal to 10, or alternatively less than or equal to 8. For example, the elements of the further vector 402 are multiplied with the prediction parameters to produce a product as an intermediate result, which is then subjected to addition or serves as an addend for the addition.

[0382] According to an embodiment, the prediction parameters include weights, each of which is associated with a corresponding matrix element of the prediction matrix. In other words, a given prediction matrix is, for example, replaced or represented by the prediction parameters. The weights are, for example, integer and / or fixed-point values.

[0383] According to an embodiment, the prediction parameters may further be scaled by one or more scaling factors, e.g., values ​​scale i,j each of which is a weight, e.g., an integer value, associated with one or more corresponding matrix elements of the predetermined prediction matrix 405.

[0384]

number

[0385] Additionally or alternatively, the prediction parameters may be associated with one or more corresponding matrix elements of a predetermined prediction matrix 405 for scaling the matrix . i,j each of which is a weight, e.g., an integer value, associated with one or more corresponding matrix elements of the predetermined prediction matrix 405.

[0386]

number

[0387] , and is associated with one or more corresponding matrix elements of a predetermined prediction matrix 405 for offsetting the .times. ...

[0388] offset i,j and scale i,j To reduce the amount of storage required for , their values ​​may be chosen to be constant for a particular set of indices i, j. For example, as shown in Figure 10, their entries may be constant for each column, or they may be constant for each row, or they may be constant for all i, j.

[0389] For example, in one preferred embodiment, as shown in FIG. i,j and scale i,j is constant for all values ​​of the matrix of one prediction mode. Thus, when there are K prediction modes, k=0...K-1, a single value o k and a single value s k Only k is needed to compute the prediction for mode k.

[0390] According to one embodiment, the offset i,j and / or scale i,j is constant, i.e., identical, for all matrix-based intra prediction modes. i,j and / or scale i,j It is possible that ∑ i = ...

[0391] offset is o k and the scale is s k , the calculation in (1) can be modified as follows: (2) final_offset = 0; for i in 0..n-1 { final_offset := final_offset - y[i];} final_offset *= final_offset * offset * scale; final_offset += 1 << (right_shift_result - 1); for i in 0..m-1 { accumulator = 0 for j in 0..n-1 { accumulator := accumulator + y[j]*C[i,j] } z[i] = (accumulator*scale + final_offset) >> right_shift_result; }

[0392] The expanded embodiments resulting from the solution The above solution suggests the following embodiments.

[0393] 1. A prediction method as in Section I, where in step 2 of Section I the following is done for integer approximation of the matrix-vector product involved: For some i0, 1≦i0≦n, the (averaged) boundary samples x=(x1,...,x n ), vector y=(y1,…,y n ) is calculated, where y for i≠i0 i =x i -mean(x),

[0394]

number

[0395] and mean(x) represents the average value of x. Then, since the vector y serves as the input for the matrix-vector product Cy (realized by an integer), the (downsampled) predicted signal pred from step 2 of section I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents a signal equal to mean(x) for each sample position in the region of the (downsampled) predicted signal (see, for example, Fig. 9b).

[0396] 2. A prediction method as in section I, wherein in step 2 of section I, the following is done for the integer approximation of the relevant matrix-vector product. For a certain i0 with 1 ≤ i0 ≤ n, from the (averaged) boundary samples x = (x1,..., x n ), a vector y = (y1,..., y n-1 ) is calculated, where for i < i0, y i = x i - mean(x), and for i ≥ i0, y i = x i+1 - mean(x), and mean(x) represents the average value of x. Then, since the vector y serves as the input for the matrix-vector product Cy (realized by an integer), the (downsampled) predicted signal pred from step 2 of section I is given by pred = Cy + meanpred(x). In this equation, meanpred(x) represents a signal equal to mean(x) for each sample position in the region of the (downsampled) predicted signal (see, for example, Fig. 9c).

[0397] 3. A prediction method as in section I, wherein the integer realization of the matrix-vector product Cy is the coefficient in the matrix-vector product z i = Σ j b i,j * y j and the coefficients

[0398]

Number

[0399] (See, for example, FIG. 10).

[0400] 4. A prediction method as in Section I, wherein step 2 uses one of K matrices, thereby assigning multiple prediction modes to different matrices, k=0...K-1.

[0401]

number

[0402] can be calculated using the matrix-vector product C k The integer realization of y is the matrix vector product

[0403]

number

[0404] Coefficient in

[0405]

number

[0406] (See, for example, FIG. 11).

[0407] That is, according to an embodiment of the present application, the encoder and decoder operate as follows to predict a given block 18 of picture 10, see FIG. 8. For the prediction, multiple reference samples are used. As outlined above, the embodiment of the present application is not limited to intra-coding, and therefore the reference samples are not limited to neighboring samples, i.e., samples of picture 10 that are near block 18. In particular, the reference samples are not limited to samples that line the outer edge of block 18, such as samples that border the outer edge of the block. However, this situation is, of course, one embodiment of the present application.

[0408] To perform the prediction, a sample value vector 400 is formed from reference samples, such as reference samples 17a and 17c. Possible formations are described above. This formation may involve averaging, and thereby reducing, the number of samples 102 or components of vector 400 compared to the reference samples 17 that contribute to the formation. This formation may also depend in some way on the dimensions or size of block 18, such as its width and height, as described above.

[0409] It is this vector 400 that should undergo an affine or linear transformation to obtain a prediction of block 18. Various terms have been used above. Using the most recent one, the aim is to perform the prediction by applying vector 400 to matrix A by matrix-vector product while performing an addition with an offset vector b. The offset vector b is optional. The affine or linear transformation determined by A or A and b can be determined by the encoder and decoder, or more precisely, for the prediction based on the size and dimensions of block 18 as already described above.

[0410] However, to achieve the computational efficiency improvements outlined above or to make prediction more effective for implementation, the affine and linear transformations are quantized, and the encoder and decoder, or their predictors, use the above-mentioned C and T to represent and perform the linear or affine transformations, which are applied in the manner described above using C and T to represent a quantized version of the affine transformation. Specifically, rather than directly applying vector 400 to matrix A, the predictors in the encoder and decoder apply vector 402, which is derived from sample vector 400 by applying a mapping via a predetermined regular linear transformation T to sample vector 400. As used herein, transform T can be the same as long as vector 400 has the same size, i.e., does not depend on the block dimensions, i.e., width and height, or is at least the same for different affine / linear transformations. Above, vector 402 is denoted as y. The exact matrix for performing the affine / linear transformation as determined by machine learning would have been B. However, rather than strictly implementing B, prediction in the encoder and decoder is performed by approximating it or a quantized version of it. Specifically, the representation is achieved via C+M representing a quantized version of B, with C appropriately represented in the manner outlined above.

[0411] Therefore, prediction in the encoder and decoder is further performed by calculating a matrix-vector product 404 of the vector 402 and a predetermined prediction matrix C, which is appropriately represented and stored in the encoder and decoder in the manner described above. A vector 406 resulting from this matrix-vector product is then used to predict the samples 104 of the block 18. As described above, for prediction, each component of the vector 406 may be summed with a parameter a, as shown at 408, to compensate for the corresponding definition of C. An optional summation of the vector 406 with an offset vector b may also be involved in deriving a prediction of the block 18 based on the vector 406. As described above, it is possible that each component of the vector 406, and therefore each component of the summation of the vector 406, the all-a vector shown at 408, and the optional vector b, may directly correspond to a sample 104 of the block 18 and thus indicate a predicted value of the sample. It is also possible that only a subset of the samples 104 of the block are predicted in this manner, with the remaining samples of the block 18, such as 108, being derived by interpolation.

[0412] As mentioned above, there are various embodiments for setting a. For example, it can be the arithmetic mean of the components of vector 400. For that case, see FIG. 9a. The regular linear transformation T 403 can be as shown in FIG. 9a. i0 is a sample value vector and a predetermined component of vector 402, which are each replaced by a. However, as also mentioned above, there are other possibilities. However, it has also been mentioned above that the same thing can be realized differently as far as the representation of C is concerned. For example, the matrix-vector product 404 can be an actual calculation of a smaller matrix-vector product with a lower dimension in its actual calculation. Specifically, as mentioned above, by the definition of C, its entire i0-th column 412 is 0, so the actual calculation of the product 404 is

[0413]

number

[0414] It is possible that this can be done by a reduced version of vector 402 obtained from vector 402 by omission of , i.e. by multiplying this reduced vector 410 with a reduced matrix C′ obtained from C by excluding the i0-th column 412.

[0415] The weights of C or C′, i.e., the elements of this matrix, may be represented and stored in a fixed-point representation. However, these weights 414 may also be stored in a manner with different scales and / or offsets, as described above. The scale and offset may be defined for the entire matrix C, i.e., equal for all weights 414 of matrix C or matrix C′, or may be defined in a manner that is constant or equal for all weights 414 in the same row or column of matrix C and matrix C′, respectively. FIG. 10 illustrates in this regard that the matrix-vector product calculation, i.e., the result of the product, may actually be performed slightly differently, i.e., by shifting the multiplication with the scale towards vector 402 or 410, thereby reducing the number of further multiplications that must be performed. FIG. 11 illustrates the case of using one scale and one offset for all weights 414 of C or C′, such as those performed in equation (2) above.

[0416] According to an embodiment, the apparatus described herein for predicting a given block of a picture may be configured to use matrix-based intra-sample prediction including the following features:

[0417] The apparatus is configured to form a sample value vector pTemp[x] 400 from a plurality of reference samples 17. Assuming pTemp[x] is 2*boundarySize, for example by direct copying, or by subsampling, or by pooling, pTemp[x] may be filled with neighboring samples located above a given block, redT[x], x=0...boundarySize-1, followed by neighboring samples located to the left of the given block, redL[x], x=0...boundarySize-1 (e.g., if Transposed=0), or vice versa in the case of a transposed operation (e.g., if Transposed=1).

[0418] An input value p[x], x=0...inSize-1, is derived, i.e. the apparatus is configured to derive from the sample value vector pTemp[x] a further vector p[x] to which the sample value vector pTemp[x] is mapped by a predetermined regular linear transformation, or more specifically by a predetermined regular affine linear transformation, as follows: - If mipSizeId is equal to 2, the following applies: p[x]=pTemp[x+1]-pTemp[0] - Otherwise (mipSizeId is less than 2), the following applies: p[0]=(1<<(BitDepth-1))-pTemp[0] For x=1…inSize-1, p[x]=pTemp[x]-pTemp[0]

[0419] Here, the variable mipSizeId indicates the size of the predetermined block, i.e., according to the present embodiment, the regular transformation with which the further vector is derived from the sample value vector depends on the size of the predetermined block.

[0420] [Table 3]

[0421] can be given according to

[0422] If predSize indicates the number of predicted samples in a given block, 2*boundarySize indicates the size of the sample value vector and is related to inSize, i.e., the size of the further vector S, according to inSize=(2*boundarySize)-(mipSizeId==2)?1:0. More precisely, inSize indicates the number of components of the further vector that actually participate in the calculation. For smaller block sizes, inSize is as large as the size of the sample value vector, and for larger block sizes, it is one component smaller than the size of the sample value vector. In the former case, one component, i.e., the component corresponding to a given component of the further vector, may be ignored, since it does not actually need to be calculated, since the contribution of the corresponding vector component in the matrix-vector product to be calculated later will be zero anyway. The dependency on the block size may be eliminated in alternative embodiments, where only one of two options is necessarily used, i.e., independently of the block size (the option corresponding to mipSizeId less than 2, or the option corresponding to mipSizeId equal to 2).

[0423] In other words, a predetermined regular linear transformation is defined such that, for example, a predetermined component of the further vector p is a, while all other components correspond to the components of the sample value vector minus a, e.g., a=pTemp[0]. In the case of the first option corresponding to mipSizeId equal to 2, this is easily clear and only differentially formed components of the further vector are taken into account further. That is, in the first option, the further vector is actually {p[0...inSize];pTemp[0]}, where pTemp[0] is a, and the matrix-vector product, i.e. the part of the matrix-vector multiplication that is actually calculated to produce the multiplication result, is limited to only the inSize components of the further vector and the corresponding columns of the matrix, since the matrix has zero columns that do not require calculation. In other cases corresponding to mipSizeId less than 2, a = pTemp[0] is chosen because all components of the further vector except p[0], i.e., each of the other components p[x] of the further vector p (for x = 1...inSize-1) except for a given component p[0], are equal to the corresponding component of the sample value vector pTemp[x] minus a, whereas p[0] is chosen to be a constant minus a. Then, the matrix-vector product is calculated. This constant is chosen to be a representable value, i.e., 2 x-1 (i.e., 1<<(BitDepth-1)), where x represents the bit depth of the computational representation used. Note that if p[0] were instead selected to be pTemp[0], the calculated product would simply deviate from that calculated using p[0] as described above by a constant vector that can be taken into account when predicting the inside of the block based on the product, i.e., the prediction vector (p[0]=(1<<(BitDepth-1))-pTemp[0]). Thus, the value a is a predetermined value, for example, pTemp[0]. The predetermined value pTemp[0] is, in this case, for example, the component of the sample value vector pTemp that corresponds to the predetermined component p[0]. It may be the neighboring sample closest to the top-left corner of the predetermined block, above or to the left of the predetermined block.

[0424] For intra sample prediction processing according to predModeIntra, which specifies eg an intra prediction mode, the apparatus is configured, for example, to apply the following steps, eg, to perform at least the first step:

[0425] 1. The matrix-based intra prediction samples predMip[x][y], where x=0...predSize-1, y=0...predSize-1, are derived as follows: - The variable modeId is set equal to predModeIntra. The weight matrix mWeight[x][y], where x=0...inSize-1, y=0...predSize*predSize-1, is derived by invoking the MIP weight matrix derivation process with mipSizeId and modeId as input. The matrix-based intra prediction samples predMip[x][y], where x=0...predSize-1, y=0...predSize-1, are derived as follows:

[0426]

number

[0427] In other words, the apparatus is configured to calculate a matrix-vector product of the further vector p[i], or {p[i]; pTemp[0]} if mipSizeId is equal to 2, and a given prediction matrix mWeight, or a prediction matrix mWeight with additional zero weight lines corresponding to omitted components of p if mipSizeId is less than 2, to obtain a prediction vector, where the prediction vector has already been assigned to an array of block positions {x, y} distributed inside the given block, to result in an array predMip[x][y]. The prediction vectors correspond to the concatenation of the rows of predMip[x][y] and the columns of predMip[x][y], respectively.

[0428] According to one embodiment, or according to a different interpretation, the component

[0429]

number

[0430] is understood as a prediction vector, and the apparatus is configured to calculate, for each component of the prediction vector, the sum of the respective component and a, e.g., pTemp[0], when predicting samples of a given block based on the prediction vector.

[0431] The device optionally receives a prediction vector, e.g., predMip or

[0432]

number

[0433] When predicting the samples of a given block based on

[0434] 2. The matrix-based intra prediction samples predMip[x][y], where x=0...predSize-1, y=0...predSize-1, are clipped, for example, as follows: predMip[x][y]=Clip1(predMip[x][y])

[0435] 3. When isTransposed is equal to TRUE, the predSize by predSize array predMip[x][y], where x=0...predSize-1, y=0...predSize-1, is transposed, for example, as follows: predTemp[y][x]=predMip[x][y] predMip=predTemp

[0436] 4. The predicted samples predSamples[x][y], where x=0...nTbW-1, y=0...nTbH-1, are derived, for example, as follows: - If nTbW, which specifies the width of the transform block, is greater than predSize, or nTbH, which specifies the height of the transform block, is greater than predSize, then the MIP prediction upsampling process is invoked with inputs: input block size predSize, matrix-based intra prediction samples predMip[x][y], x = 0... predSize-1, y = 0... predSize-1, transform block width nTbW, transform block height nTbH, top reference samples refT[x], x = 0... nTbW-1, and left reference samples refL[y], y = 0... nTbH-1; and the output is the predicted sample array predSamples. - Otherwise, predSamples[x][y], for x=0...nTbW-1, y=0...nTbH-1, is set equal to predMip[x][y].

[0437] In other words, the apparatus is configured to predict the samples predSamples of a given block based on the prediction vector predMip.

[0438] 8. Using block / matrix-based intra-prediction modes with other intra-prediction modes The following description again presents possibilities for combining block / matrix-based prediction with other intra-prediction modes, which serves as a further indication of the possibilities based on which the embodiments described in the following sections may be implemented.

[0439] Please note that in the following, the term block-based intra prediction is used to denote an intra prediction mode that may be embodied in or equivalent to that indicated by the above ALWIP.

[0440] 12 relates to a decoder and an encoder that support intra prediction for decoding / encoding a predetermined block 18, and various intra prediction modes are supported. There is an angular intra prediction mode 500 in which reference samples 17 in the neighborhood of the predetermined block 18 are used accordingly to fill the predetermined block 18 and obtain an intra prediction signal for the predetermined block 18. Specifically, the reference samples 17 aligned along the boundary of the predetermined block 18, such as along the top and left edges of the predetermined block 18, represent picture content to be extrapolated or copied into the interior of the predetermined block 18 along a predetermined direction 502. Before extrapolation or copying, the picture content represented by the neighboring samples 17 may be subjected to interpolation filtering, or in other words, may be derived from the neighboring samples 17 by interpolation filtering. The angular intra prediction modes 500 differ from each other in the intra prediction direction 502. Each angular intra-prediction mode 500 may have an associated index, and the association of the indices to the angular intra-prediction modes 500 may be such that the directions 500 rotate monotonically clockwise or counterclockwise when ordering the angular intra-prediction modes 500 according to the associated mode index.

[0441] Non-Angular intra-prediction modes are also possible. For example, 504 illustrates a Planar intra-prediction mode in FIG. 12 that may be optionally included in the set 508, according to which a two-dimensional linear function defined by a horizontal tilt, a vertical tilt, and an offset is derived based on neighboring samples 17, and this linear function defines a predicted sample value for a given block 18. The horizontal tilt, the vertical tilt, and the offset are derived based on neighboring samples 17. According to an embodiment, the first set 508 of intra-prediction modes includes the Planar intra-prediction mode 504.

[0442] Specific non-angular intra-prediction modes, i.e., DC modes, included in the set 508 are indicated at 506, where a single value, a pseudo DC value, is derived based on neighboring samples 17, and this single DC value is attributed to all samples of a given block 18 to obtain an intra-prediction signal. Two examples for non-intra-prediction modes are shown, but there may be only one or more than two.

[0443] The intra-prediction modes 500, 504, and 506 form a set 508 of intra-prediction modes supported by the encoder and decoder that compete with block-based intra-prediction modes in a rate-distortion optimization sense, generally denoted using reference numeral 510, examples of which are described above using the abbreviation ALWIP. As described above, according to these block-based intra-prediction modes 510, a matrix-vector product 520 is performed between a vector 514 derived from neighboring samples 17 of one edge and a predetermined prediction matrix 516 of the other edge. The result of the multiplication 520 is a prediction vector 518 that is used to predict samples of a given block 18. The block-based intra-prediction modes 510 differ from each other in the prediction matrix 516 associated with each mode.

[0444] Therefore, in brief summary, the encoder and decoder according to the embodiments described herein include a set of intra prediction modes 508, i.e., a first set of intra prediction modes, and a set of block-based intra prediction modes 520, i.e., a second set of matrix-based intra prediction modes, which compete with each other.

[0445] According to an embodiment of the present application, a given block 18 is coded / decoded using intra prediction in the following manner. Specifically, first, a set selection syntax element 522 indicates whether the given block 18 should be predicted using one of the set of intra prediction modes 508 or one of the modes in the set of block-based intra prediction modes 520. If the set selection syntax element indicates that the given block 18 should be predicted using one of the modes in the set 508, i.e., the first set of intra prediction modes, a list 528 of most likely candidates in the set 508 is constructed / formed in the decoder and encoder based on the intra prediction modes used to predict neighboring blocks, i.e., the neighboring blocks 18 exemplarily shown at 524 and 526. The neighboring blocks 524 and 526 may be determined relative to the position of the given block 18 in a predetermined manner, such as by determining neighboring blocks that overlap with some neighboring samples of the block 18, such as samples above the top-left border sample of the block 18, and the block 526 including a sample to the left of the immediately above-mentioned corner sample. Of course, this is just an example. The same applies to the number of neighboring blocks used for mode prediction, which is not limited to two for all embodiments. More than two or only one may be used. If any of these blocks 524 and 526 is missing, a default intra-prediction mode may be used as a substitute for the intra-prediction mode of the missing neighboring block. The same may apply if any of blocks 524 and 526 is coded / encoded using an inter-prediction mode, such as by motion-compensated prediction.

[0446] The list of modes in the set 508 of most likely intra-prediction modes, i.e., list 528, is constructed as follows. The list length of list 528, i.e., the number of most likely modes in the list, may be constant by default. This length may be four, as shown in FIG. 12, or may be different from four, such as five or six. The latter case applies to a specific example described later in this specification. An index in the data stream, which will be described later, may indicate one mode in list 528 to be used for a given block 18. The indexing is performed along a list order or ranking 530, where the list index is of a variable length, for example, coded so that the length of the index monotonically increases along the order 530. Therefore, it is worthwhile to first fill list 528 with only the most likely modes in set 508, placing more likely modes upstream along the order 530 than modes that are less likely to be appropriate for block 18. The modes in list 528 are derived based on the modes used for blocks 524 and 526, i.e., neighboring blocks near the given block 18. If any of blocks 524 and 526 are intra-predicted using a block-based mode in set 520, then the aforementioned mapping of such "ALWIP" or block-based modes 510 to modes in set 508, e.g., non-ALWIP modes, is used. The latter mapping may, for example, map a majority (i.e., more than half) of the block-based modes 510 to DC mode 506 (or either DC 506 or Planar mode 504).

[0447] According to one embodiment, the list 528 of most likely intra-prediction modes is filled with the planar intra-prediction mode 504 in a manner independent of the intra-prediction modes with which neighboring blocks are predicted. Thus, for example, only the DC intra-prediction mode 506 and the angular intra-prediction mode 500 are filled in the list 528 depending on the intra-prediction modes used for predicting the neighboring blocks 524 and 526. The planar intra-prediction mode 504 is, for example, placed in the first position in the list 528 of most likely intra-prediction modes independent of the intra-prediction modes with which the neighboring blocks 524 and 526 are predicted.

[0448] In a manner exemplarily shown in more detail below, the list 528 of most likely intra-prediction modes is constructed in such a way that the DC intra-prediction mode 506 is not present in the list 528 if the neighboring blocks 524 and 526 are predicted only by any of the angular intra-prediction modes 500. The DC intra-prediction mode 506 is not present in the list 528 of most likely intra-prediction modes if either the neighboring blocks 524 or 526 are predicted by any of the angular intra-prediction modes 500 and / or if both the neighboring blocks 524 and 526 are predicted by any of the angular intra-prediction modes 500. According to embodiments described herein below, for example, the list 528 is populated with a DC mode 506 only if one of the following conditions is true for all neighboring blocks 524 and 526: all neighboring blocks 524 and 526 are coded using one of the non-Angular intra-prediction modes 504 and 506, or all neighboring blocks 524 and 526 are intra-predicted using one of the block-based intra-prediction modes 510, which is mapped to one of the non-Angular intra-prediction modes 504 and 506 by the aforementioned mapping of the block-based intra-prediction modes 510 to modes in the set 508. Only then is the DC intra-prediction mode 506 placed in the list 528. In that case, as may be seen in subsequent examples, the DC intra-prediction mode 506 may be placed before any of the angular intra-prediction modes 500 in the order 530.

[0449] In other words, the list of most likely intra-prediction modes 528 is populated with a DC intra-prediction mode 506, for example, for each of neighboring blocks 524 and 526, only if the respective neighboring block is predicted using any of at least one non-Angular intra-prediction mode 504 and 506 together with a first set 508 that includes the DC intra-prediction mode 506, or if the respective neighboring block is predicted using any of the block-based intra-prediction modes 510 that are mapped to any of the at least one non-Angular intra-prediction mode 500 by a mapping from the second set 520 of block-based intra-prediction modes 510 to the intra-prediction modes in the first set 508 used to form the list of most likely intra-prediction modes 528. The DC intra-prediction mode 506, for example, is placed before any angular intra-prediction mode 500 in the list of most likely intra-prediction modes 528.

[0450] Thus, resuming the description of how a given block 18 is coded into data stream 12, if set selection syntax element 522 indicates that the given block 18 should be coded with any mode in first set 508, data stream 12 optionally includes an MPM syntax element 532 indicating whether the intra-prediction mode to be used for the given block 18 is in list 528, and if so, data stream 12 includes an MPM list index 534 that points into list 528, indicating the mode to be used for the given block 18 in list 528, i.e., the predetermined intra-prediction mode, by indexing the modes along order 530. However, if a mode in set 508 is not in list 528, as indicated by MPM syntax element 532, data stream 12 includes a further syntax element 536 for block 18 that indicates which mode in set 508, i.e., the predetermined intra-prediction mode, should be used for block 18. Further syntax element 536 may indicate the mode in some manner by simply distinguishing modes in set 508 that are not included in list 528 .

[0451] In other words, an apparatus for decoding a given block 18 is configured, for example, if the set selection syntax element 522 indicates that the given block 18 should be predicted using one of the first set of intra-prediction modes 508, to derive from the data stream an MPM syntax element 532 indicating whether the given intra-prediction mode of the first set of intra-prediction modes 508 is in the list of most likely intra-prediction modes 528. If the MPM syntax element 532 indicates that the given intra-prediction mode of the first set of intra-prediction modes 508 is in the list of most likely intra-prediction modes 528, the apparatus is configured, for example, to perform formation of the list of most likely intra-prediction modes 528 based on the intra-prediction modes with which nearby neighboring blocks 524, 526 of the given block 100 are predicted, and to perform derivation from the data stream 12 of an MPM list index 534 that points to the given intra-prediction mode of the list of most likely intra-prediction modes 528. If the MPM syntax element 532 from the data stream 12 indicates that the given intra-prediction mode of the first set of intra-prediction modes 508 is not in the list of maximum-likelihood intra-prediction modes 528, the apparatus is configured to derive a further list index 536 from the data stream that indicates the given intra-prediction mode from the first set of intra-prediction modes. Thus, based on the MPM syntax element 532, the data stream 12 includes either the MPM list index 534 or the further list index 536 for prediction of the given block 18.

[0452] By eliminating situations in which list 528 includes DC intra-prediction modes 506, the following advantages are achieved. Specifically, the inventors of the present application discovered that “consuming” a valuable list position in list 528 with a DC intra-prediction mode 506 in set 508 for coding / decoding a given block 18 that should use one of the intra-prediction modes in set 508, as indicated by syntax element 522, i.e., the set selection syntax element, would adversely affect coding efficiency because such a DC intra-prediction mode 506 in set 508 would anyway conflict with block-based intra-prediction mode 510. Therefore, “consuming” a list position in list 528 with such a DC intra-prediction mode 506 in set 508 increases the probability of a situation in which the intra-prediction mode that should ultimately be used for given block 18, i.e., the given intra-prediction mode, is not in list 528, and therefore syntax element 536, i.e., an additional list index, needs to be transmitted in data stream 12.

[0453] Specifically, since syntax element 522 already indicates whether block 18 should be predicted using any of the modes in set 508 or any of the block-based modes 510 of set 520, if syntax element 522 indicates that a mode in set 508 would be preferred for block 18 and therefore that block-based mode 510 should not be used for block 18, it appears that the probability that DC prediction mode 506 in set 508 may be appropriate for block 18 is so low that the presence of DC prediction modes in list 528 should be limited to a very limited set of modes used for neighboring blocks 524 and 526, i.e., the set presented above.

[0454] In other cases, i.e., if the set selection syntax element 522 indicates that a given block 18 should be predicted using one of the block-based intra-prediction modes 510, the coding of the block 18 into, and decoding therefrom, the data stream 12 may be performed in the manner presented above. To this end, indexing may be used to index a selected one of the block-based intra-prediction modes 510 within the set 520, i.e., the second set of block-based intra-prediction modes, or to indicate which of the block-based intra-prediction modes should be used. Further MPM syntax element 538 may indicate whether indexing is performed by index 540, i.e., by a further MPM list index indicating the block-based intra-prediction mode 510 to be used for block 18 in a list 542 of maximum likelihood block-based intra-prediction modes 510, i.e., by indexing along list order 544, or whether the block-based intra-prediction mode 510 to be used for block 18 is indicated by a further syntax element 546, i.e., by a yet further list index indicating the block-based intra-prediction mode 510 in set 520, which may, for example, only distinguish modes 510 in set 520 that are not already included in list 542. List construction of list 542 may be performed based on the modes with which blocks 524 and 526 are predicted. If either of blocks 524 and 526 is unavailable because it is outside the picture or because it is inter-predicted, a default intra-prediction mode, such as one in set 508, may be used instead.For each block 524 and 526 that is intra predicted using a mode in set 508 rather than set 520, the aforementioned mapping from the modes of set 508 to the modes in set 520 is used to obtain an intra prediction mode 510, i.e., a predetermined block-based intra prediction mode, for the respective block, i.e., predetermined block 18, and list 542 is interpreted based on the obtained block-based intra prediction mode for blocks 524 and 526.

[0455] According to an embodiment, an apparatus for decoding a given block 18 is configured to, if the set selection syntax element 522 indicates that the given block 18 should not be predicted using one of the first set of intra-prediction modes 508, derive from the data stream 12 a further MPM syntax element 538 indicating whether the given block-based intra-prediction mode of the second set 520 of block-based intra-prediction modes 510 is in the list 542 of maximum likelihood block-based intra-prediction modes. If the further MPM syntax element 538 indicates that the given block-based intra-prediction mode of the second set 520 of block-based intra-prediction modes 510 is in the list 542 of most likely block-based intra-prediction modes, the apparatus is configured to form the list 542 of most likely block-based intra-prediction modes based on, for example, the intra-prediction modes with which nearby neighboring blocks 524, 526 of the given block 18 are predicted, and to derive from the data stream 12 a further MPM list index 540 that points to the given block-based intra-prediction mode on the list 542 of most likely block-based intra-prediction modes. If the further MPM syntax element 538 indicates that the given block-based intra-prediction mode of the second set 520 of block-based intra-prediction modes is not in the list 542 of most likely block-based intra-prediction modes, the apparatus is configured to derive from the data stream 12 a yet further list index 546 that indicates the given block-based intra-prediction mode in the second set 520 of block-based intra-prediction modes. Thus, based on the further MPM syntax element 538 , the data stream 12 includes either a further MPM list index 540 or a further list index 546 for prediction of a given block 18 .

[0456] 12 as corresponding to MPM syntax element 532, MPM list index 534, and further list index 536, it will be apparent that either further MPM syntax element 538 and an index associated with further MPM syntax element 538, e.g., further MPM list index 540 or further list index 546, are included in data stream 12 or MPM syntax element 532 and an index associated with MPM syntax element 532, e.g., MPM list index 534 or further list index 536. Whether this syntax element or index is included in data stream 12 depends, for example, on set selection syntax element 522.

[0457] An example of a syntax element portion of data stream 12 written as pseudocode may be as shown in Figures 19a to 19d, where reference symbols indicate which syntax elements correspond to syntax elements previously discussed.

[0458] The list construction of list 528 may be defined as follows, where candIntraPredModeA / B indicates the intra-prediction mode with which either of blocks 524 and 526 was predicted, such as A for block 524 and B for block 526, or indicates which mode from set 508 the intra-prediction mode is mapped to if the corresponding block 524 or 526 was intra-predicted using one of the block-based intra-prediction modes 510. INTRA_DC is used to indicate mode 506, and angular mode 500 is indicated by INTRA_ANGULAR#, where the numbers (#) order the angular modes as exemplarily shown above, i.e., in a manner such that the angular direction 502 monotonically decreases or increases with increasing numbers. This ordering among modes within set 508 may be as defined in a subsequent table, where INTAR_PLANAR indicates mode 504.

[0459] Note that in the above example, index 534 is actually distributed into syntax elements 534' and 534''. The former 534' is specific to the first position of list 528 in order 530, where, according to this example, INTRA_PLANAR mode 504 is necessarily located. The latter 534'' points to any of the subsequent positions in list 528, and as explained, only DC mode 506 is included in this explained special situation.

[0460] Furthermore, in the above example, if syntax element 522 indicates that any of the modes in set 508 is used, then a further syntax element is included in the data stream that parameterizes the intra-prediction modes in set 500 in some way. For example, syntax element 600 parameterizes or varies the region in which reference samples 17 are located based on which modes in set 508 intra-predict the interior of block 18, such as with respect to distance to the perimeter of block 18. Additionally or alternatively, syntax element 602 parameterizes or varies whether reference samples 17 are used to intra-predict the interior of block 18 globally or en block by en block by modes in set 508, or whether intra-prediction is performed in fragments or portions into which block 18 is subdivided, and the fragments or portions are sequentially intra-predicted so that prediction residuals coded into the data stream for one portion can serve to gather new reference samples for intra-predicting the next portion. The latter coding option, controlled by a syntax element, may only be available if syntax element 600 has a predetermined state (and the corresponding syntax element may only be present in the data stream) corresponding to the region where reference sample 17 resides, for example, bordering block 18. The portions may be defined by subdividing the block along a predetermined direction, such as either horizontally, so that the portions are the same height as block 18, or vertically, so that the portions are the same width as block 18. If partitioning is signaled as enabled, syntax element 604 may be present in the data stream, which controls which division direction is used.As can be appreciated, the position in list 528 reserved for INTRA_PLANAR mode may only be available for some parameterization of the mode according to the parameterization of the syntax element mentioned immediately above, such as only if syntax element 600 has a predetermined state corresponding to an area in which reference sample 17 is located, for example, bordering block 18, and / or only if partial intra prediction mode as signaled by syntax element 602 is not enabled.

[0461] All syntax elements shown in the table but not specifically mentioned above are optional and will not be discussed further herein. - If candIntraPredModeB is equal to candIntraPredModeA and candIntraPredModeA is greater than INTRA_DC, then candModeList[x], for x=0...4, is derived as follows: candModeList[0]=candIntraPredModeA candModeList[1]=2+((candIntraPredModeA+61)%64) candModeList[2]=2+((candIntraPredModeA-1)%64) candModeList[3]=2+((candIntraPredModeA+60)%64) candModeList[4]=2+(candIntraPredModeA%64) - Otherwise, if candIntraPredModeB is not equal to candIntraPredModeA and candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC, then the following is true: The variables minAB and maxAB are derived as follows: minAB=Min(candIntraPredModeA, candIntraPredModeB) maxAB=Max(candIntraPredModeA, candIntraPredModeB) - If candIntraPredModeA and candIntraPredModeB are both greater than INTRA_DC, candModeList[x], for x=0...4, is derived as follows: candModeList[0]=candIntraPredModeA candModeList[1]=candIntraPredModeB - If maxAB-minAB equals 1, the following holds: candModeList[2]=2+((minAB+61)%64) candModeList[3]=2+((maxAB-1)%64) candModeList[4]=2+((minAB+60)%64) Otherwise, if maxAB-minAB is ≥ 62, then the following is true: candModeList[2]=2+((minAB-1)%64) candModeList[3]=2+((maxAB+61)%64) candModeList[4]=2+(minAB%64) Otherwise, if maxAB - minAB equals 2, then the following is true: candModeList[2]=2+((minAB-1)%64) candModeList[3]=2+((minAB+61)%64) candModeList[4]=2+((maxAB-1)%64) - Otherwise, the following applies: candModeList[2]=2+((minAB+61)%64) candModeList[3]=2+((minAB-1)%64) (8-36) candModeList[4]=2+(((maxAB+61))%64) - Otherwise (candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC), candModeList[x], for x=0…4, is derived as follows: candModeList[0]=maxAB candModeList[1]=2+((maxAB+61)%64) candModeList[2]=2+((maxAB-1)%64) (8-41) candModeList[3]=2+((maxAB+60)%64) candModeList[4]=2+(maxAB%64) - Otherwise, the following applies: candModeList[0]=INTRA_DC candModeList[1]=INTRA_ANGULAR50 candModeList[2]=INTRA_ANGULAR18 candModeList[3]=INTRA_ANGULAR46 candModeList[4]=INTRA_ANGULAR54

[0462] [Table 4]

[0463] 9. Embodiments utilizing block / matrix-based intra-prediction modes along with other intra-prediction modes and utilizing secondary transforms The following description presents an embodiment for combining the use of a secondary transform for coding prediction residuals and block / matrix-based prediction with other intra-prediction modes. The above presentation of possibilities for matrix-based intra-prediction (ALWIP) and its combination with other intra-prediction modes shall serve as an example for implementing the embodiments described herein below. In Figure 12, all details about MPM list construction, for example, regarding the restriction to include DC mode in the MPM list, are optional.

[0464] As mentioned above, matrix-based intra prediction (MIP), also referred to herein as block-based intra prediction and ALWIP, generates an intra prediction signal on a rectangular block by performing matrix-vector multiplication, where the output of the matrix-vector multiplication may be considered to be a prediction signal on a downsampled block, and the input of the matrix-vector multiplication may be included in the downsampled boundary samples. If the output is considered to be a prediction signal on a downsampled block, this prediction signal needs to undergo an upsampling (or linear interpolation) stage before a final prediction signal is obtained.

[0465] On the other hand, for conventional intra prediction modes such as Planar mode 504, DC mode 506, and Angular mode 500, which are also referred to as modes of set 508 in the above description, a non-separable secondary transform (LFNST) is a tool used to transform the prediction residual corresponding to these intra prediction modes. Here, a set S of transform sets is given such that each conventional intra prediction mode is associated with one of these transform sets. Then, in a decoder, whether an LFNST should be applied to a given block can be extracted from the bitstream. If it should be applied, one transform set from set S is given depending on the intra prediction mode used on the current block, and if this transform set consists of more than one transform, which transform T of this set should be used can be extracted from the bitstream. Then, in a decoder, transform T is applied as a secondary transform T, which means that it is applied to a subset 622 of the residual transform coefficients 620 of the separable primary transform T, as shown in FIG. 14, for example.

[0466] The problem is that the aforementioned secondary transform Ts is defined a priori only for conventional intra-prediction modes: Providing a specific secondary transform Ts for each MIP mode 510 may be too expensive in terms of memory requirements for additionally storing the extra transforms.

[0467] Figure 13 shows a decoder that solves this problem. The decoder decodes a given block 18 of a picture using intra prediction. According to one embodiment, the encoder includes similar features and / or functionality to the decoder.

[0468] The decoder / encoder is configured to select (602) a predetermined intra-prediction mode 604 from a plurality of intra-prediction modes 600 including a first set of intra-prediction modes 508 and a second set 520 of matrix-based intra-prediction modes 510. This intra-mode selection 602 is performed by the decoder based on data stream 12, and the encoder is configured to signal the predetermined intra-prediction mode 604 in data stream 12. The intra-mode selection 602 may be performed as described with respect to FIG. 12 .

[0469] The first set 508 of intra-prediction modes includes a DC intra-prediction mode 506, an Angular prediction mode 500, and optionally a Planar intra-prediction mode 504. If the given intra-prediction mode 604 is a matrix-based intra-prediction mode 510 in the second set 520, the decoder / encoder is configured to use a matrix-vector product 512 of a vector 514 derived from reference samples 17 in the neighborhood of the given block 18 and a prediction matrix 516 associated with the respective matrix-based intra-prediction mode 510 to obtain a prediction vector 518, based on which samples of the given block 18 are predicted. Prediction of the given block 18 using the matrix-based intra-prediction mode 510 as the given intra-prediction mode 604 may be performed by the decoder / encoder according to the embodiments of Figures 6 to 11. The decoder / encoder is configured to derive a prediction signal 606 for the given block 18 using the given intra-prediction mode 604.

[0470] The decoder / encoder selects a set 612 of secondary transforms Ts, e.g., Ts, in a manner that depends on the given intra-prediction mode 604, such that the subset 610 is non-empty if the given intra-prediction mode 604 is included in the first set 508 of intra-prediction modes and if the given intra-prediction mode 604 is included in the second set 520 of matrix-based intra-prediction modes 510. (1) -Ts (N) from one or more quadratic transformations Ts(i1) -Ts (in) where i1 ranges from 1 to N and i n is in the range of i1 to N.

[0471] According to an embodiment, the decoder / encoder is configured to select (608) the subset 610 such that each secondary transform Ts of the set 612 of secondary transforms Ts is included in the subset 610 of one or more secondary transforms Ts selected for at least one of the intra prediction modes in the first set 508 and the second set 520. Thus, the subset 610 may be equal to the set 612 of secondary transforms. Such a subset may be selected for one or more matrix-based intra prediction modes 510. For one or more matrix-based intra prediction modes 510, it may be possible to select all secondary transforms Ts of the set of secondary transforms 612, where the subset 610 for the one or more matrix-based intra prediction modes 510 includes the secondary transforms Ts selectable for the intra prediction modes in the first set 508 of intra prediction modes. Thus, for at least one of the matrix-based intra prediction modes 510, no specific additional secondary transform is required in the set of secondary transforms 612.

[0472] According to an embodiment, the decoder / encoder may select a subset 610 of the secondary transform Ts selected for any matrix-based intra-prediction mode 510. matrix a subset of secondary transforms Ts, e.g., 610, selected for at least one intra prediction mode in the first set 508 that does not belong to the angular prediction mode 500; DC and / or 610 planar15 illustrates various possible subsets of secondary transforms Ts selected for any matrix-based intra-prediction mode 510. The decoder / encoder is configured to select 608 the subset 610 of secondary transforms Ts in such a manner that the subset 610 is included in the matrix-based intra-prediction mode 510. For each matrix-based intra-prediction mode 510, the decoder / encoder performs a first combination 611 of the subset 610 of secondary transforms for the matrix-based intra-prediction mode 510. matrix 610.

[0473] As shown in FIG. 15 , the selectable subsets 610 for one or more matrix-based intra-prediction modes 510 may be different from the selectable subsets for the DC intra-prediction mode 506, e.g., 610 DC1 Or 610 DC2 or a selectable subset for the Planar intra-prediction mode 504, e.g., 610 planar1 Or 610 planar2 The selectable subsets 610 for one or more matrix-based intra-prediction modes 510 may be equal to the subsets 610 matrix4 610, a subset of secondary transforms Ts selected for one intra prediction mode in the first set 508 that does not belong to the angular prediction mode 500, as indicated by DC One or more secondary transformations of, say, Ts (ax) From Ts (aY) It is possible to include only

[0474] The selectable subsets 610 for the matrix-based intra prediction modes 510 are the subsets 610 matrix1 Two or more subsets selectable for the DC intra prediction mode 506, such as 610 DC1 and 610 DC2 or a subset 610 matrix3 Two or more subsets selectable for the Planar intra prediction mode 504, such as 610 planar1 and 610planar2 may include one or more secondary transformations in

[0475] Further possible subsets 610 that can be selected for the matrix-based intra-prediction mode 510 are the subsets 610 matrix2 One or more subsets selectable for DC intra-prediction mode 506, such as 610 DC2 and one or more subsets selectable for the Planar intra-prediction mode 504, e.g., 610 planar1 may include one or more secondary transformations from

[0476] According to an embodiment, the decoder / encoder performs a first union 611 of the subset 610 of secondary transforms selected for the matrix-based intra-prediction mode 510. matrix and a subset of quadratic transforms selected for all angular intra prediction modes 610 angular The Second Union 611 angular The third union 611 is configured to select (608) a subset 610 in such a way that the intersection of DC is all the subsets 610 selected for the DC intra prediction mode 506 DC Including the Fourth Union 611 planar is all the subsets 610 selected for the Planar intra prediction mode 504 planar For example, because MIP mode 510 is non-directional like planar mode 504 and DC mode 506 due to the downsampling optionally applied to the reduced prediction signal 606, their prediction residuals 618 have more statistical similarity to the prediction residuals of DC mode 506 and planar mode 504 than to the prediction residuals of angular mode 500.

[0477] Further, as shown in Figure 13, the decoder is configured to derive (614) a transformed version 616 of the prediction residual for a given block 18 from the data stream encoded by the encoder into the data stream, where this transformed version is related to the spatial domain version 618 of the prediction residual for the given block 18 via a transform T defined by the concatenation of a primary transform Tp and a given secondary transform Ts from the subset 610 of secondary transforms. As shown in Figure 14, the encoder may be configured to apply a transform T to the subset 622 of coefficients 620 of the primary transform Tp when the given intra prediction mode is included in the first set 508 of intra prediction modes and when the given intra prediction mode is included in the second set 520 of matrix-based intra prediction modes 510. The decoder may apply an inverse function T of the transform T to obtain the spatial domain version 618 of the prediction residual for the given block 18. -1 The primary transformation may be, for example, a separable 2D transformation, and the secondary transformation may be, for example, a non-separable 2D transformation.

[0478] The decoder is configured to reconstruct (624) a given block 18 using the prediction signal 606 and the prediction residual 618 for the given block 18.

[0479] If the subset 610 of one or more secondary transforms Ts includes more than one secondary transform Ts, the decoder may be configured to select a given secondary transform Ts from the subset 610 of one or more secondary transforms in response to a secondary transform indication syntax element transmitted in data stream 12 for a given block. The secondary transform indication syntax element may be an index that points into the selected subset 610 of one or more secondary transforms. In this case, the encoder may be configured to transmit the secondary transform indication syntax element in data stream 12.

[0480] According to one embodiment, if the dimensions of a given block 18 satisfy a predetermined criterion, the decoder is configured to infer that the transform T, by which the transformed version 616 of the prediction residual for the given block 18 is associated with the spatial domain version 618 of the prediction residual for the given block 18, is the primary transform Tp. Otherwise, the transform T may be a concatenation of the primary transform Tp and a predetermined secondary transform Ts. The decoder makes available the second set 520 of matrix-based intra-prediction modes 510 for the selection 602 of a given intra-prediction mode 604, independent of whether the dimensions of the given block 18 satisfy the predetermined criterion. The predetermined criterion for the dimensions of the given block 18 may relate only to the selection 608 of the subset and not to the selection 602 of the intra-mode. For example, there are block dimensions for which MIP / ALWIP 510 is available but LFNST, i.e., the selection 608 of the subset, is not available. For blocks 18 of such dimensions, the secondary transform indication syntax element does not need to be transmitted in or read from the data stream. The predetermined criterion is met, for example, if the dimension is below a predetermined threshold. For certain small blocks 18, it may not be beneficial, which means that for the latter shapes, the cost of additional signaling to signal whether LFNST should be applied to the block in MIP mode is, on average, higher than the benefit obtained by allowing LFNST transforms for MIP mode.

[0481] According to an embodiment, as shown in FIG. 16 , a decoder is configured to read a non-zero zone indication transmitted in the data stream 12 for each given block 18, which indicates a non-zero transform domain area 623 within the transformed version 616 of the prediction residual for the given block 18. All non-zero coefficients are located exclusively in the non-zero transform domain area 623. The non-zero zone indication is transmitted in the data stream 12, for example, by an encoder. The decoder / encoder is configured to decode / encode coefficients within the non-zero transform domain area 623 from / to the data stream 12. An LP syntax element may operate as the non-zero zone indication. The last non-zero coefficient position along the scan path from the DC coefficient position to the opposite, or highest frequency, coefficient position is indicated by the LP syntax element herein. The LP syntax element may essentially be a measure of the expected count of non-zero coefficients within the non-zero transform domain area 623.

[0482] According to an embodiment, the decoder is configured to infer that the transformation T by which the transformed version 616 of the prediction residual for a given block 18 is associated with the spatial domain version 618 of the prediction residual of the given block 18 is a linear transformation Tp depending on the extension and / or position of the non-zero transformation domain area 623 that meets a first predetermined criterion and / or the number of non-zero coefficients within the non-zero transformation domain area 623 that meets a second predetermined criterion.

[0483] The first predetermined criterion is, for example, satisfied when the non-zero transform domain area 623 does not exclusively include the subset 622 of coefficients of the primary transform Tp to which the secondary transform Tp is applied by concatenation. This is based on the idea that the secondary transform Ts should include all non-zero coefficients of the transform coefficients of the primary transform Tp. If a non-zero coefficient is outside the subset 622 of coefficients of the primary transform Tp, it may not be advantageous to apply the predetermined secondary transform, and for this reason, the decoder infers that the transform T is the primary transform Tp. The decoder performs subset selection 608, and if the non-zero transform domain area 623 is entirely located within the subset 622 of coefficients of the primary transform Tp to which the secondary transform Tp is applied by concatenation, the decoder can apply a transform T defined by concatenation of the primary transform Tp and the predetermined secondary transform Ts in the subset 610 of the secondary transform applied to the subset 622 of coefficients of the primary transform Tp to obtain a spatial domain version 618 of the prediction residual of the given block 18; see, for example, FIG. 16 .

[0484] The second predetermined criterion may be satisfied, for example, when the number of non-zero coefficients in the non-zero transform domain area 623 is below a predetermined threshold. This is based on the idea that if the number of non-zero coefficients in the non-zero transform domain area 623 is below the predetermined threshold, there is no need to further reduce it by the secondary transform Ts. If the primary transform Tp has a small number of non-zero coefficients, it may not be advantageous to additionally apply the predetermined secondary transform, and for this reason, the decoder infers that the transform T is the primary transform Tp. If the number of non-zero coefficients is below the predetermined threshold, the cost of additional signaling for the predetermined secondary transform is greater than the improvement in coding efficiency achieved by the predetermined secondary transform. The decoder may perform subset selection 608 and, if the number of non-zero coefficients in the non-zero transform domain area 623 is equal to or greater than the predetermined threshold, apply a transform T defined by the concatenation of the primary transform Tp and the predetermined secondary transform Ts in the subset 610 of secondary transforms applied to the subset 622 of coefficients of the primary transform Tp to obtain a spatial domain version 618 of the prediction residual of the given block 18.

[0485] According to an embodiment, as described with respect to, for example, Figure 12, the decoder / encoder is configured to derive from / encode into data stream 12 a set selection syntax element 522 that indicates whether a given block 18 should be predicted using one of the first set of intra-prediction modes 508. If the set selection syntax element 522 indicates that the given block 18 should be predicted using one of the first set of intra-prediction modes 508, the decoder / encoder is configured to form a list 528 of most likely intra-prediction modes based on the intra-prediction modes with which nearby neighboring blocks 524 and 526 of the given block 18 are predicted, and to derive from / signal into data stream 12 an MPM list index 534 that points to the given intra-prediction mode 604 on the list 528 of most likely intra-prediction modes. If the set selection syntax element 522 indicates that a given block 18 should be predicted using one of the first set 508 of intra-prediction modes, the decoder / encoder is configured to derive from / encode into the data stream a further index 540 and / or 546 indicating a given intra-prediction mode 604 within the second set 520 of matrix-based intra-prediction modes 510.

[0486] The decoder and encoder described with respect to FIG. 13 may include additional features and / or functionality as described with respect to FIG.

[0487] When the predetermined intra prediction mode 604 for a given block 18 is a matrix-based intra prediction mode 510 in the second set 520, an apparatus for decoding the given block 18, i.e., a decoder according to FIG. 13, and / or an apparatus for encoding the given block 18, i.e., an encoder according to FIG. 13, may include one or more of the following features.

[0488] According to an embodiment, the device is configured to form a sample value vector from among a plurality of reference samples 17, for example the sample value vector 400 as described with respect to one of the embodiments of Figures 6 to 9, and to derive the vector 514 from the sample value vector, such that the sample value vector is mapped by a predetermined regular linear transformation to the vector 514. In this case, the vector 514 may be understood as a further vector. The vector 514 is determined and / or defined, for example, as described for the further vector 402 with respect to one of the embodiments of Figures 8 to 11.

[0489] According to an embodiment, the apparatus is configured to form a sample value vector from the plurality of reference samples 17 by, for each component of the sample value vector, employing one reference sample from the plurality of reference samples as the respective component of the sample value vector, and / or by averaging two or more components of the sample value vector to obtain the respective component of the sample value vector.

[0490] A number of reference samples 17 are arranged in the picture, for example along the outer edge of a given block 18 .

[0491] The regular linear transformation is defined, for example, such that a predetermined component of vector 514, e.g., a further vector, is a, and each of the other components of vector 514, excluding the predetermined component, is equal to the corresponding component of the sample value vector minus a. The value a is, for example, a predetermined value 1400.

[0492] According to one embodiment, the predetermined value 1400 is one of an average, such as an arithmetic mean or a weighted mean, of the components of the sample value vector, a default value, a value signaled in the data stream in which the picture is coded, and a component of the sample value vector corresponding to the predetermined component.

[0493] A regular linear transformation is defined, for example, such that a predetermined component of vector 514, e.g., a further vector, is a, and each of the other components of vector 514 except for the predetermined component is equal to the corresponding component of the sampled vector minus a, where a is the arithmetic mean of the components of the sampled vector.

[0494] The regular linear transform is defined, for example, such that a predetermined component of vector 514, e.g., a further vector, is a, and each of the other components of vector 514 excluding the predetermined component is equal to a corresponding component of the sample-valued vector minus a, where a is the component of the sample-valued vector that corresponds to the predetermined component. The apparatus is configured, for example, to include a plurality of regular linear transforms, each associated with one component of vector 514, to select the predetermined component from the components of the sample-valued vector, and to use the regular linear transform from the plurality of regular linear transforms associated with the predetermined component as the predetermined regular linear transform.

[0495] According to an embodiment, the matrix elements of the prediction matrix 516 in a column of the prediction matrix 516 that corresponds to a predetermined element of the vector 514, e.g. the further vector, are all 0. The apparatus is configured to calculate the matrix-vector product 512 by performing a multiplication by calculating the matrix-vector product 512 of a reduced prediction matrix obtained from the prediction matrix 516 by excluding that column and the further vector 410 obtained from the vector 514 by excluding the predetermined element, as shown in Figure 9c.

[0496] According to an embodiment, the device is configured to, when predicting samples of a given block 18 based on the prediction vector 518, calculate for each component of the prediction vector 518 the sum of the respective component and a.

[0497] The matrix obtained by adding each matrix element of the prediction matrix 516 in a column of the prediction matrix 516 that has a value of 1 corresponding to a given element of the vector 514, e.g., a further vector (i.e., adding matrix C405 and matrix M1300 shown in FIG. 9a, resulting in matrix B in FIG. 8), multiplied by a regular linear transformation, corresponds, for example, to a quantized version of the machine learning prediction matrix (i.e., prediction matrix A1100 shown in FIG. 8).

[0498] According to one embodiment, the apparatus is configured to calculate the matrix vector product 512 using fixed-point arithmetic.

[0499] According to one embodiment, the apparatus is configured to calculate the matrix vector product 512 without floating-point arithmetic operations.

[0500] According to one embodiment, the device is configured to store a fixed-point representation of the prediction matrix 516 .

[0501] According to an embodiment, the apparatus is configured to represent a prediction matrix 516 using the prediction parameters and to calculate a matrix-vector product 512 by performing multiplications and additions on the components of the vector 514, e.g., further vectors, and the prediction parameters and intermediate results obtained therefrom, where the absolute values ​​of the prediction parameters are representable by an n-bit fixed-point representation, where n is less than or equal to 14, or alternatively less than or equal to 10, or alternatively less than or equal to 8. This can be performed similarly or as described in Figure 10 or 11.

[0502] The prediction parameters include, for example, weights each associated with a corresponding matrix element of the prediction matrix 516 .

[0503] The prediction parameters further include, for example, one or more scaling coefficients each associated with one or more corresponding matrix elements of the prediction matrix 516 for scaling weights associated with the one or more corresponding matrix elements of the prediction matrix 516, and / or one or more offsets each associated with one or more corresponding matrix elements of the prediction matrix 516 for offsetting weights associated with the one or more corresponding matrix elements of the prediction matrix 516.

[0504] According to an embodiment, when predicting samples of a given block 18 based on a prediction vector 518, the device is configured to use interpolation to calculate at least one sample position of the given block 18 based on the prediction vector 518 with which each of its components is associated with a corresponding position within the given block 18.

[0505] The decoder and encoder described with respect to FIG. 13 may include additional features and / or functionality as described with respect to one or more of the embodiments of FIGS.

[0506] Therefore, the solution provided by the present invention is to associate one specific set of transforms 160 in a given transform set S612 with each MIP mode 510, the latter originally defined for conventional intra-prediction modes, i.e., the intra-prediction modes in the first set 508. One specific way to do this is for all MIP modes 510 to use the LFNST transform, i.e., the secondary transform, originally designed for planar mode 504 and DC mode 506. For example, because MIP modes 510 are non-directional like planar mode 504 and DC mode 506, due to the downsampling optionally applied to the downscaled prediction signal, their prediction residuals will have a greater statistical similarity to the prediction residuals of DC mode 506 and planar mode 504 than to the prediction residuals of angular mode 500.

[0507] Allowing LFNST for MIP 510 in the manner described above, i.e., allowing the use of a secondary transform, may prove beneficial in terms of coding efficiency for some block shapes and not for other block shapes, meaning that for the latter shapes, the additional signaling cost of signaling whether LFNST should be applied for block 18 in MIP mode 510 is, on average, greater than the benefit obtained by allowing the LFNST transform for MIP mode 510. Therefore, in one embodiment of the present invention, the above-described combination of LFNST and MIP 510 is allowed only for a subset of all block shapes for which such a combination is, in principle, possible.

[0508] FIG. 17 shows a method 6000 for decoding a predetermined block (18) of a picture using intra prediction, which includes a step of selecting (602) a predetermined intra prediction mode (604) from a plurality of intra prediction modes (600) including a first set (508) of intra prediction modes including a DC intra prediction mode (506), an Angular prediction mode (500), and optionally a Planar intra prediction mode (504), based on a data stream, and a second set (520) of matrix-based intra prediction modes (510), in which, according to each of the matrix-based intra prediction modes, a matrix-vector product (512) of a vector (514) derived from reference samples (17) in the vicinity of the predetermined block and a prediction matrix (516) associated with the respective matrix-based intra prediction mode is used to obtain a prediction vector (518), and samples of the predetermined block are predicted based on the prediction vector. A prediction signal (606) for a given block is derived (6100) using a given intra-prediction mode, and a subset (610) of one or more secondary transforms in a set (612) of secondary transforms is selected (608) in a manner dependent on the given intra-prediction mode such that the subset (610) is non-empty when the given intra-prediction mode is included in a first set (508) of intra-prediction modes and when the given intra-prediction mode is included in a second set (520) of matrix-based intra-prediction modes (510). The method 6000 includes a step of deriving (614) a transformed version (616) of the prediction residual for a given block (18) related to a spatial domain version (618) of the prediction residual for the given block via a transform (T) defined by concatenation of a primary transform (Tp) and a given secondary transform (Ts) from a subset (610) of secondary transforms applied to a subset (622) of coefficients of the primary transform (620) from the data stream when the given intra prediction mode is included in a first set (508) of intra prediction modes and when the given intra prediction mode is included in a second set (520) of matrix-based intra prediction modes (510).Additionally, the method 6000 includes reconstructing (624) a given block (18) using the prediction signal and the prediction residual for the given block.

[0509] Figure 18 shows a method 7000 for encoding a given block (18) of a picture using intra prediction, comprising a step (602) of selecting a given intra prediction mode (604) from a plurality of intra prediction modes (600) including a first set (508) of intra prediction modes including a DC intra prediction mode (506) and an Angular prediction mode (500), and a second set (520) of matrix-based intra prediction modes (510), according to each of which a matrix-based intra prediction mode is used to obtain a prediction vector (518) based on which samples of the given block are predicted. In addition, the method 7000 includes a step (7100) of signaling a predetermined intra-prediction mode (604) in a data stream, a step (7200) of deriving a prediction signal (606) for a predetermined block using the predetermined intra-prediction mode, and a step (608) of selecting a subset (610) of one or more secondary transforms from a set (612) of secondary transforms in a manner dependent on the predetermined intra-prediction mode such that the subset (610) is non-empty when the predetermined intra-prediction mode is included in a first set (508) of intra-prediction modes and when the predetermined intra-prediction mode is included in a second set (520) of matrix-based intra-prediction modes (510).The method includes a step of encoding (614) into a data stream a transformed version (616) of a prediction residual for a given block (18) that is associated with a spatial domain version (618) of a prediction residual for the given block via a transform (T) defined by a concatenation of a primary transform (Tp) and a given secondary transform (Ts) from a subset (610) of secondary transforms applied to a subset (622) of coefficients (620) of the primary transform, when the given intra prediction mode is included in a first set (508) of intra prediction modes and when the given intra prediction mode is included in a second set (520) of matrix-based intra prediction modes (510), the transformed version being related to a spatial domain version (618) of a prediction residual for the given block, the given block being reconstructable using a prediction signal and the prediction residual for the given block (18) (624).

[0510] (References)

[0511] Further Embodiments and Examples In general, the examples may be implemented as a computer program product with program instructions that are operable to perform one of the methods when the computer program product is executed on a computer. The program instructions may be stored on, for example, a machine-readable medium.

[0512] Other examples comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.

[0513] In other words, an example of a method is therefore a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.

[0514] A further example of a method is therefore a data carrier medium (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-transitory, rather than a signal, which is intangible and transitory.

[0515] A further example of a method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. A data stream or a sequence of signals may for example be transmitted via a data communication connection, for example via the Internet.

[0516] Further examples include a processing means, for example a computer, or a programmable logic device adapted to perform one of the methods described herein.

[0517] A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0518] Further examples include an apparatus or system that transfers (e.g., electrically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.

[0519] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods may be performed by any suitable hardware apparatus.

[0520] The above examples are merely illustrative of the principles discussed above. It will be understood that modifications and variations of the arrangements and details described herein will be apparent. It is therefore the intention to be limited only by the scope of the following claims and not by the specific details presented for purposes of illustration and description of the examples herein.

[0521] One or more elements having equal or equivalent functions are designated in the following description by equal or equivalent reference numerals, even if they are present in different drawings. [Explanation of symbols]

[0522] 10 Pictures 12 Data Stream 14 Encoder 16 videos 17 neighboring blocks, neighboring samples 18 blocks, predetermined blocks 19 ALWIP Conversion 20 Coding Order 22 Subtractor 24 Predictive Signals 26 Prediction residual signal 28 Prediction Residual Encoder 30 Quantizer 32 conversion stages 34 Quantized prediction residual signal 36 Prediction residual reconstruction stage 38 Inverse quantizer 40 Inverse converter 42 Adder 44 Predictor 46 In-loop filter 54 Decoder 56 Entropy Decoder 102 Reduced Set 104 sample values, samples 108 samples 110 Group 112 samples 118 samples 119 samples 120 Groups 150 prescribed ingredients 156 Residual Provider 400 sample value vector 402 More Vectors 403 Regular Linear Transformations 404 Matrix Vector Multiplication 405 Predetermined prediction matrix 406 Prediction Vector 407 Matrix Vector Product 408 Further Offsets 409 Vector 410 Yet Another Vector Column 412 414 Weights, Matrix Elements 500 Angular Intra Prediction Mode 502 Angle Direction 504 INTRA_PLANAR mode, Planar intra prediction mode 506 DC mode, DC intra prediction mode 508 First Set 510 Matrix-based intra prediction modes 512 Matrix Vector Product 514 Vector 516 Prediction Matrix 518 Prediction Vectors 520 Second Set, Multiplication 522 Set Selection Syntax Elements 524 neighboring blocks 526 neighboring blocks 528 List 530 order 532 MPM Syntax Elements 534 MPM List Index 536 Further syntax elements 538 Further MPM syntax elements 540 Index, MPM List Index 542 List 544 List Order 546 Further syntax elements 600 Syntax Elements 602 Syntax Elements 604 Predetermined intra prediction mode 606 Predictive Signal 610 subset 611 First Union, Second Union 612 Gathering 616 converted version 618 Spatial Domain Version 620 coefficient 622 subset 623 Non-zero Transform Domain Area 1100 Machine Learning Prediction Matrix 1110 Offset B 1200 Further Matrix B 1300 Integer matrix, matrix M(i0) 1310 Matrix Vector Product 1400 given value 1500 prescribed ingredients

Claims

[Claim 1] 1. A method for decoding pictures from a data stream, comprising: selecting, for a block of the picture, an intra-prediction mode from a first set of intra-prediction modes or a second set of intra-prediction modes based on an instruction included in a data stream, wherein the first set of intra-prediction modes comprises at least one of an angular prediction mode, a planar intra-prediction mode, and a DC intra-prediction mode, and the second set of intra-prediction modes comprises at least one matrix-based intra-prediction mode; deriving a prediction signal for the block using the selected intra-prediction mode; and selecting a subset of secondary transforms from a set of secondary transforms comprising a plurality of low frequency non-separable secondary transforms (LFNSTs) based on the selected intra prediction mode, wherein the selected subset of secondary transforms comprises at least one of the plurality of LFNSTs; deriving a prediction residual for the block from the data stream; transforming the prediction residual using an LFNST from the selected subset; reconstructing the block using the prediction signal and the transformed prediction residual for the block; Equipped with The method of claim 1, wherein the same subset of secondary transforms is selected for both a Planar intra prediction mode and the at least one matrix-based intra prediction mode.

Citation Information

Patent Citations

  • Encoding device, decoding device, encoding method, and decoding method

    WO2020213677A1