Coding using intra-prediction
By forming intra prediction mode lists based on adjacent blocks and separating modes, the method optimizes coding efficiency by excluding unlikely modes and reducing unnecessary syntax element transmission, improving prediction accuracy.
Patent Information
- Application Number
- JP2025064240
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing intra prediction methods often require additional syntax elements to be transmitted due to unlikely prediction modes not being included in the most likely mode list, leading to inefficiencies in coding efficiency.
A method for forming a list of most likely intra prediction modes based on adjacent blocks, excluding unlikely modes like DC when angular modes are used, and separating modes into distinct sets to optimize coding efficiency.
Improves coding efficiency by reducing the need for additional syntax elements and ensuring that the intra prediction mode used is within the most likely mode list, enhancing prediction accuracy and reducing data transmission requirements.
Smart Images

Figure 2025106466000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intra prediction. Embodiments relate to advantageous methods for generating the most likely mode list.
[0002] Today, there are various methods for generating the most likely mode list. However, there is still a high likelihood of situations where additional syntax elements need to be transmitted because the intra prediction mode ultimately used is not within this list. Therefore, there is a problem of optimizing the generation of the most likely mode list and / or improving coding efficiency. This is achieved by the subject matter of the independent claims of this application. Further embodiments according to the invention are defined by the subject matter of the dependent claims of this application.
Summary of the Invention
[0003] According to a first aspect of the present invention, the inventors of the present application recognized that one problem encountered when forming a list of the most likely intra prediction modes is that unlikely prediction modes have an adverse effect on coding efficiency and inherit valuable list positions that increase the likelihood of a situation where the intra prediction mode ultimately used to predict a given block is not within this list. According to a first aspect of the present application, this difficulty is overcome by forming a list of the most likely intra prediction modes based on already predicted adjacent blocks adjacent to a given block. Thus, unlikely intra prediction modes can be omitted. A high probability of an intra prediction mode of a given block similar to that of the intra prediction mode of the adjacent blocks can be expected. In particular, when at least one of the adjacent blocks is predicted by any angular intra prediction mode, the list does not include the DC intra prediction mode. This enables a list of the most likely intra prediction modes having a diversity of angular intra prediction modes that increases the likelihood that an intra prediction mode will be used for a given block within the list. Further, matrix-based intra prediction modes are, for example, not considered for the list of the most likely intra prediction modes, and thus form a separate second set of intra prediction modes that do not conflict with the intra prediction modes of the first set of intra prediction modes for positions within the list of the most likely intra prediction modes.
[0004] Accordingly, according to a first aspect of the present application, an apparatus for decoding a predetermined block of an image using intra prediction is configured to derive from a data stream a set selection syntax element indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. Optionally, the first set of intra prediction modes may include a planar intra prediction mode in addition to or instead of the DC intra prediction mode. When the set selection syntax element indicates that a predetermined block is to be predicted using one of the first set of intra prediction modes, the apparatus forms a list of the most likely intra prediction modes based on the intra prediction mode when adjacent blocks adjacent to the predetermined block are predicted, derives from the data stream an MPM (i.e., most probable mode) list index that directs the list of the most likely intra prediction modes onto a predetermined intra prediction mode, and is configured to intra predict the predetermined block using the predetermined intra prediction mode. In other words, in this case, the apparatus is configured to form a list of the most likely intra prediction modes based on the intra prediction mode used for predicting adjacent blocks adjacent to the predetermined block. When the set selection syntax element indicates that a predetermined block is not to be predicted using one of the first set of intra prediction modes, the apparatus calculates a matrix-vector product between a vector derived from reference samples within the neighborhood of the predetermined block to obtain a prediction vector and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode, and by predicting samples of the predetermined block based on the prediction vector, further derives from the data stream an index indicating a predetermined matrix-based intra prediction mode of a second set of intra prediction modes, i.e., a second set of intra prediction modes including the matrix-based intra prediction mode. In this case, the prediction is similar or equal to the ALWIP prediction described with respect to the embodiments of FIGS. 5 to 11, for example.When an adjacent block is predicted by any of the angular intra prediction modes, the apparatus is configured to perform the formation of a list of the most likely intra prediction modes such that the list does not include the DC intra prediction mode, using when an adjacent block adjacent to a given block is predicted. In other words, the apparatus is configured to perform the formation of a list of the most likely intra prediction modes such that the list does not include the DC intra prediction mode when an adjacent block is exclusively predicted by any of the angular intra prediction modes, using when an adjacent block adjacent to a given block is predicted. Thus, the DC intra prediction mode does not occupy a position within the list of the most likely intra prediction modes when the likelihood that the DC intra prediction mode is selected for a given block is low.
[0005] This apparatus introduces an advantageous and efficient method for determining the intra prediction mode of a given block. In particular, an advantageous analysis of the prediction of adjacent blocks adjacent to a given block for forming a list of the most likely intra prediction modes is presented, where the adjacent blocks have already been predicted.
[0006] According to one embodiment, the apparatus is configured to perform the formation of a list of the most likely intra prediction modes based on the intra prediction mode such that, when an adjacent block adjacent to a given block is predicted, the list of the most likely intra prediction modes is filled with the DC intra prediction mode only if the adjacent block is predicted using any one of at least one non-angular intra prediction mode having a first set including the DC intra prediction mode for each of the adjacent blocks, or mapped to any one of at least one non-angular intra prediction mode by mapping to an intra prediction mode within the first set used for the formation of the list of the most likely intra prediction modes from a second set of block-based intra prediction modes. In other words, the list of the most likely intra prediction modes includes the DC intra prediction mode in the case of prediction of all adjacent blocks, for example both adjacent blocks, using any one of at least one non-angular intra prediction mode of the first set of intra prediction modes. Alternatively, the list of the most likely intra prediction modes includes the DC intra prediction mode in the case of prediction of all adjacent blocks, for example both adjacent blocks, using any one of the block-based intra prediction modes of the second set of intra prediction modes, and the block-based intra prediction modes are mapped to non-angular intra prediction modes within the first set from the second set of block-based intra prediction modes. According to one embodiment, the apparatus is configured to position the DC intra prediction mode before any angular intra prediction mode in the list of the most likely intra prediction modes. This is based on the idea that in the above case, the DC intra prediction mode is the most likely mode of a given block, thereby enhancing the coding efficiency of this positioning.
[0007] According to one embodiment, the apparatus is configured to derive an MPM syntax element from the data stream and form a list of the most likely intra prediction modes only in the case of an MPM syntax element indicating that a given intra prediction mode of the first set of intra prediction modes is within the list of the most likely intra prediction modes. By this feature, the list of the most likely intra prediction modes is formed only when necessary or advantageous, thus improving the coding efficiency.
[0008] When a given block is to be predicted using one of the second set of intra prediction modes, the apparatus is configured, according to one embodiment, to form a list of the most likely block-based intra prediction modes. In this case, the apparatus is configured to derive, for example, an additional MPM list index from the data stream of the given matrix-based intra prediction modes, i.e., the given block-based intra prediction modes, which point to the most likely block-based intra prediction modes. Optionally, this list of the most likely block-based intra prediction modes is formed only when an additional MPM syntax element derived from the data stream indicates that the given block-based intra prediction mode is within the list of the most likely block-based intra prediction modes.
[0009] Thus, the apparatus is configured to form different MPM lists for, for example, a first set of intra prediction modes and a second set of intra prediction modes. The list of most likely intra prediction modes comprises, for example, the intra prediction modes of the first set of intra prediction modes, and the list of most likely block-based intra prediction modes comprises, for example, the second set of intra prediction modes, i.e., the intra prediction modes of the second set of block-based intra prediction modes. Thereby, it becomes possible that the block-based intra prediction modes do not need to compete with the intra prediction modes of the first set of intra prediction modes, such as the DC intra prediction mode and the angular prediction mode, for a position within the overall MPM list. With this separation, it is likely that the intra prediction modes of a given block are actually within their respective MPM lists.
[0010] One embodiment relates to an apparatus for encoding a predetermined block of an image using intra prediction, the apparatus being configured to signal a set selection syntax element in a data stream indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. Optionally, the first set of intra prediction modes may include a planar intra prediction mode in addition to or instead of the DC intra prediction mode. When the set selection syntax element indicates that a predetermined block is to be predicted using one of the first set of intra prediction modes, the apparatus forms a list of the most likely intra prediction modes based on the intra prediction mode when adjacent blocks adjacent to the predetermined block are predicted, signals an MPM list index in the data stream that points to the list of the most likely intra prediction modes over a predetermined intra prediction mode, and is configured to intra predict the predetermined block using the predetermined intra prediction mode. In other words, in this case, the apparatus is configured to form a list of the most likely intra prediction modes based on the intra prediction mode used for prediction of adjacent blocks adjacent to the predetermined block. When the set selection syntax element indicates that a predetermined block is not to be predicted using one of the first set of intra prediction modes, the apparatus calculates a matrix-vector product between a vector derived from reference samples in the neighborhood of the predetermined block to obtain a prediction vector and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode, and signals a further index in the data stream indicating a predetermined matrix-based intra prediction mode of a second set of matrix-based intra prediction modes, i.e., a second set of intra prediction modes including the matrix-based intra prediction mode, by predicting samples of the predetermined block based on the prediction vector. In this case, the prediction is similar or equal to, for example, the ALWIP prediction described with respect to the embodiments of FIGS. 5 to 11.The list of the most likely intra prediction modes is formed based on the intra prediction modes when adjacent blocks to a given block are predicted, such that when the adjacent blocks are predicted by any of the angular intra prediction modes, the list of the most likely intra prediction modes does not include the DC intra prediction mode. In other words, the list of the most likely intra prediction modes is formed based on the intra prediction modes used when adjacent blocks to a given block are predicted, such that when the adjacent blocks are exclusively predicted by any of the angular intra prediction modes, the list of the most likely intra prediction modes does not include the DC intra prediction mode.
[0011] One embodiment is a method for decoding a predetermined block of an image using intra prediction, the method including deriving from a data stream a set selection syntax element indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. When the set selection syntax element indicates that a predetermined block is to be predicted using one of the first set of intra prediction modes, the method includes forming a list of the most likely intra prediction modes based on the intra prediction mode using when adjacent blocks adjacent to the predetermined block are predicted, and deriving from the data stream an MPM list index indicating the list of the most likely intra prediction modes on a predetermined intra prediction mode, and intra predicting the predetermined block using the predetermined intra prediction mode. When the set selection syntax element indicates that a predetermined block is not to be predicted using one of the first set of intra prediction modes, the method includes calculating a matrix-vector product between a vector derived from reference samples in the vicinity of the predetermined block and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode to obtain a prediction vector, and deriving from the data stream a further index indicating the predetermined matrix-based intra prediction mode of a second set of matrix-based intra prediction modes by predicting samples of the predetermined block based on the prediction vector. The list of the most likely intra prediction modes is formed based on the intra prediction mode using when adjacent blocks adjacent to the predetermined block are predicted such that the list of the most likely intra prediction modes does not include the DC intra prediction mode when the adjacent blocks are predicted by any of the angular intra prediction modes.In other words, the list of the most likely intra prediction modes is formed based on the intra prediction modes used when adjacent blocks are predicted when an adjacent block adjacent to a given block is predicted such that the list of the most likely intra prediction modes does not include the DC intra prediction mode when the adjacent blocks are exclusively predicted by any of the angular intra prediction modes.
[0012] One embodiment is a method for encoding a predetermined block of an image using intra prediction, the method including the step of signaling a set-selection syntax element in a data stream indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. When the set-selection syntax element indicates that a predetermined block is to be predicted using one of the first set of intra prediction modes, the method includes the steps of forming a list of the most likely intra prediction modes based on the intra prediction mode using when adjacent blocks adjacent to the predetermined block are predicted; signaling an MPM list index in the data stream directing the list of the most likely intra prediction modes onto a predetermined intra prediction mode; and intra predicting the predetermined block using the predetermined intra prediction mode. When the set-selection syntax element indicates that a predetermined block is not to be predicted using one of the first set of intra prediction modes, the method includes the steps of calculating a matrix-vector product between a vector derived from reference samples in the neighborhood of the predetermined block and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode to obtain a prediction vector; and predicting samples of the predetermined block based on the prediction vector, and signaling a further index in the data stream indicating the predetermined matrix-based intra prediction mode of a second set of matrix-based intra prediction modes. The list of the most likely intra prediction modes is formed based on the intra prediction mode using when adjacent blocks adjacent to the predetermined block are predicted such that the list of the most likely intra prediction modes does not include the DC intra prediction mode when the adjacent blocks are predicted by any of the angular intra prediction modes.In other words, the list of the most likely intra prediction modes is formed based on the intra prediction modes used when adjacent blocks adjacent to a given block are predicted such that when the adjacent blocks are exclusively predicted by any of the angular intra prediction modes, the list of the most likely intra prediction modes does not include the DC intra prediction mode. One embodiment relates to a data stream having an image encoded using the encoding method described herein. One embodiment relates to a computer program having program code for performing the methods described herein when executed on a computer.
Brief Description of the Drawings
[0013] The drawings are not necessarily to scale and instead generally focus on explaining the principles of the present invention. In the following description, various embodiments of the present invention are described with reference to the following drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7.1
Figure 7.2
Figure 7.3
Figure 7.4
Figure 8
Figure 9a
Figure 9b-9c
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
DETAILED DESCRIPTION OF THE INVENTION
[0014] Elements that are the same or equivalent, or have the same or equivalent functions, are denoted by the same or equivalent reference numerals in the following description even if they occur in different figures.
[0015] In the following description, in order to provide a more overall description of embodiments of the present invention, a plurality of details are set forth. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present invention. Furthermore, features of different embodiments described later in this specification can be combined with each other unless otherwise specified. 1 Introduction
[0016] In the following, different examples, embodiments and aspects of the present invention will be described. At least some of these examples, embodiments, and aspects relate, inter alia, to methods and / or apparatuses for video coding and / or for performing intra prediction using, for example, linear or affine transforms with adjacent sample reduction and / or for optimizing video delivery (such as broadcasting, streaming, file playback, etc.) for video applications and / or virtual reality applications.
[0017] Furthermore, examples, embodiments, and aspects may refer to High Efficiency Video Coding (HEVC) or successors. Also, further embodiments, examples and aspects are defined by the appended claims.
[0018] Note that any embodiment, example and aspect defined by the claims can be supplemented by any of the details (features and functions) described in the following chapters.
[0019] Also, the embodiments, examples, and aspects described in the following chapters can be used individually and can also be supplemented by any of the features in another chapter or any feature included in the claims.
[0020] Note also that the individual examples, embodiments, and aspects described herein can be used individually or in combination. Thus, details can be added to each of the individual aspects without adding details to another one of the above examples, embodiments, and aspects. It should also be noted that the present disclosure explicitly or implicitly describes features of decoding and / or encoding systems and / or methods.
[0021] Furthermore, the features and functions disclosed herein in connection with the method can also be used in an apparatus. Further, any features and functions disclosed herein in connection with the apparatus can also be used in a corresponding method. In other words, the methods disclosed herein can be complemented by any of the features and functions described in connection with the apparatus.
[0022] Also, any of the features and functions described herein can be implemented in hardware or software, or a combination of hardware and software, as described in the "Implementation Options" section.
[0023] Furthermore, any of the features described within parentheses ("(...) " or "[...]") can be considered optional in some examples, embodiments, or aspects. 2 Encoder, Decoder
[0024] In the following, various examples that can assist in achieving more effective compression when using block-based prediction are described. Some examples achieve high compression efficiency by expending a set of intra prediction modes. The latter may be added to, or provided exclusively in addition to, other intra prediction modes designed heuristically, for example. Also, in still other examples, both of the aforementioned expertise areas are utilized. However, as a variation of these embodiments, intra prediction can be converted to inter prediction by instead using reference samples within another image.
[0025] To facilitate understanding of the following examples of this application, the description begins with the presentation of a possible encoder and a decoder that fits it, which can construct the examples of this application outlined subsequently. FIG. 1 shows an apparatus for encoding an image 10 into a data stream 12 in block units. The apparatus is shown using reference numeral 14 and may be a still image encoder or a video encoder. In other words, if the encoder 14 is configured to encode a video 16 including the image 10 into the data stream 12, the image 10 may be the current image of the video 16, or the encoder 14 may encode the image 10 into the data stream 12 exclusively.
[0026] As described above, the encoder 14 performs encoding in block units or on a block basis. For this, the encoder 14 subdivides the image 10 into blocks, and in units of blocks, the encoder 14 encodes the image 10 into the data stream 12. Examples of possible subdivisions of the image 10 into blocks 18 are described in more detail below. Generally, the subdivision can end with an array of blocks 18 of a certain size, such as an array of blocks arranged in rows and columns, or blocks 18 of different block sizes, starting from the entire image area of the image 10 or using hierarchical multi-tree subdivision that starts from a pre-division of the image 10 into an array of tree blocks, and these examples are not treated as excluding other possible ways of subdividing the image 10 into blocks 18.
[0027] Furthermore, the encoder 14 is a predictive encoder configured to predictively encode the image 10 into the data stream 12. In the case of a particular block 18, this means that the encoder 14 determines a prediction signal for the block 18 and encodes into the data stream 12 a prediction residue, i.e., a prediction error in which the prediction signal deviates from the actual image content within the block 18.
[0028] Symbolizer 14 can support different prediction modes in order to derive the prediction signal of a certain block 18. The prediction mode that is important in the following example is the intra prediction mode, according to which the inside of block 18 is spatially predicted from adjacent already encoded samples of image 10. The encoding of the data stream 12 of image 10, and thus the corresponding decoding procedure, can be based on a specific encoding order 20 defined between blocks 18. For example, the encoding order 20 can traverse the blocks 18 in a raster scan order such as row by row from top to bottom while traversing each row from left to right. In the case of hierarchical multi-tree based subdivision, the raster scan ordering may be applied within each hierarchical level, or a depth-first traversal order may be applied, that is, the leaf nodes within a block of a specific hierarchical level may precede the blocks of the same hierarchical level having the same parent block according to the encoding order 20. Depending on the encoding order 20, the adjacent already encoded samples of block 18 can usually be located on one or more sides of block 18. In the case of the example presented herein, for example, the adjacent already encoded samples of block 18 are located at the top and left side of block 18.
[0029] The intra prediction mode may not be the only mode supported by symbolizer 14. For example, if symbolizer 14 is a video encoder, symbolizer 14 may also support an inter prediction mode in which block 18 is temporarily predicted from a previously encoded image of video 16. Such an inter prediction mode may be a motion compensated prediction mode, according to which a motion vector indicating the relative spatial offset of the portion where the prediction signal of block 18 is derived as a copy is signaled for such a block 18. In addition or alternatively, other non-intra prediction modes such as an inter prediction mode when symbolizer 14 is a multi-view encoder, or a non-prediction mode in which the inside of block 18 is encoded as it is, i.e., without prediction, may also be available.
[0030] Before beginning to focus the description of this application on the intra prediction mode, as described with respect to FIG. 2, a more specific example of a possible block-based encoder, i.e., a possible implementation form of encoder 14, is shown, and then two corresponding examples of decoders that conform to FIGS. 1 and 2, respectively, are shown.
[0031] FIG. 2 shows a possible implementation of the encoder 14 of FIG. 1, namely, the encoder is configured to use transform coding to encode the prediction residual, which is only an example, and the present application is not limited to that kind of prediction residual coding. According to FIG. 2, the encoder 14 is configured to subtract a prediction signal 24 corresponding to an inbound signal, i.e., the image 10, or the current block 18 on a block-by-block basis from the current block 18 to obtain a prediction residual signal 26 that is encoded into the data stream 12 by the prediction residual encoder 28, and includes a subtractor 22. The prediction residual encoder 28 is composed of an irreversible coding stage 28a and a reversible coding stage 28b. The irreversible stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes the samples of the prediction residual signal 26. As already described above, in this example, transform coding of the prediction residual signal 26 is used. Therefore, the lossy coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such spectrally decomposed prediction residual 26, and the quantization by the quantizer 30 is performed on the transform coefficients presenting the residual signal 26. The transform may be DCT, DST, FFT, Hadamard transform, etc. Next, the transformed and quantized prediction residual signal 34 is reversibly encoded by the reversible coding stage 28b, which is the entropy encoding of the entropy encoder for the quantized prediction residual signal 34, to become the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal in a manner available to the decoder from the transformed and quantized prediction residual signal 34, i.e., considering the coding loss in the quantizer 30. For this purpose, the prediction residual reconstruction stage 36 includes an inverse quantizer 38 that performs the inverse of the quantization by the quantizer 30, and then an inverse transform device 40 that performs an inverse transform with respect to the transform performed by the transform device 32, such as the inverse of the spectral decomposition for any of the above specific transform examples. The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transform device 40 and the prediction signal 24 to output a reconstructed signal, i.e., a reconstructed sample. This output is supplied to the predictor 44 of the encoder 14, and the predictor 44 determines the prediction signal 24 based thereon.This is the predictor 44 that supports all the prediction modes already described above with respect to FIG. 1. FIG. 2 also shows that when the encoder 14 is a video encoder, the encoder 14 may include an in-loop filter 46 that filters the fully reconstructed image to form a reference image for the predictor 44 for the inter prediction block after filtering.
[0032] As already described above, the coder 14 operates on a block-by-block basis. For the following description, the block of interest is one sub-divided image 10 in which an intra prediction mode is selected from a set or plurality of intra prediction modes respectively supported by the predictor 44 or the coder 14, and the selected intra prediction mode is executed individually. However, there may be other types of blocks into which the image 10 is subdivided. For example, the above determination of whether the image 10 is inter-coded or intra-coded can be made at a granularity, or in units of blocks deviating from the block 18. For example, the inter / intra mode determination may be performed at the level of the coding blocks in which the image 10 is re-divided and each coding block is re-divided into prediction blocks. Prediction blocks having coding blocks determined to use intra prediction are each subdivided into intra prediction mode determinations. In contrast, for each of these prediction blocks, it is determined which supported intra prediction mode should be used for each prediction block. These prediction blocks form the block 18 of interest here. Prediction blocks within coding blocks associated with inter prediction are treated differently by the predictor 44. They are inter-predicted from a reference image by determining a motion vector and copying the prediction signal of this block from the position in the reference image indicated by the motion vector. Another block subdivision relates to the subdivision into transform blocks at the unit where the transformation by the transformer 32 and the inverse transformer 40 is executed. The transformed block can be, for example, the result of further subdividing the coding block. Of course, the embodiments described herein should not be treated as limiting, and there are other embodiments. For the sake of completeness, it should be noted that the subdivision into coding blocks may use, for example, multi-tree subdivision, and the prediction blocks and / or the transform blocks may be obtained by further subdividing the coding blocks using multi-tree subdivision.
[0033] A decoder 54 or apparatus for block-based decoding adapted to the encoder 14 of FIG. 1 is shown in FIG. 3. This decoder 54 does the opposite of the encoder 14, i.e., decodes the data stream 12 image 10 in block units, and for this purpose, supports a plurality of intra prediction modes. The decoder 54 can comprise, for example, a residual provider 156. All of the other possibilities described above with respect to FIG. 1 are also valid for the decoder 54. In contrast, the decoder 54 may be a still image decoder or a video decoder, and all prediction modes and prediction possibilities are also supported by the decoder 54. The difference between the encoder 14 and the decoder 54 lies mainly in the fact that the encoder 14 selects or chooses encoding decisions according to some optimization in order to minimize some cost function that may depend, for example, on the encoding rate and / or encoding distortion. One of these encoding options or encoding parameters may include the selection of the intra prediction mode used for the current block 18 from among the available or supported intra prediction modes. The selected intra prediction mode can then be signaled by the encoder 14 of the current block 18 in the data stream 12 by the decoder 54 re-executing the selection using this signaling within the data stream 12 of block 18. Similarly, the subdivision of the image 10 into blocks 18 can be the subject of optimization within the encoder 14, and the corresponding subdivision information can be transmitted within the data stream 12 by the decoder 54 restoring the subdivision of the image 10 into blocks 18 based on the subdivision information. Summarizing the above, the decoder 54 may be a predictive decoder operating on a block basis, and in addition to the intra prediction mode, the decoder 54 may support other prediction modes such as the inter prediction mode, for example, if the decoder 54 is a video decoder. In decoding, the decoder 54 can also use the encoding order 20 described with respect to FIG. 1, and since this encoding order 20 is observed by both the encoder 14 and the decoder 54, the same adjacent samples are available for the current block 18 in both the encoder 14 and the decoder 54.Therefore, to avoid unnecessary repetition, the description of the operation mode of the encoder 14 shall also apply to the decoder 54 as far as the subdivision of the image 10 into blocks is concerned, for example, as far as prediction is concerned, and as far as the encoding of the prediction residue is concerned. The difference is that the encoder 14, by optimization, selects some encoding options or encoding parameters and then signals and inserts the encoding parameters derived from the data stream 12 into the data stream 12 for re-executing prediction, subdivision, etc. by the decoder 54.
[0034] Figure 4 shows a possible implementation of the decoder 54 of FIG. 3, i.e., one fitting to the implementation of the encoder 14 of FIG. 1 shown in FIG. 2. Since many elements of the encoder 54 of FIG. 4 occur in the corresponding encoder of FIG. 2, the same reference signs with apostrophes are used in FIG. 4 to denote these elements. In particular, the adder 42', the optional in-loop filter 46' and the predictor 44' are connected to the prediction loop in the same way as the encoder of FIG. 2. The reconstructed, i.e., inverse quantized and inverse transformed, prediction residue signal applied to the adder 42' is derived by a sequence of an entropy decoder 56 that reverses the entropy encoding of the entropy encoder 28b, followed by a residue signal reconstruction stage 36' composed of an inverse quantizer 38' and an inverse transformer 40' as in the case of the encoding side. The output of the decoder is the reconstruction of the image 10. The reconstruction of the image 10 may be directly available at the output of the adder 42' or may be available at the output of the in-loop filter 46'. To subject the reconstruction of the image 10 to some post-filtering to improve the image quality, some post-filters may be arranged at the output of the decoder, but this option is not shown in FIG. 4.
[0035] Here too, with respect to FIG. 4, the description given above with respect to FIG. 2 is equally valid for FIG. 4, except that only the encoder performs the relevant decisions regarding the optimization task and the coding options. However, all the explanations regarding block subdivision, prediction, inverse quantization, and retransformation are also valid for the decoder 54 of FIG. 4. 3 ALWIP (Affine Linear Weighted Intra Predictor)
[0036] Some non-limiting examples regarding ALWIP are described herein, even in cases where ALWIP is not necessarily required to implement the techniques described herein.
[0037] This application relates, inter alia, to an improved block-based prediction mode concept for block-based image coding such as can be used in HEVC or any subsequent video codec of HEVC. The prediction mode may be an intra prediction mode, but in theory, the concepts described herein can also be transferred to an inter prediction mode where the reference samples are part of another image. There is a need for block-based prediction concepts that enable efficient implementations such as hardware-friendly implementations. This object is achieved by the subject matter of the independent claims of this application.
[0038] The intra prediction mode is widely used in image and video coding. In video coding, the intra prediction mode competes with other prediction modes such as the inter prediction mode like the motion compensation prediction mode. In the intra prediction mode, the current block is predicted based on adjacent samples, i.e., samples that have already been encoded on the encoder side and already decoded on the decoder side. The adjacent sample values are extrapolated to the current block to form the prediction signal of the current block, and the prediction residual is transmitted in the data stream of the current block. The better the prediction signal, the smaller the prediction residual, and thus the fewer bits required to encode the prediction residual.
[0039] To be effective, several aspects should be considered to form an effective framework for intra prediction in a block - wise image coding environment. For example, the more intra prediction modes supported by a codec, the greater the side - information rate consumption for signaling the selection to the decoder. On the other hand, the set of supported intra prediction modes must be able to provide a good prediction signal, i.e., a prediction signal with a low prediction residual.
[0040] In the following, as a comparative embodiment or a basic example, there is disclosed an apparatus (encoder or decoder) for decoding an image from a data stream in block units, wherein the intra - prediction signal for a block of a predetermined size of the image is determined by applying a first template of samples adjacent to the current block to an affine - linear predictor, ultimately called an Affine Linear Weighted Intra Predictor (ALWIP), and supporting at least one intra - prediction mode.
[0041] The apparatus may have at least one of the following characteristics (the same may be implemented in a non - transitory storage unit that stores instructions for causing a processor to implement the method and / or operate as the apparatus when executed, for example, by the processor, and may be applied to a method or another technology). 3.1 The predictor may be complementary to other predictors.
[0042] The intra prediction mode that can form the subject of the implementation improvement to be further described below can be complementary to other intra prediction modes of the codec. Therefore, they can complement the DC prediction mode, planar prediction mode, or angular prediction mode defined in the HEVC codec and JEM reference software. Hereinafter, the latter three intra prediction modes are referred to as conventional intra prediction modes. Therefore, for a given block of the intra mode, a decoder needs to parse a flag indicating whether one of the intra prediction modes supported by the device should be used. 3.2 Two or more proposed prediction modes
[0043] The device can include two or more ALWIP modes. Therefore, if the decoder knows that one of the ALWIP modes supported by the device should be used, the decoder needs to parse additional information indicating which of the ALWIP modes supported by the device should be used.
[0044] The signaling of the supported modes can have the characteristic that the encoding of some ALWIP modes may require fewer bins than other ALWIP modes. Which of these modes require fewer bins and which modes require more bins can depend on information that can be extracted from the already decoded bitstream or can be fixed in advance. 4 Some aspects
[0045] FIG. 2 shows a decoder 54 for decoding an image from a data stream 12. The decoder 54 can be configured to decode a predetermined block 18 of the image. In particular, the predictor 44 can be configured to map a set of P adjacent samples adjacent to the predetermined block 18 to a set of Q predicted values of the samples of the predetermined block using a linear or affine linear transformation [e.g., ALWIP].
[0046] As shown in FIG. 5, a predetermined block 18 includes Q predicted values (which will become “predicted values” at the end of the operation). If block 18 has M rows and N columns, then Q = M·N. The Q values of block 18 can be in the spatial domain (e.g., pixels) or the transform domain (e.g., DCT, discrete wavelet transform, etc.). The Q values of block 18 can generally be predicted based on P values obtained from adjacent blocks 17a - 17c adjacent to block 18. The P values of adjacent blocks 17a - 17c may be in the position closest to block 18 (e.g., adjacent). The P values of adjacent blocks 17a - 17c have already been processed and predicted. The P values are shown as values in portions 17’a - 17’c to distinguish them from the blocks of which they are a part (in some examples, 17’b is not used).
[0047] As shown in FIG. 6, to perform the prediction, it is possible to operate with a first vector 17P having P entries (each entry is associated with a specific position in adjacent portions 17’a - 17’c), a second vector 18Q having Q entries (each entry is associated with a specific position in block 18), and a mapping matrix 17M (each row is associated with a specific position within block 18 and each column is associated with a specific position within adjacent portions 17’a - 17’c). Thus, mapping matrix 17M performs the prediction of the P values of adjacent portions 17’a - 17’c to the values of block 18 according to a predetermined mode. Thus, the entries within mapping matrix 17M can be understood as weight coefficients. In the following sections, the adjacent portions of the boundary are referred to using symbols 17a - 17c instead of 17’a - 17’c.
[0048] In the relevant art, several conventional modes such as DC mode, planar mode, and 65 - direction prediction mode are known. For example, 67 modes are known.
[0049] However, it should be noted that it is also possible here to utilize a different mode called linear or affine-linear transformation. The linear or affine-linear transformation includes P·Q weight coefficients, at least 1 / 4 P·Q of which are non-zero weight values, and for each of the Q predicted values, it includes a series of P weight coefficients regarding that respective predicted value. When the series is arranged vertically in raster scan order among the samples of a given block, it forms an envelope that is non-linear in all directions.
[0050] It is possible to map the P positions of the adjacent values 17’a~17’c (templates), the Q positions of the adjacent samples 17’a~17’c, and the values of the P*Q weight coefficients of the matrix 17M. The plane is an example of the envelope of the series for DC conversion (it is the plane for DC conversion). The envelope is clearly a plane and is thus excluded by the definition of linear or affine-linear transformation (ALWIP). Another example is a matrix that results in the emulation of the angular mode, and the envelope is excluded from the ALWIP definition and, plainly speaking, looks like a hill slanting diagonally from top to bottom along a direction in the P / Q plane. The plane mode and the 65-direction prediction mode have different envelopes, but these are linear in at least one direction, that is, for example, all directions of the illustrated DC, and in the direction of the hill of the angular mode, for example.
[0051] Conversely, the envelope of the linear or affine transformation is not linear in all directions. It is understood that such a type of transformation may be optimal for performing the prediction of block 18 depending on the situation. It should be noted that it is preferable that at least 1 / 4 of the weight coefficients are different from 0 (that is, at least 25% of the P*Q weight coefficients are different from 0).
[0052] The weight coefficients may be independent of each other according to any regular mapping rule. Thus, the matrix 17M may be such that the values of its entries do not have an obvious recognizable relationship. For example, the weight coefficients cannot be described by any analytical or differential function.
[0053] In the example, the ALWIP conversion is the average of the maximum values of the cross-correlations between the first series of weight coefficients associated with each predicted value and the second series of weight coefficients associated with predicted values other than each predicted value, or the inverted version of the latter series may be lower than a predetermined threshold value (for example, 0.2 or 0.3 or 0.35 or 0.1, for example, a threshold value within the range between 0.05 and 0.035) even though the maximum value is higher. For example, for each combination (i1, i2) of rows of the ALWIP matrix 17M, the cross-correlation may be calculated by multiplying the P value of the i1-th row by the P value of the i2-th row. For each obtained cross-correlation, the maximum value can be obtained. Therefore, an average (mean) can be obtained for the entire matrix 17M (that is, the maximum values of the cross-correlations in all combinations are averaged). Then, the threshold value may be, for example, 0.2 or 0.3 or 0.35 or 0.1, for example, a threshold value in the range between 0.05 and 0.035.
[0054] The P adjacent samples of the blocks 17a to 17c may be arranged along a one-dimensional path extending along the boundary of a predetermined block 18 (for example, 18c, 18a). For each of the Q predicted values of the predetermined block 18, a series of P weight coefficients associated with each predicted value may be ordered so as to cross the one-dimensional path in a predetermined direction (for example, from left to right, from top to bottom, etc.). In the example, the ALWIP matrix 17M may be non-diagonal or non-block diagonal. An example of the ALWIP matrix 17M for predicting the 4×4 block 18 from four already predicted adjacent samples may be as follows. { { 37, 59, 77, 28}, { 32, 92, 85, 25}, { 31, 69, 100, 24}, { 33, 36, 106, 29}, { 24, 49, 104, 48}, { 24, 21, 94, 59}, { 29, 0, 80, 72}, {35, 2, 66, 84}, {32, 13, 35, 99}, {39, 11, 34, 103}, {45, 21, 34, 106}, {51, 24, 40, 105}, {50, 28, 43, 101}, {56, 32, 49, 101}, {61, 31, 53, 102}, {61, 32, 54, 100} [[ID=18}]}.
[0055] (Here, {37, 59, 77, 28} is the first row of matrix 17M, {32, 92, 85, 25} is the second row, and {61, 32, 54, 100} is the 16th row). Matrix 17M has dimensions 16×4 and contains 64 weight coefficients (as a result of 16*4 = 64). This is because matrix 17M has dimensions Q×P, where Q = M*N is the number of samples of the prediction target block 18 (block 18 is a 4×4 block), and P is the number of samples of the already predicted samples. Here, M = 4, N = 4, Q = 16 (as a result of M*N = 4*4 = 16), and P = 4. The matrix is non - diagonal and non - block - diagonal and is not described by a specific rule.
[0056] As can be seen, less than 1 / 4 of the weight coefficients are 0 (in the case of the above matrix, one of the 64 weight coefficients is 0). The envelope formed by these values forms an envelope that is non - linear in all directions when arranged one below the other in raster scan order. Even if the above explanation is mainly described with reference to a decoder (e.g., decoder 54), the same may be executed in an encoder (e.g., encoder 14).
[0057] In some examples, for each block size (in a set of block sizes), the ALWIP transforms of the intra prediction modes within a second set of intra prediction modes for each respective block size are different from each other. In addition or alternatively, the cardinality of the second set of intra prediction modes for block sizes in a set of block sizes may match, but the associated linear transforms or affine linear transforms of the intra prediction modes within the second set of intra prediction modes for different block sizes may be non-transferable to each other by scaling.
[0058] In some examples, the ALWIP transform may be defined to "have nothing" to share with the corresponding conventional transform (e.g., the ALWIP transform may have "nothing" to share with the corresponding conventional transform even though it is mapped via one of the above mappings).
[0059] In some examples, the ALWIP mode is used for both the luma and chroma components, while in other examples, the ALWIP mode is used for the luma component but not for the chroma component. 5 Acceleration of the Encoder with Affine Linear Weighted Intra Prediction Mode (e.g., Test CE3-1.2.1) 5.1 Description of the Method or Apparatus
[0060] The Affine Linear Weighted Intra Prediction (ALWIP) mode tested in CE3-1.2.1 may be the same as that proposed in JVET-L0199 under Test CE3-2.2.2, except for the following changes.
[0061] · Multiple Reference Line (MRL) intra prediction, particularly in terms of encoder estimation and signaling compatibility, i.e., MRL is not combined with ALWIP, and the transmission of the MRL index is restricted to non-ALWIP blocks.
[0062] · Subsampling is mandatory for all blocks with W×H≧32×32 (previously optional for 32×32). Therefore, additional tests in the encoder and transmission of subsampling flags are removed.
[0063] · The ALWIP for 64×N and N×64 blocks (N≦32) is added by downsampling to 32×N and N×32 respectively and applying the corresponding ALWIP mode. Furthermore, test CE3-1.2.1 includes the following encoder optimizations for ALWIP.
[0064] · Combined mode estimation: The conventional mode and the ALWIP mode use a shared Hadamard candidate list for full RD estimation, i.e., the ALWIP mode candidates are added to the same list as the conventional (and MRL) mode candidates based on the Hadamard cost.
[0065] · EMT intra fast and PB intra fast are supported for the combined mode list and there are additional optimizations to reduce the number of full RD checks. · Only the MPM of the available left and upper blocks is added to the list for full RD estimation of ALWIP following the same method as in the case of the conventional mode. 5.2 Complexity Evaluation
[0066] Excluding the calculations that call the discrete cosine transform, in test CE3-1.2.1, a maximum of 12 multiplications per sample were required to generate the predicted signal. Furthermore, a total of 136 out of 492 parameters, each 16 bits, were required. This corresponds to 0.273 megabytes of memory. 5.3 Experimental Results
[0067] The test evaluation was carried out according to the common test conditions JVET-J1010[2] for the intra-only (AI) and random access (RA) configurations using the VTM software version 3.0.1. The corresponding simulation was executed on an Intel Xeon cluster (E5-2697A v4, AVX2 on, turbo boost off) equipped with the Linux (registered trademark) OS and the GCC 7.2.1 compiler.
[0068]
Table 1
[0069]
Table 2
[0070] The technique tested in CE2 is related to the "Affine Linear Intra Prediction" described in JVET-L0199[1], but simplifies it in terms of memory requirements and computational complexity.
[0071] · There can only be three different sets of prediction matrices (e.g., S0, S1, S2, see also below) and bias vectors (e.g., to provide offset values) that cover all block shapes. As a result, the number of parameters is reduced to 14,400 10-bit values, which is
Number
[0072] · The input and output sizes of the predictor are further reduced. Further, instead of converting the boundaries via the DCT, averaging or downsampling can be performed on the boundary samples, and linear interpolation can be used to generate the prediction signal instead of the inverse DCT. As a result, up to 4 multiplications per sample may be required to generate the prediction signal. 6 Embodiments Here, a method of performing several predictions (e.g., as shown in FIG. 6) using ALWIP prediction will be described.
[0073] In principle, referring to FIG. 6, to obtain the Q = M*N values of the M×N block 18 to be predicted, the multiplication of the Q*P samples of the Q×P ALWIP prediction matrix 17M and the P samples of the P×1 adjacent vector 17P should be performed. Therefore, generally, to obtain each of the Q = M*N values of the M×N block 18 to be predicted, at least P = M + N value multiplications are required.
[0074] These multiplications have highly undesirable effects. The dimension P of the boundary vector 17P generally depends on the number M + N of boundary samples (bins or pixels) 17a, 17c adjacent to the M×N block 18 to be predicted (e.g., adjacent). This means that as the size of the block 18 to be predicted increases, the number of boundary pixels M + N (17a, 17c) correspondingly increases, so the dimension P = M + N of the P×1 boundary vector 17P, the length of each row of the Q×P ALWIP prediction matrix 17M, and thus the required number of multiplications (generally speaking, Q = M*N = W*H, where W (Width) is another symbol for N and H (Height) is another symbol for M, and P = M + N = H + W when the boundary vector is formed by only one row and / or one column of samples) increases.
[0075] This problem is generally exacerbated in microprocessor-based systems (or other digital processing systems) by the fact that multiplication is generally a power-consuming operation. It can be inferred that a large number of multiplications carried out on a very large number of samples of a large number of blocks generally causes an undesirable waste of computing power. Therefore, it is preferable to reduce the number of multiplications Q*P required to predict the M×N block 18.
[0076] It is understood that it is possible to reduce in some way the computing power required for each intra-prediction of each predicted block 18 by intelligently selecting an operation that is easier to process instead of multiplication.
[0077] In particular, referring to FIGS. 7.1 to 7.4, an encoder or decoder can predict a predetermined block (e.g., 18) of an image by using a plurality of adjacent samples (e.g., 17a, 17c) to obtain a set of reduced sample values with a lower number of samples compared to the plurality of adjacent samples (e.g., by averaging or downsampling at step 811), and subjecting the set of reduced sample values to a linear or affine-linear transformation (e.g., at step 812) to obtain a predicted value for a predetermined sample of the predetermined block.
[0078] In some cases, the decoder or encoder can also derive predicted values for additional samples of a predetermined block based on the predicted values of the predetermined samples and a plurality of adjacent samples, e.g., by interpolation. Thus, an upsampling strategy can be obtained.
[0079] In an example, it is possible to perform several averages on the samples of boundary 17 (e.g., in step 811) so as to reach a reduced set 102 of samples with a reduced number of samples (Figures 7.1 to 7.4) (at least one of the samples of the reduced number of samples 102 may be the average of two samples of the original boundary samples or a selection of the original boundary samples). For example, if the original boundary has P = M + N samples, the reduced set of samples can have M red <M and N red <at least one of N, with P red <such that P red = M red + N red Thus, the boundary vector 17P actually used for prediction (e.g., in step 812b) does not have P × 1 entries, but has P red <P with P red × 1 entries. Similarly, the ALWIP prediction matrix 17M selected for prediction does not have a Q × P dimension, but has at least P red <P for (M red <M and N red <at least one of N) the number of matrix elements reduced for Q × P red (or Q red xP red , see below) for.
[0080] In some examples (e.g., Figures 7.2, 7.3), the block obtained by ALWIP (in step 812) is
Number
Number
Number
Number
Number
Number
[0081] These techniques can be advantageous because while the matrix multiplication involves a reduced number of multiplications (Q red *P red Or Q*P red ), both the initial reduction (e.g., averaging or downsampling) and the final transformation (e.g., interpolation) can be performed by reducing (or avoiding) multiplication. For example, downsampling, averaging, and / or interpolation can be performed (e.g., in steps 811 and / or 813) by employing binary operations that require non-computation power such as addition and shift. Also, addition is a very simple operation that can be easily performed without much computational effort.
[0082] This shift operation can be used, for example, to average two boundary samples, and / or to interpolate two samples (support values) of a reduced prediction block (or obtained from the boundary) in order to obtain a final prediction block. (For interpolation, two sample values are required. Within the block, there are always two predetermined values, but along the left and upper boundaries of the block, there is only one predetermined value as shown in Figure 7.2, and thus boundary samples are used as support values for interpolation.) The following, A step of first summing the values of two samples, and Then, a step of halving the sum value (e.g., by a right shift), and A two-step procedure such as this can be used. Alternatively, the following, A step of first halving each of the samples (e.g., by a left shift), and Then, a step of summing the values of the two halved samples, and Is possible.
[0083] Since it is only necessary to select one sample amount and a group of samples (e.g., adjacent samples to each other), easier operations can be performed during downsampling (e.g., in step 811).
[0084] Therefore, here, it is possible to define techniques (multiple possible) for reducing the number of multiplications to be performed. Some of these techniques may be based, among other things, on at least one of the following principles.
[0085] Even if the size of the actually predicted block 18 is M×N, the block is reduced (in at least one of the two dimensions), and the size is reduced to Q red xP red (Here,[[]] [Number][[]] P red =Nred +M red 、 [Number] and / or [Number] M red <M and / or N red <An ALWIP matrix of (M and / or N) may be applied. Thus, the boundary vector 17P has size P red ×1 and P red <means only P multiplications (P red =M red +N red and P = M + N). P red ×1 boundary vector 17P can be easily obtained from the original boundary 17, for example, by downsampling (e.g., by selecting only some samples of the boundary), and / or by averaging multiple samples of the boundary (which can be easily obtained by addition and shift without multiplication). can be easily obtained from the original boundary 17.
[0086] In addition, or alternatively, instead of predicting all Q = M * N values of the block 18 to be predicted by multiplication, a reduced block with reduced dimensions (e.g., [Number] 、 [Number] and / or [Number] ) can be predicted. The remaining samples of the block 18 to be predicted are obtained, for example, by interpolation using Q red samples as support values for the remaining Q - Q red values.
[0087] According to the example shown in FIG. 7.1, a 4×4 block 18 (M = 4, N = 4, Q = M*N = 16) would be predicted, and the neighborhoods 17 (a vertical matrix with four already predicted samples) and 17c (a horizontal row with four already predicted samples) of sample 17a have already been predicted in the previous iteration (neighborhoods 17a and 17c can be collectively denoted as 17). Previously, by using the formula shown in FIG. 6, the prediction matrix 17M should be a Q×P = 16×8 matrix (due to Q = M*N = 4*4 and P = M+N = 4+4 = 8), and the boundary vector 17P should have an 8×1 dimension (due to P = 8). However, this leads to the need to perform eight multiplications for each of the 16 samples of the 4×4 block 18 to be predicted, and thus to the need to perform a total of 16*8 = 128 multiplications. (Note that the average number of multiplications per sample is a good measure of computational complexity. In conventional intra prediction, four multiplications per sample are required, which increases the computational effort involved. Therefore, it is possible to use this as the upper limit of ALWIP, and it is guaranteed that the complexity is reasonable and does not exceed that of conventional intra prediction.)
[0088] Nevertheless, by using this technique, in step 811, the number of samples 17a and 17c adjacent to the predicted block 18 is changed from P to P redIt is understood that it is possible to reduce to . In particular, in order to obtain a reduced boundary 102 having two horizontal rows and two vertical columns, it is possible to average adjacent boundary samples (17a, 17c) with each other (for example, at 100 in FIG. 7.1), and thus it is understood that operating as block 18 was a 2×2 block (the reduced boundary is formed by the average value). Alternatively, it is possible to perform downsampling, and thus select two samples for row 17c and two samples for column 17a. Thus, the horizontal row 17c is processed as having two samples (for example, averaged samples) instead of having four original samples, and the vertical column 17a, which originally had four samples, is processed as having two samples (for example, averaged samples). After subdividing row 17c and column 17a into groups 110 of two samples each, it can also be understood that a single sample is maintained (for example, the average of the samples in group 110 or a simple selection between the samples in group 110). Thus, by having only the set 102 with four samples (M red =2, N red =2, P red =M red +N red =4, P red <P), a so-called reduced set 102 of sample values is obtained.
[0089] It is understood that it is possible to perform operations (such as averaging or downsampling 100) without performing an excessive amount of multiplication at the processor level. That is, the averaging or downsampling 100 performed in step 811 can be easily obtained by simple and computationally non-power-consuming operations such as addition and shift.
[0090] At this point, it is understood that it is possible to apply the set 102 of reduced sample values to a linear or affine linear (ALWIP) transform 19 (e.g., using a prediction matrix such as matrix 17M in FIG. 6). In this case, the ALWIP transform 19 directly maps four samples 102 to the sample values 104 of block 18. In this case, interpolation is not required.
[0091] In this case, the ALWIP matrix 17M has dimensions Q×P red = 16×4. This follows from the fact that all Q = 16 samples of block 18 to be predicted are obtained directly by ALWIP multiplication (interpolation is not required).
[0092] Accordingly, in step 812a, an appropriate ALWIP matrix 17M having dimensions Q×P red is selected. The selection may be based at least in part on signaling from data stream 12, for example. The selected ALWIP matrix 17M may also be denoted as A k where k is an index that may be signaled in data stream 12 (optionally, the matrix may also be denoted as follows
Number
[0093] In step 812b, the selected Q×P red ALWIP matrix 17M (also denoted as A k ) is multiplied by the P red ×1 boundary vector 17P.
[0094] In step 812c, an offset value (e.g., b k ) can be added to all the acquired values 104 of the vector 18Q acquired, for example, by ALWIP. The value of the offset (b k , or possibly even further in some cases [Number] (see below) may be associated with a specific selected ALWIP matrix (A k ) or may be based on an index (which may be signaled, for example, in data stream 12). Therefore, the comparison between using this technology and not using this technology resumes here. When not using this technology: A prediction target block 18 with dimensions M = 4 and N = 4, Q = M * N = 4 * 4 = 16 predicted values, P = M + N = 4 + 4 = 8 boundary samples P = 8 multiplications for each of the Q = 16 predicted values, Total number of P * Q = 8 * 16 = 128 times, In this technology, A prediction target block 18 with dimensions M = 4 and N = 4, Q = M * N = 4 * 4 = 16 finally predicted values, Reduced dimension of the boundary vector: P red = M red + N red = 2 + 2 = 4, P red = 4 multiplications for each of the Q = 16 values to be predicted by ALWIP, P red * Q total number = 4 * 16 = 64 times (half of 128!) The ratio of the number of multiplications to the number of finally acquired values is P red * Q / Q = 4, that is, it is half of the P = 8 multiplications for each predicted sample!
[0095] As can be appreciated, appropriate values can be obtained in step 812 by relying on operations that are direct and do not require computationally intensive power, such as averaging (optionally with addition and / or shifting and / or downsampling).
[0096] Referring to FIG. 7.2, the predicted block 18 is here an 8×8 block of 64 samples (M = 8, N = 8). Here, a priori, the prediction matrix 17M should have a size Q×P = 64×16 (Q = M*N = 8*8 = 64, Q = 64 due to M = 8 and N = 8, and P = M + N = 8 + 8 = 16). Thus, a priori, for each of the Q = 64 samples of the predicted 8×8 block 18, P = 16 multiplications are required, reaching 64*16 = 1024 multiplications for the entire 8×8 block 18!
[0097] However, as seen in FIG. 7.2, instead of using all 16 samples at the boundary, a method 820 can be provided where only 8 values (e.g., 4 of the horizontal boundary row 17c and 4 of the vertical boundary column 17a between the original samples at the boundary) are used. From the boundary column 17c, instead of 8 samples, 4 samples can be used (e.g., they can be a 2×2 average and / or a selection of 1 out of 2 samples). Thus, the boundary vector becomes not a P×1 = 16×1 vector, but only a P red ×1 = 8×1 vector (P red = M red + N red = 4 + 4). Instead of the original P = 16 samples, it is understood that the samples of the horizontal row 17c and the vertical column 17a can be selected or averaged (e.g., 2×2) to have only P red = 8 boundary values, forming a reduced set 102 of sample values. This reduced set 102 makes it possible to obtain a reduced version of the block 18, where the reduced version has Q red = M red *N red = 4*4 = 16 samples. The size Mred xN red An ALWIP matrix for predicting a 4×4 block can be applied. The reduced version of block 18 includes the samples shown in gray in scheme 106 of FIG. 7.2, and the samples shown by the gray squares (including samples 118’ and 118’’) are the Q obtained in the target step 812 red = forms a 4×4 reduced block having 16 values. The 4×4 reduced block is obtained by applying the linear transformation 19 in the target step 812. After obtaining the values of the 4×4 reduced block, it is possible to obtain the values of the remaining samples (the samples shown as white samples in scheme 106), for example, by interpolation.
[0098] Regarding the method 810 of FIG. 7.1, this method 820 further includes a step 813 of deriving predicted values of the remaining Q - Q of the predicted M×N = 8×8 block 18, i.e., 64 - 16 = 48 samples (white squares), for example, by interpolation. red = The remaining Q - Q red = 64 - 16 = 48 samples can be obtained from the directly obtained Q by interpolation (the interpolation can also utilize the values of the boundary samples, for example). red = 16 samples. As can be seen in FIG. 7.2, samples 118’ and 118’’ are obtained in step 812 (as shown by the gray squares), while sample 108’ (which is in the middle of samples 118’ and 118’’ and is shown by the white square) is obtained by interpolation between samples 118’ and 118’’ in step 813. It is understood that the interpolation can also be obtained by an operation similar to that for averaging such as shifting and adding. Thus, in FIG. 7.2, the value 108’ can generally be determined (and can be an average) as the value intermediate between the value of sample 118’ and the value of sample 118’’.
[0099] By performing interpolation, it is also possible to reach the final version of the M×N = 8×8 block 18 based on the plurality of sample values shown in 104 in step 813. Therefore, comparing using this technology with not using it, When not using this technology: The dimensions within the prediction target block 18 are M = 8, N = 8, Q = M * N = 8 * 8 = 64 samples in the prediction target block 18, P = M + N = 8 + 8 = 16 samples within the boundary 17, P = 16 multiplications for each of the Q = 64 values to be predicted, The total number of P * Q = 16 * 64 = 1028 multiplications The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 16. In this technology, The block 18 to be predicted has dimensions M = 8, N = 8 Q = M * N = 8 * 8 = 64 values to be finally predicted,
[0100] However, the Q used red xP red The ALWIP matrix has P red = M red + N red , Q red = M red * N red , M red = 4, N red = 4, and P red = M red + N red = 4 + 4 = 8 samples within the boundary, P red < P P red = Q of the 4×4 reduced block to be predicted red = 8 multiplications for each of the 16 values (formed by the gray squares in scheme 106), P red * Q red The total number of = 8 * 16 = 128 times (far less than 1024!) The ratio of the number of multiplications to the number of final values to be obtained is P red * Q red / Q = 128 / 64 = 2 (much less than 16 obtained without using this technology!) Therefore, the technology presented in this specification requires 8 times less power than the previous technology.
[0101] Figure 7.3 shows another example (which can be obtained based on method 820), where the predicted block 18 is a rectangular 4×8 block (M = 8, N = 4) with Q = 4 * 8 = 32 predicted samples. The boundary 17 is formed by a row 17c of N = 8 samples and a column 17a of M = 4 samples. Thus, beforehand, the boundary vector 17P has dimension P×1 = 12×1, but the predicted ALWIP matrix should be a Q×P = 32×12 matrix, and thus Q * P = 32 * 12 = 384 multiplications are required.
[0102] However, for example, it is possible to average or downsample at least 8 samples of the horizontal row 17c to obtain a reduced horizontal row of only 4 samples (e.g., averaged samples). In some examples, the vertical column 17a remains as it is (e.g., without averaging). In total, the reduced boundary has dimension P red = 8, and P red < P. Thus, the boundary vector 17P has dimension P red ×1 = 8×1. The ALWIP prediction matrix 17M is a matrix with dimension M * N red * P red = 4 * 4 * 8 = 64. The 4×4 reduced block directly obtained in target step 812 (formed by the gray columns of schema 107) has size Q red = M * N red = 4 * 4 = 16 samples (instead of Q = 4 * 8 = 32 of the original 4×8 block 18 to be predicted). When the reduced 4×4 block is obtained by ALWIP, an offset value b k is added (step 812c), and interpolation can be performed in step 813. As can be seen in step 813 of Figure 7.3, the reduced 4×4 block is expanded to a 4×8 block 18, and the value 108' not obtained in step 812 is obtained in step 813 by interpolating the values 118' and 118'' (gray squares) obtained in step 812. Therefore, comparing using this technique with not using it, When not using this technology: Prediction target block 18 with dimensions M = 4 and N = 8 Value Q = M * N = 4 * 8 = 32 predicted, P = M + N = 4 + 8 = 12 samples at the boundary, P = 12 multiplications for each of the Q = 32 predicted values, Total number of P * Q = 12 * 32 = 384 multiplications The ratio of the number of multiplications to the number of final values obtained is P * Q / Q = 12. In this technology, Prediction target block 18 with dimensions M = 4 and N = 8 Q = M * N = 4 * 8 = 32 values finally predicted,
[0103] However, Q red xP red = Can use 16×8 ALWIP matrix, M = 4, N red = 4, Q red = M * N red = 16, P red = M + N red = 4 + 4 = 8, and P red = M + N red = 4 + 4 = 8 samples within the boundary, P red <P P red = Q of the reduced block to be predicted red = 8 multiplications for each of the 16 values Q red *P red Total number = 16 * 8 = 128 times (less than 384!) The ratio of the number of multiplications to the number of final values to be obtained is P red *Q red / Q = 128 / 32 = 4 (much less than 12 obtained without using this technology!) Therefore, in this technology, the computational effort is reduced to 1 / 3.
[0104] FIG. 7.4 shows an example of block 18 to be predicted with dimensions M×N = 16×16, where the finally predicted value is Q = M*N = 16*16 = 256 and P = M+N = 16+16 = 32 boundary samples. This results in a prediction matrix having dimensions Q×P = 256×32, which means 256*32 = 8192 multiplications!
[0105] However, by applying method 820, in step 811, it is possible to reduce the number of boundary samples, for example, from 32 to 8 (e.g., by averaging or downsampling). For example, for each group 120 of four consecutive samples in row 17a, a single sample (e.g., selected from among the four samples or the average of the samples) remains. Also, for each group of four consecutive samples in column 17c, one sample (e.g., selected from among the four samples or the average of the samples) remains.
[0106] Here, the ALWIP matrix 17M is red xP red = a 64×8 matrix. This is due to the fact that it is selected to P red = 8 (by using the eight averaged samples or the samples selected from 32 boundaries), and the fact that the reduced block predicted in step 812 is an 8×8 block (in scheme 109, the gray square is 64).
[0107] Thus, when 64 samples of the reduced 8×8 block are obtained in step 812, in step 813, it is possible to derive the remaining Q - Q red = 256 - 64 = 192 values 104 of the predicted block 18.
[0108] In this case, in order to perform interpolation, it has been selected to use only all the samples of boundary column 17a and the alternative samples of boundary row 17c. Other selections may be made.
[0109] In this method, the ratio between the number of multiplications and the number of finally obtained values is Q red *P red / Q = 8 * 64 / 256 = 2, which is much less than the 32 multiplications for each value without using this technology! The comparison between using and not using this technology is as follows. When not using this technology: The prediction target block 18 with dimensions M = 16, N = 16 Q = M * N = 16 * 16 = 256 predicted values, P = M + N = 16 * 16 = 32 samples within the boundary, P = 32 multiplications for each of the Q = 256 predicted values, The total number of P * Q = 32 * 256 = 8192 multiplications, The ratio between the number of multiplications and the number of finally obtained values is P * Q / Q = 32. In this technology, The prediction target block 18 with dimensions M = 16, N = 16 Q = M * N = 16 * 16 = 256 finally predicted values,
[0110] However, the Q used red xP red = 64 × 8 The ALWIP matrix has M red = 4, N red = 4, Q red = 8 * 8 = 64 samples predicted by ALWIP, P red = M red +N red = 4 + 4 = 8. P red = M red +N red = 4 + 4 = 8 samples within the boundary, P red <P P red = The Q of the reduced block to be predicted red = 8 multiplications for each of the 64 values Q red *P red The total number of = 64 * 4 = 256 times (less than 8192!) The ratio between the number of multiplications and the number of finally obtained values is P red *Qred / Q = 8 * 64 / 256 = 2 (much less than 32 obtained without using this technology!) Therefore, the computing power required by this technique is 16 times smaller than that of conventional techniques. Therefore,
[0111] Reducing a plurality of adjacent samples (100, 813), comparing with a plurality of adjacent samples (17), and obtaining a set of reduced sample values (102) with fewer samples;
[0112] Subjecting the set of reduced sample values (102) to a linear or affine linear transformation (19, 17M) to obtain predicted values of predetermined samples (104, 118’, 188’’) of a predetermined block (18) (step 812); thereby A predetermined block (18) of an image can be predicted using a plurality of adjacent samples (17).
[0113] In particular, reduction (100, 813) can be performed by downsampling a plurality of adjacent samples to obtain a reduced set of sample values (102) with fewer samples compared to the plurality of adjacent samples (17).
[0114] Alternatively, reduction (100, 813) can be performed by averaging a plurality of adjacent samples to obtain a reduced set of sample values (102) with fewer samples compared to the plurality of adjacent samples (17).
[0115] Furthermore, based on the predicted values of predetermined samples (104, 118’, 118’’) and a plurality of adjacent samples (17), it is possible to derive predicted values of further samples (108, 108’) of a predetermined block (18) by interpolation (813).
[0116] A plurality of adjacent samples (17a, 17c) may extend one-dimensionally along two sides of a predetermined block (18) (e.g., toward the right and bottom in FIGS. 7.1 to 7.4). A predetermined sample (e.g., the one obtained by ALWIP in step 812) may also be arranged in rows and columns, and may be arranged at every nth position from the samples (112) of the predetermined samples 112 adjacent to both sides of the predetermined block 18 along at least one of the rows and columns.
[0117] Based on a plurality of adjacent samples (17), it is possible to determine the support value (118) of one (118) of a plurality of adjacent positions aligned with each of at least one of the rows and columns. Based on the predicted value of a predetermined sample (108, 108') and the support value of the adjacent samples (118) aligned with at least one of the rows and columns, it is also possible to derive the predicted value 118 of additional samples (104, 118', 118'') of the predetermined block (18) by interpolation.
[0118] A predetermined sample (104) may be arranged at every nth position from the samples (112) of the samples (112) adjacent to two sides of the predetermined block 18 along the row, and the predetermined sample may be arranged at every mth position from the samples (112) of the predetermined samples (112) adjacent to two sides of the predetermined block (18) along the column, where n, m > 1. In some cases, n = m (e.g., in FIGS. 7.2 and 7.3, the samples 104, 118', 118'' directly obtained by ALWIP in 812 and shown by gray squares are alternated with the samples 108, 108' subsequently obtained in step 813 along the rows and columns).
[0119] It may be possible to perform a step of determining a support value by downsampling or averaging (122) a group (120) of adjacent samples within a plurality of adjacent samples (118) each including an adjacent sample for which a respective support value is determined, along at least one of a row (17c) and a column (17a), for each support value. Thus, in FIG. 7.4, in step 813, it is possible to obtain the value of sample 119 by using the value of a predetermined sample 118''' (previously obtained in step 812) and the values of adjacent samples 118 as support values.
[0120] The plurality of adjacent samples may extend one-dimensionally along two sides of a predetermined block (18). It may be possible to perform reduction (811) by grouping a plurality of adjacent samples (17) into one or more groups (110) of consecutive adjacent samples and performing downsampling or averaging on each of the one or more groups (110) of adjacent samples having two or three or more adjacent samples.
[0121] In an example, a linear or affine-linear transformation can include P red *Q red or P red *Q weight coefficients, where P red is the number (102) of sample values within a reduced set of sample values, and Q red or Q is the number (18) of predetermined samples within a predetermined block. At least 1 / 4 of P red *Q red or 1 / 4 of P red *Q weight coefficients are non-zero weight values. P red *Q red or P red *Q weight coefficients can include, for each of Q or Q red predetermined samples, a series of P red weight coefficients for the respective predetermined sample, and a series of P redWhen the weight coefficients are arranged vertically according to the raster scan order between predetermined samples of a predetermined block (18), they form an envelope that is non-linear in all directions. P red *Either Q or P red *Q red The weight coefficients may be independent of each other via a regular mapping rule. The maximum value of the cross-correlation between a first series of weight coefficients associated with each predetermined sample and a second series of weight coefficients associated with predetermined samples other than each predetermined sample, or the reversed version of the latter series, is lower than a predetermined threshold value even though the maximum value becomes higher. The predetermined threshold value may be 0.3 [or in some cases 0.2 or 0.1]. P red The adjacent samples (17) may be arranged along a one-dimensional path extending along both sides of a predetermined block (18), either Q or Q red For each of Q or Q predetermined samples, a series of P associated with each predetermined sample red The weight coefficients are ordered so as to cross the one-dimensional path in a predetermined direction. 6.1 Description of the method and apparatus
[0122] Width
Number
Number
Number
[0123] 1. Among the boundary samples 17, the sample 102 (e.g., 4 samples when W = H = 4 and / or 8 samples in other cases) can be extracted by averaging or downsampling (e.g., step 811).
[0124] 2. Matrix-vector multiplication followed by addition of an offset can be performed with the averaged samples (or the samples remaining from downsampling) as input. The result can be a reduced predicted signal (e.g., step 812) on a set of subsampled samples within the original block.
[0125] 3. The predicted signals at the remaining positions can be generated, for example, from the predicted signals on the subsampled set, for example, by linear interpolation (e.g., step 813), for example, by upsampling.
[0126] Thanks to step 1. (811) and / or 3. (813), the total number of multiplications required for the calculation of the matrix-vector product can be made to always be
Number
[0127] In some examples, the matrix (e.g., 17M) and offset vector (e.g., b k ) required to generate the predicted signal can be a set of matrices (e.g., 3 sets) stored, for example, in the storage units of the decoder and encoder, for example
Number
Number
[0128] In some examples, the set
Number
Number
Number
Number
Number
Number
Number
[0129] In some examples,
Number
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
Mathematics
[0130] In addition to, or alternatively to, this
Number
Number
Number
Number
Number
Number
[0131] As described above, the boundary samples (17a, 17c) can be averaged and / or downsampled (e.g., from P samples to P red <P samples).
[0132] In the first step, the input boundaries
Number
Number
Number
Number
Number
Number
Number
[0133] Similarly
Number
Number
Number
[0134] In all other cases (for example, for blocks with a width or height of withers different from 4), the block width W [Number] is W [Number] when given as [Number] . Similarly [Number] is defined as
[0135] In still other cases, the boundary (for example, by selecting one specific boundary sample from a group of boundary samples) can be downsampled to reduce the number of samples. For example, [Number] is [Number] and [Number] may be selected from among [Number] is
Number
Number
Number
[0136] Two reduced boundaries
Number
Number
Number
Number
Number
Number
Number
Mathematics
Mathematics
Mathematics
[0137] Therefore, according to a specific state (
Mathematics
Mathematics
[0138] Other strategies may be executed. In other examples, the mode index "mode" is not necessarily within the range from 0 to 35 (other ranges may be defined). Furthermore, each of the three sets S0, S1, S2 does not necessarily have 18 matrices (therefore,
Mathematics
Mathematics
Mathematics
Mathematics
[0139] The mode and transposition information are not necessarily stored and / or transmitted as a single combined mode index "mode". That is, in some examples, it may be explicitly signaled as a transposition flag and a matrix index (0 to 15 for S0, 0 to 7 for S1, 0 to 5 for S2).
[0140] In some cases, the combination of the transposition flag and the matrix index may be interpreted as a set index. For example, there may be a 1-bit operating as the transposition flag and several bits indicating the matrix index, which are collectively shown as the "set index". 6.3 Generation of Reduced Prediction Signals by Matrix-Vector Multiplication Here, features are provided regarding step 812.
[0141] Reduced input vector [Number] One of (the boundary vector 17P) can be the reduced prediction signal [Number] can generate. The latter signal can be a signal on a downsampled block with width [Number] and height [Number] Here, [Number] and [Number] can be defined as follows.
Number
Number
Number
Number
Number
[0142] Here,
Number
Number
Number
Number
Number
Number
Number
Number
[0143] Matrix A and vector
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0144] Other strategies may be implemented. In other examples, the mode index "mode" is not necessarily within the range from 0 to 35 (other ranges may be defined). Further, instead of expressions such as
Number
Number
Number
Number
[0145] For interpolation of the subsampled prediction signal, in a large block, a second version of the averaged boundary may be required. That is,
Number
[0146] In addition, or alternatively, it is possible to have "hard downsampling", [Number] is [Number] equal to
[0147] Also, [Number] can be defined similarly. [Number] At the sample positions excluded in the generation of [Number] It can be generated by linear interpolation from (for example, step 813 in the examples of FIGS. 7.2 to 7.4). In some examples,
Number
[0148] Linear interpolation may be given as follows (other examples are also possible).
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0149] Here, [Number] and [Number] it is. The latter process can be executed k times until H [Number] H red = H. Therefore, H [Number] or H [Number] in the case of, it can be executed at most once. H [Number] in the case of, it may be done twice. H [Number] if it is, it may be done three times. Next, the horizontal upsampling operation can be applied to the result of the vertical upsampling. The latter upsampling operation can use the entire left boundary of the prediction signal. Finally, [Number] in the case of, it can be similarly proceeded by first upsampling horizontally (if necessary) and then upsampling vertically.
[0150] This is an example of interpolation that uses the boundary samples reduced by the first interpolation (horizontal or vertical) and the original boundary samples for the second interpolation (vertical or horizontal). Depending on the block size, only the second interpolation or no interpolation may be required. If both horizontal and vertical interpolations are needed, the order depends on the width and height of the block. However, different techniques may be implemented. For example, the original boundary samples may be used for both the first and second interpolations, and the order may be fixed. For example, it may be horizontal first and then vertical (or vertical first and then horizontal in other cases). Therefore, the interpolation order (horizontal / vertical) and the use of reduced / original boundary samples can be changed.
[0151] 6.5 Description of an Example of the Entire ALWIP Process The entire process of averaging, matrix-vector multiplication, and linear interpolation is shown for different shapes in FIGS. 7.1 to 7.4. Note that the remaining shapes are treated like one of the illustrated cases.
[0152] 1.
Number
Number
Number
[0153] 2.
Number
Number
Number
[0154] 3.
Number
Number
Number
[0155] 4.
Number
Number
Number
Number
Number
Number
[0156] 6.6 Evaluation of the Number and Complexity of Required Parameters The parameters required for all possible proposed intra prediction modes can be composed of matrices and offset vectors belonging to the set
Number
Number
Number
[0157] 6.7 Signaling of the Proposed Intra Prediction Mode For the luma block, for example, 35 ALWIP modes are proposed (other numbers of modes may be used). For each coding unit (CU) in the intra-mode, a flag indicating whether the ALWIP mode should be applied to the corresponding prediction unit (PU) is transmitted in the bitstream. The signaling of the latter indicator can be coordinated with the MRL in the same way as in the first CE test. When the ALWIP mode is applied, the index of the ALWIP mode
Number
[0158] Here, the derivation of the MPM may be performed using the intra-modes of the above and left PUs as follows. Each conventional intra prediction mode
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0159] In all other cases
Number
Number
Number
Number
Number
Number
Number
[0160] The embodiments described in this specification are not limited by the above-described signaling of the proposed intra prediction mode. According to an alternative embodiment, the MPM and / or the mapping table are not used for MIP (ALWIP).
[0161] 6.8 Adaptive MPM List Derivation for Conventional Luma Intra Prediction Mode and Chroma Intra Prediction Mode The proposed ALWIP mode can be reconciled with the MPM-based coding of the conventional intra prediction mode as follows. The luma and chroma MPM list derivation processes of the conventional intra prediction mode use fixed tables
Number
Number
Number
Number
Number
Number
[0162] 7 Efficient Embodiments To form a basis for further expanding the embodiments described below, the above examples are briefly summarized. To predict a predetermined block 18 of the image 10, a plurality of adjacent samples 17a, 17c are used. Reduction 100 by averaging a plurality of adjacent samples is performed to obtain a set 102 of reduced sample values having fewer samples compared to the plurality of adjacent samples. This reduction is optional in the embodiments of this specification and results in a so-called sample value vector referred to below. The reduced set of sample values undergoes a linear or affine linear transformation 19 to obtain a predicted value of a given sample 104 of a given block. This transformation is later shown using a matrix A and an offset vector b that are obtained by machine learning (ML) and should be implemented efficiently.
[0163] By 2 interpolation, a predicted value of a further sample 108 of a given block is derived based on the predicted values of the given sample and the plurality of adjacent samples. Theoretically, it can be said that the result of the affine / linear transformation can be associated with the non-full sample positions of block 18 such that all samples of block 18 can be obtained by interpolation according to alternative embodiments. Interpolation may not be necessary at all.
[0164] A plurality of adjacent samples can extend one-dimensionally along two sides of a predetermined block, a predetermined sample is arranged in rows and columns and along at least one of the rows and columns, and the predetermined sample can be arranged at every nth position from a sample (112) of a predetermined sample adjacent to two sides of the predetermined block. Based on the plurality of adjacent samples, for each of at least one of the rows and columns, a support value of one (118) of the plurality of adjacent positions can be determined, which is aligned with each of at least one of the rows and columns, and by interpolation, a predicted value of a further sample 108 of the predetermined block can be derived based on the predicted value of the predetermined sample and the support values of the adjacent samples aligned with at least one of the rows and columns. The predetermined sample may be arranged at every nth position from a sample 112 of a predetermined sample adjacent to both sides of the predetermined block along a row, and the predetermined sample may be arranged at every mth position from a sample 112 of a predetermined sample adjacent to both sides of the predetermined block along a column, where n, m>1. It may be that n = m. Along at least one of the rows and columns, the determination of the support value can be performed by averaging (122) a group 120 of adjacent samples within the plurality of adjacent samples including the adjacent samples 118 for which each support value is determined. The plurality of adjacent samples may extend one-dimensionally along two sides of the predetermined block, and the reduction may be performed by grouping the plurality of adjacent samples into one or more groups 110 of consecutive adjacent samples and performing averaging for each of the one or more groups of adjacent samples having three or more adjacent samples.
[0165] For a predetermined block, a prediction residual can be transmitted in a data stream. This may be derived therefrom in a decoder, and the predetermined block is reconstructed using the prediction residual and predicted value of the predetermined sample. In an encoder, the prediction residual is encoded into the data stream by the encoder. The image can be subdivided into a plurality of blocks of different block sizes. These plurality of blocks comprise a predetermined block. Next, the linear or affine-linear transformation of block 18 is selected according to the width W and height H of the predetermined block, and as long as the width W and height H of the predetermined block are within a first set of width / height pairs, the linear or affine-linear transformation selected for the predetermined block is selected from a first set of linear or affine-linear transformations, and as long as the width W and height H of the predetermined block are within a second set of width / height pairs different from the first set of width / height pairs, a second set of linear or affine-linear transformations is selected. Again, later it becomes apparent that the affine / linear transformation is represented by other parameters, namely the weights of C, and optionally offset and scale parameters.
[0166] The decoder and the encoder subdivide the image into a plurality of blocks of different block sizes including a predetermined block, and the linear or affine-linear transformation selected for the predetermined block is selected from a first set of linear or affine-linear transformations as long as the width W and height H of the predetermined block are within a first set of width / height pairs, selected from a second set of linear or affine-linear transformations as long as the width W and the height H of the predetermined block are within a second set of width / height pairs not common to the first set of width / height pairs, configured to select a linear or affine-linear transformation according to the width W and height H of the predetermined block such that, as long as the width W and height H of the predetermined block are within a third set of one or more width / height pairs and are independent of the first and second sets of width / height pairs, it is selected from a third set of linear or affine-linear transformations. The third set of one or more width / height pairs comprises simply one width / height pair W', H', and each linear or affine-linear transformation within the first set of linear or affine-linear transformations is for converting N' sample values into W'*H' predicted values of a W'xH' array of sample positions.
[0167] Each of the first and second sets of width / height pairs is W p is H p a first width / height pair W that is not equal to H p H p and H q =W p and Wq = H p a second width / height pair W that is q H q can include. Each of the first and second sets of width / height pairs can further include a third width / height pair W p H p where W p is equal to H p and H p >H q is. For a given block, a set index indicating which linear or affine-linear transformation of a given set of linear or affine-linear transformations should be selected for block 18 can be transmitted in the data stream.
[0168] A plurality of adjacent samples may extend one-dimensionally along two sides of a predetermined block, and the reduction is for a first subset of a plurality of adjacent samples adjacent to a first side of the predetermined block, the first subset is grouped into a first group 110 of one or more consecutive adjacent samples, for a second subset of a plurality of adjacent samples adjacent to a second side of the predetermined block, the second subset is grouped into a second group 110 of one or more consecutive adjacent samples, and for obtaining a first sample value from the first group and a second sample value of the second group, the second subset may be averaged for each of the first and second groups of one or more adjacent samples having two or more adjacent samples. Next, the linear or affine-linear transformation may be selected according to a set index from a predetermined set of linear or affine-linear transformations, so that two different states of the set index result in one selection of the linear or affine-linear transformations of the predetermined set of linear or affine-linear transformations, and the reduced set of sample values, in the case of a set index assuming the first state of the two different states in the form of a first vector, undergoes a predetermined linear or affine-linear transformation to generate an output vector of predicted values, and the predicted values of the output vector are distributed to predetermined samples of the predetermined block along a first scan order, in the case of a set index assuming the second state of the two different states in the form of a second vector, the first and second vectors are such that one of the sample values of the first sample value in the first vector and the second vector is the value of one of the sample values of the second sample value in the second vector, generate an output vector of predicted values, and distribute the predicted values of the output vector along a second scan order over predetermined samples of the predetermined block transposed with respect to the first scan order, such that a component populated by one of the second sample values in the first vector is populated by one of the first sample values in the second vector.
[0169] Each linear or affine linear transformation within the first set of linear or affine linear transformations may be for transforming N1 sample values into w1*h1 predicted values for a w1×h1 array of sample positions, and each linear or affine linear transformation within the second set of linear or affine linear transformations is for transforming N2 sample values into w2*h2 predicted values for a w2×h2 array of sample positions, where for a first predetermined one of the width / height pairs, w1 can exceed the width of the first predetermined width / height pair, or h1 can exceed the height of the first predetermined width / height pair, and for a second predetermined one of the width / height pairs, w1 cannot exceed the width of the second predetermined width / height pair. By averaging, the step (100) of reducing a plurality of adjacent samples to obtain a reduced set (102) of sample values may be performed such that the reduced set 102 of sample values has N1 sample values when the predetermined block is of the first predetermined width / height pair and when the predetermined block is of the second predetermined width / height pair. Applying the selected linear or affine linear transformation to the selected reduced set of sample values may be performed using only the first sub-part of the selected linear or affine linear transformation associated with the subsampling of the w1×h1 array of sample positions, where when w1 exceeds the width of one width / height pair, along the width dimension, or when h1 exceeds the height of one width / height pair, along the height dimension, the selected linear or affine linear transformation is fully performed when the predetermined block is of the first predetermined width / height pair and when the predetermined block is of the second predetermined width / height pair.
[0170] Each linear or affine linear transformation within the first set of linear or affine linear transformations can be for transforming N1 sample values into w1*h1 predicted values for a w1×h1 array of sample positions where w1 = h1, and each linear or affine linear transformation within the second set of linear or affine linear transformations is for transforming N2 sample values into w2*h2 predicted values for a w2×h2 array of sample positions where w2 = h2. All of the above embodiments are merely examples in that they can form the basis of the embodiments described below in this specification. That is, the above concepts and details are useful for understanding the following embodiments and serve as a reservoir for possible extensions and modifications of the embodiments described later in this specification. In particular, many of the above details are optional, such as the averaging of adjacent samples and the fact that adjacent samples are used as reference samples. More generally, the embodiments described in this specification assume that the prediction signal on the rectangular block is generated from already reconstructed samples, such as the intra prediction signal on the rectangular block being generated from the adjacent already reconstructed samples on the left and above the block. The generation of the prediction signal is based on the following steps.
[0171] 1. However, without excluding the possibility of transferring the description to reference samples placed elsewhere, samples can be extracted by averaging from reference samples called boundary samples. Here, the averaging is performed only on the boundary samples on both the left and above the block, or only on the boundary samples on one of the two sides. If averaging is not performed on one side, the samples on that side remain unchanged. 2. Matrix-vector multiplication is performed, and optionally, an offset is added later. The input vector of the matrix-vector multiplication is, when averaging is applied only to the left side, the concatenation of the left of the averaged boundary samples of the block and the original boundary samples above the block, or when averaging is applied only to the upper side, the concatenation of the left of the original boundary samples of the block and the averaged boundary samples above the block, or when averaging is applied to both sides of the block, any of the concatenations of the left of the averaged boundary samples of the block and the averaged boundary samples above the block. Again, there are alternatives such as cases where averaging is not used at all.
[0172] 3. The result of the matrix-vector multiplication and the optional offset addition may optionally be a reduced prediction signal on a set of subsampled samples within the original block. The prediction signal at the remaining positions may be generated from the prediction signal on the subsampled set by linear interpolation. The calculation of the matrix-vector product in step 2 should preferably be performed with integer arithmetic. Thus,
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0173] Figure 8 shows an improved ALWIP prediction. The samples of a given block can be predicted based on a first matrix-vector product between a matrix A1100 derived by some machine learning-based training algorithm and a sample value vector 400. Optionally, an offset b1110 can be added. To achieve an integer approximation or fixed-point approximation of this first matrix-vector product, the sample value vector can undergo a reversible linear transformation 403 to determine a further vector 402. A second matrix-vector product between a further matrix B1200 and the further vector 402 can be made equal to the result of the first matrix-vector product. For further features of the further vector 402, the second matrix-vector product can be approximated as an integer by a matrix-vector product 404 between a predetermined prediction matrix C405, the further vector 402, and a further offset 408. The further vector 402 and the further offset 408 can consist of integer values or fixed-point values. All components of the further offset are, for example, the same. The predetermined prediction matrix 405 may be a quantized matrix or a matrix to be quantized. The result of the matrix-vector product 404 between the predetermined prediction matrix 405 and the further vector 402 can be understood as a prediction vector 406.
[0174] Further details regarding this integer approximation are provided below. Possible solution according to Example I: Subtraction and addition of average values Equations usable in the above scenario
Number
Number
Number
Number
Number
Number
Count
Count
Count
Count
Count
Count
Count
Count
Count
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0175] . The integer approximation described in the specification of Equation
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0176] According to one embodiment, the apparatus 1000 is configured to include a plurality of reversible linear transformations 403, each of which is associated with a component of a further vector 402. Further, the apparatus is configured to select, for example, a predetermined component 1500 among the components of the sample value vector 400, and use the reversible linear transformation 403 among the plurality of reversible linear transformations associated with the predetermined component 1500 as the predetermined reversible linear transformation. This is due to, for example, different positions of the rows of the reversible linear transformation 403 corresponding to the predetermined component, which depend on the position of the predetermined component in the i0-th row, that is, in the further vector. For example, if the first component of the further vector 402, that is, y1, is the predetermined component, the i0-th row replaces the first row of the reversible linear transformation.
[0177] As shown in FIG. 9b, the matrix components 414 of the predetermined prediction matrix 405 in the column 412 corresponding to the predetermined component 1500 of the further vector 402, that is, in the i0-th column, are, for example, all 0. In this case, the apparatus is configured to calculate the matrix-vector product 404 by performing a multiplication, for example, by calculating the matrix-vector product 407 between the reduced prediction matrix C'405 resulting from the predetermined prediction matrix C405 by leaving the column 412 and the still further vector 410 resulting from the further vector 402 by leaving the predetermined component 1500. Therefore, the prediction vector 406 can be calculated with fewer multiplications.
[0178] As shown in FIGS. 8, 9b, and 9c, when predicting samples of a given block based on the prediction vector 406, the apparatus 1000 can be configured to calculate, for each component of the prediction vector 406, the sum of that component and a, i.e., a predetermined value 1400. This sum can be represented, as shown in FIGS. 8 and 9c, by the sum of the prediction vector 406 and the vector 409, where all components of the vector 409 are equal to the predetermined value 1400. Alternatively, the addition can be represented, as shown in FIG. 9b, by the sum of the prediction vector 406 and the matrix-vector product 1310 between the integer matrix M1300 and a further vector 402, where the matrix components of the integer matrix 1300 are 1 within the column corresponding to the predetermined component 1500 of the further vector 402, i.e., the i0 column, and all other components are, for example, 0.
[0179] The result of the sum of the given prediction matrix 405 and the integer matrix 1300 is equal to or approximates, for example, a further matrix 1200 shown in FIG. 8. In other words, each matrix component of the column 412 of the given prediction matrix 405 corresponding to the predetermined component 1500 of the further vector 402, i.e., the i0-th column of the given prediction matrix C405, summed with the (i.e., matrix B) multiple of the invertible linear transformation 403, i.e., a further matrix B1200, corresponds to, for example, the quantized version of the machine learning prediction matrix A1100 as shown in FIGS. 8, 9a, and 9b. As shown in FIG. 9b, summing each matrix component of the given prediction matrix C405 within the i0-th column 412 can correspond to the sum of the given prediction matrix 405 and the integer matrix 1300. As shown in FIG. 8, the machine learning prediction matrix A1100 can be made equal to 1200 times the result of a further matrix by the invertible linear transformation 403. This is
Number
[0180] Matrix multiplication using only integer arithmetic For low complexity implementations (in terms of the complexity of adding and multiplying scalar values, as well as the storage required for the entries of the participating matrices), it is desirable to perform matrix multiplication 404 using only integer arithmetic.
Number
Number
[0181] According to one embodiment, operations on integers only are used to map a real value
Number
Number
Number
[0182] (1) final_offset = 1 << (right_shift_result - 1); for i in 0…m - 1 { accumulator = 0 for j in 0…n - 1 { accumulator := accumulator + y[j] * C[i,j] } z[i] = (accumulator + final_offset) >> right_shift_result; } Here, the array C, i.e., the predetermined prediction matrix 405, stores fixed-point numbers as integers, for example. The final addition of final_offset and the right shift operation by right_shift_result reduce the precision by rounding to obtain the fixed-point format required in the output.
[0183] To enable an increase in the range of real values representable by the integers of C, as shown in the embodiments of FIGS. 10 and 11, two additional matrices
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0184] According to one embodiment, the prediction parameters include weights each associated with a corresponding matrix component of the prediction matrix. That is, a predetermined prediction matrix is, for example, replaced by or represented by the prediction parameters. The weights are, for example, integers and / or fixed-point values. According to one embodiment, the prediction parameters further include one or more scaling factors, such as the value scale i , j each of which is associated with a weight, such as an integer value, associated with one or more corresponding matrix components of a predetermined prediction matrix 405
Number
Number
Number
Number
[0185] For example, in a preferred embodiment,
Number
Number
Number
Number
Number
Number
Number
Number
[0186] Offset representation
Number
Number
[0187] A wide range of embodiments resulting from that solution The above solution means the following embodiments. 1. A prediction method as described in Section I, wherein in step 2 of Section I, the following is performed on the integer approximation of the involved matrix vector product, namely, the (averaged) boundary samples
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0188] 2. A prediction method as described in section I, wherein in step 2 of section I, the following is performed for the integer approximation of the involved matrix-vector product, that is, the (averaged) boundary samples
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0189] 4. The prediction method described in Section I, wherein in step 2, one of the K matrices is used so that a plurality of prediction modes can be calculated, each being a different matrix with k = 0
Number
Number
Number
Number
[0190] To perform the prediction, a sample value vector 400 is formed from reference samples such as reference samples 17a and 17c. The possible formations have been described above. The formation can include averaging, thereby reducing the number of samples 102 or the number of components of vector 400 compared to the reference samples 17 contributing to the formation. The formation can also depend in some way on the dimensions or size of the block, such as the width and height of block 18, as described above. It is this vector 400 that should undergo an affine or linear transformation to obtain the prediction of block 18. Different nomenclatures have been used above. Using the latest one, it is the object to perform the prediction by applying vector 400 to matrix A by means of a matrix-vector product within the range of performing addition with offset vector b. The offset vector b is optional. The affine or linear transformation determined by A or A and b can be determined by the encoder and decoder for prediction, or more precisely, based on the size and dimensions of block 18 as already described above.
[0191] However, to achieve the improvement in computational efficiency outlined above or to make the prediction more effective with respect to implementation, an affine or linear transformation is quantized and the encoder and decoder, or its predictor, uses the C and T applied in the above method to represent and perform a linear or affine transformation using a quantized version of the affine transformation. In particular, instead of applying the vector 400 directly to the matrix A, the predictor of the encoder and decoder applies the vector 402 obtained from the sample value vector 400 by subjecting it to a mapping via a predetermined invertible linear transformation T. The transformation T used here is the same as long as the vector 400 has the same size, i.e., is the same regardless of the block dimensions, i.e., width and height, or is at least the same for different affine / linear transformations. Above, the vector 402 is denoted by y. The exact matrix for performing the affine / linear transformation determined by machine learning was B. However, instead of exactly performing B, the prediction in the encoder and decoder is performed by its approximate or quantized version. In particular, the representation is done by appropriately representing C in the manner outlined above, and C + M represents the quantized version of B.
[0192] Thus, the prediction in the coder and decoder is further accomplished by calculating the matrix-vector product 404 between the vector 402 and a predetermined prediction matrix C properly represented and stored in the coder and decoder in the above-described manner. Next, the vector 406 resulting from this matrix-vector product is used to predict the samples 104 of block 18. As described above, for prediction, each component of the vector 406 may receive a sum with the parameter a as shown at 408 to compensate for the corresponding definition of C. An optional sum of the vector 406 with the offset vector b may also be included in the derivation of the prediction of block 18 based on the vector 406. As described above, each component of the vector 406, thus each component of the sum of the vector 406, all the vectors of a shown at 408, and the optional vector b directly correspond to the samples 104 of block 18 and can thus indicate the predicted values of the samples. It is also possible that only a subset of the samples 104 of the block is predicted in this way and the remaining samples of block 18, such as 108, are derived by interpolation.
[0193] As described above, there are different embodiments for setting a. For example, it may be the arithmetic mean of the components of the vector 400. In that case, refer to FIG. 9. The invertible linear transformation T may be as shown in FIG. 9. i0 is a predetermined component of the sample value vector and the vector 402 respectively and is replaced by a. However, as also shown above, there are other possibilities. However, as far as the representation of C is concerned, it has also been shown above that it can be embodied differently. For example, the matrix-vector product 404 may, in its actual calculation, be the actual calculation of a smaller matrix-vector product having a lower dimension. In particular, as described above, due to the definition of C, the entire i0-th column 412 thereof becomes 0, so the actual calculation of the product 404 is
Number
[0194] According to one embodiment, the apparatus described herein for predicting a predetermined block of an image can be configured to use matrix-based intra prediction including the following features. The device is configured to form a sample value vector pTemp[x]400 from a plurality of reference samples 17. Assuming that pTemp[x] is 2*boundarySize, pTemp[x] may be populated, for example, by adjacent samples redT[x] located at the top of a given block, x = 0..boundarySize - 1, by direct copy or by subsampling or pooling, and subsequently by adjacent samples redL[x] located to the left of the given block, x = 0..boundarySize - 1 (for example, when isTransposed = 0), or vice versa in the case of a transposed process (for example, when isTransposed = 1).
[0195] An input value p[x] for x = 0..inSize - 1 is derived, that is, the device is configured to derive from the sample value vector pTemp[x] a further vector p[x] such that the sample value vector pTemp[x] is mapped by a given invertible linear transformation, or is a more specific given invertible affine linear transformation, as follows. - When -mipSizeId is equal to 2, the following applies. p[x]=pTemp[x + 1]-pTemp[0] - Otherwise (when mipSizeId is less than 2), the following applies. p[0]=(1<<(BitDepth - 1))-pTemp[0] p[x]=pTemp[x]-pTemp[0] (when x = 1) ··· inSize - 1 Here, the variable mipSizeId indicates the size of a given block. That is, according to this embodiment, the invertible transformation for deriving a further vector from the sample value vector depends on the size of a given block. The dependency may be given as follows.
[0196]
Table 3
[0197] The predSize indicates the number of predicted samples within a predetermined block, 2 * bondarySize indicates the size of the sample value vector, and inSize is related to the S size of the further vector, i.e., inSize, according to inSize = (2 * boundarySize) - (mipSizeId == 2)? 1: 0. More precisely, inSize indicates the number of components of the further vector that are actually involved in the calculation. inSize is the same size as the sample value vector for small block sizes and one component smaller for large block sizes. In the former case, one component, i.e., the component corresponding to a predetermined component of the further vector, such as in the matrix-vector product calculated later, may be decomposed, and the contribution of the corresponding vector component will be 0 anyway and thus does not actually need to be calculated. The dependence on the block size can be excluded in the case of alternative embodiments where only one of two alternatives is necessarily used, i.e., used regardless of the block size (the option corresponding to mipSizeId is less than 2 or the option corresponding to mipSizeId is equal to 2).
[0198] In other words, a given reversible linear transformation is defined such that, for example, a given component of a further vector p becomes a, and all other components correspond to the components obtained by subtracting a from the sample value vector. For example, a = pTemp[0]. In the case of the first option corresponding to mipSizeId equal to 2, this is easily seen, and only the separately formed components of the further vector are further considered. That is, in the case of the first option, the further vector is actually {p[0...inSize], pTemp[0]}, where pTemp[0] is such that the actually calculated part of the matrix-vector multiplication, i.e., the result of the multiplication, is limited to only the inSize components of the corresponding columns of the further vector and the matrix since the matrix has a column of 0s that does not require calculation. In other cases, corresponding to mipSizeId smaller than 2, each of all components of the further vector except p[0], i.e., the other components p[x] (x = 1...inSize - 1) of the further vector p except the given component p[0], is equal to the corresponding component - a of the sample value vector pTemp[x], but a = pTemp[0] is selected such that p[0] is the constant - a. Then, the matrix-vector product is calculated. The constant is the average of the representable values, i.e., 2 x-1 (i.e., 1 << (bit depth - 1)), where x indicates the bit depth of the computational representation used. It should be noted that if p[0] is selected to be pTemp[0] instead, the calculated product deviates from the product calculated using p[0] as above (p[0] = (1 << (BitDepth - 1)) - pTemp[0]) by only a constant vector that can be considered when predicting inside the block based on the product, i.e., the prediction vector. Thus, the value a is a given value, for example, pTemp[0]. The given value pTemp[0] is, in this case, for example, the component of the sample value vector pTemp corresponding to the given component p[0]. It can be the sample adjacent to the top of the given block or to the left of the given block that is closest to the upper left corner of the given block.
[0199] For example, in the case of the in-sample prediction process by predModeIntra, such as specifying the intra prediction mode, the apparatus is configured to apply, for example, the following steps and, for example, execute at least the first step. 1. For the matrix-based intra prediction sample predMip[x][y] where x = 0..predSize - 1 and y = 0..predSize - 1, it is derived as follows. - The variable modeId is set equal to predModeIntra. - The weight matrix mWeight[x][y] where x = 0..inSize - 1 and y = 0..predSize * predSize - 1 is derived by calling the MIP weight matrix derivation process with mipSizeId and modeId as inputs. - For the matrix-based intra prediction sample predMip[x][y] where x = 0..predSize - 1 and y = 0..predSize - 1, it is derived as follows. oW = 32 - 32 * (
Number
Number
[0200] In other words, the apparatus performs a matrix-vector product between additional vectors p[i], or when mipSizeId is equal to 2, {p[i], pTemp[0]} and a predetermined prediction matrix mWeight, or when mipSizeId is less than 2, the prediction matrix mWeight already has additional zero-weight lines corresponding to the omitted components of p in order to obtain the prediction vectors assigned to the array of block positions {x, y} distributed inside a predetermined block so as to result in the array predMip[x][y] here. The prediction vectors correspond respectively to the concatenation of the rows of predMip[x][y] or the columns of predMip[x][y]. According to one embodiment, or according to a different interpretation, the component ((
Number
Number
[0201] For the matrix-based intra prediction sample predMip[x][y] where 2.x = 0..predSize - 1 and y = 0..predSize - 1, it is clipped as follows, for example. predMip[x][y] = Clip1(predMip[x][y]) When isTransposed is equal to TRUE, the predSize×predSize array predMip[x][y] where x = 0..predSize - 1 and y = 0..predSize - 1 is transposed as follows, for example. predTemp[y][x] = predMip[x][y] predMip = predTemp For the prediction sample predSamples[x][y] where x = 0..nTbW - 1 and y = 0..nTbH - 1, it is derived as follows, for example. - If nTbW specifying the transform block width is greater than predSize or nTbH specifying the transform block height is greater than predSize, the MIP prediction upsampling process is called with as input a matrix-based intra prediction sample predMip[x][y] having an input block size predSize, where x = 0..predSize - 1, y = 0..predSize - 1, a transform block width nTbW, a transform block height nTbH, an upper reference sample refT[x] having x = 0..nTbW - 1, and a left reference sample refL[y] having y = 0..nTbH - 1, and the output is a prediction sample array predSamples. - Otherwise, predSamples[x][y], where x = 0..nTbW - 1, y = 0..nTbH - 1, is set equal to predMip[x][y]. In other words, the apparatus is configured to predict samples predSamples of a given block based on a prediction vector predMip.
[0202] 8 Embodiments using a block-based intra prediction mode together with other intra prediction modes All of the above descriptions shall be considered as optional implementation details of the embodiments described herein. Note that hereinafter, the term block-based intra prediction is used to indicate an intra prediction mode that can be embodied as or equal to that indicated by the above ALWIP.
[0203] FIG. 12 shows an embodiment of an apparatus 3000 for decoding a given block 18 of an image 10 using intra prediction. The apparatus 3000 is configured to derive a set selection syntax element 522 from a data stream 12 indicating whether a given block 18 is to be predicted using one of a first set 508 of intra prediction modes including a DC intra prediction mode 506 and an angular prediction mode 500. The data stream 12 can include different syntax elements and / or indices indicating the functionality of the apparatus 3000.
[0204] When the set-selective syntax element 522 indicates that a given block 18 is to be predicted using one of the first set 508 of intra prediction modes, the apparatus 3000 is configured to form a list 528 of the most likely intra prediction modes based on the intra prediction mode 3050 when adjacent blocks 524, 526 adjacent to the given block 18 are predicted. In other words, the intra prediction modes 506, 500 of the first set 508 of intra prediction modes are positioned / placed in the list 528 of the most likely intra prediction modes based on the intra prediction mode 3050 used for predicting the adjacent blocks 524 and 526. The apparatus 3000 is configured to, for example, save the prediction modes used for the already predicted blocks and obtain the prediction modes 3050 used for the adjacent blocks 524 and 526 among the saved prediction modes, or analyze the adjacent blocks 524 and 526 to obtain the prediction modes 3050 used for the adjacent blocks 524 and 526. According to one embodiment, the apparatus 3000 is configured to search for intra prediction modes equal to or similar to the prediction modes 3050 used for the adjacent blocks 524 and 526 within the first set 508 of intra prediction modes, and form a list 528 of the most likely intra prediction modes among these equal or similar intra prediction modes.
[0205] The list 528 of the most likely intra prediction modes is formed such that when the adjacent blocks 524 and 526 are exclusively predicted by any of the angular intra prediction modes 506, the list 528 of the most likely intra prediction modes does not include the DC intra prediction mode 500. Thus, the availability of the DC intra prediction mode 506 depends only on the adjacent blocks 524 and 526 of a given block 18 and not on other blocks of the image 10. The DC intra prediction mode 506 is not placed / arranged in the list 528 of the most likely intra prediction modes, for example, when at least one of the adjacent blocks 524 or 526 is predicted using the angular intra prediction mode 500. The list 528 of the most likely intra prediction modes may also not include the DC intra prediction mode 506 when both of the adjacent blocks 524 and 526 are predicted using the angular intra prediction mode 500.
[0206] Furthermore, the apparatus is configured to derive an MPM list index 534 from a data stream when the set selective syntax element 522 indicates that a given block 18 is to be predicted using one of the first set 508 of intra prediction modes. The MPM list index 534 directs the list 528 of the most likely intra prediction modes to a given intra prediction mode. The apparatus 3000 is configured to intra predict a given block 18 using a given intra prediction mode 3100.
[0207] When the set - selective syntax element 522 indicates that a given block 18 should not be predicted using one of the first set 508 of intra - prediction modes, the apparatus 3000 is configured to derive from the data stream 12 a further index 540 that indicates a second set 520 of matrix - based intra - prediction modes, i.e., a given matrix - based intra - prediction mode among the block - based intra - prediction modes 510, i.e., a given block - based intra - prediction mode 3200. In other words, a given block - based intra - prediction mode 3200 among the second set 520 of block - based intra - prediction modes is selected based on the further index 540 for the prediction of the given block 18. The apparatus 3000 is configured to calculate a matrix - vector product 512 between a vector 514 derived from reference samples 17 in the neighborhood of the given block 18 and a given prediction matrix 516 associated with the given matrix - based intra - prediction mode 3200 to obtain a prediction vector 518, and when the set - selective syntax element 522 indicates that a given block 18 is not predicted using one of the first set 508 of intra - prediction modes, the apparatus 3000 is configured to predict samples of the given block 18 based on the prediction vector 518. The data stream 12 includes either an MPM list index 534 or a further index 540 based on the set - selective syntax element 522. The apparatus 3000 can include features and / or functions as described with respect to FIG. 13.
[0208] Accordingly, the embodiments described below with respect to FIG. 13 relate to a decoder and an encoder that support intra prediction for decoding / encoding a predetermined block 18 for which different intra prediction modes are supported. To obtain an intra prediction signal for the predetermined block 18, there is an angular intra prediction mode 500 in which reference samples 17 adjacent to the predetermined block 18 are used to satisfy the predetermined block 18. In particular, the reference samples 17 arranged along the boundary of the predetermined block 18, such as along the upper and left ends of the predetermined block 18, represent image content that is extrapolated or copied inside the predetermined block 18 along a predetermined direction 502. Before extrapolation or copying, the image content represented by the adjacent samples 17 may be subject to interpolation filtering, or in other words, in the case of interpolation filtering, may be derived from the adjacent samples 17. The angular intra prediction modes 500 have different intra prediction directions 502 from each other. Each angular intra prediction mode 500 can have an associated index, and the association of the index to the angular intra prediction mode 500 can be such that the direction 500 monotonically rotates clockwise or counterclockwise when ordering the angular intra prediction modes 500 according to the associated mode index.
[0209] There can also be a non-angular intra prediction mode. For example, in FIG. 13, 504 shows a planar intra prediction mode optionally included in the set 508, according to which a two-dimensional linear function defined by a horizontal gradient, a vertical gradient, and an offset is derived based on the adjacent samples 17, and this linear function defines the predicted sample values of the predetermined block 18. The horizontal gradient, the vertical gradient, and the offset are derived based on the adjacent samples 17. According to one embodiment, the first set 508 of intra prediction modes includes the planar intra prediction mode 504. The DC mode, which is a specific non-angular intra prediction mode included in set 508, is shown at 506. Here, based on neighboring sample 17, one value that is pseudo-DC value is derived, and this one DC value is for all samples of a given block 18 to obtain an intra prediction signal. Two examples of non-intra prediction modes are shown, but there may be only one or three or more.
[0210] Intra prediction modes 500, 504, and 506 form a set 508 of intra prediction modes supported by an encoder and a decoder, and the encoder and the decoder compete with a block-based intra prediction mode generally denoted using reference numeral 510 in terms of rate / distortion optimization, and in that example, it was described above using the abbreviation ALWIP. As described above, according to these block-based intra prediction modes 510, on the one hand, a matrix-vector product 520 is performed between a vector 514 derived from adjacent samples 17 and, on the other hand, a given prediction matrix 516. The result of multiplication 512 is a prediction vector 518 used to predict the samples of a given block 18. The block-based intra prediction modes 510 are different from each other in the prediction matrices 516 associated with their respective modes. Thus, to summarize briefly, the encoder and decoder according to the embodiments described herein include a set 508 of intra prediction modes, i.e., a first set of intra prediction modes, and a set 520 of block-based intra prediction modes, i.e., a second set of matrix-based intra prediction modes, and they compete with each other.
[0211] According to an embodiment of the present application, a predetermined block 18 is encoded / decoded using intra prediction in the following manner. In particular, first, a set selection syntax element 522 that determines whether the predetermined block 18 is predicted using any of a set 508 of intra prediction modes or any of a set 520 of block-based intra prediction modes. If the set-selected syntax element indicates that the predetermined block 18 should be predicted using any mode of set 508, i.e., the first set of intra prediction modes, a list 528 of the most likely candidates from set 508 is interpreted / formed in the decoder and encoder based on which intra prediction mode the adjacent blocks 18, and the adjacent blocks exemplarily shown at 524 and 526, were predicted using. The adjacent blocks 524 and 526 can be determined relative to the position of the predetermined block 18 in a predetermined manner, such as by overlaying specific adjacent samples of block 18, such as samples, over the top left sample of block 18, and determining block 526 that includes samples to the left of the above-described corner sample. Of course, this is just an example. The same applies to the number of adjacent blocks used for mode prediction, which is not limited to two for all embodiments. There may be more than two or just one. If any of these blocks 524 and 526 are missing, a default intra prediction mode can be used by default in place of the intra prediction mode of the missing adjacent block. The same may be true if either of blocks 524 and 526 is encoded / decoded using an inter prediction mode, such as by motion compensation prediction.
[0212] The configuration of the list of modes from set 508, i.e., the list 528 of the most likely intra prediction modes, is as follows. The list length of list 528, i.e., the number of the most likely modes therein, can be fixed by default. The length may be 4 as shown in FIG. 13, or may be different, such as 5 or 6. In the latter case, it applies to the specific examples described below. The index in the data stream described later may indicate one of the modes of list 528 used for a given block 18. The indexing is performed along the list order or ranking 530, and the list index is a variable length encoded such that, for example, the length of the index monotonically increases along the order 530. Thus, first, only the most likely modes from set 508 are added to list 528, and it is valuable to place the more likely modes upstream along the order 530 compared to the modes less likely to be suitable for block 18. The modes of list 528 are derived based on the modes used for blocks 524 and 526, i.e., the adjacent blocks adjacent to a given block 18. If either of blocks 524 and 526 is intra predicted using the block-based mode 510 of set 520, the aforementioned mapping from such an "ALWIP" or block-based mode 510 to a mode within set 508, e.g., a non-ALWIP mode, is used. The latter mapping can map, for example, most (i.e., more than half) of the block-based mode 510 to the DC mode 506 (or either DC506 or planar mode 504).
[0213] According to one embodiment, the list 528 of the most likely intra prediction modes is filled with the planar intra prediction mode 504 in a way independent of the intra prediction mode using when neighboring blocks are predicted. Thus, for example, only the DC intra prediction mode 506 and the angular intra prediction mode 500 are incorporated into the list 528 depending on the intra prediction modes used for predicting the neighboring blocks 524 and 526. The planar intra prediction mode 504 is placed at the first position of the list 528 of the most likely intra prediction modes independent of the intra prediction mode, for example, using when the adjacent blocks 524 and 526 are predicted.
[0214] As will be illustrated in more detail below, the list construction of the list 528 of the most likely intra prediction modes is performed such that the list 528 does not include the DC intra prediction mode 500 when the neighboring blocks 524 and 526 are exclusively predicted by the angular intra prediction mode 506. The DC intra prediction mode 506 is not in the list 528 of the most likely intra prediction modes when one of the adjacent blocks 524 or 526 is predicted by any of the angular intra prediction modes 500, and / or when both of the adjacent blocks 524 and 526 are predicted by any of the angular intra prediction modes 500. According to the embodiments shown below in this specification, for example, the list 528 is filled with the DC mode 506 only when all of the following situations apply to all of the adjacent blocks 524 and 526: either is encoded using either of the non-angular intra prediction modes 504 and 506, or either is intra-predicted using the block-based intra prediction mode 510, which is mapped to either of the non-angular intra prediction modes 504 and 506 by the aforementioned mapping from the block-based intra prediction mode 510 to the modes within the set 508. Only in that case is the DC intra prediction mode 506 positioned in the list 528. In that case, as can be seen from the subsequent examples, it can be placed before any of the angular intra prediction modes 500 in the order 530. In other words, the list 528 of the most likely intra prediction modes is predicted using, for each of the adjacent blocks 524 and 526, either at least one non-angular intra prediction mode 504 or 506 having a first set 508 including the DC intra prediction mode 506, or is mapped to any of the non-angular intra prediction modes 500 by mapping to the intra prediction modes within the first set 506 used to form the list 528 of the most likely intra prediction modes from a second set 520 of block-based intra prediction modes 510, and is filled with the DC intra prediction mode 508 only in the case of each adjacent block predicted using any of the block-based intra prediction modes 510 mapped to any of the non-angular intra prediction modes 500.
[0215] Accordingly, resuming the description of how a given block 18 is encoded in the data stream 12, the data stream 12 optionally includes an MPM syntax element 532 indicating whether the intra prediction mode to be used for the given block 18 is in the list 528 when the set selection syntax element 522 indicates that the given block 18 should be encoded by any of the modes in the first set 508. If so, the data stream 12 includes an MPM list index 534 indicating the mode to be used for the given block 18 in the list 528, i.e., the given intra prediction mode, by indexing the same along the order 530. However, if the mode from the set 508 indicated by the MPM syntax element 532 is not in the list 528, the data stream 12 includes a further syntax element 536 indicating which mode, i.e., the given intra prediction mode, should be used for the block 18 from the set 508. The further syntax element 536 can indicate the mode by simply distinguishing the mode from the set 508 not included in the list 528. In other words, the apparatus 3000 is configured to derive from the data stream an MPM syntax element 532 indicating whether a given intra prediction mode of the first set 508 of intra prediction modes is within a list 528 of the most likely intra prediction modes, and the set selection syntax element 522 indicates that a given block 18 is predicted using one of the first set 508 of intra prediction modes. When the MPM syntax element 532 indicates that a given intra prediction mode of the first set 508 of intra prediction modes is within the list 528 of the most likely intra prediction modes, the apparatus 3000 is configured to, for example, perform formation of the list 528 of the most likely intra prediction modes using when adjacent blocks 524, 526 adjacent to a given block 100 are predicted, and perform derivation of an MPM list index 534 from the data stream 12 indicating the given intra prediction mode of the list 528 of the most likely intra prediction modes. When the MPM syntax element 532 from the data stream 12 indicates that a given intra prediction mode of the first set 508 of intra prediction modes is not within the list 528 of the most likely intra prediction modes, the apparatus 3000 is configured to derive a further list index 536 from the data stream indicating the given intra prediction mode of the first set of intra prediction modes. Thus, based on the MPM syntax element 532, the data stream 12 includes either an MPM list index 534 or a further list index 536 for prediction of a given block 18.
[0216] By removing the situation where list 528 includes the DC intra prediction mode 506, the following advantages are achieved. In particular, the inventors of the present application have found that in order to encode / decrypt a predetermined block 18 that should use any of the intra prediction modes of set 508 indicated by syntax element 522, i.e., the set-selective syntax element, "consuming" the valuable list positions of list 528 having the DC intra prediction mode 506 in set 508 will, in any case, have an adverse effect on the encoding efficiency because such a DC intra prediction mode 506 in set 508 competes with the block-based intra prediction mode 510. Therefore, "consuming" the list positions of list 528 using such a DC intra prediction mode 506 from set 508 increases the possibility of a situation where the intra prediction mode that should ultimately be used for the predetermined block 18, i.e., the predetermined intra prediction mode, is not within list 528, and as a result, syntax element 536, i.e., an additional list index, needs to be transmitted within data stream 12. In particular, since syntax element 522 already indicates whether block 18 should be predicted using any of the modes in set 508 or any of the block-based modes 510 of set 520, if syntax element 522 indicates that the mode in set 508 is preferable for block 18 and thus the block-based mode 510 is not used for block 18, the possibility that the DC prediction mode 506 of set 508 is suitable for block 18 is very low. Therefore, its occurrence in list 528 should be restricted to a very limited set of constellations of the modes used for adjacent blocks 524 and 526, i.e., the constellations set above.
[0217] In the case of the other party, that is, when the set selective syntax element 522 indicates that a predetermined block 18 is predicted using any of the block-based intra prediction modes 510, the encoding of the block 18 into the data stream 12 and its decoding therefrom can be performed in the above-described manner. For this purpose, indexing can be used to index or indicate the selected one of the block-based intra prediction modes 510 in the set 520, that is, the second set of block-based intra prediction modes. A further MPM syntax element 538 can indicate whether the indexing is performed by the index 540, that is, by a further MPM list index indicating the block-based intra prediction mode 510 to be used for the block 18 in the list 542 of the most likely block-based intra prediction modes 510, that is, by indexing along the list order 544, or whether the block-based intra prediction mode 510 to be used for the block 18 is indicated by a further syntax element 546, that is, by a still further list index indicating the block-based intra prediction mode 510 in the set 520, and the latter syntax element 546 can distinguish only between those modes 510 in the set 520 that are not yet included in the list 542, for example. The list construction of the list 542 can be performed based on the modes in which the blocks 524 and 526 are predicted. If either of the blocks 524 and 526 is outside the image or not available because it is inter-predicted, a default intra prediction mode such as one of the set 508 can be used instead.For each of the blocks 524 and 526 intra-predicted using a mode outside of set 508 rather than set 520, the aforementioned mapping from the mode of set 508 to a mode outside of set 520 is used to obtain the intra-prediction mode 510 of a respective block, i.e., a predetermined block 18, i.e., a block-based intra-prediction mode based on a predetermined block, and list 542 is interpreted based on the block-based intra-prediction mode resulting from blocks 524 and 526.
[0218] According to one embodiment, the apparatus 3000 is configured to derive from the data stream 12 a further MPM syntax element 538 indicating whether a predetermined block-based intra prediction mode of a second set 520 of block-based intra prediction modes 510 is within a list 542 of most likely block-based intra prediction modes when a set-selective syntax element 522 indicates that a predetermined block 18 is not predicted using one of a first set 508 of intra prediction modes. When the further MPM syntax element 538 indicates that a predetermined block-based intra prediction mode of the second set 520 of block-based intra prediction modes 510 is within the list 542 of most likely block-based intra prediction modes, the apparatus 3000 forms, for example, the list 542 of most likely block-based intra prediction modes using when adjacent blocks 524, 526 adjacent to the predetermined block 18 are predicted, and is configured to derive from the data stream 12 a further MPM list index 540 indicating the list 542 of most likely block-based intra prediction modes over the predetermined block-based intra prediction mode. When the further MPM syntax element 538 indicates that a predetermined block-based intra prediction mode of the second set 520 of block-based intra prediction modes 510 is not within the list 542 of most likely block-based intra prediction modes, the apparatus 3000 is configured to derive from the data stream 12 a yet further list index 546 indicating the predetermined block-based intra prediction mode of the second set 520 of block-based intra prediction modes. Thus, based on the further MPM syntax element 538, the data stream 12 includes either a further MPM list index 540 or a yet further list index 546 for prediction of the predetermined block 18.
[0219] A further MPM syntax element 538, a further MPM list index 540, and yet another list index 546 are represented in data stream 12 of FIG. 13 as being parallel to MPM syntax element 532, MPM list index 534, and further list index 536, but it is clear that either the further MPM syntax element 538, or an index associated with the further MPM syntax element 538, such as either the further MPM list index 540 or yet another list index 546, is constituted by data stream 12 or MPM syntax element 532, and an index associated with MPM syntax element 532, such as MPM list index 534 or further list index 536. Which of these syntax elements and indexes are included in data stream 12 depends, for example, on set-selective syntax element 522. An example of the syntax element portion of data stream 12 written as pseudocode might look like the following, where the reference numerals indicate which syntax elements correspond to the foregoing syntax elements.
[0220] [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5]
[0221] The list construction of list 528 may be defined as follows, where candIntraPredModeA / B indicates the intra prediction mode in which either of blocks 524 and 526, such as A of block 524 and B of block 526, is predicted if the corresponding block 524 or 526 is intra predicted using any of the block-based intra prediction modes 510, or the mode to which the intra prediction mode is mapped from set 508. INTRA_DC is used to indicate mode 506, and the angular mode 500 is indicated by INTRA_ANGULAR#, and the angular modes are ordered and numbered (#) such that, as exemplified above, i.e., the angular direction 502 monotonically decreases or increases with the increase of the number. The order among the modes within set 508 may be as defined in the subsequent table where INTRA_PLANAR indicates mode 504.
[0222] Note that in the above example, the index 534 is actually distributed over the syntax elements 534’ and 534’’. That is, the former 534’ is unique to the first, ordered 530 position of list 528 where, according to this example, the INTRA_PLANAR mode 504 is necessarily positioned. The latter 534’’ refers to any of the subsequent positions of list 528, and as explained, the DC mode 506 is included only in the special situation explained.
[0223] Furthermore, in the above example, in the case of a syntax element 522 indicating that any one of the modes within set 508 is used, a further syntax element is included in data stream 12, and this further syntax element parameterizes the intra prediction mode within set 508 in some way. For example, syntax element 600 parameterizes or varies the region where reference sample 17 is located, and based on that, the mode within set 508 intra-predicts the inside of block 18 with respect to, for example, the distance towards the outer perimeter of block 18. In addition to or instead of this, syntax element 602 parameterizes or changes whether reference sample 17 is used by a mode within set 508 to globally or within-block intra-predict the inside of block 18, or whether the intra-prediction is performed on a subdivided part or parts of block 18 and sequentially intra-predicted, so that the prediction residual encoded in data stream 12 for one part can help to adopt a new reference sample for intra-predicting a subsequent part. The latter encoding option controlled by the syntax element may only be available (and the corresponding syntax element may only exist in the data stream) when syntax element 600 has a predetermined state corresponding to, for example, the region where reference sample 17 adjacent to block 18 exists. The parts may be defined by subdividing along a predetermined direction such as either horizontally, which results in parts of the same height as block 18, or vertically, which results in parts of the same width as block 18. If the splitting is signaled to be active, a syntax element 604 may exist in the data stream to control which splitting direction is used.As can be seen, the positions within list 528 reserved for the INTRA_PLANAR mode are only available in the case of specific parameterizations of the mode by the above-described parameterized syntax elements, such as when the syntax element 600 has a predetermined state corresponding to a region where a reference sample 17 adjacent to block 18 exists, and / or when the partial-wise intra prediction mode is not active as notified by the syntax element 602.
[0224] All syntax elements shown in the table and not specifically described above are optional and will not be further described herein. - If candIntraPredModeB is equal to candIntraPredModeA and candIntraPredModeA is greater than INTRA_DC, candModeList[x] for x = 0..4 is derived as follows. candModeList[0]=candIntraPredModeA candModeList[1]=2+((candIntraPredModeA+61)% 64) candModeList[2]=2+((candIntraPredModeA-1)% 64) candModeList[3]=2+((candIntraPredModeA+60)% 64) candModeList[4]=2+(candIntraPredModeA % 64) - Otherwise, if candIntraPredModeB is not equal to candIntraPredModeA and either candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC, the following applies. - The variables minAB and maxAB are derived as follows.
[0225] minAB=Min(candIntraPredModeA,candIntraPredModeB) maxAB = Max(candIntraPredModeA, candIntraPredModeB) - When both candIntraPredModeA and candIntraPredModeB are greater than INTRA_DC, candModeList[x] where x = 0..4 is derived as follows. candModeList[0] = candIntraPredModeA candModeList[1] = candIntraPredModeB - When maxAB - minAB is equal to 1, the following applies. candModeList[2] = 2 + ((minAB + 61) % 64) candModeList[3] = 2 + ((maxAB - 1) % 64) candModeList[4] = 2 + ((minAB + 60) % 64) - Otherwise, when maxAB - minAB is 62 or more, the following applies. candModeList[2] = 2 + ((minAB - 1) % 64) candModeList[3] = 2 + ((maxAB + 61) % 64) candModeList[4] = 2 + (minAB % 64) - Otherwise, when maxAB - minAB is equal to 2, the following applies. candModeList[2] = 2 + ((minAB - 1) % 64) candModeList[3] = 2 + ((minAB + 61) % 64) candModeList[4] = 2 + ((maxAB - 1) % 64) - Otherwise, the following applies. candModeList[2] = 2 + ((minAB + 61) % 64) candModeList[3] = 2 + ((minAB - 1) % 64)(8 - 36) candModeList[4] = 2 + (((maxAB + 61)) % 64) - Otherwise (if candIntraPredModeA or candIntraPredModeB is greater than INTRA_DC), candModeList[x] for x = 0..4 is derived as follows. candModeList[0]=maxAB candModeList[1]=2+((maxAB+61)% 64) candModeList[2]=2+((maxAB-1)% 64)(8-41) candModeList[3]=2+((maxAB+60)% 64) candModeList[4]=2+(maxAB % 64) - Otherwise, the following applies. candModeList[0]=INTRA_DC candModeList[1]=INTRA_ANGULAR50( candModeList[2]=INTRA_ANGULAR18 candModeList[3]=INTRA_ANGULAR46 candModeList[4]=INTRA_ANGULAR54
[0226]
Table 5
[0227] In the case of a set selection syntax element 522 indicating that a predetermined block 18 should not be predicted using one of the first set 508 of intra prediction modes, the apparatus for decoding the predetermined block 18 and / or the apparatus for encoding the predetermined block 18 may include one or more of the following features. According to one embodiment, the apparatus is configured to form a sample value vector from a plurality of reference samples 17, for example, the sample value vector 400 described with respect to one of the embodiments of FIGS. 6 to 9, and to derive a vector 514 from the sample value vector such that the sample value vector is mapped to the vector 514 by a predetermined reversible linear transformation. In this case, the vector 514 can be understood as a further vector. The vector 514 is determined and / or defined as described for a further vector 402 with respect to one of the embodiments of FIGS. 8 to 11, for example.
[0228] According to one embodiment, the apparatus is configured to form a sample value vector from a plurality of reference samples 17 by adopting, for each component of the sample value vector, one of the plurality of reference samples as each respective component of the sample value vector and / or by averaging two or more components of the sample value vector to obtain each respective component of the sample value vector. The plurality of reference samples 17 are arranged, for example, within the image along the outer edge of a predetermined block 18. The reversible linear transformation is defined such that, for example, a predetermined component of the vector 514, for example a predetermined component of a further vector, becomes a, and each of the other components of the vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a. The value a is, for example, a predetermined value.
[0229] According to one embodiment, the predetermined value is an average such as an arithmetic mean or a weighted mean of the components of the sample value vector, a default value, a value signaled within a data stream in which the image is encoded, and one of the components of the sample value vector corresponding to the predetermined component. The reversible linear transformation is defined such that, for example, a predetermined component of the vector 514, for example a predetermined component of a further vector, becomes a, and each of the other components of the vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the arithmetic mean of the components of the sample value vector. A reversible linear transformation is defined such that, for example, a predetermined component of vector 514, such as a predetermined component of a further vector, becomes a, and each of the other components of vector 514 excluding the predetermined component is equal to the corresponding component of the sample value vector minus a, where a is the component of the sample value vector corresponding to the predetermined component. The apparatus comprises, for example, a plurality of reversible linear transformations each associated with one component of vector 514, and is configured to select a predetermined component from the components of the sample value vector and use the reversible linear transformation from among the plurality of reversible linear transformations associated with the predetermined component as the predetermined reversible linear transformation.
[0230] According to one embodiment, for example, all of the matrix components of prediction matrix 516 within the column of prediction matrix 516 corresponding to the predetermined component of vector 514 of a further vector are zero. The apparatus is configured to perform multiplication by calculating matrix-vector product 512 between a reduced prediction matrix resulting from excluding the column from prediction matrix 516 and yet another vector resulting from excluding the predetermined component from vector 514. According to one embodiment, when predicting samples of a predetermined block 18 based on prediction vector 518, the apparatus is configured to calculate, for each component of prediction vector 518, the sum of the respective component and a. The matrix obtained by summing, for example, the multiple (i.e., matrix B in FIG. 8) of each matrix component of prediction matrix 516 within the column of prediction matrix 516 corresponding to the predetermined component of vector 514 of a further vector, corresponds to, for example, a quantized version of a machine learning prediction matrix.
[0231] According to one embodiment, the apparatus is configured to calculate matrix-vector product 512 using fixed-point arithmetic. According to one embodiment, the apparatus is configured to calculate matrix-vector product 512 without using floating-point arithmetic. According to one embodiment, the apparatus is configured to store a fixed-point number representation of prediction matrix 516.
[0232] According to one embodiment, the apparatus is configured to represent the prediction matrix 516 using prediction parameters, and to calculate a matrix-vector product 512 by performing multiplication and addition on, for example, the components of the vector of vectors 514 of further vectors, and the prediction parameters and the intermediate results resulting therefrom. The absolute value of the prediction parameters can be represented by an n-bit fixed-point decimal representation, where n is 14 or less, or alternatively 10, or alternatively 8. This can be performed in a similar manner, or as described in FIGS. 10 or 11. The prediction parameters include, for example, weights each associated with a corresponding matrix component of the prediction matrix 516. The prediction parameters further include, for example, one or more scaling coefficients each associated with a corresponding matrix component of the prediction matrix 516 for scaling the weights each associated with a corresponding matrix component of the prediction matrix 516, and / or one or more offsets each associated with a corresponding matrix component of the prediction matrix 516 for offsetting the weights each associated with a corresponding matrix component of the prediction matrix 516.
[0233] According to one embodiment, when predicting samples of a predetermined block 18 based on a prediction vector 518, the apparatus is configured to use interpolation to calculate at least one sample position of the predetermined block 18 based on the prediction vector 518 in which each component is associated with a corresponding position within the predetermined block 18.
[0234] FIG. 14 shows an apparatus 6000 for encoding a predetermined block 18 of an image 10 using intra prediction. The apparatus is configured to signal a set selective syntax element 522 in a data stream 12 indicating whether the predetermined block 18 is to be predicted using one of a first set 508 of intra prediction modes including a DC intra prediction mode 506 and an angular prediction mode 500. When the set selective syntax element 522 indicates that the predetermined block 18 is to be predicted using one of the first set 508 of intra prediction modes, the apparatus 6000 forms a list 528 of the most likely intra prediction modes based on the intra prediction modes in which adjacent blocks 524, 526 adjacent to the predetermined block 18 are predicted, signals an MPM list index 534 in the data stream 12 that points to the list 528 of the most likely intra prediction modes on a predetermined intra prediction mode 3100, and is configured to intra predict the predetermined block 18 using the predetermined intra prediction mode 3100. When the set selective syntax element 522 indicates that the predetermined block 18 is not to be predicted using one of the first set 508 of intra prediction modes, the apparatus 6000 calculates a matrix-vector product 512 between a vector 514 derived from reference samples 17 in the vicinity of the predetermined block 18 and a predetermined prediction matrix 516 associated with a predetermined matrix-based intra prediction mode 3200 to obtain a prediction vector 518, and signals a further index 540 in the data stream 12 indicating the predetermined matrix-based intra prediction mode 3200 of a second set 520 of matrix-based intra prediction modes 510 by predicting samples of the predetermined block 18 based on the prediction vector 518.
[0235] The list 528 of the most likely intra prediction modes is formed based on the intra prediction modes used by the adjacent blocks 524, 526 adjacent to a given block 18 such that the list of the most likely intra prediction modes does not include the DC intra prediction mode 506 when at least one of the adjacent blocks 524, 526 is predicted by any of the angular intra prediction modes 500. The first set 508 of intra prediction modes further includes, for example, the planar intra prediction mode 504.
[0236] According to one embodiment, the apparatus 6000 can include features and / or functions similar to those described with respect to the apparatus 3000 of FIGS. 12 and / or 13. Similarly, the apparatus 3000 can include features and / or functions similar to those described with respect to the apparatus 6000 of FIG. 14. When the set - selective syntax element 522 indicates that a given block 18 is to be predicted using one of the first set 508 of intra - prediction modes, the apparatus 6000 is configured to signal, for example, the MPM syntax element 532 in the data stream 12 indicating whether a given intra - prediction mode 3100 of the first set 508 of intra - prediction modes is within the list 528 of most - likely intra - prediction modes. When the MPM syntax element 532 indicates that a given intra - prediction mode of the first set 508 of intra - prediction modes is within the list 528 of most - likely intra - prediction modes, the apparatus 6000 is configured to perform, for example, the formation of the list 528 of most - likely intra - prediction modes based on the intra - prediction modes in which the adjacent blocks 524, 526 adjacent to the given block 18 are predicted, and to signal the MPM list index 534 in the data stream 12 that points to the given intra - prediction mode 3100 of the list 528 of most - likely intra - prediction modes. When the MPM syntax element 532 in the data stream 12 indicates that a given intra - prediction mode 3100 of the first set 508 of intra - prediction modes is not within the list 528 of most - likely intra - prediction modes, the apparatus 6000 is configured to signal, for example, an additional list index 536 in the data stream 12 that indicates the given intra - prediction mode 3100 of the first set 508 of intra - prediction modes.
[0237] If the set selective syntax element 522 indicates that a given block 18 should not be predicted using one of the first set 508 of intra prediction modes, the apparatus 6000 is configured to signal, for example, an additional MPM syntax element 538 within the data stream 12 indicating whether a given block - based intra prediction mode 3200 of the second set 520 of block - based intra prediction modes is within the list 542 of most likely block - based intra prediction modes, i.e., within the second list of most likely block - based intra prediction modes, for example, the list 542 of most likely block - based intra prediction modes of the second set 520 of block - based intra prediction modes. If the additional MPM syntax element 538 indicates that a given block - based intra prediction mode 3200 of the second set 520 of block - based intra prediction modes is within the list 542 of most likely block - based intra prediction modes, the apparatus 6000 is configured to signal, for example, an additional MPM list index 540 within the data stream 12 that points to the given block - based intra prediction mode 3200 on the list 542 of most likely block - based intra prediction modes, based on the intra prediction mode 3050 in which the adjacent blocks 524, 526 adjacent to the given block 18 are predicted. If the additional MPM syntax element 538 indicates that a given block - based intra prediction mode 3200 of the second set 520 of block - based intra prediction modes is not within the list 542 of most likely block - based intra prediction modes, the apparatus 6000 is configured to signal, for example, an additional list index 546 within the data stream 12 indicating the given block - based intra prediction mode 3200 of the second set 520 of block - based intra prediction modes.
[0238] According to one embodiment, for each of the adjacent blocks 524 and 526, the apparatus 6000 is predicted using either at least one non-angular intra prediction mode 504 and / or 506 having a first set 508 including the DC intra prediction mode 506, or is predicted by mapping to an intra prediction mode within the first set 508 from a second set 520 of block-based intra prediction modes used to form the list 528 of the most likely intra prediction modes. The list 528 incorporates the DC intra prediction mode 528 only if each of the adjacent blocks predicted using any of the block-based intra prediction modes 510 is predicted using the DC intra prediction mode 506. The apparatus 6000 is configured to execute the formation of the list 528 of the most likely intra prediction modes of the first set 508 of intra prediction modes using which of the adjacent blocks 524, 526 is adjacent to a given block 18, and is mapped to either at least one non-angular intra prediction mode 504 and / or 506.
[0239] According to one embodiment, for each of the adjacent blocks 524 and 526, the apparatus 6000 is predicted using either at least one non-angular intra prediction mode 504 and / or 506 having a first set 508 including the DC intra prediction mode 506, or is predicted by mapping from a second set 520 of block-based intra prediction modes used to form the list 528 of the most likely intra prediction modes to an intra prediction mode within the first set 508, to either at least one non-angular intra prediction mode 504 and / or 506. The apparatus 6000 is configured to execute the formation of the list 528 of the most likely intra prediction modes based on the intra prediction mode by which the adjacent blocks 524, 526 predicted to be adjacent to a given block 18 using it are predicted. The DC intra prediction mode 506 is placed before any angular intra prediction mode 500 within the list 528 of the most likely intra prediction modes.
[0240] According to one embodiment, the apparatus 6000 is configured to perform the formation of a list 528 of the most likely intra prediction modes based on an intra prediction mode in which adjacent blocks 524, 526 adjacent to a predetermined block 18 are predicted. As a result, the list 528 is populated with a planar intra prediction mode 504 in a manner independent of the intra prediction mode in which the adjacent blocks 524, 526 are predicted. According to one embodiment, the apparatus 6000 is configured to perform the formation of a list 528 of the most likely intra prediction modes based on an intra prediction mode in which adjacent blocks 524, 526 adjacent to a predetermined block 18 are predicted, whereby the planar intra prediction mode 504 is placed in a first position of the list 528 of the most likely intra prediction modes independent of the intra prediction mode in which the adjacent blocks 524, 526 are predicted.
[0241] FIG. 15 shows a block diagram of a method 4000 for decoding a predetermined block 18 of an image using intra prediction, and includes a step 4100 of deriving from a data stream a set selection syntax element indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. If the set selection syntax element indicates 4150 that a predetermined block is to be predicted using one of the first set of intra prediction modes, method 4000 includes a step 4200 of forming a list of the most likely intra prediction modes based on the intra prediction mode in which adjacent blocks adjacent to the predetermined block are predicted, a step 4300 of deriving an MPM list index from a data stream that directs the list of the most likely intra prediction modes to a predetermined intra prediction mode, and a step 4400 of intra predicting the predetermined block using the predetermined intra prediction mode. If the set selection syntax element indicates 4155 that a predetermined block is not to be predicted using one of the first set of intra prediction modes, method 4000 includes a step 4350 of calculating a matrix-vector product between a vector derived from reference samples in the vicinity of the predetermined block and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode to obtain a prediction vector, and a step of predicting 4450 samples of the predetermined block based on the prediction vector, and includes a step 4250 of deriving from the data stream a further index indicating a predetermined matrix-based intra prediction mode of a second set of matrix-based intra prediction modes. The list of the most likely intra prediction modes is formed 4200 based on the intra prediction mode in which adjacent blocks adjacent to the predetermined block are predicted such that the list of the most likely intra prediction modes does not include the DC intra prediction mode when at least one of the adjacent blocks is predicted by any of the angular intra prediction modes.
[0242] FIG. 16 shows a block diagram of a method 5000 for encoding a predetermined block of an image using intra prediction, and includes signaling 5100 a set selection syntax element in a data stream indicating whether a predetermined block is to be predicted using one of a first set of intra prediction modes including a DC intra prediction mode and an angular prediction mode. When the set selection syntax element indicates 5150 that a predetermined block is to be predicted using one of the first set of intra prediction modes, method 5000 includes forming 5200 a list of the most likely intra prediction modes based on the intra prediction mode in which adjacent blocks adjacent to the predetermined block are predicted, signaling 5300 an MPM list index in a data stream that directs the list of the most likely intra prediction modes to a predetermined intra prediction mode, and intra predicting 5400 the predetermined block using the predetermined intra prediction mode. When the set selection syntax element indicates 5155 that a predetermined block is not to be predicted using one of the first set of intra prediction modes, method 5000 includes calculating 5350 a matrix-vector product between a vector derived from reference samples in the vicinity of the predetermined block and a predetermined prediction matrix associated with a predetermined matrix-based intra prediction mode to obtain a prediction vector, and signaling 5250 an additional index in a data stream indicating the predetermined matrix-based intra prediction mode of a second set of matrix-based intra prediction modes by predicting 5450 samples of the predetermined block based on the prediction vector. The list of the most likely intra prediction modes is formed 5200 based on the intra prediction mode in which adjacent blocks adjacent to the predetermined block are predicted such that the list of the most likely intra prediction modes does not include the DC intra prediction mode when at least one of the adjacent blocks is predicted by any of the angular intra prediction modes.
[0243] References [1] P. Helle et al., "Non-linear weighted intra prediction", JVET-L0199, Macau, China, October 2018. [2] F. Bossen, J. Boyce, K. Suehring, X. Li, V. Seregin, "JVET common test conditions and software reference configurations for SDR video", JVET-K 1010, Ljubljana, SI, July 2018.
[0244] Further embodiments and examples In general, an example may be implemented as a computer program product having program instructions, the program instructions operative to perform one of the methods when the computer program product is executed on a computer. The program instructions may be stored, for example, on a machine-readable medium. Another example includes a computer program for performing one of the methods described herein, stored on a machine-readable carrier. In other words, thus, an example of a method is a computer program having program instructions for performing one of the methods described herein when the computer program is executed on a computer.
[0245] Thus, a further example of a method is a data carrier medium (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier medium, digital storage medium, or recording medium is tangible and / or non-transitory, not an intangible and transitory signal. Thus, a further example of a method is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may be transferred, for example, via a data communication connection, such as via the Internet.
[0246] Further examples include a processing means, such as a computer, or a programmable logic device that executes one of the methods described herein. Further examples include a computer having installed thereon a computer program for executing one of the methods described herein. Further examples include an apparatus or system that transfers (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0247] In some examples, a programmable logic device (e.g., a field programmable gate array) can be used to execute some or all of the functions of the methods described herein. In some examples, the field programmable gate array can cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods may be executed by any suitable hardware device.
[0248] The above examples are merely illustrative of the above principles. It will be understood that modifications and variations of the configurations and details described herein are apparent. Accordingly, it is intended to be limited not by the descriptions and specific details presented as examples herein, but by the appended claims. Elements that are the same or equivalent, or that have the same or equivalent functions, will be denoted by the same or equivalent reference numerals in the following description, even if they occur in different figures.
Claims
Claims 1 An apparatus (3000) for decoding a predetermined block (18) of an image (10) using intra prediction, decoding from a data stream (12) a set selection syntax element (522) indicating whether the predetermined block (18), which is a luma block, should be predicted using an intra prediction mode from a first set (508) of intra prediction modes including a DC intra prediction mode (506) and an angular intra prediction mode (500), when the set selection syntax element (522) indicates that the predetermined block (18) should be predicted using one of the first set (508) of intra prediction modes, forming a list (528) of the most likely intra prediction modes having an intra prediction mode (3050) used to predict adjacent blocks (524, 526) of the predetermined block (18), wherein the list does not include the DC intra prediction mode (506) when one of the adjacent blocks (524, 526) is predicted by any of the angular intra prediction modes (500) and the other of the adjacent blocks (524, 526) is predicted by a non-angular intra prediction mode, deriving from the data stream (12) an MPM list index (534) indicating a predetermined intra prediction mode (3100) from the list (528) of the most likely intra prediction modes, when the set selection syntax element (522) indicates that the predetermined block (18) is not to be predicted using one of the first set (508) of intra prediction modes, deriving from the data stream (12) a further index (540, 546) indicating a predetermined matrix-based intra prediction mode (510) from a second set (520) of matrix-based intra prediction modes (3200), obtaining a reference sample in the vicinity of the predetermined block (18) to obtain a prediction vector (518), predicting samples of the predetermined block (18) based on the prediction vector (518), apparatus (3000).
Citation Information
Patent Citations
Coding using intra prediction
JP7665822B2
Most probable mode list construction for matrix-based intra prediction
WO2020207502A1
Image encoding / decoding method and device for utilizing simplified mpm list generation method, and method for transmitting bitstream
WO2020251330A1