Device and method for decoding / encoding a predetermined block of a picture using intra prediction
By using the prediction matrix represented by fixed points and the unified prelocal depth, the deviation problem of the intra prediction mode based on the matrix in matrix-vector product approximation is solved, and the quality and computational efficiency of the prediction signal are improved.
Patent Information
- Application Number
- CN202080095826.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-12-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-12-04
AI Technical Summary
The matrix-based intra prediction mode may have a large deviation in the integer operation approximation of matrix-vector product, affecting the quality of the predicted signal.
Effective calculation of matrix-vector product is achieved by representing all terms of the prediction matrix using a fixed point representation with a pre-local depth and applying the same pre-local depth for all matrix-based intra prediction modes.
This method effectively overcomes the deviation problem in matrix-vector product approximation, improves the quality of intra-prediction signals, and optimizes storage and computing efficiency.
Smart Images

Figure CN115066896B_ABST
Abstract
Description
Technical Field
[0001] Embodiments according to the present invention relate to matrix-based intra prediction with mode global settings for picture and video encoding / decoding. Background Art
[0002] Typical block-based image or video codecs generally operate by predictive coding. Thus, when a receiver of an encoded image or video signal generates the signal for a given block, he constructs a prediction signal from the information already available from the encoded data. This prediction signal serves as a first approximation of the signal on that block. In a second step, the prediction residual is decoded from the bitstream and added to the prediction signal. The better the prediction signal, the smaller the number of bits needed to transmit the prediction residual. Thus, the quality of the prediction signal greatly affects the overall efficiency of the codec.
[0003] Generally, there are two ways to generate a prediction signal. The first method, used only in video codecs, is inter prediction. Here, the prediction signal is generated from reconstructed samples belonging to a frame different from the current frame. The second method is intra prediction. Here, the prediction signal is generated from reconstructed samples belonging to the same frame and typically spatially adjacent to the given block.
[0004] In classical codecs, intra prediction is performed using angular prediction modes or DC and planar modes. The angular prediction modes replicate the reconstructed samples to the left and above the block along a specific direction defined by an angular parameter, where for fractional angular positions, an interpolation filter is used. The DC mode generates the prediction signal as the average sample value of the neighboring samples to the left and above the block. Finally, the planar mode generates the prediction signal as a linear combination of predictions along the horizontal and vertical directions. Optionally, post-filtering of the prediction signal or pre-smoothing of the reference samples can be applied for any of the aforementioned prediction techniques.
[0005] Different from the classical intra prediction methods described above, matrix-based intra prediction (MIP) is introduced as a new technique for generating intra prediction signals. It is part of the current draft of the Versatile Video Coding (VVC) standard [1]. MIP can be regarded as a low-complexity variant of a more general data-driven neural-network-based intra prediction mode. Each MIP mode generates an intra prediction signal by multiplying a predefined matrix depending on the prediction mode with a downsampled version of the top and left border samples and then upsampling the result. For more details, we refer to the chapter review on matrix-based intra prediction.
[0006] The key property of MIP is to determine the matrices for various MIP patterns via a training algorithm using a large training data set. In this training algorithm, an attempt is made to find matrices such that they minimize a predefined loss function with respect to the training data. Here, the stochastic gradient descent method is used, where the matrix entries are updated iteratively. Such methods for determining matrix entries require computations in floating point arithmetic, and thus, the resulting matrix entries are given as floating point numbers. Thus, after training, for each MIP pattern i, a matrix with floating point entries is obtained such that in floating point, for MIP pattern i, the reduced prediction signal is given as
[0007]
[0008] where r red represents a downsampled version of the boundary of a given block, and where · denotes matrix-vector multiplication.
[0009] On the other hand, for applications in the final standard, each matrix-vector multiplication (1) needs to be approximated by rules specified in integer arithmetic. This means that for each MIP pattern i, a matrix A i with integral entries and positive integers c i and d i must be specified such that the computation of the reduced prediction signal pred red is specified as
[0010] pred red = ((A i - c i ) · r red + (1 << (d i - 1))) >> d i . (2)
[0011] Here, A i - c i represents the matrix resulting when c i is subtracted from each entry of A i . Finally, if v and w are vectors where w has integral entries, then v + (1 << (d i - 1)) represents the vector resulting by adding 1 << (d i - 1) to each entry of v, and w >> d i represents the vector resulting by shifting each entry of w to the right by d i . By the basic idea of MIP, (2) must approximate (1) for all possible input vectors r red .
[0012] Thus, it is desired to obtain a matrix A with integral entriesi (For its equation (2) for the variable input vector r red is moderately well approximated by equation (1)). Otherwise, the MIP prediction mode specified in the codec and requiring the use of matrix A i to perform the matrix-vector product in (2) may deviate significantly from the "true" behavior (i.e., the MIP mode trained using the matrix-vector product of the matrix ), see equation (1). Thus, the entire concept of a data-driven approach to in-frame prediction supporting MIP is violated.
[0013] Accordingly, there is a desire to provide a concept for more effectively supporting matrix-based in-frame prediction in picture encoding and / or video encoding. Additionally or alternatively, there is a desire to reduce the bitstream and thus the signaling cost.
[0014] This is achieved by the subject matter of the independent claims of the present application. Summary of the Invention
[0015] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application.
[0016] According to a first aspect of the present invention, the inventors of the present application have recognized that a problem encountered when attempting to use a matrix-based intra prediction mode (MIP mode) to predict samples of a predetermined block of a picture stems from the fact that the matrix-vector product (i.e., matrix-vector multiplication) performed in the MIP mode needs to be approximated by integer operations, whereby a large deviation between the approximated and the non-approximated matrix-vector product (i.e., the "true" matrix-vector product) may occur. In the following, the approximated matrix-vector product may be understood as the matrix-vector product of the corresponding MIP mode, because for each MIP mode, only this approximated matrix-vector product is calculated and the "true" matrix-vector product is not calculated to determine the prediction signal of the predetermined block. According to a first aspect of the present application, this difficulty is overcome by imposing a constraint on the calculation of the prediction signal by the matrix-vector product. The inventors have found it advantageous to represent all terms of the prediction matrix associated with the MIP mode by a fixed-point representation with a predetermined bit depth, and to apply the same predetermined bit depth for all matrix-based intra prediction modes (e.g., at least for modes related to the same block size, but optionally possibly for prediction matrices of all block sizes). This enables the efficient implementation of the matrix-vector product, because if all terms of all prediction matrices have a common fixed predetermined bit depth, a specific multiplier adapted to this predetermined bit depth can be used, and it is possible to share this specific multiplier across all MIP modes for calculating the matrix-vector product. In addition, if all terms of all prediction matrices can be stored in fixed-point representation (i.e., with fixed precision), efficient memory management (when dealing with the prediction matrices) is enabled. Furthermore, the inventors have found it advantageous to calculate the matrix-vector product between the input vector and the prediction matrix associated with the corresponding MIP mode by performing a right shift with an equal number of bits for each component of the output vector for all MIP modes (e.g., at least for MIP modes related to the same block size, but optionally possibly for prediction matrices of all block sizes). This enables the efficient implementation of the shift in the matrix-vector product based on a fixed right shift for all MIP modes, because if the shift value does not depend on the MIP mode, table lookups are saved, and a single fixed shift operation can be implemented for the MIP, which is beneficial for a compact SIMD implementation of the matrix-vector product and reduces the case-dependent implementation of the matrix-vector product in hardware implementation.
[0017] Accordingly, according to a first aspect of the present application, a device for decoding a predetermined block of a picture using intra prediction is configured to read a mode index from a data stream, and a device for encoding a predetermined block of a picture using intra prediction is configured to insert the mode index into the data stream. For example, a device for encoding, i.e., an encoder, may have selected this mode from a mode list through rate-distortion optimization, and optionally, further modes such as an inter prediction mode. The mode index points to one in a matrix-based intra prediction mode list. Further, the device (i.e., the device for decoding and / or the device for encoding) is configured to predict the samples of the predetermined block by calculating a matrix-vector product between an input vector derived from reference samples in the neighborhood of the predetermined block and a prediction matrix associated with the matrix-based intra prediction mode pointed to by the mode index, and by associating the components of the output vector obtained through the matrix-vector product to the sample positions of the predetermined block. For each matrix-based intra prediction mode, all terms of the prediction matrix associated with the corresponding matrix-based intra prediction mode are represented by fixed-point representations having a predetermined bit depth, and the predetermined bit depth is equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes related to the same block size, but optionally may be for matrices of all block sizes). Further, the device is configured to calculate the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode for each matrix-based intra prediction mode by performing a right shift with a number of bits equal for the matrix-based intra prediction mode (e.g., at least for matrix-based intra prediction modes related to the same block size, but optionally may be for matrices of all block sizes) for each component of the output vector.
[0018] According to an embodiment, the number of matrix-based intra prediction modes in the matrix-based intra prediction mode list is 12, 16, or 32.
[0019] According to an embodiment, the device for decoding and / or the device for encoding is configured such that the matrix-based intra prediction modes in the matrix-based intra prediction mode list have 6, 8, or 16 different matrices associated therewith. For example, there may be different lists for mutually exclusive block size sets, one list having 6 different matrices associated with 12 modes for block sizes within a first block size set, one list having 8 different matrices associated with 16 modes for smaller block sizes within a second block size set, and one list having 16 different matrices associated with 32 modes for even smaller block sizes within a third block size set.
[0020] According to an embodiment, a device for decoding and / or a device for encoding is configured to calculate the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode in fixed-point arithmetic by applying the right shift to an intermediate result obtained by the matrix-vector product for each component of the output vector. For example, the intermediate result is obtained by the matrix-vector product between the input vector and the prediction matrix or by the matrix-vector product between the input vector and the prediction matrix, the intermediate result is offset by a positive integer, e.g., subtracting a positive integer from each entry of the prediction matrix / adding a positive integer to each entry of the prediction matrix results in an intermediate matrix, and the intermediate result is obtained by the matrix-vector product between the input vector and the intermediate matrix.
[0021] According to an embodiment, a device for decoding and / or a device for encoding is configured to, before calculating the matrix-vector product, for each matrix-based intra prediction mode, offset all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode by an offset value equal for the matrix-based intra prediction mode (e.g., at least for matrix-based intra prediction modes related to the same block size but optionally possibly for matrices of all block sizes) by adding or subtracting. This enables an efficient implementation of the matrix-vector product by saving table lookups based on a fixed offset value for all MIP modes.
[0022] According to an embodiment, a device for decoding and / or a device for encoding is configured to store, for each entry of the prediction matrix associated with the corresponding matrix-based intra prediction mode, the fixed-point representation in the predetermined bit depth, for each matrix-based intra prediction mode.
[0023] According to an embodiment, a device for decoding / a device for encoding is configured to: decode / encode the picture with 10-bit resolution; store the magnitude of the entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode with 7-bit precision, for each matrix-based intra prediction mode, and use 6 bits as the number of bits for the right shift.
[0024] According to an embodiment, a device for decoding and / or a device for encoding is configured to represent, for each matrix-based intra prediction mode, the entries of the prediction matrix associated with the respective matrix-based intra prediction mode as 8-bit signed magnitudes. Alternatively, in a case where the entries of the prediction matrix associated with the respective matrix-based intra prediction mode have the same sign, the device for decoding and / or the device for encoding is configured to, before calculating the matrix-vector product, for each matrix-based intra prediction mode, offset all the entries of the prediction matrix associated with the respective matrix-based intra prediction mode by an offset value equal for the matrix-based intra prediction mode, for example, by adding or by subtracting, wherein all the entries of the prediction matrix associated with the respective matrix-based intra prediction mode are representable by an 8-bit signed representation. Accordingly, according to this alternative, since indication of the sign is not necessary, only 7-bit magnitudes can be stored for each matrix entry.
[0025] According to an embodiment, a device for decoding and / or a device for encoding is configured to calculate the matrix-vector product between the input vector and the prediction matrix associated with the respective matrix-based intra prediction mode in fixed-point arithmetic by applying the right shift to an intermediate result obtained by the matrix-vector product for each component of the output vector and represented with a bit precision twice as high as the bit precision with which the entries of the prediction matrix associated with the matrix-based intra prediction mode are stored. For example, the matrix-vector product calculated in such a way is represented with a bit precision twice as high as the bit precision with which the entries of the prediction matrix associated with the matrix-based intra prediction mode are stored, for example, a prediction signal obtained by offsetting the prediction matrix by a positive integer (resulting in an intermediate matrix), calculating the matrix-vector product between the input vector and the intermediate matrix (resulting in an intermediate result), and performing a right shift on the intermediate result.
[0026] According to an embodiment, the matrix-based intra prediction mode list includes one or more matrix-based intra prediction mode pairs. Note that the matrix-based intra prediction mode list may not be specifically composed of such mode pairs, but there may also be other modes that are specifically applied using the transpose option or the non-transpose option. For each matrix-based intra prediction mode pair, the prediction matrix associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair is equal to the prediction matrix associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair. The device (i.e., the device for decoding and / or the device for encoding) is configured such that if the matrix-based intra prediction mode pointed to by the mode index is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair (e.g., the mode with an odd mode index), then the association of the reference samples in the neighborhood of the predetermined block with the components of the input vector and the association of the sample positions of the predetermined block with the components of the output vector are transposed relative to the case where the matrix-based intra prediction mode pointed to by the mode index is the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair (e.g., the mode with an even mode index). That is, if in the former case a certain component of the input vector is associated with the position (x, y) (where (0, 0) represents the top-left sample of the predetermined block), then in the latter case it is associated with (y, x). The same applies to the components of the output vector.
[0027] According to an embodiment, the device for decoding and / or the device for encoding is configured to use the matrix-based intra prediction mode list for multiple block sizes.
[0028] According to an embodiment, the device for decoding and / or the device for encoding is configured to predict the samples of the predetermined block by upsampling and / or interpolating based on the output vector or based on the output vector and the reference samples in the neighborhood of the predetermined block, and the samples are offset from the sample positions associated with the components of the output vector.
[0029] According to an embodiment, the device for decoding and / or the device for encoding is configured to derive the input vector from the reference samples in the neighborhood of the predetermined block by downsampling and / or pooling.
[0030] According to an embodiment, the reference samples in the neighborhood of the predetermined block include a first reference sample above the predetermined block and a second reference sample to the left of the predetermined block. The device (i.e., a device for decoding and / or a device for encoding) is configured to derive the input vector from the reference samples in the neighborhood of the predetermined block by: deriving a first intermediate component from the first reference sample by downsampling and / or pooling; deriving a second intermediate component from the second reference sample by downsampling and / or pooling; concatenating the first intermediate component and the second intermediate component to derive a preliminary input vector, and forming the input vector from the preliminary input vector.
[0031] According to an embodiment, a device for decoding / a device for encoding is configured to decode / encode the picture at a B-bit resolution. The device is configured to form the input vector from the preliminary input vector by: subtracting 2 from a first component of the preliminary input vector B-1 to obtain the first component of the input vector, and subtracting the first component of the preliminary input vector from further components of the preliminary input vector to obtain the further components of the input vector, or subtracting the first component of the preliminary input vector from further components of the preliminary input vector such that the input vector is formed from the further components. Further, the device is configured to correct the output vector by a component-wise addition of the first component of the preliminary input vector.
[0032] According to an embodiment, the entries of the prediction matrix of the matrix-based intra prediction mode in the matrix-based intra prediction mode list correspond to the entries in Table 2 (shown below), but note that another shift value may be selected to list the values in the table, and the values in the table may be represented in another scale.
[0033] According to an embodiment, a device for decoding and / or a device for encoding is configured to use a trained prediction matrix selected for a predetermined block to calculate a matrix-vector product between an input vector derived from reference samples in the neighborhood of the predetermined block and the trained prediction matrix associated with the matrix-based intra prediction mode selected for the predetermined block, and to predict the samples of the predetermined block by associating the components of the output vector obtained by the matrix-vector product with the sample positions of the predetermined block. For example, the trained prediction matrix is trained by a device for training the prediction matrix, e.g., by a device according to the second aspect.
[0034] According to a second aspect of the present invention, the inventors of the present application have recognized that a problem encountered when attempting to use a matrix-based intra prediction mode (MIP mode) to predict samples of a predetermined block of a picture stems from the fact that the matrix-vector product (i.e., matrix-vector multiplication) performed in the MIP mode needs to be approximated by integer operations, and the fact that the training prediction matrix for such matrix-vector multiplication is obtained with floating-point precision. According to the second aspect of the present application, this difficulty is overcome by imposing a constraint on the calculation of the prediction signal through the matrix-vector product when training the prediction matrix for such matrix-vector products. The inventors have found it advantageous to optimize the terms of the prediction matrix associated with the MIP mode by using a cost function that depends on a prediction distortion measure, the prediction distortion measure being associated with setting the terms of the prediction matrix to intermediate values, and using a differentiable function to map representative values onto the intermediate values. By this method, it is possible to limit the range of all prediction matrix terms and avoid the possibility that some terms may not be updated during training. During training, the terms are represented in floating-point notation, and then the intermediate values are quantized to a fixed-point notation with a predetermined bit depth that is equal for all MIP modes. This enables the matrix-vector product to be efficiently implemented (since if all terms of all prediction matrices have a common fixed predetermined bit depth). Such prediction matrices enable a video or picture encoder / decoder to use a specific multiplier adapted to that predetermined bit depth and share that specific multiplier across all MIP modes for calculating the matrix-vector product. Additionally, if all terms of all prediction matrices can be stored in fixed-point notation (i.e., with fixed precision), efficient memory management (when dealing with the prediction matrices) is enabled.
[0035] Accordingly, in a second aspect of the present application, there is provided an apparatus for training a prediction matrix of a matrix-based intra prediction mode list by calculating a matrix-vector product between an input vector derived from reference samples in a neighborhood of a predetermined block and one of the prediction matrices associated with a matrix-based intra prediction mode selected for the predetermined block, and associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block, where one of the matrix-based intra prediction modes should be selected for the predetermined block to predict samples of the predetermined block. The apparatus is configured to optimize, by using a gradient descent method, a representative value of the terms of the prediction matrix of the matrix-based intra prediction mode list, represented in floating-point representation, by using a cost function that depends on a prediction distortion measure associated with setting the terms of the prediction matrix to intermediate values (e.g., by using a set of training of predetermined blocks of known (e.g., original) samples and their corresponding neighborhoods), and to map the representative value to the intermediate value by using a differentiable function. The prediction distortion measure defines, for example, an increasing cost with decreasing prediction quality, as caused by applying the differentiable function to the prediction matrix under training (meaning applying the differentiable function to each term of the prediction matrix under training). The domain and co-domain of the differentiable function are defined by the floating-point representation, the image of the differentiable function has a predetermined dynamic range, and the differentiable function is equal for the matrix-based intra prediction modes. Further, the apparatus is configured to, for example, after training, quantize the intermediate value to a fixed-point representation such that for each matrix-based intra prediction mode, the prediction matrix associated with the corresponding matrix-based intra prediction mode has all terms represented by a fixed-point representation having a predetermined bit depth, such that the predetermined bit depth is equal for the matrix-based intra prediction modes, and such that the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode is computable by performing a right shift by a number of bits equal for each component of the output vector for the matrix-based intra prediction modes.
[0036] According to an embodiment, the differentiable function (i.e., the clipping function) has a slope of 1 at the origin of the image, is strictly monotonically increasing, and has horizontal asymptotes at the upper and lower bounds of the image. The horizontal asymptotes at the upper and lower bounds of the image of the differentiable function may define the predetermined dynamic range.
[0037] According to an embodiment, the differentiable function is represented / defined by the following formula:
[0038]
[0039] where α, β, γ, and δ are real numbers that depend on the said predetermined dynamic range (i.e., the clipping range), and λ is a non - negative integer.
[0040] According to an embodiment, the differentiable function (i.e., the clipping function) is parameterizable by a shift parameter (e.g., δ) in terms of the shift of the image within the co - domain. The device is configured to subject the shift parameter to optimization using the gradient descent method (which can be, but does not have to be), and to derive from the shift parameter offset values equal for the matrix - based intra - prediction modes, for use before the calculation of the matrix - vector product for each matrix - based intra - prediction mode, e.g., by adding or by subtracting, an offset from all terms of the prediction matrix associated with the corresponding matrix - based intra - prediction mode.
[0041] An embodiment related to a method for decoding a picture of a predetermined block using intra - prediction, the method comprising: reading a mode index from a data stream, the mode index pointing to one in a list of matrix - based intra - prediction modes; and predicting the samples of the predetermined block by calculating a matrix - vector product between an input vector derived from reference samples in the neighborhood of the predetermined block and a prediction matrix associated with the matrix - based intra - prediction mode pointed to by the mode index, and by associating the components of the output vector obtained by the matrix - vector product to the sample positions of the predetermined block. For each matrix - based intra - prediction mode, all terms of the prediction matrix associated with the corresponding matrix - based intra - prediction mode are represented by a fixed - point representation with a predetermined bit depth, the predetermined bit depth being equal for the matrix - based intra - prediction modes (e.g., at least for matrix - based intra - prediction modes related to the same block size but optionally possibly for matrices of all block sizes). Further, the method comprises, for each matrix - based intra - prediction mode, calculating the matrix - vector product between the input vector and the prediction matrix associated with the corresponding matrix - based intra - prediction mode by performing a right - shift with a number of bits equal for the matrix - based intra - prediction modes (e.g., at least for matrix - based intra - prediction modes related to the same block size but optionally possibly for matrices of all block sizes) for each component of the output vector.
[0042] Embodiments related to a method for encoding a picture using intra prediction of a predetermined block, the method comprising: inserting a mode index into the data stream, the mode index pointing to one of a matrix-based intra prediction mode list, e.g., this mode may have been selected from the matrix-based intra prediction mode list (and optionally also from further modes such as inter prediction modes) by rate distortion optimization. The method further comprises predicting the samples of the predetermined block by computing a matrix-vector product between an input vector derived from reference samples in a neighborhood of the predetermined block and a prediction matrix associated with the matrix-based intra prediction mode pointed to by the mode index, and by associating components of the output vector obtained by the matrix-vector product to sample positions of the predetermined block. For each matrix-based intra prediction mode, all terms of the prediction matrix associated with the corresponding matrix-based intra prediction mode are represented by fixed-point representations having a predetermined bit depth, the predetermined bit depth being equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes related to the same block size but optionally may be for matrices of all block sizes). Further, the method comprises, for each matrix-based intra prediction mode, computing the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode by performing a right shift by a number of bits equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes related to the same block size but optionally may be for matrices of all block sizes) for each component of the output vector.
[0043] The method as described above is based on the same considerations as the encoder / decoder described above. Incidentally, the method can be accomplished with all the features and functionality that are also described with respect to the encoder / decoder.
[0044] Embodiments related to a method for training a prediction matrix of a matrix-based intra prediction mode list by calculating a matrix-vector product between an input vector derived from reference samples in a neighborhood of a predetermined block and one of the prediction matrices associated with a matrix-based intra prediction mode selected for the predetermined block, and associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block, wherein one of the matrix-based intra prediction modes should be selected for the predetermined block to predict samples of the predetermined block. The method includes optimizing, e.g., by using a gradient descent method, a representative value of entries of the prediction matrix of the matrix-based intra prediction mode list, represented in floating-point representation, using a cost function depending on a prediction distortion measure associated with setting the entries of the prediction matrix to intermediate values, by using a set of training of predetermined blocks of known (e.g., original) samples and their corresponding neighborhoods, mapping the representative value to the intermediate value using a differentiable function, the domain and co-domain of the differentiable function being defined by the floating-point representation, the image of the differentiable function having a predetermined dynamic range, and the differentiable function being equal for the matrix-based intra prediction modes. Further, the method includes quantizing the intermediate value to a fixed-point representation, e.g., after training, such that for each matrix-based intra prediction mode, the prediction matrix associated with the corresponding matrix-based intra prediction mode has all entries represented by a fixed-point representation having a predetermined bit depth, such that the predetermined bit depth is equal for the matrix-based intra prediction modes, and such that for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode is computable by performing a right shift by a number of bits equal for the matrix-based intra prediction modes for each component of the output vector.
[0045] The method as described above is based on the same considerations as the device described above for training the prediction matrix. Incidentally, the method can be accomplished with all features and functionalities that are also described with respect to the device for training the prediction matrix.
[0046] Embodiments relate to a data stream having a picture or video encoded therein using the method for encoding described herein.
[0047] Embodiments relate to a computer program having program code for performing the method described herein when run on a computer. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings need not be to scale; instead, the emphasis is generally on showing the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0049] Figure 1 An embodiment encoded into a data stream is shown;
[0050] Figure 2 An embodiment of an encoder is shown;
[0051] Figure 3 An embodiment of picture reconstruction is shown;
[0052] Figure 4 An embodiment of a decoder is shown;
[0053] Figure 5.1 Prediction of a block with a reduced sample value vector according to an embodiment is shown;
[0054] Figure 5.2 Prediction of a block using sample interpolation according to an embodiment is shown;
[0055] Figure 5.3 Prediction of a block with a reduced sample value vector according to an embodiment, where only some of the boundary samples are averaged;
[0056] Figure 5.4 Prediction of a block with a reduced sample value vector according to an embodiment, where a group of four boundary samples are averaged;
[0057] Figure 6.1 Matrix-based intra prediction of a predetermined block of a picture based on a mode index is shown;
[0058] Figure 6.2 The relationship between the application of a matrix-based intra prediction mode pair and a sample-to-sample distance setting is shown;
[0059] Figure 7 A device for decoding using a MIP mode for prediction according to an embodiment is shown;
[0060] FIG. 8 shows a device for training a prediction matrix according to an embodiment; and
[0061] Figure 9 A demonstration differentiable function is shown. DETAILED DESCRIPTION
[0062] In the following description, equal or equivalent elements or elements having the same or equivalent functionality are denoted by equal or equivalent reference numerals (even if they appear in different figures).
[0063] In the following description, numerous specific details are set forth to provide a more thorough explanation of embodiments of the present invention. It will be apparent, however, to one skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present invention. Additionally, unless otherwise specifically noted, the features of the different embodiments described hereinafter may be combined with each other.
[0064] In the following, various examples are described that can help achieve more efficient compression when using matrix-based intra prediction. For example, matrix-based intra prediction can be added to other heuristically designed intra prediction modes or can be provided specifically.
[0065] To facilitate understanding of the following examples, the presentation begins with a description of possible encoders and decoders (into which the examples outlined above in this application can be incorporated) that are suitable therefor. Figure 1 A device for encoding picture 10 block by block into data stream 12 is shown. The device is designated by reference numeral 14 and can be a still picture encoder or a video encoder. In other words, when encoder 14 is configured to encode video 16 including picture 10 into data stream 12, picture 10 can be the current picture among videos 16, or encoder 14 can specifically encode picture 10 into data stream 12.
[0066] As mentioned, encoder 14 performs encoding in a block-by-block or block-based manner. To this end, encoder 14 subdivides picture 10 into blocks, and the units of encoder 14 encode picture 10 into data stream 12. Examples of possible subdivisions of subdividing picture 10 into blocks 18 are elaborated in more detail below. Generally, the subdivision can end at blocks 18 of a constant size, such as an array of blocks arranged in rows and columns, or at blocks 18 of different block sizes, such as by using hierarchical multi-tree subdivision, where the multi-tree subdivision starts from the entire picture area of picture 10 or from a pre-segmentation of picture 10 into an array of tree blocks, and these examples should not be considered as excluding other possible ways of subdividing picture 10 into blocks 18.
[0067] Furthermore, encoder 14 is a predictive encoder that is configured to predictively encode picture 10 into data stream 12. For a certain block 18, this means that encoder 14 determines the prediction signal of block 18 and encodes the prediction residual (i.e., the prediction error, by which the prediction signal deviates from the actual picture content within block 18) into data stream 12.
[0068] The encoder 14 may support different prediction modes in order to derive a prediction signal for a certain block 18. The important prediction mode in the following examples is the intra prediction mode, according to which the interior of block 18 is predicted spatially from adjacent, already encoded samples of picture 10. Picture 10 is encoded into data stream 12, and correspondingly, the corresponding decoding process may be based on a certain encoding order 20 defined between blocks 18. For example, the encoding order 20 may traverse blocks 18 in raster scan order, such as row by row from top to bottom, and for example, traverse each row from left to right. In the case of a hierarchical multi-tree based subdivision, raster scan sorting may be applied within each hierarchical level, where a depth-first traversal order may be applied, i.e., according to the encoding order 20, the leaf notes within a block of a certain hierarchical level may be before the blocks of the same hierarchical level having the same parent block. Depending on the encoding order 20, the adjacent, already encoded samples of block 18 may typically be located on one or more sides of block 18. In the case of the example presented herein, for example, the adjacent, already encoded samples of block 18 are located at the top and left of block 18.
[0069] The intra prediction mode may not be the only mode supported by encoder 14. For example, in the case where encoder 14 is a video encoder, encoder 14 may also support an intra prediction mode according to which block 18 is temporarily predicted from the aforementioned encoded pictures of video 16. Such an intra prediction mode may be a motion compensated prediction mode, according to which a motion vector is signaled for such a block 18, the motion vector indicating the relative spatial offset of a portion (from which the prediction signal for block 18 is to be derived as a copy). Additionally or alternatively, other non-intra prediction modes may also be available, such as an inter-view prediction mode in the case where encoder 14 is a multi-view encoder, or a non-predictive mode, according to which the interior of block 18 is encoded as it is (i.e., without any prediction).
[0070] Before starting by concentrating the description of the present application on the intra prediction mode, for a more specific example of a possible block-based encoder (i.e., for a possible implementation of encoder 14 as described with respect to Figure 2 ), two corresponding examples of decoders suitable for Figure 1 and 2 are presented respectively.
[0071] Figure 2 Shows a possible implementation of encoder 14 of Figure 1 , that is, an implementation in which the encoder is configured to use transform coding to encode prediction residuals, although this is almost an example and the present application is not limited to this type of prediction residual coding. According to Figure 2, the encoder 14 includes a subtractor 22 configured to subtract a corresponding prediction signal 24 from an inbound signal (i.e., picture 10) or, on a block basis, from the current block 18 to obtain a prediction residual signal 26, which is then encoded by a prediction residual encoder 28 into a data stream 12. The prediction residual encoder 28 consists of a lossy coding stage 28a and a lossless coding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes the samples of the prediction residual signal 26. As already mentioned above, this example uses transform coding of the prediction residual signal 26, and accordingly the lossy coding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such spectrally decomposed prediction residuals 26, where the quantization of the quantizer 30 occurs with respect to the transform coefficients in which the residual signal 26 is presented. The transform can be a DCT, DST, FFT, Hadamard transform, etc. Then, the transformed and quantized prediction residual signal 34 undergoes lossless coding by the lossless coding stage 28b, which is an entropy encoder that entropy-codes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further includes a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner that is also available at the decoder (i.e., taking into account the coding loss in the quantizer 30). To this end, the prediction residual reconstruction stage 36 includes a dequantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs an inverse transform with respect to the transform performed by the transformer 32 (e.g., the inverse of the spectral decomposition, e.g., the inverse of any of the specific transform examples mentioned above). The encoder 14 includes an adder 42 that adds the reconstructed prediction residual signal output by the inverse transformer 40 to the prediction signal 24 to output a reconstructed signal, i.e., reconstructed samples. This output is fed into a predictor 44 of the encoder 14, which then determines the prediction signal 24 based on it. It is the predictor 44 that supports all the prediction modes already discussed with respect to Figure 1 discussed above. Figure 2 It is also shown that in the case where the encoder 14 is a video encoder, the encoder 14 may further include an in-loop filter 46 that filters a fully reconstructed picture, which, after being filtered, forms a reference picture for the inter-frame prediction block of the predictor 44.
[0072] As already mentioned above, the encoder 14 operates on a block-by-block basis. For the following description, the block-based operation of interest is an operation that subdivides the picture 10 into blocks, for which an intra prediction mode is selected from a set or plurality of intra prediction modes supported by the predictor 44 or the encoder 14 respectively, and the selected intra prediction mode is performed individually. However, there may also be other kinds of blocks into which the picture 10 is subdivided. For example, the decision as to whether the above-mentioned picture 10 is inter-coded or intra-coded may be made in units of granularity or blocks offset from the block 18. For example, the inter / intra mode decision may be performed at the level of the coding blocks into which the picture 10 is subdivided, and each coding block is subdivided into prediction blocks. Prediction blocks having coding blocks (for which intra prediction has been decided) are each subdivided into intra prediction mode decisions. For this purpose, for each of these prediction blocks, it is decided which supported intra prediction mode should be used for the corresponding prediction block. These prediction blocks will form the block 18 of interest here. Prediction blocks within coding blocks associated with inter prediction are handled differently by the predictor 44. They are inter-predicted from a reference picture by determining a motion vector and copying the prediction signal of this block from the position in the reference picture pointed to by the motion vector. Another kind of block subdivision involves subdivision into transform blocks, and the transform by the transformer 32 and the inverse transformer 40 is performed in units of the transform blocks. The transform blocks may be, for example, the result of further subdividing the coding blocks. Naturally, the examples set forth herein should not be considered restrictive, and there are also other examples. For completeness only, note that, for example, the subdivision into coding blocks may use a multi-tree subdivision, and the prediction blocks and / or transform blocks may also be obtained by further subdividing the coding blocks in one step using a multi-tree subdivision.
[0073] Figure 3 depicted in is suitable for Figure 1 the encoder 14 for block-by-block decoding. This decoder 54 performs the opposite operation of the encoder 14, i.e., it decodes the picture 10 from the data stream 12 in a block-by-block manner and, for this purpose, supports a plurality of intra prediction modes. The decoder 54 may include, for example, a residual provider 156. As described above with respect to Figure 1All other possibilities discussed are also valid for decoder 54. To this end, decoder 54 can be a still picture decoder or a video decoder, and all prediction modes and prediction possibilities are also supported by decoder 54. The difference between encoder 14 and decoder 54 mainly lies in the fact that encoder 14 selects or chooses encoding decisions according to a certain optimization, for example, in order to minimize a certain cost function that can depend on the encoding rate and / or encoding distortion. One of these encoding options or encoding parameters can involve selecting an intra prediction mode to be used for current block 18 among the available or supported intra prediction modes. Then, the selected intra prediction mode can be signaled in data stream 12 by encoder 14 for current block 18, where decoder 54 uses this signaling in data stream 12 to re - make the selection for block 18. Similarly, the subdivision of picture 10 into blocks 18 can be subject to optimization within encoder 14, and the corresponding subdivision information can be conveyed in data stream 12, where decoder 54 restores the subdivision of picture 10 into blocks 18 based on the subdivision information. In summary, decoder 54 can be a predictive decoder operating on a block - by - block basis, and in addition to the intra prediction mode, decoder 54 can support other prediction modes (e.g., inter prediction mode) if decoder 54 is a video decoder, for example. In decoding, decoder 54 can also use the Figure 1 discussed encoding order 20, and since this encoding order 20 is followed at both encoder 14 and decoder 54, the same neighboring samples are available for current block 18 at both encoder 14 and decoder 54. Accordingly, to avoid unnecessary repetition, the description of the operating mode of encoder 14 should also apply to decoder 54 as long as it involves the subdivision of picture 10 into blocks, for example, as long as it involves prediction and as long as it involves the encoding of prediction residuals. The difference lies in the fact that encoder 14 selects some encoding options or encoding parameters and signals within data stream 12 through optimization, or inserts the encoding parameters into data stream 12, and then decoder 54 derives the encoding parameters from data stream 12 in order to re - make predictions, subdivisions, etc.
[0074] Figure 4 is shown Figure 3 a possible implementation of decoder 54, that is, an implementation suitable for the Figure 2 shown in Figure 1 encoder 14 as shown in Figure 4 Many elements of encoder 54 are the same as those that appear in the Figure 2 corresponding encoder, so the same reference numerals provided with an apostrophe are used in Figure 4 to indicate these elements. In particular, adder 42’, optional in - loop filter 46’ and predictor 44’ are in the same way as they are in Figure 2in the encoder in the same way is connected into the prediction loop. The reconstructed (i.e., dequantized and inverse-transformed) prediction residual signal applied to the adder 42’ is derived from the entropy decoder 56 of the entropy encoder 28b that performs inverse entropy coding, followed by a sequence of the residual signal reconstruction stage 36’ composed of the dequantizer 38’ and the inverse transformer 40’, just as in the case of the encoding side. The output of the decoder is the reconstruction of the picture 10. The reconstruction of the picture 10 can be directly available at the output of the adder 42’, or alternatively, can be available at the output of the in-loop filter 46’. A certain post-filter can be arranged at the output of the decoder to subject the reconstruction of the picture 10 to a certain post-filtering to improve the picture quality, but this option is not depicted in Figure 4 this.
[0075] Again, regarding Figure 4 , the description presented above regarding Figure 2 should also be valid for Figure 4 , except that only the encoder performs the optimization tasks and the associated decisions regarding the encoding options. However, all the descriptions regarding block subdivision, prediction, dequantization, and inverse transformation are also valid for the decoder 54 of Figure 4 .
[0076] The embodiments described herein use the so-called matrix-based intra prediction. The general concept is outlined below.
[0077] Review of Matrix-Based Intra Prediction
[0078] To keep this application self-contained, in this section, the main steps of the current matrix-based intra prediction (MIP) method included in the working draft 7 of the general video coding [1] are described. For more details, refer to [1].
[0079] Matrix-based intra prediction (MIP) is a method for generating an intra prediction signal on a rectangular block with width W and height H. The input for the MIP prediction process is the reconstructed samples r top from one row above the block and the reconstructed samples r left from one column to the left of the block, the MIP mode index i, and the information on whether the MIP mode is to be transposed. Then, the MIP prediction signal is generated using the following three steps:
[0080] 1. For specified natural numbers w in,red ≤W and h in,red ≤H that depend on W and H and are satisfied, one of r in,red and h in,red is downsampled / averaged to generate a reduced top input r top of size w in,ted , and r top,red , and rleft One of them generates a size h through downsampling / averaging in,red is the reduced left input r left,red .
[0081] Then, connect r top,red and r left,red to the reduced input r red,full . If the MIP mode is not to be transposed, then r red,full is defined as
[0082] r red,full = [r top,red , r left,red .
[0083] If the MIP mode is to be transposed, it is defined as
[0084] r red,full = [r left,red , r top,red .
[0085] Next, one of the r red,full defines the reduced input r red . Here, r red has the same size w red,full as r in,red + h in,red or has a size of w in,red + h in,red - 1. In the first case, r red is defined as
[0086] r red [0] = r red,full [0] - 2 B-1 ,
[0087] where B is the bit depth and is defined as
[0088] r red [i] = r red,full [i] - r red,full [0], i > 0.
[0089] In the second case, r red is defined as
[0090] r red [i] = r red,full [i + 1] - r red,full [0].[[]END]]
[0091] 2. For specified natural numbers w out,red ≤ W and h out,red ≤ H that depend on W and H out,red and hout,red , on a block having a width w out,red and a height h out,red the reduced prediction signal pred red is generated as
[0092] pred red = ((A i - c i ) · r red + (1 << (d i - 1))) >> d i .
[0093] Here A i is a matrix depending on W and H and depending on the MIP mode index i and c i and d i are non - negative integers depending on the MIP mode index i, where this dependency is to be removed by the present invention. Further, if the MIP mode does not need to be transposed, pred red is a vector of size w out,red · h out,red - identified by signals on a block having a width w out,red and a height h out,red in row - major order, and if the MIP mode needs to be transposed, it is in column - major order.
[0094] Then, r red,full [0] is added to pred red .
[0095] Finally, the result is clipped to a given bit range [0, 2 B ).
[0096] 3. If w out,red < W or h out,red < H, then upsampling / linear interpolation is applied to generate a full MIP prediction signal from the reduced prediction signal obtained at the end of the foregoing step. Here, the reconstructed samples are included in the linear interpolation.
[0097] Presentation of implementation examples
[0098] The general concepts were outlined above. The concepts are sometimes referred to hereinafter as ALWIP (Affine Linear Weighted Intra - Prediction) (as an alternative synonym for MIP (Matrix - based Intra - Prediction)) in order to explain the use of these modes in more detail again.
[0099] In the following Figures 5.1 - 5.4 , the entire process of filling the input vector, calculating the matrix - vector multiplication, and linear interpolation on a neighborhood basis is shown for different block shapes. Note that the remaining shapes are considered to be one of the depicted cases.
[0100] 1. Given a 4×4 block, ALWIP (or MIP) can take two averages along each axis of the boundary, see Figure 5.1 . As an alternative to taking the average, every other sample in the neighborhood is taken, or more generally and precisely, each component of the input vector of the matrix - vector - multiplication 19 is taken from exactly one sample in the neighborhood. The four resulting input samples enter the matrix - vector multiplication. The matrix is taken from the set S o matrix, where the set is a set of matrices for a nearby block size. After adding the offset, this can produce 16 final prediction samples. For generating the prediction signal, linear interpolation is not required. Thus, a total of (4 * 16) / (4 * 4) = 4 multiplications are performed per sample. For example, see the ALWIP of the 4×4 block shown in Figure 5.1 . The exact calculation has been explained above.
[0101] 2. Given an 8×8 block, ALWIP can take four averages along each axis of the boundary, see Figure 5.2 . The eight resulting input samples enter the matrix - vector multiplication 19. The matrix is taken from the set S1. This produces 16 samples at the odd positions of the prediction block. Thus, a total of (8 * 16) / (8 * 8) = 2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically using a reduced top boundary. The original left boundary is used after horizontal interpolation. For example, see the ALWIP of the 8×8 block shown in Figure 5.2 .
[0102] 3. Given an 8×4 block, ALWIP can take four averages along the horizontal axis of the boundary and four original boundary values on the left boundary, see Figure 5.3 . The eight resulting input samples enter the matrix - vector multiplication. The matrix is taken from the set S1. This produces 16 samples at the odd horizontal and each vertical position of the prediction block. Thus, a total of (8 * 16) / (8 * 4) = 4 multiplications are performed per sample. After adding the offset, these samples are interpolated horizontally using the original left boundary. For example, see the ALWIP of the 8×4 block shown in Figure 5.3 .
[0103] The transposed cases are handled accordingly.
[0104] 4. Given a 16×16 block, ALWIP can take four averages along each axis of the boundary. The eight resulting input samples enter a matrix-vector multiplication. The matrix is taken from the set S2. This produces 64 samples at the odd positions of the prediction block. Thus, a total of (8 * 64) / (16 * 16) = 2 multiplications are performed per sample. After adding the offset, these samples are interpolated vertically using the eight averages of the top boundary. The original left boundary is used after horizontal interpolation. For example, see the Figure 5.4 .
[0105] For larger shapes, the process can be substantially the same, and it is easy to check that the number of multiplications per sample is less than two.
[0106] For a W×8 block, since samples are given at the odd horizontal and each vertical position, only horizontal interpolation is required. Thus, at most (8 * 64) / (16 * 8) = 4 multiplications are performed per sample in these cases.
[0107] Finally, for a W×4 block (where W > 8), let A k be the matrix produced by omitting every row corresponding to the odd terms along the horizontal axis of the downsampled block. Thus, the output size can be 32, and again, only horizontal interpolation remains to be performed. At most (8 * 32) / (16 * 4) = 4 multiplications can be performed per sample.
[0108] The transposed cases can be handled accordingly. This is shown in the subsequent figures.
[0109] Figure 6.1 Device 54 for using intra prediction to decode a picture for a predetermined block 18 is shown.
[0110] Device 54 is configured to read a mode index 200 from a data stream 12 using a binary code 202, the mode index pointing to one of a matrix-based intra prediction mode list 204. The matrix-based intra prediction mode list 204 consists of an even number of matrix-based intra prediction modes, where the matrix-based intra prediction modes of the list 204 are grouped into matrix-based intra prediction mode pairs 212. Each pair 212 consists of a first matrix-based intra prediction mode and a second matrix-based intra prediction mode. Device 54 is configured to read the mode index 200 from the data stream 12 using the binary code 202 in such a way that for each matrix-based intra prediction mode pair 212, the first matrix-based intra prediction mode is assigned a first codeword and the second matrix-based intra prediction mode is assigned a second codeword and the lengths of the two codewords are equal.
[0111] Optionally, the binarization code 202 is a variable length code that includes codewords of different lengths. Alternatively, the binarization code may be a truncated binary code, and the number of matrix-based intra prediction modes is not a power of two, such that the truncated binary code has codewords of different lengths. The matrix-based intra prediction mode associated with the first matrix-based intra prediction mode pair 212 may be assigned a codeword that has a different length from the codeword assigned to the matrix-based intra prediction mode associated with the second matrix-based intra prediction mode pair 212. However, the two codeword lengths of the matrix-based intra prediction mode pair 212 are equal.
[0112] According to an embodiment, the device 54 may be configured to read the mode index 200 from the data stream 12 using the equiprobable bypass mode of a context adaptive binary arithmetic decoder.
[0113] Similar to the device 54 (i.e., the decoder) for using intra prediction to decode a predetermined block 18 of a picture, a device (i.e., the encoder) for using intra prediction to encode a predetermined block 18 of a picture may be configured to encode the mode index 200 into the data stream 12 using the binarization code 202 and optionally using the equiprobable bypass mode of a context adaptive binary arithmetic encoder.
[0114] The decoder and the encoder are configured to predict the samples 108 of the predetermined block 18 by computing the matrix-vector product 206 between the input vector 102 derived from the reference samples 17 in the neighborhood of the predetermined block 18 and the prediction matrix 19 associated with the matrix-based intra prediction mode k pointed to by the mode index 200. The computation of the matrix-vector product 206 results in an output vector 208. Further, the samples 108 of the predetermined block 18 are predicted by associating the components 210 of the output vector 208 obtained by the matrix-vector product 206 to the sample positions 104 of the predetermined block 18. This prediction of the samples 108 of the predetermined block 18 may be performed as described with respect to Figures 5.1 to 5.4 is performed.
[0115] For each matrix-based intra prediction mode pair 212, the prediction matrix 19 associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212 is equal to the prediction matrix 19 associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212. Thus, for matrix-based intra prediction modes 2k and 2k + 1, the same prediction matrix 19 is used. For each matrix-based intra prediction mode pair 212, the encoder and decoder are configured such that if the matrix-based intra prediction mode pointed to by the mode index 200 is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212 (e.g., the mode with an odd mode index 2k + 1), the association of the reference samples 17 in the neighborhood of the predetermined block 18 with the components 214 of the input vector 112 and the association of the sample positions 104 of the predetermined block 18 with the components 210 of the output vector 208 are transposed with respect to the associations in the case where the matrix-based intra prediction mode pointed to by the mode index 200 is the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212 (e.g., the mode with an even mode index 2k).
[0116] The decoder / encoder can be configured to determine, based on the parity of the mode index 200, whether the matrix-based intra prediction mode pointed to by the mode index 200 is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair or the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212. The parity of the mode index 200 can indicate whether the input vector 102 and the output vector 208 are used in a transposed manner for the prediction of the samples 108 of the predetermined block 18. That is, as Figure 6.2 shown, if in the former case a certain component of components 1 to n of the input vector 102 is associated with the position (x, y) (where (0, 0) represents the top-left sample AA of the predetermined block 18), then in the latter case it is associated with (y, x). The same applies to the components of the output vector 208 (AA, AB, AC, BA, CA,...).
[0117] Each pair 212 consists of a first matrix-based intra prediction mode and a second matrix-based intra prediction mode, which are related to each other by the same prediction matrix 19 and differ from each other only in terms of whether the input vector 102 and the output vector 208 are transposed or not. This is advantageous because only the mode index 200 is required in the data stream 12 to indicate the matrix-based intra prediction mode and whether the matrix-based intra prediction mode is used in a transposed manner. No additional index or flag is required to indicate that the input vector 102 and the output vector 208 are to be used in a transposed manner for the matrix vector product 206.
[0118] According to an embodiment, the decoder / encoder is configured to index prediction matrix 19 from a plurality of prediction matrices using the integer part of dividing mode index 200 by 2. This is based on the idea that two matrix-based intra prediction modes of a pair 212 use the same prediction matrix 19 to predict samples 108 of a predetermined block 18, and for this reason, the prediction matrix 19 has been sufficiently indicated by pointing to the relevant pair 212 in list 204 with mode index 200.
[0119] As Figure 6.1 and 6.2 shown in, the decoder / encoder may be configured to horizontally set 217 the inter-sample distance 216 of the sample positions 104 of the samples of the predetermined block 18 and the inter-sample distance 218 of the reference samples 17 in the neighborhood of the predetermined block 18 according to a first ratio of the horizontal size 220 of the predetermined block 18 relative to a horizontal default size and / or vertically according to a second ratio of the vertical size 222 of the predetermined block 18 relative to a vertical default size. This enables the use of the matrix-based intra prediction mode list 204 for multiple block sizes. The device may fill the space between the predicted samples by interpolation. The setting 217 of the inter-sample distances 216 of the sample positions 104 of the samples of the predetermined block 18 and the inter-sample distances 218 of the reference samples 17 in the neighborhood of the predetermined block 18 enables an improved distribution of the predicted samples 108 in the predetermined block 18 and the reference samples 17 in the neighborhood of the predetermined block 18. Therefore, the predicted samples can be equally distributed, enabling improved interpolation of the samples of the predetermined block 18.
[0120] According to an embodiment, the decoder / encoder is configured to equally sort the matrix-based intra prediction modes in the matrix-based intra prediction mode list 204 for multiple block sizes. Alternatively, the order may be adapted to, for example, blocks with a width greater than the height, and vice versa, i.e., height greater than width, or quadratic blocks. This sorting can increase the coding efficiency and reduce the bitstream because the matrix-based intra prediction modes for a common block size can be associated with short codewords, and the matrix-based intra prediction modes for rare block sizes can be associated with longer codewords.
[0121] Optionally, the multiple block sizes include at least one block size corresponding to an aspect ratio greater than 4. The matrix-based intra prediction can be optimized such that a predetermined block 18 having an aspect ratio of the horizontal size 220 to the vertical size 222 greater than 4. That is, the multiple block sizes include a predetermined block having a horizontal size 220 that is at least four times greater than the vertical size 222 and / or a predetermined block having a vertical size 222 that is at least four times greater than the horizontal size 220. Figure 6.2 A predetermined block 18 having a block size corresponding to an aspect ratio greater than 4 can be shown.
[0122] According to the embodiments presented below, the MIP mode is applied in a way that makes the use of MIP even more efficient compared to the use expected so far in the current VVC version.
[0123] The embodiments below will mainly show in view of the characteristics and functionality of the decoder. However, it is clear that the encoder may include the same or similar characteristics and functionality. For example, the decoding performed by the decoder may correspond to the encoding by the encoder. In addition, the encoder may include the same characteristics as those described regarding the decoder in the feedback loop (e.g., in the prediction stage 36).
[0124] Figure 7 An apparatus 54 for decoding a picture of a predetermined block 18 using intra prediction is shown.
[0125] The apparatus 54 is configured to read a mode index 200 from a data stream 12. The mode index 200 points to one of the matrix-based intra prediction modes 2051 - 205 n in a list 204, i.e., the MIP mode. The number n of the matrix-based intra prediction modes 2051 - 205 in the matrix-based intra prediction mode list 204 is, for example, 12, 16, or 32. Moreover, the embodiments focus on the intra prediction using the MIP modes 2051 - 205 n It is clear that the mode index 200 can also be used to indicate further modes, such as further intra prediction modes and / or inter prediction modes. The mode index 200 may be inserted into the data stream 12 by an apparatus 14 for encoding a picture of a predetermined block 18 using intra prediction. n
[0126] According to an embodiment, the matrix-based intra prediction modes 2051 - 205 in the matrix-based intra prediction mode list 204 n are associated with 6, 8, or 16 different prediction matrices 19.
[0127] According to an embodiment, the matrix-based intra prediction mode list 204 includes MIP modes for multiple block sizes.
[0128] According to an embodiment, there may be two or more MIP mode lists 204, where the two or more MIP mode lists 204 differ from each other in terms of the block size of a predetermined block 18 associated with the MIP mode. MIP modes associated with the same or similar block sizes are included in the same list of the two or more MIP mode lists 204. For example, there may be different lists of mutually exclusive block size sets, one having 6 different matrices associated with 12 modes of block sizes within a first block size set, one having 8 different matrices associated with 16 modes of smaller block sizes within a second block size set, and one having 16 different matrices associated with 32 modes of even smaller block sizes within a third block size set. This is merely an example, and it is clear that different numbers of lists 204 are possible and each list may include MIP modes associated with a block size set different from the block size sets described above.
[0129] Device 54 is configured to predict samples 108 of a predetermined block 18 by calculating a matrix-vector product 206 between an input vector 102 derived from reference samples 17 in the neighborhood of the predetermined block 18 and a prediction matrix 19 associated with a matrix-based intra prediction mode 205 pointed to by a mode index 200, and by associating components 210 of an output vector 208 obtained through the matrix-vector product 206 to sample positions 104 of the predetermined block 18.
[0130] For each matrix-based intra prediction mode 2051 - 205 n , all terms of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 are represented by a fixed-point representation 190 having a predetermined bit depth 192. The fixed-point representation 190 may have a fractional part 194 (having n bits), an integer part 196 (having m bits), and an optional sign bit 198. As Figure 7 shown, the predetermined bit depth 192 is equal for all matrix-based intra prediction modes 2051 - 205 in the MIP mode list 204, see Constraint 1 below. In the case where there are two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, for example, the predetermined bit depth 192 is equal for all matrix-based intra prediction modes 2051 - 205 n in the corresponding MIP mode list. In other words, for example, the predetermined bit depth 192 is at least the same for MIP modes related to the same block size set. n
[0131] According to an embodiment, device 54 is configured for each matrix-based intra prediction mode 2051 - 205 n, for each term in the prediction matrix 19 associated with a corresponding matrix-based intra prediction mode 205, it is stored in a fixed-point representation with a predefined bit depth. For example, this is shown in the following representative examples, see List 1 to List 4.
[0132] For each matrix-based intra prediction mode 2051 - 205 n , the device 54 is configured to calculate the matrix-vector product 206 between the input vector 102 and the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 by performing a right shift 209 with an equal number of bits 211 for each component 210 of the output vector for all matrix-based intra prediction modes 2051 - 205 in the MIP mode list 204, as n shown. The number of bits 211 for the right shift is, for example, x bits. In the case where there are two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, the number of bits 211 is, for example, equal for all matrix-based intra prediction modes 2051 - 205 in the corresponding MIP mode list. In other words, the number of bits 211, for example, is at least the same for MIP modes associated with the same set of block sizes. The right shift 209 is, for example, indicated by >> in the equation (2) described above, and for the number of bits 211, compare d Figure 7 as shown. In the case where there are two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, the number of bits 211 is, for example, equal for all matrix-based intra prediction modes 2051 - 205 in the corresponding MIP mode list. That is to say, the number of bits 211, for example, is at least the same for MIP modes associated with the same set of block sizes. The right shift 209 is, for example, indicated by >> in the equation (2) described above, and for the number of bits 211, compare d n is equal. In other words, the number of bits 211, for example, for MIP modes related to the same set of block sizes is at least the same. The right shift 209, for example, is indicated by >> in the equation (2) described above, and for the number of bits 211, compare d i , where the device 54 applies the constraint that the number of bits 211 is equal for matrix-based intra prediction modes 2051 - 205 n and compare the following constraint 2.
[0133] For the matrix A i (i.e., the prediction matrix 19) and the parameters c i and d i in the equation (2), the following constraints are desirable.
[0134] 1. The range of matrix terms for MIP is fixed. For example, the terms of the prediction matrix 19 are represented by a fixed-point representation 190 with a predefined bit depth 192. Thus, there are predefined non-negative integers μ 1,low , μ 1,up and μ 2,low , μ 2,up such that for each MIP mode i 2051 - 205 n and for each matrix term a i of the matrix A k,l (i.e., the prediction matrix 19), there is
[0135]
[0136] and
[0137]
[0138] Two specific examples of such a constraint are given below.
[0139] The first example is that for a fixed positive integer μ, for all matrix entries a of all matrices A used in MIP prediction i of all matrices k,l , there is
[0140] 0 ≤ a k,l ≤ 2 μ -1
[0141] and
[0142] -2 μ ≤ a k,l - c i ≤ 2 μ -1.
[0143] According to the first example, the entries of prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 have the same sign, e.g., positive, and device 54 is configured to, before computing matrix-vector product 206, for each matrix-based intra prediction mode 2051 - 205 n , shift all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 by an offset value c n equal for the matrix-based intra prediction modes 2051 - 205 i , where, for each matrix-based intra prediction mode 2051 - 205 n , all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 can be represented by an 8-bit representation of the sign. Accordingly, according to this alternative, only a 7-bit magnitude can be stored for each matrix entry. This is due to the fact that the sign bit does not have to be stored.
[0144] The second example is that for each MIP mode i 2051 - 205 n , there is c i = 0, and there exists a fixed positive integer v such that for each MIP mode i 2051 - 205 n and each matrix entry a of the matrix A i k,l , there is
[0145] -2 v ≤ a k,l ≤ 2 v -1.
[0146] According to the second example, the device 54 is configured for each matrix-based intra prediction mode 2051-205 n to store the entries of the prediction matrix 19 associated with the respective matrix-based intra prediction mode 205 in an 8-bit signed magnitude representation. In this second example, the device 54 does not offset the entries of the prediction matrix 19 associated with the respective matrix-based intra prediction mode 205 by an offset value c before computing the matrix-vector product 206 i .
[0147] 2. Shift value d i , i.e., the number of bits 211 for the right shift 209, is independent of the MIP mode i2051-205 n . Thus, there exists a positive integer d such that for each MIP mode i2051-205 n , there is
[0148] d i = d.
[0149] Optionally, the following constraints may also be desirable:
[0150] 3. Value c i , i.e., the offset of the entries of the prediction matrix 19, is independent of the MIP mode i2051-205 n . Thus, there exists a positive integer c such that for each MIP mode i2051-205 n , there is
[0151] c i = c.
[0152] Thus, the device 54 may be configured to, before computing the matrix-vector product 206, for each matrix-based intra prediction mode 2051-205 n , e.g., by adding or by subtracting an offset value that offsets all the entries of the prediction matrix 19 associated with the respective matrix-based intra prediction mode 205 by an amount equal for all the matrix-based intra prediction modes 2051-205 of the MIP mode list 204 n , compare c.
[0153] In the case of two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, the offset value c is, e.g., equal for all the matrix-based intra prediction modes 2051-205 in the respective MIP mode list n . In other words, the offset value c is, e.g., the same for at least the MIP modes associated with the same set of block sizes.
[0154] The reasons for imposing these constraints are as follows. Constraint 1 enables an efficient implementation of the matrix-vector multiplication 206 of equation (2) (Ai -c i )·r red , because if all the terms of all matrices (A i -c i ) have a common fixed bit depth (i.e., a predetermined bit depth of 192), then a specific multiplier suitable for that bit depth can be used, and it can be shared across all MIP modes 2051 - 205 n for calculating the matrix - vector product 206. Additionally, if all the terms of all matrices A i can be stored with fixed precision, i.e., the terms of the prediction matrix 19 are represented in fixed - point notation 190, then efficient memory management is enabled (when dealing with matrix A i ). Here, an important example is that all the terms of all matrices A i can be stored with 8 - bit precision (i.e., in one byte). In other words, the terms of the prediction matrix 19 can be represented in fixed - point notation 190 (with a predetermined bit depth 192 of 8 bits).
[0155] Constraint 2 enables efficient implementation of the shift in expression (2) because if the shift value, i.e., the number of bits 211 for the right shift, does not depend on the MIP mode i, then table lookups are saved, and a single fixed shift operation can be implemented for MIP, which is beneficial for a compact SIMD implementation of equation (2), and it reduces the case - dependent implementation of equation (2) in hardware implementation, i.e., the case - dependent implementation of the matrix - vector product 206 of the matrix - vector product with a prediction matrix approximately in floating - point precision. Here, a particularly important example is that the value 6 is used as the fixed shift, i.e., as the number of bits 211. The reason is that for 10 - bit content, clipping to a 10 - bit range is applied to pred during the MIP prediction process red . Thus, before shifting down by 6, i.e., before performing the right shift 209 (with the number of bits 211 being 6) in equation (2), the term (A i -c i )·r red +(1 << (d i -1)) can be stored in 16 bits (i.e., 2 bytes), where the term (A i -c i )·r red +(1 << (d i -1)) represents the intermediate result 108’ obtained from the matrix - vector product 206 for each component 210 of the output vector 208.
[0156] According to an embodiment, the device 54 is configured to apply a right shift 209 to an intermediate result 108’ obtained, for example, from the matrix - vector product 206 for each component 210 of the output vector 208 (e.g., (A i -ci )·r red +(1 << (d i -1))) to compute the matrix-vector product 206 between the input vector 102 and the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 in fixed-point arithmetic. Optionally, the intermediate result 108' is represented with a bit precision at least twice as high as the bit precision (with which the entries of the prediction matrix 19 associated with the matrix-based intra prediction mode are stored), e.g., the intermediate result 108' can be stored in 16 bits, and the entries of the prediction matrix 19 can be stored as 7-bit magnitudes or as a signed 8-bit representation.
[0157] According to an embodiment, the device is configured to decode the picture at 10-bit resolution, for each matrix-based intra prediction mode 2051-205 n store the magnitudes of the entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode with 7-bit precision, and use 6 bits as the number of bits 211.
[0158] Similar to constraint 2), constraint 3) enables a more efficient implementation of equation (2) (again saving table lookups).
[0159] The problem that this application intends to solve is that it is not obvious how constraint 2 or constraint 3 should be satisfied together with constraint 1 such that equation (2) is used as an approximation of equation (1).
[0160] For example, assume that by constraint 1, all matrix entries of each MIP matrix A i must be stored with 8-bit precision, and by constraint 2, a fixed shift d i = 6 must be used for all MIP modes i in equation (2). Also, assume for simplicity that ci = 0. This would mean that if equation (2) should approximate equation (1), then for each floating-point matrix that is the result of the training algorithm for MIP there must exist a non-negative integer c i such that each entry of must satisfy
[0161] where ε is moderately small so that
[0162] equation (2) can be executed to moderately approximate equation (1). Here, rounding is applied to each entry of the matrix. Additionally, for a real number x,
[0163] clip(x, 2 7 ) := min(max(x, -2 7 ), 2 7-1)
[0164] and is denoted by clip(A, 2 7 ), where A is a matrix that is produced by applying clip(-, 2 7 ) to each entry of A. Note that, if it is assumed that there exists a matrix A′ i with integer entries in the 8-bit range (for which equation (2) is an approximate equation (1) for a variable input vector r red with a fixed shift of 6), then the matrix A i defined in assignment (4) must approximate
[0165] On the other hand, a priori, there is no reason why a training algorithm (whose output is a MIP matrix ) should produce a matrix that is approximated by the corresponding matrix A i defined in (4). The problem is that may contain matrix entries with absolute values greater than , and thus the clipping in equation (4) introduces a substantial difference between (1) and (2) by discarding parts of the most significant matrix entries of the matrix . Therefore, applying the posterior (4) to the trained matrix may result in a significant deviation of the MIP prediction mode specified in the codec and requiring the use of the matrix A to perform the matrix-vector product in equation (2) from the "true" behavior (i.e., the trained MIP mode using the matrix-vector product with the matrix i ). Thus, the whole concept of the data-driven approach for in-frame prediction supporting MIP is violated.
[0166] In fact, it can be observed that applying equation (4) to the training matrix that is the basis of the MIP mode used in the current VVC draft [1] significantly changes the behavior of some MIP modes when compared to the basic training mode, because some of the entries in the matrix
[0167] are much greater than 2. Finally, note that it is trivial to solve only constraint 1 without solving constraint 2 as long as each entry of each matrix lies between -2 7 and 2 7 -1, which is the case for the matrices supporting the MIP mode of the current VVC draft [1]. Here, assuming that currently c i = 0, the shift value d i is simply defined such that
[0168]
[0169] For each matrix entry of it holds, and such that (5) does not hold for any d′ > di. i >di does not hold.
[0170] Furthermore, it should be noted that the device 54 may include features and / or functionality as described with respect to Figure 6.1 and 6.2 This means, for example, that the matrix-based intra prediction mode list 2051 - 205 n list 204 includes one or more matrix-based intra prediction mode pairs 212. Note that the list 204 may not be specifically composed of such MIP mode pairs 212 (since they are depicted as being present in the list 204 in Figure 6.1 ), but there may also be other MIP modes that are specifically applied using the transpose option or the non-transpose option. For each matrix-based intra prediction mode 2051 - 205 n pair 212, the prediction matrix 19 associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212 is equal to the prediction matrix 19 associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair, for example, for modes 2k and 2k + 1, the same matrix 19 is used. The device is configured such that if the matrix-based intra prediction mode 205 pointed to by the mode index 200 is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212, the association of the reference samples 17 in the neighborhood of the predetermined block with the components 214 of the input vector 112 and the association of the sample positions 104 of the predetermined block 18 with the components 210 of the output vector 208 are transposed with respect to the case where the matrix-based intra prediction mode 205 pointed to by the mode index 200 is the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212. That is, if in the former case a certain component of the input vector 102 is associated with the position (x, y) (where (0, 0) represents the upper left corner sample of the predetermined block 18), then in the latter case it is associated with (y, x). The same applies to the components of the output vector 208. For more details, see the descriptions of Figure 6.1 and 6.2 .
[0171] Furthermore, it should be noted that the device 54 may include features and / or functionality as described with respect to Figures 5.1 to 5.4 .
[0172] According to an embodiment, the device 54 is configured to predict samples of a predetermined block 18 offset from a sample position associated with a component 210 of the output vector 208 by upsampling and / or interpolation based on the output vector 208 or based on the output vector 208 and reference samples 17 in a neighborhood of the predetermined block 18, as shown in, for example Figures 5.1 to 5.4 one of those shown in
[0173] According to an embodiment, the device 54 is configured to derive an input vector 102 from reference samples 17 in a neighborhood of the predetermined block 18 by downsampling and / or pooling, as shown in, for example Figures 5.1 to 5.4 one of those shown in
[0174] According to an embodiment, the reference samples 17 in a neighborhood of the predetermined block 18 include a first reference sample 17c above the predetermined block 18 and a second reference sample 17a to the left of the predetermined block 18. The device 54 is configured to derive an input vector 102 from the reference samples 17 in a neighborhood of the predetermined block 18 by: deriving a first intermediate component from the first reference sample 17c by downsampling and / or pooling, deriving a second intermediate component from the second reference sample 17a by downsampling and / or pooling, concatenating the first intermediate component and the second intermediate component to derive a preliminary input vector, and forming an input vector from the preliminary input vector
[0175] According to an embodiment, the device 54 is configured to decode a picture at a B-bit resolution. The device 54 may be configured to form the input vector 102 from the preliminary input vector by subtracting 2 from a first component of the preliminary input vector B-1 so as to obtain a first component of the input vector 102, and subtracting the first component of the preliminary input vector from further components of the preliminary input vector so as to obtain further components of the input vector 102. Alternatively, the device 54 may be configured to form the input vector 102 from the preliminary input vector by subtracting the first component of the preliminary input vector from further components of the preliminary input vector such that the input vector 102 is formed from the further components. Additionally, the device 54 is configured to correct the output vector 208 by component-wise addition of the first component of the preliminary input vector
[0176] According to an embodiment, entries of a prediction matrix of a matrix-based intra prediction mode 2051 - 205 in a matrix-based intra prediction mode list 204 correspond to entries in Table 2 below, see, for example, List 2. However, note that another shift value (i.e., another number of bits 211) may be selected for listing the values in the table, and possibly, the values in the table may be represented in another scale n
[0177] The following embodiments will focus on data-driven training of matrix-based intra prediction modes with predefined fixed coefficient ranges and predefined fixed shifts and its application in codecs.
[0178] The solution presented in the present invention to the problem of obtaining a fixed bit depth, a fixed shift and a fixed offset is to already in the training of the MIP prediction mode, i.e. in the matrix In the derivation of , constraints 1, 2 and 3 are included. Therefore, the range of all matrix entries has been restricted during training, where a gradient descent algorithm is applied to continuously guide the matrix towards a (local) optimum for a predefined loss function on a large training data set.
[0179] The simplest way to do this would be to multiply each matrix by 2 d , d as in constraint 2, then add the offset c from constraint 3 (if desired), then clip the result to the desired range of constraint 1, then subtract the offset c, and finally divide the result by 2 d However, this is not feasible because the clipping function has a gradient of zero outside the clipping range, and therefore, in such methods, every weight that falls outside the clipping range at some point in stochastic gradient descent will never be updated from then on.
[0180] FIG. 8 shows a method for training matrix-based intra prediction models 2051-205 n List 204 Prediction Matrix 19 i In an embodiment of the device 310, the matrix-based intra prediction mode 205 i The matrix-based intra prediction mode 205 selected for the predetermined block 18 should be selected for the predetermined block 18 by calculating the input vector 102 derived from the reference samples 17 in the neighborhood of the predetermined block 18 and the matrix-based intra prediction mode 205 selected for the predetermined block 18. i Correlated prediction matrix 19 i , and associates components 210 of an output vector 208 obtained by the matrix-vector product 206 to sample positions 104 of a predetermined block 18 to predict samples 108 of the predetermined block 18 .
[0181] The device 310 is configured to train 320 the matrix-based intra prediction modes 2051-205 using a gradient descent method 322 n List 204 Prediction Matrix 19 i . Prediction Matrix 19 i The training 320 is performed, for example, by using a training set of predetermined blocks 18 of known original samples and their corresponding neighborhoods 17. The matrix-based intra prediction modes 2051-205 represented in floating point representation are optimized using a cost function 324. nThe prediction matrix 19 of list 204 i The representative values of the items of, compare To train the 320 prediction matrix 19 i , the cost function 324 depends on the prediction distortion measure 326, and the prediction distortion measure 326 is related to setting the items of the prediction matrix 19 i To the intermediate value, using the differentiable function 328, compare f(x), and map the representative value to the intermediate value. The cost function 324 depends on the prediction distortion measure 326 such that the cost 325 increases with the decreasing prediction quality, as from Obtained, which means that f(x) is applied to Each item of. The prediction distortion measure 326 can be defined as the deviation between the prediction signal that can be obtained using the prediction matrix including the intermediate value And the original signal associated with the predetermined block of the training set. As can be seen in FIG. 8, in the case where the prediction signal Is equal to the original signal, the prediction distortion measure 326 is zero. The matrix The items of can be the representative values mentioned above, and the matrix The items of can be the intermediate values mentioned above. The matrix Can represent the prediction matrix under training. It should be noted that the gradient descent method 322 with the correlation of the cost function 324 and the cost 325 and the prediction distortion measure 326 is only schematically shown to illustrate the basic principle of the basic device 310
[0182] The domain of the differentiable function 328 (e.g., Figure 9 The x-axis 302 in), and the codomain (e.g., Figure 9 The y-axis 304 in) are defined by the floating-point representation. The graph 300 of the differentiable function 328 has a predetermined dynamic range, and the differentiable function 328 is equal for the matrix-based intra prediction modes 2051-205 n Is equal. The predetermined dynamic range is defined by max(image 300) / min(image300) for example. According to an embodiment, max(image 300) is α + δ and min(image 300) is -α + δ, see equation (7) below. It should be noted that Figure 9 Only shows the curve graph of the exemplary differentiable function f(x) 328
[0183] In addition, the device 310 is configured to quantize the intermediate value 330 to the fixed-point representation 190 after training 320 for example, such that for each matrix-based intra prediction mode 2051-205 n , the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 i i All terms are represented by a fixed-point representation 190 having a predetermined bit depth 192, such that the predetermined bit depth 192 is equal for the matrix-based intra prediction modes 2051 - 205 n and such that for each matrix-based intra prediction mode 2051 - 205 n , the matrix-vector product 206 between the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 i and the input vector 102 is computable by performing a right shift 209 by a number of bits 211 that is equal for the matrix-based intra prediction modes 2051 - 205 n . For example, this means that the non-shifted-out part b x+1 to b y+x of the fixed-point representation is sufficient to represent α, see equation (7) below. It should be noted that FIG. 8 shows the fixed-point representation 190 as an (x + y + 1)-bit signed magnitude representation. However, for example, in the case where all intermediate values have the same sign (e.g., all intermediate values may be positive), it is also possible that the fixed-point representation 190 is an (x + y)-bit magnitude representation.
[0184] Thus, as a solution, in the present invention, the clipping operation is approximated by a smoothing function (e.g., a differentiable function 328). More precisely, from among Constraint 1, Constraint 2, and optionally Constraint 3, the range of the unscaled matrix terms is calculated, i.e., the range of the representative values of the prediction matrix terms, and compared such that if applied during training
[0185]
[0186] where represents the current matrix during the training process, the result lies within the range of Constraint 1. Then, during training, each is clipped to this range by applying a smoothed approximation f of the clipping function implemented as follows, i.e., the differentiable function 328:
[0187]
[0188] where α, β, γ, and δ are real numbers depending on the clipping range (i.e., the predetermined dynamic range). In addition, λ is a non-negative integer that can be chosen experimentally. Figure 9 FIG. shows an example of the clipping function f(x), i.e., an example of the differentiable function 328.
[0189] According to an embodiment, the differentiable function 328 has a slope of 1 at the origin, is strictly monotonically increasing, and has horizontal asymptotes at the upper and lower bounds of the image 300.
[0190] According to an embodiment, the differentiable function 328 can be parameterized by a shift parameter (e.g., δ in terms of the shift of the image 300 within the co-domain 304). Further, the device 310 can be configured to subject the shift parameter to optimization using the gradient descent method 322. Further, the device 310 can be configured to derive an offset value, c, from the shift parameter for comparison, such that, prior to computing the matrix-vector product 206, it is used to offset all terms of the prediction matrix 19 associated with the respective matrix-based intra prediction mode, e.g., by addition or by subtraction, for each matrix-based intra prediction mode. The derived offset value c is equal for the matrix-based intra prediction modes 2051 - 205 n is equal.
[0191] In summary, the invention of the present application is an implementation of MIP having a portion given by Equation (2) that satisfies Constraint 1 and Constraint 2, or an implementation of MIP having a portion given by Equation (2) that satisfies Constraint 1, Constraint 2, and Constraint 3, and in both cases, in the training algorithm for the floating-point matrix which is then quantized into an integer matrix, Constraint 1 and Constraint 2 are employed as described in this section, and if desired, Constraint 3 is also employed.
[0192] The following embodiments describe examples of the storage representation of the prediction matrix.
[0193] The following shows the floating-point matrix obtained from the training for the MIP mode with mipSizeId = 2, [1] using the techniques provided in the present application Specifically, the parameters of the clipping function (i.e., the differentiable function) are selected in such a way that the training generates matrix coefficients that can be represented using 7-bit unsigned integer numbers with a fixed shift of 6 and a fixed offset of 32.
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220] List 1: Floating-point matrix coefficients obtained from training
[0221] The matrix coefficients shown in List 1 satisfy the requirements presented in the foregoing section. To show this, List 2 shows the matrix coefficients after multiplying by 2 6 The matrix coefficients after multiplication are offset from the final range by only a fixed offset of 32. Thus, the following values are the stored matrix terms in fixed-point representation according to the example. The matrix is a 7×64 matrix (7-component input vector and 64-component output vector). There are 6 matrices for 6 modes. According to an embodiment, the terms may deviate from the values shown below. For example, multiplying by 2 is chosen for illustrative purposes only 6 , and accordingly, when another factor is chosen, the terms of the matrix (as they are shown below) may look different.
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238]
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247] List 2: Floating-point matrix coefficients in proportion
[0248] Now, a fixed offset of 32 is added to these matrix coefficients to obtain a set of matrices in which all coefficients are equal to or greater than -0.5 and are thus rounded to non-negative values. This set is shown in List 3.
[0249]
[0250]
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258]
[0259]
[0260]
[0261]
[0262]
[0263]
[0264]
[0265]
[0266]
[0267]
[0268]
[0269]
[0270]
[0271]
[0272]
[0273]
[0274]
[0275] List 3: Proportional floating-point matrix coefficients after adding a constant offset
[0276] Finally, the above matrix coefficients are rounded to integer precision, i.e., the intermediate values are quantized from 330 to the fixed-point representation 190. Since the smallest coefficient of the above set is -0.5, the resulting integer coefficients are non-negative and can therefore be represented by unsigned integer numbers. The largest coefficient of the above matrix set is 127.5. These coefficients would typically be rounded to 128, but here they are rounded to 127, introducing the same absolute rounding error. Therefore, the resulting integer coefficients shown in List 4 are from the unsigned 7-bit range. Additionally, since the smallest rounded coefficient is 0 and the largest rounded coefficient is 127, the resulting integer coefficients fully utilize this range.
[0277]
[0278]
[0279]
[0280]
[0281]
[0282]
[0283]
[0284]
[0285]
[0286]
[0287]
[0288]
[0289]
[0290]
[0291] List 4: Unsigned 7-bit integer coefficients
[0292] Although some aspects have been described in the context of devices, these aspects also represent a clear description of the corresponding method, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of the corresponding blocks or items or features of the corresponding device. Some or all of the method steps can be performed by (or with) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0293] The data stream of the present invention can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0294] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be performed using a digital storage medium on which an electronically readable control signal is stored, such as a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, which cooperates (or is capable of cooperating) with a programmable computer system such that the corresponding method is executed. Thus, the digital storage medium can be computer-readable.
[0295] Some embodiments according to the present invention include a data carrier having an electronically readable control signal, which is capable of cooperating with a programmable computer system in order to execute one of the methods described herein.
[0296] In general, embodiments of the present invention can be implemented as a computer program product having program code, which, when the computer program product runs on a computer, the program code is operative to execute one of the methods. The program code can be stored, for example, on a machine-readable carrier.
[0297] Other embodiments include a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0298] Thus, in other words, embodiments of the method of the present invention are computer programs having program code, which, when the computer program runs on a computer, the program code is for executing one of the methods described herein.
[0299] Thus, a further embodiment of the method of the present invention is a data carrier (or a digital storage medium, or a computer-readable medium), the data carrier including a computer program recorded thereon for executing one of the methods described herein. The data carrier, the digital storage medium, or the recording medium is generally tangible and / or non-transitory.
[0300] Accordingly, a further embodiment of the method of the present invention is a data stream or signal sequence representing a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0301] A further embodiment includes a processing component, such as a computer or a programmable logic device, which is configured to or adapted to perform one of the methods described herein.
[0302] A further embodiment includes a computer on which a computer program for performing one of the methods described herein is installed.
[0303] A further embodiment according to the present invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0304] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0305] The apparatus described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0306] The apparatus described herein or any component of the apparatus described herein may be implemented at least in part in hardware and / or in software.
[0307] The methods described herein may be performed using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0308] The methods described herein or any component of the apparatus described herein may be performed at least in part by hardware and / or by software.
[0309] The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to other technicians in the art. Therefore, it is intended to be limited only by the scope of the appended patent claims and not by the specific details presented by the description and explanation of the embodiments herein.
[0310] References
[0311] [1]B. Bross et al., Versatile Video Coding (Draft 7), document JVET-P2001, Geneva, October 2019
Claims
1. An apparatus (54) for decoding a predetermined block (18) of a picture using intra prediction, configured to read a mode index (200) from a data stream (12), the mode index pointing to one in a matrix-based intra prediction mode list (204), Predicting the samples (108) of the predetermined block (18) by calculating a matrix-vector product (206) between an input vector (102) derived from reference samples (17) in the neighborhood of the predetermined block (18) and a prediction matrix (19) associated with the matrix-based intra prediction mode (k) pointed to by the mode index (200), and by associating components (210) of an output vector (208) obtained through the matrix-vector product (206) to sample positions (104) of the predetermined block, wherein for each matrix-based intra prediction mode (2051-205 n ), all terms of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ) are represented by a fixed-point representation (190) having a predetermined bit depth (192), and the predetermined bit depth (192) is equal for the matrix-based intra prediction modes (2051-205 n ). wherein the device (54) is configured to, for each matrix-based intra prediction mode (2051-205 n ), compute the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 n ) by performing a right shift (209) with an equal number of bits (211) for each component (210) of the output vector (208) for the matrix-based intra prediction mode (2051-205 i ). For each matrix-based intra prediction mode (2051-205 n ), store, with 7-bit precision, the magnitudes of the entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ), and using 6 bits as the number of bits (211).
2. The apparatus (54) according to claim 1, wherein at least one of the following is satisfied: The number of the matrix-based intra prediction modes (2051-205 n ) in the matrix-based intra prediction mode list is 12, 16, or 32; or The device (54) is configured to use the matrix-based intra prediction mode (2051-205 n ) list for multiple block sizes.
3. The apparatus (54) according to claim 1, configured such that the matrix-based intra prediction mode (2051-205 n ) in the matrix-based intra prediction mode list has 6, 8, or 16 different matrices associated therewith.
4. The apparatus (54) according to claim 1, wherein the apparatus (54) is configured to calculate the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with a respective matrix-based intra prediction mode (205; 205 i ) by applying the right shift (209) to an intermediate result (108’) obtained from the matrix-vector product (206) for each component (210) of the output vector (208), the intermediate result (108’) being represented with a bit-precision that is twice as high as the bit-precision in which the entries of the prediction matrix (19) associated with the matrix-based intra prediction mode (2051-205 n ) are stored.
5. The apparatus (54) according to claim 1, wherein the apparatus (54) is configured to, before calculating the matrix-vector product (206), for each matrix-based intra prediction mode (2051-205 n ), offset all terms of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ) by an offset value equal for the matrix-based intra prediction modes (2051-205 n ).
6. The apparatus (54) according to claim 1, wherein the apparatus (54) is configured to, for each matrix-based intra prediction mode (2051-205 n ), store the fixed-point representation (190) at the predetermined bit depth (192) for each entry of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ).
7. The apparatus (54) according to claim 1, wherein the apparatus (54) is configured to decode the picture at 10-bit resolution.
8. The apparatus (54) according to claim 1, wherein the matrix-based intra prediction mode (2051 - 205 n ) list (204) includes one or more matrix-based intra prediction mode pairs (212) consisting of a first matrix-based intra prediction mode and a second matrix-based intra prediction mode, and for each matrix-based intra prediction mode pair, the prediction matrix (19) associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair is equal to the prediction matrix (19) associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair, and the apparatus (54) is configured such that if the matrix-based intra prediction mode pointed to by the mode index is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair, the association of the reference samples (17) in the neighborhood of the predetermined block with the components (214) of the input vector (102) and the association of the sample positions (104) of the predetermined block (18) with the components (210) of the output vector (208) are transposed relative to the association in the case where the matrix-based intra prediction mode pointed to by the mode index is the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair.
9. An apparatus (14) for encoding a predetermined block (18) of a picture using intra prediction, configured to Insert a mode index (200) into the data stream (12), the mode index pointing to one of the matrix-based intra prediction mode (2051-205 n ) lists (204), Predicting the samples (108) of the predetermined block (18) by calculating a matrix-vector product (206) between an input vector (102) derived from reference samples (17) in the neighborhood of the predetermined block (18) and a prediction matrix (19) associated with the matrix-based intra prediction mode (k) pointed to by the pattern index (200), and by associating components (210) of an output vector (208) obtained by the matrix-vector product (206) to sample positions (104) of the predetermined block, wherein for each matrix-based intra prediction mode (2051-205 n ), all terms of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ) are represented by a fixed-point representation (190) having a predetermined bit depth (192), and the predetermined bit depth (192) is equal for the matrix-based intra prediction modes (2051-205 n ). wherein the device (14) is configured to, for each matrix-based intra prediction mode (2051 - 205 n ), calculate the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 n ) by performing a right shift (209) with an equal number of bits (211) for each component of the output vector for the matrix-based intra prediction mode (2051 - 205 i ). For each matrix-based intra prediction mode (2051-205 n ), store, with 7-bit precision, the magnitudes of the entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ), and using 6 bits as the number of bits (211).
10. The apparatus (14) according to claim 9, wherein at least one of the following is satisfied: The number of the matrix-based intra prediction modes (2051-205 n ) in the matrix-based intra prediction mode list is 12, 16, or 32; or The device (14) is configured to use the matrix-based intra prediction mode (2051-205 n ) list for multiple block sizes.
11. The apparatus (14) according to claim 9, configured such that the matrix-based intra prediction mode (2051-205 n ) in the matrix-based intra prediction mode list has 6, 8 or 16 different matrices associated therewith.
12. The apparatus (14) according to claim 9, wherein the apparatus (14) is configured to calculate the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with a respective matrix-based intra prediction mode (205; 205 i ) by applying the right shift (209) to an intermediate result (108') obtained from the matrix-vector product (206) for each component (210) of the output vector (208), the intermediate result (108') being represented with a bit precision twice as high as the bit precision in which the entries of the prediction matrix (19) associated with the matrix-based intra prediction mode (2051-205 n ) are stored.
13. The apparatus (14) according to claim 9, wherein the apparatus (14) is configured to at least one of the following: Before calculating the matrix-vector product (206), for each matrix-based intra prediction mode (2051-205 n ), all terms of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ) are offset by an offset value equal for the matrix-based intra prediction modes (2051-205 n ); or For each matrix-based intra prediction mode (2051-205 n ), for each entry of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ), store the fixed-point representation (190) at the predetermined bit depth (192).
14. The apparatus (14) according to claim 9, wherein the apparatus (14) is configured to encode the picture at 10-bit resolution.
15. The apparatus (14) according to claim 9, wherein the matrix-based intra prediction mode (2051-205 n ) list (204) includes one or more matrix-based intra prediction mode pairs (212) consisting of a first matrix-based intra prediction mode and a second matrix-based intra prediction mode, and for each matrix-based intra prediction mode pair, the prediction matrix (19) associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair is equal to the prediction matrix (19) associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair, and the apparatus (14) is configured such that if the matrix-based intra prediction mode pointed to by the mode index is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair, the association of the reference samples (17) in the neighborhood of the predetermined block with the components (214) of the input vector (102) and the association of the sample positions (104) of the predetermined block (18) with the components (210) of the output vector (208) are transposed relative to the association in the case where the matrix-based intra prediction mode pointed to by the mode index is the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair.
16. A method for decoding a predetermined block (18) of a picture using intra prediction, comprising reading a mode index (200) from a data stream (12), the mode index pointing to one in a matrix-based intra prediction mode list (204), predicting samples (108) of the predetermined block (18) by calculating a matrix-vector product (206) between an input vector (102) derived from reference samples (17) in a neighborhood of the predetermined block (18) and a prediction matrix (19) associated with the matrix-based intra prediction mode (k) pointed to by the mode index (200), and by associating components (210) of an output vector (208) obtained by the matrix-vector product (206) to sample positions (104) of the predetermined block, wherein for each matrix-based intra prediction mode, all entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode are represented by a fixed-point representation (190) having a predetermined bit depth (192), the predetermined bit depth (192) being equal for the matrix-based intra prediction modes, wherein the method includes, for each matrix-based intra prediction mode, calculating the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (k) by performing a right shift (209) with an equal number of bits (211) for each component of the output vector, For each matrix-based intra prediction mode (2051-205 n ), store, with 7-bit precision, the magnitudes of the entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ), and using 6 bits as the number of bits (211).
17. A method for encoding a predetermined block (18) of a picture using intra prediction, comprising inserting a mode index (200) into a data stream (12), the mode index pointing to one in a matrix-based intra prediction mode list (204), Predicting samples (108) of the predetermined block (18) by calculating a matrix-vector product (206) between an input vector (102) derived from reference samples (17) in the neighborhood of the predetermined block (18) and a prediction matrix (19) associated with the matrix-based intra prediction mode (k) pointed to by the pattern index (200), and associating components (210) of an output vector (208) obtained by the matrix-vector product (206) to sample positions (104) of the predetermined block, wherein for each matrix-based intra prediction mode, all entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode are represented by fixed-point representation (190) having a predetermined bit depth (192), and the predetermined bit depth is equal for the matrix-based intra prediction modes. Wherein the method includes, for each matrix-based intra prediction mode, calculating the matrix-vector product (206) between the input vector (102) and the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (k) by performing a right shift with a number of bits equal for the matrix-based intra prediction mode for each component of the output vector. For each matrix-based intra prediction mode (2051-205 n ), store, with 7-bit precision, the magnitudes of the entries of the prediction matrix (19) associated with the corresponding matrix-based intra prediction mode (205; 205 i ), and Using 6 bits as the number of bits (211).
18. A computer program product having a computer program which, when run on a computer, is configured to perform the method according to claim 16 or 17.
19. A computer-readable medium having instructions stored thereon which, when executed, cause a computing device to perform the method according to claim 16 or 17.