Apparatus and method for decoding / encoding predetermined block of picture using intra prediction
By using fixed point representation and right shift operations in matrix-based intra prediction mode, the matrix-vector product approximation problem is solved, improving the efficiency of video encoding and storage management, and reducing the bitstream size.
Patent Information
- Application Number
- CN202510820947.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-12-04
- Publication Date
- 2025-08-12
AI Technical Summary
The existing matrix-based intra prediction mode has a matrix-vector product in video encoding that requires approximation through integer operations, causing the prediction signal to deviate from the real behavior and affect the encoding and decoding efficiency.
Matrix-vector product is calculated by using fixed point representations with pre-located depths and fixed right shift operations, ensuring that all prediction matrix terms have the same bit depth, and optimizing the prediction matrix as fixed point representation during training, reducing storage and computational complexity.
More efficient matrix-vector product calculation is implemented, which improves coding efficiency, reduces bitstream size, and improves the performance of the codec.
Smart Images

Figure CN120475145A_ABST
Abstract
Description
Technical Field
[0001] Embodiments according to the present invention relate to matrix-based intra prediction with global setting of modes for picture and video encoding / decoding. Background Art
[0002] Typical block-based image or video codecs typically operate through predictive coding. Therefore, when a receiver of a coded image or video signal generates this signal for a given block, it constructs a prediction signal from information already available from the coded data. This prediction signal serves as a first approximation of the signal for that block. In a second step, a prediction residual is decoded from the bitstream and added to the prediction signal. The better the prediction signal, the smaller the number of bits required to transmit the prediction residual. Therefore, the quality of the prediction signal significantly impacts the overall codec efficiency.
[0003] Generally, there are two methods for generating prediction signals. The first method, used exclusively in video codecs, is inter-frame prediction. Here, the prediction signal is generated from reconstructed samples belonging to a different frame than the current one. The second method is intra-frame prediction. Here, the prediction signal is generated from reconstructed samples belonging to the same frame and typically spatially adjacent to a given block.
[0004] In classic codecs, intra-frame prediction is performed using either angular prediction mode or DC and planar modes. Angular prediction mode replicates the reconstructed samples to the left and above the block in a specific direction defined by an angle parameter, using an interpolation filter for fractional angular positions. DC mode generates the prediction signal as the average sample value of neighboring samples to the left and above the block. Finally, planar mode generates the prediction signal as a linear combination of predictions in the horizontal and vertical directions. Optionally, post-filtering of the prediction signal or pre-smoothing of the reference samples can be applied for any of the aforementioned prediction techniques.
[0005] Different from the classical intra prediction methods described above, matrix-based intra prediction (MIP) is introduced as a new technique for generating intra prediction signals. It is part of the current draft of the Evolved Versatile Video Coding (VVC) standard [1]. MIP can be seen as a low-complexity variant of the more general data-driven neural network-based intra prediction mode. Each MIP mode generates an intra prediction signal by multiplying a predefined matrix depending on the prediction mode with the downsampled versions of the top and left boundary samples and then upsampling the result. For more details, we refer to the section review on matrix-based intra prediction.
[0006] A key property of MIP is that the matrices for various MIP patterns are determined via a training algorithm using a large set of training data. In this training algorithm, an attempt is made to find matrices such that they minimize a predefined loss function with respect to the training data. Here, a stochastic gradient descent method is used, in which the matrix entries are updated iteratively. Such methods for determining the matrix entries require calculations in floating-point operations, and therefore, the resulting matrix entries are given as floating-point numbers. Thus, after training, for each MIP pattern i, a matrix with floating-point entries is obtained. So that in floating point, for MIP mode i, the reduced prediction signal Given as
[0007]
[0008] where r red represents a downsampled version of the boundary of a given block, and where · represents a matrix-vector multiplication.
[0009] On the other hand, for application in the final standard, each matrix-vector multiplication (1) needs to be approximated by the rules specified in integer operations. This means that for each MIP mode i, there is an integral term and a positive integer c i and d i The matrix A i must be specified so that the predicted signal pred is reduced red The calculation is specified as
[0010] pred red =((A i -c i )·r red +(1<<(d i -1)))>>d i (2)
[0011] Here, A i -c i Indicates that when A i Subtract c from each term i Finally, if v and w are vectors, where w has integer entries, then v+(1<<(d i -1)) means that by changing 1<<(d i -1) is added to each term of v, and w>>d i represents the sum of the terms of w by shifting them right by d. i The vector generated. According to the basic idea of MIP, (2) must be for all possible input vectors r red Approximately (1).
[0012] Therefore, we expect to obtain a matrix A with integer entriesi (For its equation (2) for the variable input vector r red A reasonably good approximation of equation (1). Otherwise, the matrix A specified in the codec and required to be used i The MIP prediction mode that performs the matrix-vector product in (2) may deviate significantly from the “real” behavior (i.e., using the same matrix as Thus, the whole concept of a data-driven approach to MIP-enabled intra prediction would be violated.
[0013] Hence, it is desirable to provide concepts for making picture coding and / or video coding more efficient in supporting matrix-based intra prediction. Additionally or alternatively, it is desirable to reduce the bitstream and, therefore, the signaling cost.
[0014] This is achieved by the subject matter of the independent claims of the present application. Summary of the Invention
[0015] Further embodiments according to the invention are defined by the subject matter of the dependent claims of the present application.
[0016] According to a first aspect of the present invention, the inventors of the present application have recognized that a problem encountered when attempting to predict samples of a predetermined block of a picture using a matrix-based intra prediction mode (MIP mode) stems from the fact that the matrix-vector product (i.e., matrix-vector multiplication) performed in the MIP mode needs to be approximated by integer operations, resulting in a large deviation between the approximated and non-approximated matrix-vector products (i.e., the "true" matrix-vector product). Hereinafter, the approximated matrix-vector product may be understood as the matrix-vector product of the corresponding MIP mode, because for each MIP mode, only this approximate matrix-vector product is calculated, and the "true" matrix-vector product is not calculated, to determine the prediction signal for the predetermined block. According to the first aspect of the present application, this difficulty is overcome by constraining the calculation of the prediction signal by the matrix-vector product. The inventors have found that it is advantageous to represent all entries of the prediction matrix associated with the MIP mode using a fixed-point representation having a predetermined bit depth, and to apply the same predetermined bit depth to all matrix-based intra prediction modes (e.g., at least for modes associated with the same block size, but optionally for prediction matrices of all block sizes). This enables an efficient implementation of the matrix-vector product, since if the entries of all prediction matrices have a common fixed predetermined bit depth, it is possible to use a specific multiplier adapted to the predetermined bit depth and to share this specific multiplier across all MIP modes for computing the matrix-vector product. Furthermore, if all entries of all prediction matrices can be stored in a fixed point representation (i.e. with a fixed precision), this enables efficient memory management (when handling the prediction matrices). Furthermore, the inventors have found it advantageous to compute, for each MIP mode, the matrix-vector product between the input vector and the prediction matrix associated with the respective MIP mode by performing a right shift for each component of the output vector by a number of bits that is equal for all MIP modes (e.g. at least for MIP modes associated with the same block size, but optionally for prediction matrices of all block sizes). This is based on the idea that a fixed right shift for all MIP modes enables efficient implementation of shifts in the matrix-vector product, because if the shift value does not depend on the MIP mode, a table lookup is saved and a single fixed shift operation can be implemented for the MIP, which benefits from a compact SIMD implementation of the matrix-vector product and reduces the case-dependent implementation of the matrix-vector product in hardware implementation.
[0017] Accordingly, according to a first aspect of the present application, a device for decoding a predetermined block of a picture using intra-frame prediction is configured to read a mode index from a data stream, and a device for encoding a predetermined block of a picture using intra-frame prediction is configured to insert the mode index into the data stream, for example, the device for encoding, i.e., the encoder, may have selected this mode from a mode list by rate-distortion optimization, and optionally, selected a further mode such as an inter-frame prediction mode. The mode index points to one of the matrix-based intra-frame prediction mode lists. In addition, the device (i.e., the device for decoding and / or the device for encoding) is configured to predict the samples of the predetermined block by calculating the matrix-vector product between an input vector derived from a reference sample in a neighborhood of the predetermined block and a prediction matrix associated with the matrix-based intra-frame prediction mode pointed to by the mode index, and by associating the components of the output vector obtained by the matrix-vector product to the sample positions of the predetermined block. For each matrix-based intra prediction mode, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode are represented by a fixed-point representation having a predetermined bit depth, wherein the predetermined bit depth is equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes associated with the same block size, but optionally for matrices of all block sizes). In addition, the device is configured to calculate, for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode by performing a right shift on each component of the output vector by a number of bits that is equal for the matrix-based intra prediction mode (e.g., at least for matrix-based intra prediction modes associated with the same block size, but optionally for matrices of all block sizes).
[0018] According to an embodiment, the number of matrix-based intra prediction modes in the matrix-based intra prediction mode list is 12, 16 or 32.
[0019] According to an embodiment, the device for decoding and / or the device for encoding is configured such that the matrix-based intra prediction mode in the matrix-based intra prediction mode list has 6, 8 or 16 different matrices associated therewith. For example, there may be different lists for mutually exclusive sets of block sizes, one list having 6 different matrices associated with 12 modes for block sizes within a first set of block sizes, one list having 8 different matrices associated with 16 modes for smaller block sizes within a second set of block sizes, and one list having 16 different matrices associated with 32 modes for even smaller block sizes within a third set of block sizes.
[0020] According to an embodiment, the device for decoding and / or the device for encoding is configured to calculate the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode in fixed-point arithmetic by applying the right shift to an intermediate result obtained by the matrix-vector product for each component of the output vector. For example, the intermediate result is obtained by the matrix-vector product between the input vector and the prediction matrix or by the matrix-vector product between the input vector and the prediction matrix, the intermediate result being offset by a positive integer, for example, an intermediate matrix resulting from the subtraction of a positive integer from each entry of the prediction matrix / the addition of a positive integer to each entry of the prediction matrix, and the intermediate result is obtained by the matrix-vector product between the input vector and the intermediate matrix.
[0021] According to an embodiment, the device for decoding and / or the device for encoding is configured to, before calculating the matrix-vector product, for example for each matrix-based intra prediction mode, offset all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode by adding or subtracting an offset value that is equal for the matrix-based intra prediction mode (e.g. at least for matrix-based intra prediction modes associated with the same block size but optionally for matrices of all block sizes). This is based on the idea that fixed offset values for all MIP modes enable efficient implementation of the matrix-vector product by saving table lookups.
[0022] According to an embodiment, the device for decoding and / or the device for encoding is configured to store, for each matrix-based intra prediction mode, for each entry of the prediction matrix associated with the respective matrix-based intra prediction mode, the fixed point representation at the predetermined bit depth.
[0023] According to an embodiment, the device for decoding / the device for encoding is configured to: decode / encode the picture with 10-bit resolution; store, for each matrix-based intra-frame prediction mode, the magnitude of the entry of the prediction matrix associated with the corresponding matrix-based intra-frame prediction mode with 7-bit precision, and use 6 bits as the number of bits for the right shift.
[0024] According to an embodiment, the decoding device and / or the encoding device is configured to store, for each matrix-based intra prediction mode, the entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode in an 8-bit labeled quantity representation. Alternatively, in the case where the entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode have the same label, the decoding device and / or the encoding device is configured to offset, for each matrix-based intra prediction mode, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode by an offset value that is equal to that of the matrix-based intra prediction mode, for example by addition or by subtraction, before calculating the matrix-vector product, wherein, for each matrix-based intra prediction mode, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode are representable by the labeled 8-bit representation. Accordingly, according to this alternative, since the indicator label is not required, only a 7-bit quantity can be stored for each matrix entry.
[0025] According to an embodiment, the device for decoding and / or the device for encoding is configured to calculate the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode in fixed-point arithmetic by applying the right shift to an intermediate result obtained by the matrix-vector product for each component of the output vector and represented with a bit precision as high as twice the bit precision at which the entries of the prediction matrix associated with the matrix-based intra prediction mode are stored. For example, the matrix-vector product thus calculated is represented with a bit precision as high as twice the bit precision at which the entries of the prediction matrix associated with the matrix-based intra prediction mode are stored, such as a prediction signal obtained by shifting the prediction matrix by a positive integer (obtaining an intermediate matrix), calculating the matrix-vector product between the input vector and the intermediate matrix (obtaining an intermediate result), and performing a right shift on the intermediate result.
[0026] According to an embodiment, the matrix-based intra-frame prediction mode list includes one or more matrix-based intra-frame prediction mode pairs. Note that the matrix-based intra-frame prediction mode list may not exclusively consist of such mode pairs, but may also include other modes that are exclusively applied using the transposed option or the non-transposed option. For each matrix-based intra-frame prediction mode pair, the prediction matrix associated with the first matrix-based intra-frame prediction mode in the corresponding matrix-based intra-frame prediction mode pair is equal to the prediction matrix associated with the second matrix-based intra-frame prediction mode in the corresponding matrix-based intra-frame prediction mode pair. The device (i.e., a device for decoding and / or a device for encoding) is configured such that, if the matrix-based intra-frame prediction mode pointed to by the mode index is the first matrix-based intra-frame prediction mode of the corresponding matrix-based intra-frame prediction mode pair (e.g., the mode with an odd mode index), the association of the reference samples in the neighborhood of the predetermined block with the components of the input vector and the association of the sample positions of the predetermined block with the components of the output vector are transposed relative to the association when the matrix-based intra-frame prediction mode pointed to by the mode index is the second matrix-based intra-frame prediction mode of the corresponding matrix-based intra-frame prediction mode pair (e.g., the mode with an even mode index). That is, if in the former case a component of the input vector is associated with position (x, y) (where (0, 0) represents the top left sample of the predetermined block), then in the latter case it is associated with (y, x). The same applies to the components of the output vector.
[0027] According to an embodiment, the device for decoding and / or the device for encoding is configured to use the matrix-based intra prediction mode list for multiple block sizes.
[0028] According to an embodiment, the device for decoding and / or the device for encoding is configured to predict samples of the predetermined block by upsampling and / or interpolating on the basis of the output vector or on the basis of the output vector and the reference samples in the neighborhood of the predetermined block, the samples being offset from the sample positions associated with the components of the output vector.
[0029] According to an embodiment, the device for decoding and / or the device for encoding is configured to derive said input vector from said reference samples in said neighborhood of said predetermined block by downsampling and / or pooling.
[0030] According to an embodiment, the reference samples in the neighborhood of the predetermined block include a first reference sample above the predetermined block and a second reference sample to the left of the predetermined block. The device (i.e., the device for decoding and / or the device for encoding) is configured to derive the input vector from the reference samples in the neighborhood of the predetermined block by: deriving a first intermediate component from the first reference sample by downsampling and / or pooling; deriving a second intermediate component from the second reference sample by downsampling and / or pooling; concatenating the first intermediate component and the second intermediate component to derive a preliminary input vector, and forming the input vector from the preliminary input vector.
[0031] According to an embodiment, the device for decoding / the device for encoding is configured to decode / encode the picture with a B-bit resolution. The device is configured to form the input vector from the preliminary input vector by: subtracting 2 from the first component of the preliminary input vector B-1 In order to obtain a first component of the input vector, the device is configured to subtract the first component of the preliminary input vector from a further component of the preliminary input vector to obtain the further component of the input vector, or to subtract the first component of the preliminary input vector from the further component of the preliminary input vector so that the input vector is formed from the further component. Furthermore, the device is configured to correct the output vector by component-by-component addition of the first component of the preliminary input vector.
[0032] According to an embodiment, the entries of the prediction matrix of the matrix-based intra-frame prediction mode in the matrix-based intra-frame prediction mode list correspond to the entries in Table 2 (shown below), but note that another shift value may be selected to list the values in the table, and the values in the table may be expressed in another scale.
[0033] According to an embodiment, a device for decoding and / or a device for encoding is configured to use a trained prediction matrix selected for a predetermined block to predict samples of the predetermined block by calculating a matrix-vector product between an input vector derived from reference samples in a neighborhood of the predetermined block and a trained prediction matrix associated with a matrix-based intra prediction mode selected for the predetermined block, and by associating components of an output vector obtained by the matrix-vector product with sample positions of the predetermined block. For example, the trained prediction matrix is trained by a device for training a prediction matrix, for example, by a device according to the second aspect.
[0034] According to a second aspect of the present invention, the inventors of the present application have recognized that one problem encountered when attempting to predict samples of a predetermined block of a picture using a matrix-based intra prediction mode (MIP mode) stems from the fact that the matrix-vector products (i.e., matrix-vector multiplications) performed in MIP modes need to be approximated by integer operations, and that the training prediction matrices for such matrix-vector multiplications are obtained in floating-point precision. According to the second aspect of the present application, this difficulty is overcome by already constraining the computation of the prediction signal when training the prediction matrix for such matrix-vector products. The inventors have found that it is advantageous to optimize the entries of the prediction matrix associated with the MIP mode using a cost function that depends on a prediction distortion measure, the prediction distortion measure being associated with setting the entries of the prediction matrix to intermediate values onto which the representative values are mapped using a differentiable function. This approach makes it possible to limit the range of all prediction matrix entries and avoid the possibility that some entries will not be updated during training. During training, the entries are represented in floating-point representation, and the intermediate values are then quantized to a fixed-point representation having a predetermined bit depth that is equal for all MIP modes. This enables efficient implementation of matrix-vector products (as if the entries of all prediction matrices have a common fixed predetermined bit depth). Such prediction matrices enable the video or picture encoder / decoder to use a specific multiplier suitable for that predetermined bit depth and to share that specific multiplier across all MIP modes for computing the matrix-vector product. Furthermore, if all entries of all prediction matrices can be stored in a fixed-point representation (i.e., with a fixed precision), this enables efficient memory management (when handling the prediction matrices).
[0035] Accordingly, according to a second aspect of the present application, there is provided a device for training a prediction matrix for a matrix-based intra prediction mode list by computing a matrix-vector product between an input vector derived from reference samples in a neighborhood of a predetermined block and one of the prediction matrices associated with a matrix-based intra prediction mode selected for the predetermined block, and associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block, wherein one of the matrix-based intra prediction modes should be selected for the predetermined block for predicting samples of the predetermined block. The device is configured to train the prediction matrix for the matrix-based intra prediction mode list by optimizing representative values represented in floating point representation of entries of the prediction matrix for the matrix-based intra prediction mode list using a gradient descent method (e.g., by using a training set of predetermined blocks of known (e.g., original) samples and their corresponding neighborhoods), using a cost function that depends on a prediction distortion measure associated with setting the entries of the prediction matrix to intermediate values, the representative values being mapped to the intermediate values using a differentiable function. The prediction distortion measure defines, for example, an increasing cost with decreasing prediction quality, such as that resulting from applying a differentiable function to the prediction matrix under training (meaning applying the differentiable function to each entry of the prediction matrix under training), wherein the domain and codomain of the differentiable function are defined by the floating point representation, the image of the differentiable function has a predetermined dynamic range, and the differentiable function is equal for the matrix-based intra prediction mode. Furthermore, the apparatus is configured to quantize the intermediate values to a fixed point representation, e.g. after training, such that for each matrix-based intra prediction mode, the prediction matrix associated with the corresponding matrix-based intra prediction mode has all entries represented by a fixed point representation having a predetermined bit depth, such that the predetermined bit depth is equal for the matrix-based intra prediction modes, and such that for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode is computable by performing a right shift for each component of the output vector by a number of bits that is equal for the matrix-based intra prediction mode.
[0036] According to an embodiment, the differentiable function (i.e., the clipping function) has a slope of 1 at the origin of the image, is strictly monotonically increasing, and has horizontal asymptotes at the upper and lower bounds of the image. The horizontal asymptotes at the upper and lower bounds of the image of the differentiable function may define a predetermined dynamic range.
[0037] According to an embodiment, the differentiable function is represented / defined by the following formula:
[0038]
[0039] Wherein α, β, γ and δ are real numbers depending on the predetermined dynamic range (ie, the clipping range), and λ is a non-negative integer.
[0040] According to an embodiment, the differentiable function (i.e., the clipping function) is parameterizable by a shift parameter (e.g., δ) in terms of the shift of the image within the common domain. The device is configured to subject the shift parameter to optimization using the gradient descent method (which may be, but not necessarily be), and to derive from the shift parameter an offset value that is equal for the matrix-based intra prediction modes, so as to be used for offsetting, for example by addition or by subtraction, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode for each matrix-based intra prediction mode before the calculation of the matrix-vector product.
[0041] An embodiment relates to a method for decoding a predetermined block of a picture using intra prediction, the method comprising: reading a mode index from a data stream, the mode index pointing to one of a list of matrix-based intra prediction modes; and predicting samples of the predetermined block by computing a matrix-vector product between an input vector derived from reference samples in a neighborhood of the predetermined block and a prediction matrix associated with the matrix-based intra prediction mode pointed to by the mode index, and by associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block. For each matrix-based intra prediction mode, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode are represented by a fixed-point representation having a predetermined bit depth, the predetermined bit depth being equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes associated with the same block size, but optionally for matrices of all block sizes). Furthermore, the method comprises calculating, for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode by performing a right shift on each component of the output vector by a number of bits that is equal for the matrix-based intra prediction mode (e.g. at least for matrix-based intra prediction modes associated with the same block size but optionally possibly for matrices of all block sizes).
[0042] An embodiment relates to a method for encoding a predetermined block of a picture using intra prediction, the method comprising: inserting a mode index into the data stream, the mode index pointing to one of a list of matrix-based intra prediction modes, for example, this mode may have been selected from the list of matrix-based intra prediction modes (and optionally also from further modes such as inter prediction modes) using rate-distortion optimization. The method further comprises predicting samples of the predetermined block by computing a matrix-vector product between an input vector derived from reference samples in a neighborhood of the predetermined block and a prediction matrix associated with the matrix-based intra prediction mode pointed to by the mode index, and by associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block. For each matrix-based intra prediction mode, all entries of the prediction matrix associated with the corresponding matrix-based intra prediction mode are represented by a fixed-point representation having a predetermined bit depth, the predetermined bit depth being equal for the matrix-based intra prediction modes (e.g., at least for matrix-based intra prediction modes associated with the same block size, but optionally for matrices of all block sizes). Furthermore, the method comprises calculating, for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode by performing a right shift on each component of the output vector by a number of bits that is equal for the matrix-based intra prediction mode (e.g. at least for matrix-based intra prediction modes associated with the same block size but optionally possibly for matrices of all block sizes).
[0043] The method as described above is based on the same considerations as the encoder / decoder described above. Incidentally, the method can be accomplished by all features and functionalities which were also described with respect to the encoder / decoder.
[0044] An embodiment relates to a method for training a prediction matrix of a matrix-based intra-frame prediction mode list by computing a matrix-vector product between an input vector derived from reference samples in a neighborhood of a predetermined block and one of the prediction matrices associated with the matrix-based intra-frame prediction mode selected for the predetermined block, and by associating components of an output vector obtained by the matrix-vector product to sample positions of the predetermined block, one of the matrix-based intra-frame prediction modes that should be selected for the predetermined block for predicting samples of the predetermined block. The method comprises training the prediction matrix for the matrix-based intra-prediction mode list by optimizing representative values of the entries of the prediction matrix for the matrix-based intra-prediction mode list represented in floating point representation using a gradient descent method (e.g., by using a training set of predetermined blocks of known (e.g., original) samples and their corresponding neighborhoods), mapping the representative values to the intermediate values using a differentiable function, the domain and codomain of the differentiable function being defined by the floating point representation, an image of the differentiable function having a predetermined dynamic range, and the differentiable function being equal for the matrix-based intra-prediction modes. Furthermore, the method comprises quantizing the intermediate values to a fixed point representation, e.g. after training, such that for each matrix-based intra prediction mode, the prediction matrix associated with the corresponding matrix-based intra prediction mode has all entries represented by a fixed point representation having a predetermined bit depth, such that the predetermined bit depth is equal for the matrix-based intra prediction modes, and such that for each matrix-based intra prediction mode, the matrix-vector product between the input vector and the prediction matrix associated with the corresponding matrix-based intra prediction mode is computable by performing a right shift for each component of the output vector by a number of bits that is equal for the matrix-based intra prediction mode.
[0045] The method as described above is based on the same considerations as the above-described device for training a prediction matrix. Incidentally, the method can be accomplished by all features and functionalities that were also described with respect to the device for training a prediction matrix.
[0046] Embodiments relate to a data stream having pictures or video encoded therein using the method for encoding described herein.
[0047] An embodiment relates to a computer program with a program code for performing the method described herein when the computer program runs on a computer. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings are not necessarily to scale; instead, emphasis is generally placed on illustrating the principles of the invention. In the following description, various embodiments of the invention are described with reference to the following drawings, in which:
[0049] Figure 1 An embodiment of encoding into a data stream is shown;
[0050] Figure 2 An embodiment of an encoder is shown;
[0051] Figure 3 An embodiment of picture reconstruction is shown;
[0052] Figure 4 An embodiment of a decoder is shown;
[0053] Figure 5.1 shows the prediction of a block with a reduced sample value vector according to an embodiment;
[0054] Figure 5.2 shows prediction of a block using sample interpolation according to an embodiment;
[0055] Figure 5.3 shows the prediction of a block with a reduced sample value vector according to an embodiment, where only some boundary samples are averaged;
[0056] Figure 5.4 shows a prediction of a block with a reduced sample value vector according to an embodiment, wherein groups of four boundary samples are averaged;
[0057] Figure 6.1 Matrix-based intra prediction of a predetermined block of a picture based on a mode index is shown;
[0058] Figure 6.2 shows the relationship between matrix-based intra prediction mode pairs and the application of inter-sample distance settings;
[0059] Figure 7 An apparatus for decoding using MIP mode for prediction according to an embodiment is shown;
[0060] FIG8 shows an apparatus for training a prediction matrix according to an embodiment; and
[0061] Figure 9 An exemplary differentiable function is shown. DETAILED DESCRIPTION
[0062] In the following description, equal or equivalent elements or elements having the same or equivalent functionality are denoted by equal or equivalent reference numerals even though they are shown in different drawings.
[0063] In the following description, a number of details are set forth to provide a more comprehensive explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be implemented without these specific details. In other instances, in order to avoid obscuring embodiments of the present invention, well-known structures and devices are shown in block diagram form rather than in detail. In addition, unless otherwise specifically noted, the features of the different embodiments described hereinafter may be combined with each other.
[0064] In the following, various examples are described that can help achieve more efficient compression when using matrix-based intra prediction.For example, matrix-based intra prediction can be added to other heuristically designed intra prediction modes, or can be provided specifically.
[0065] In order to facilitate understanding of the following examples, the description begins with the presentation of possible encoders and decoders suitable therefor, into which the above-outlined examples of the present application may be built. Figure 1 An apparatus is shown for encoding a picture 10 block by block into a data stream 12. The apparatus is indicated using reference numeral 14 and can be a still picture encoder or a video encoder. In other words, when the encoder 14 is configured to encode a video 16 including the picture 10 into the data stream 12, the picture 10 can be the current picture in the video 16, or the encoder 14 can specifically encode the picture 10 into the data stream 12.
[0066] As mentioned, the encoder 14 performs encoding in a block-by-block manner or on a block basis. To this end, the encoder 14 subdivides the picture 10 into blocks, the units of which encode the picture 10 into a data stream 12. Examples of possible subdivisions of the picture 10 into blocks 18 are explained in more detail below. In general, the subdivision can end up with blocks 18 of constant size, such as an array of blocks arranged in rows and columns, or with blocks 18 of different block sizes, such as by using a hierarchical multi-tree subdivision, where the multi-tree subdivision starts from the entire picture area of the picture 10, or starts from a pre-division of the picture 10 into an array of tree blocks, wherein these examples should not be regarded as excluding other possible ways of subdividing the picture 10 into blocks 18.
[0067] Furthermore, the encoder 14 is a predictive encoder configured to predictively encode the picture 10 into the data stream 12. For a certain block 18, this means that the encoder 14 determines a prediction signal for the block 18 and encodes a prediction residual (i.e., a prediction error by which the prediction signal deviates from the actual picture content within the block 18) into the data stream 12.
[0068] The encoder 14 can support different prediction modes to derive a prediction signal for a block 18. The prediction mode of interest in the following example is the intra-frame prediction mode, according to which the interior of the block 18 is spatially predicted from neighboring, already coded samples of the picture 10. The encoding of the picture 10 into the data stream 12, and accordingly, the corresponding decoding process, can be based on a coding order 20 defined between the blocks 18. For example, the coding order 20 can be in a raster scan order, such as traversing the blocks 18 row by row from top to bottom, for example, traversing each row from left to right. In the case of a hierarchical multi-tree based subdivision, a raster scan order can be applied within each hierarchical level, wherein a depth-first traversal order can be applied, i.e., according to the coding order 20, leaf notes within a block at a certain hierarchical level can precede blocks at the same hierarchical level with the same parent block. Depending on the coding order 20, the neighboring, already coded samples of the block 18 can generally be located on one or more sides of the block 18. In the case of the examples presented herein, for example, the adjacent, already encoded samples of block 18 are located at the top and to the left of block 18 .
[0069] Intra-prediction mode may not be the only mode supported by encoder 14. For example, in the case where encoder 14 is a video encoder, encoder 14 may also support intra-prediction mode according to which block 18 is temporarily predicted from a previously encoded picture of video 16. Such intra-prediction mode may be a motion-compensated prediction mode according to which a motion vector for such block 18 is signaled, the motion vector indicating the relative spatial offset of the portion from which the prediction signal for block 18 is to be derived as a copy. Additionally or alternatively, other non-intra-prediction modes may also be available, such as inter-view prediction mode in the case where encoder 14 is a multi-view encoder, or a non-predictive mode according to which the interior of block 18 is encoded as is (i.e., without any prediction).
[0070] Before starting by focusing the description of this application on intra prediction modes, a more specific example of a possible block-based encoder (i.e., for Figure 2 The possible implementations of the described encoder 14) are respectively presented as suitable for Figure 1 and 2 Two corresponding examples of decoders for .
[0071] Figure 2 Shown Figure 1 A possible implementation of the encoder 14 is one in which the encoder is configured to encode the prediction residual using transform coding, although this is merely an example and the application is not limited to this kind of prediction residual coding. Figure 2, the encoder 14 includes a subtractor 22 configured to subtract the corresponding prediction signal 24 from the incoming signal (i.e., picture 10) or from the current block 18 on a block basis to obtain a prediction residual signal 26, which is then encoded into the data stream 12 by a prediction residual encoder 28. The prediction residual encoder 28 consists of a lossy encoding stage 28a and a lossless encoding stage 28b. The lossy stage 28a receives the prediction residual signal 26 and includes a quantizer 30 that quantizes the samples of the prediction residual signal 26. As mentioned above, this example uses transform coding of the prediction residual signal 26, and accordingly, the lossy encoding stage 28a includes a transform stage 32 connected between the subtractor 22 and the quantizer 30 to transform such spectrally decomposed prediction residual 26, wherein the quantization of the quantizer 30 occurs with respect to the transform coefficients (in which the residual signal 26 is present). The transform can be a DCT, a DST, an FFT, a Hadamard transform, or the like. The transformed and quantized prediction residual signal 34 then undergoes lossless encoding by a lossless encoding stage 28b, which is an entropy encoder that entropy encodes the quantized prediction residual signal 34 into the data stream 12. The encoder 14 further comprises a prediction residual signal reconstruction stage 36 connected to the output of the quantizer 30 in order to reconstruct the prediction residual signal from the transformed and quantized prediction residual signal 34 in a manner that is also usable at the decoder (i.e., taking into account the coding losses in the quantizer 30). To this end, the prediction residual reconstruction stage 36 comprises a dequantizer 38 that performs the inverse of the quantization of the quantizer 30, followed by an inverse transformer 40 that performs the inverse transform relative to the transform performed by the transformer 32 (e.g., the inverse of the spectral decomposition, such as the inverse of any of the specific transform examples mentioned above). The encoder 14 comprises an adder 42 that adds the reconstructed prediction residual signal as output by the inverse transformer 40 to the prediction signal 24 in order to output a reconstructed signal, i.e., reconstructed samples. This output is fed into the predictor 44 of the encoder 14, which then determines the prediction signal 24 based thereon. It is the predictor 44 that supports the above-mentioned Figure 1 All forecast models discussed. Figure 2 It is also shown that in case the encoder 14 is a video encoder, the encoder 14 may also comprise an in-loop filter 46 that filters the fully reconstructed picture which, after having been filtered, forms a reference picture for the predictor 44 with respect to the inter-prediction blocks.
[0072] As mentioned above, encoder 14 operates on a block-based basis. For the following description, the block-based operation of interest is one in which picture 10 is subdivided into blocks, for which an intra-prediction mode is selected from a set or multiple intra-prediction modes supported by predictor 44 or encoder 14, respectively, and the selected intra-prediction mode is individually performed. However, other types of blocks into which picture 10 is subdivided are also possible. For example, the aforementioned decision as to whether picture 10 is inter-coded or intra-coded can be made at a granular level, or in units of blocks offset from block 18. For example, the inter / intra mode decision can be made at the level of the coding blocks into which picture 10 is subdivided, with each coding block being subdivided into prediction blocks. Prediction blocks for each coding block for which intra-prediction has been decided are each subdivided into an intra-prediction mode decision. To this end, for each of these prediction blocks, a decision is made as to which supported intra-prediction mode should be used for the corresponding prediction block. These prediction blocks will form block 18 of interest here. Prediction blocks within coding blocks associated with inter-prediction are processed differently by predictor 44. They are inter-predicted from a reference picture by determining a motion vector and copying the prediction signal of this block from the position in the reference picture pointed to by the motion vector. Another block subdivision involves subdivision into transform blocks, the transform by the transformer 32 and the inverse transformer 40 being performed in units of said transform blocks. The transform blocks can, for example, be the result of a further subdivision of the coding blocks. Naturally, the examples set forth herein should not be considered limiting, and other examples exist. For the sake of completeness only, it is noted that, for example, the subdivision into coding blocks can use a multi-tree subdivision, and that prediction blocks and / or transform blocks can also be obtained by subdividing the coding blocks in one step using a multi-tree subdivision.
[0073] Figure 3 Depicted in Figure 1 The decoder 54 is a block-by-block decoding device or decoder 54 of the encoder 14. This decoder 54 performs the reverse operation of the encoder 14, i.e. it decodes the picture 10 from the data stream 12 in a block-by-block manner and supports multiple intra prediction modes for this purpose. The decoder 54 may include, for example, a residual provider 156. Figure 1All other possibilities discussed are also valid for decoder 54. To this end, decoder 54 can be a still picture decoder or a video decoder, and all prediction modes and prediction possibilities are also supported by decoder 54. The difference between encoder 14 and decoder 54 lies primarily in the fact that encoder 14 chooses or selects coding decisions based on some optimization, such as, for example, minimizing some cost function that may depend on coding rate and / or coding distortion. One of these coding options or coding parameters may involve selecting an intra-prediction mode to be used for current block 18 among available or supported intra-prediction modes. The selected intra-prediction mode may then be signaled by encoder 14 within data stream 12 for current block 18, with decoder 54 using this signaling in data stream 12 to redo the selection for block 18. Similarly, the subdivision of picture 10 into blocks 18 may be optimized within encoder 14, and corresponding subdivision information may be communicated within data stream 12, with decoder 54 resuming the subdivision of picture 10 into blocks 18 based on the subdivision information. In summary, the decoder 54 may be a predictive decoder that operates on a block basis, and in addition to the intra prediction mode, the decoder 54 may support other prediction modes (e.g., inter prediction mode), for example, if the decoder 54 is a video decoder. In decoding, the decoder 54 may also use information about Figure 1 The coding order 20 discussed is followed, and since this coding order 20 is followed at both the encoder 14 and the decoder 54, the same adjacent samples are available for the current block 18 at both the encoder 14 and the decoder 54. Accordingly, in order to avoid unnecessary repetition, the description of the mode of operation of the encoder 14 should also apply to the decoder 54 as far as the subdivision of the picture 10 into blocks is concerned, for example, as far as prediction is concerned and as far as the coding of the prediction residual is concerned. The difference lies in the fact that the encoder 14 selects some coding options or coding parameters and signals within the data stream 12 by optimization, or inserts the coding parameters into the data stream 12, which are then derived from the data stream 12 by the decoder 54 in order to re-perform the prediction, subdivision, etc.
[0074] Figure 4 Shown Figure 3 A possible implementation of the decoder 54 is, that is, suitable for Figure 2 Shown in Figure 1 One implementation of the encoder 14. Figure 4 Many components of the encoder 54 are related to Figure 2 The same elements appear in the corresponding encoder, so Figure 4 The same reference numerals provided with a prime are used in order to indicate these elements. In particular, the adder 42', the optional in-loop filter 46' and the predictor 44' are shown in the same manner as they are in FIG. Figure 2is connected to the prediction loop in the same way as in the encoder of . The reconstructed (i.e. dequantized and retransformed) prediction residual signal applied to the adder 42' is derived by an entropy decoder 56 for entropy coding of the inverse entropy encoder 28b, followed by a sequence of residual signal reconstruction stages 36' consisting of a dequantizer 38' and an inverse transformer 40', just as in the case on the encoding side. The output of the decoder is a reconstruction of the picture 10. The reconstruction of the picture 10 may be available directly at the output of the adder 42' or, alternatively, at the output of the in-loop filter 46'. A certain postfilter may be arranged at the output of the decoder in order to subject the reconstruction of the picture 10 to a certain postfiltering in order to improve the picture quality, but in Figure 4 This option is not depicted in .
[0075] Again, about Figure 4 , the above about Figure 2 The description presented is for Figure 4 It should also be efficient, except that only the encoder performs the optimization tasks and the associated decisions about coding options. However, all descriptions of block subdivision, prediction, dequantization and retransformation are for Figure 4 The decoder 54 is also effective.
[0076] The embodiments described herein use so-called matrix-based intra prediction.The general concept is outlined below.
[0077] A Review of Matrix-Based Intra Prediction
[0078] In order to keep this application independent, in this section we describe the main steps of the current matrix-based intra prediction (MIP) method included in Working Draft 7 of General Video Coding [1]. For more details, reference is made to [1].
[0079] Matrix-based intra prediction (MIP) is a method for generating intra prediction signals on rectangular blocks of width W and height H. The input to the MIP prediction process is the reconstructed samples r from one row above the block. top and the reconstructed sample r of the left column of the block left The reconstructed sample r, the MIP mode index i, and the information of whether the MIP mode is to be transposed. Then, the MIP prediction signal is generated using the following three steps:
[0080] 1. For depends on W and H, and satisfies w in,red ≤W and h in,red ≤H specified natural number w in,red and h in,red , r top One of them is generated into size w by downsampling / averaging in,red The reduced top input r top,red , and rleft One of them is generated into size h by downsampling / averaging in,red is the reduced left input r left,red .
[0081] Then, r top,red and r left,red Connect to the reduced input r red,full , if the MIP mode is not to be transposed, then r red,full is defined as
[0082] r red,full =[r top,red , r left,red ],
[0083] If the MIP mode is to be transposed it is defined as
[0084] r red,full =[r left,red , r top,red ],
[0085] Next, r red,full One of the definitions reduces the input r red Here, r red With r red,full Same size w in,red +h in,red or with size w in,red +h in,red -1. In the first case, r red Defined as
[0086] r red [0] = r red,full [0]-2 B-1 ,
[0087] where B is the bit depth and is defined as
[0088] r red [i]=r red,full [i]-r red,full [0],i>0.
[0089] In the second case, r red Defined as
[0090] r red [i]=r red,full [i+1]-r red,full [0].
[0091] 2. For depends on W and H, and satisfies w out,red ≤W and h out,red ≤H specified natural number w out,red and hout,red , with width w out,red and height h out,red The prediction signal pred will be reduced on the block red Generated as
[0092] pred red =((A i -c i )·r red +(1<<(d i -1)))>>d i .
[0093] Here A i is a matrix that depends on W and H and on the MIP mode index i and c i and d i is a non-negative integer that depends on the MIP mode index i, wherein this dependency is to be removed by the present invention. In addition, if the MIP mode does not need to be transposed, then pred red It is w out,red ·h out,red - A row-major order of width w out,red and height h out,red A vector of sizes identifying the blocks of signals and in column-major order if the MIP mode needs to be transposed.
[0094] Then, r red,full [0] Add to pred red .
[0095] Finally, the result is clipped to the given bit range [0, 2 B ).
[0096] 3. If w out,red <W or h out,red < H, then upsampling / linear interpolation is applied to generate the full MIP prediction signal from the reduced prediction signal obtained at the end of the previous steps. Here, the reconstructed samples are included in the linear interpolation.
[0097] Implementation example presentation
[0098] The general concept is outlined above. The concept is sometimes referred to below as ALWIP (Affine Linearly Weighted Intra Prediction) (as an alternative synonym for MIP (Matrix-based Intra Prediction)) in order to explain the use of these modes again in more detail.
[0099] In the subsequent Figures 5.1-5.4 In FIG, the whole process of filling the input vector, computing the matrix-vector multiplication and linear interpolation on a neighborhood basis is shown for different block shapes. Note that the remaining shapes are considered as one of the depicted cases.
[0100] 1. Given a 4×4 block, ALWIP (or MIP) can take two averages along each axis of the boundary, see Figure 5.1 . As an alternative to averaging, take every other sample of the neighborhood, or more generally and accurately, take each component of the input vector of the matrix-vector-multiplication 19 from exactly one sample in the neighborhood. The resulting four input samples go into the matrix-vector multiplication. Take a matrix from the set S0 matrices, which is the set of matrices for the nearby block sizes. After adding the offset, this can produce 16 final prediction samples. Linear interpolation is not necessary for generating the prediction signal. Therefore, a total of (4*16) / (4*4)=4 multiplications per sample are performed. For example, see ALWIP showing a 4×4 block Figure 5.1 The exact calculation has been explained above.
[0101] 2. Given an 8×8 block, ALWIP can take four averages along each axis of the boundary, see Figure 5.2 The resulting eight input samples enter the matrix-vector multiplication 19. The matrix is taken from set S1. This produces 16 samples at odd positions in the prediction block. Therefore, a total of (8*16) / (8*8)=2 multiplications per sample are performed. After adding the offset, these samples are interpolated vertically by using the reduced top boundary. The original left boundary is used after horizontal interpolation. For example, see the ALWIP diagram showing an 8×8 block. Figure 5.2 .
[0102] 3. Given an 8×4 block, ALWIP can take four average values along the horizontal axis of the boundary and four original boundary values on the left boundary, see Figure 5.3 The resulting eight input samples enter the matrix vector multiplication. The matrix is taken from set S1. This produces 16 samples at odd horizontal and vertical positions of the prediction block. Therefore, a total of (8*16) / (8*4)=4 multiplications per sample are performed. After adding the offset, these samples are horizontally interpolated using the original left boundary. For example, see the ALWIP diagram showing an 8×4 block. Figure 5.3 .
[0103] The transposed case is handled accordingly.
[0104] 4. Given a 16x16 block, ALWIP can take four averages along each axis of the boundary. The resulting eight input samples go into a matrix-vector multiplication. The matrix is taken from set S2. This produces 64 samples at odd positions in the prediction block. Therefore, a total of (8*64) / (16*16)=2 multiplications per sample are performed. After adding the offset, these samples are interpolated vertically by using the eight averages of the top boundary. The original left boundary is then used after horizontal interpolation. For example, see the example showing ALWIP for a 16x16 block. Figure 5.4 .
[0105] For larger shapes, the process can be essentially the same, and it is easy to check that the number of multiplications per sample is less than two.
[0106] For Wx8 blocks, only horizontal interpolation is necessary since samples are given at odd horizontal and every vertical position.Thus, at most (8*64) / (16*8)=4 multiplications per sample are performed in these cases.
[0107] Finally, for a W×4 block (where W>8), let A k is a matrix produced by omitting every row corresponding to an odd-numbered entry along the horizontal axis of the downsampled block. Thus, the output size can be 32, and again, only horizontal interpolation remains to be performed. At most (8*32) / (16*4)=4 multiplications per sample can be performed.
[0108] The transposed case can be handled accordingly. This is illustrated in the subsequent figures.
[0109] Figure 6.1 A device 54 for decoding a predetermined block 18 of a picture using intra prediction is shown.
[0110] The device 54 is configured to read a mode index 200 from the data stream 12 using a binarized code 202, the mode index pointing to one of a matrix-based intra-prediction mode list 204. The matrix-based intra-prediction mode list 204 consists of an even number of matrix-based intra-prediction modes, wherein the matrix-based intra-prediction modes of the list 204 are grouped into matrix-based intra-prediction mode pairs 212. Each pair 212 consists of a first matrix-based intra-prediction mode and a second matrix-based intra-prediction mode. The device 54 is configured to read the mode index 200 from the data stream 12 using the binarized code 202 such that, for each matrix-based intra-prediction mode pair 212, the first matrix-based intra-prediction mode is assigned a first codeword and the second matrix-based intra-prediction mode is assigned a second codeword, and the lengths of the two codewords are equal.
[0111] Optionally, the binarized code 202 is a variable length code comprising codewords of different lengths. Alternatively, the binarized code may be a truncated binary code, and the number of matrix-based intra-frame prediction modes is not a power of two, such that the truncated binary code has codewords of different lengths. The matrix-based intra-frame prediction mode associated with the first matrix-based intra-frame prediction mode pair 212 may be assigned a codeword having a length different from the codeword assigned to the matrix-based intra-frame prediction mode associated with the second matrix-based intra-frame prediction mode pair 212. However, the two codewords of the matrix-based intra-frame prediction mode pair 212 are equal in length.
[0112] According to an embodiment, the device 54 may be configured to read the pattern index 200 from the data stream 12 using an equiprobable bypass mode of a context-adaptive binary arithmetic decoder.
[0113] Similar to the device 54 (i.e., a decoder) for decoding a predetermined block 18 of a picture using intra-frame prediction, the device for encoding a predetermined block 18 of a picture using intra-frame prediction (i.e., an encoder) can be configured to encode a mode index 200 into the data stream 12 using a binarization code 202 and, optionally, using an equiprobable bypass mode of a context-adaptive binary operation encoder.
[0114] The decoder and the encoder are configured to predict samples 108 of the predetermined block 18 by computing a matrix-vector product 206 between an input vector 102 derived from reference samples 17 in a neighborhood of the predetermined block 18 and a prediction matrix 19 associated with the matrix-based intra prediction mode k pointed to by the mode index 200. The computation of the matrix-vector product 206 results in an output vector 208. Furthermore, the samples 108 of the predetermined block 18 are predicted by associating components 210 of the output vector 208 obtained by the matrix-vector product 206 to the sample positions 104 of the predetermined block 18. This prediction of the samples 108 of the predetermined block 18 may be as described with respect to Figures 5.1 to 5.4 is performed as described.
[0115] For each matrix-based intra prediction mode pair 212, the prediction matrix 19 associated with the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212 is equal to the prediction matrix 19 associated with the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212. Therefore, the same prediction matrix 19 is used for matrix-based intra prediction modes 2k and 2k+1. For each matrix-based intra-frame prediction mode pair 212, the encoder and the decoder are configured such that, if the matrix-based intra-frame prediction mode pointed to by the mode index 200 is the first matrix-based intra-frame prediction mode in the corresponding matrix-based intra-frame prediction mode pair 212 (e.g., the mode with the odd mode index 2k+1), the association of the reference sample 17 in the neighborhood of the predetermined block 18 with the component 214 of the input vector 112 and the association of the sample position 104 of the predetermined block 18 with the component 210 of the output vector 208 are transposed relative to the associations when the matrix-based intra-frame prediction mode pointed to by the mode index 200 is the second matrix-based intra-frame prediction mode in the corresponding matrix-based intra-frame prediction mode pair 212 (e.g., the mode with the even mode index 2k).
[0116] The decoder / encoder may be configured to determine, based on the parity of the mode index 200, whether the matrix-based intra prediction mode pointed to by the mode index 200 is the first matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair or the second matrix-based intra prediction mode in the corresponding matrix-based intra prediction mode pair 212. The parity of the mode index 200 may indicate whether the input vector 102 and the output vector 208 are used in a transposed manner for the prediction of the samples 108 of the predetermined block 18. That is, as Figure 6.2 As shown in FIG, if in the former case a certain component of the components 1 to n of the input vector 102 is associated with the position (x, y) (where (0, 0) represents the top left corner sample AA of the predetermined block 18), then in the latter case it is associated with (y, x). The same applies to the components (AA, AB, AC, BA, CA, ...) of the output vector 208.
[0117] Each pair 212 consists of a first matrix-based intra prediction mode and a second matrix-based intra prediction mode, the modes being related to each other by the same prediction matrix 19 and differing from each other only in whether the input vector 102 and the output vector 208 are transposed or not transposed. This is advantageous because only a mode index 200 is required in the data stream 12 to indicate the matrix-based intra prediction mode and whether the matrix-based intra prediction mode is to be used in a transposed manner. No additional index or flag is required to indicate for the matrix-vector product 206 that the input vector 102 and the output vector 208 are to be used in a transposed manner.
[0118] According to an embodiment, the decoder / encoder is configured to index a prediction matrix 19 from a plurality of prediction matrices using the integer part of the mode index 200 divided by 2. This is based on the idea that two matrix-based intra prediction modes of a pair 212 use the same prediction matrix 19 to predict the samples 108 of the predetermined block 18, for which reason the prediction matrix 19 is already sufficiently indicated by pointing to the relevant pair 212 in the list 204 with the mode index 200.
[0119] like Figure 6.1 and 6.2As shown in FIG. 2 , the decoder / encoder can be configured to horizontally set 217 the inter-sample distances 216 of the sample positions 104 of the predetermined block 18 and the inter-sample distances 218 of the reference samples 17 in the neighborhood of the predetermined block 18 according to a first ratio of the horizontal size 220 of the predetermined block 18 to the horizontal default size, and / or vertically set 217 according to a second ratio of the vertical size 222 of the predetermined block 18 to the vertical default size. This enables the use of the matrix-based intra prediction mode list 204 for multiple block sizes. The device can fill the spaces between the prediction samples by interpolation. The inter-sample distance setting 217 of the inter-sample distances 216 of the sample positions 104 of the predetermined block 18 and the inter-sample distances 218 of the reference samples 17 in the neighborhood of the predetermined block 18 enables improved distribution of the prediction samples 108 in the predetermined block 18 and the reference samples 17 in the neighborhood of the predetermined block 18. As a result, the prediction samples can be equally distributed, enabling improved interpolation of the samples of the predetermined block 18.
[0120] According to an embodiment, the decoder / encoder is configured to sort the matrix-based intra prediction modes in the matrix-based intra prediction mode list 204 equally for multiple block sizes. Alternatively, the order may be suitable for blocks that are wider than they are tall, or vice versa, i.e., taller than they are wide, or quadratic blocks. This ordering may increase coding efficiency and reduce bitstream, because matrix-based intra prediction modes for common block sizes may be associated with short codewords, and matrix-based intra prediction modes for rare block sizes may be associated with longer codewords.
[0121] Optionally, the plurality of block sizes includes at least one block size corresponding to an aspect ratio greater than 4. The matrix-based intra prediction may be optimized such that the predetermined block 18 has an aspect ratio of the horizontal dimension 220 to the vertical dimension 222 greater than 4. That is, the plurality of block sizes includes a predetermined block having a horizontal dimension 220 that is at least four times greater than the vertical dimension 222 and / or a predetermined block having a vertical dimension 222 that is at least four times greater than the horizontal dimension 220. Figure 6.2 Predetermined blocks 18 having a block size corresponding to an aspect ratio greater than 4 may be shown.
[0122] According to the embodiments presented below, MIP mode is applied in a way that makes the use of MIP even more efficient than and what was hitherto expected in current VVC versions.
[0123] The embodiments below will primarily illustrate features and functionality with respect to the decoder. However, it is clear that the encoder may include the same or similar features and functionality, e.g., the decoding performed by the decoder may correspond to the encoding performed by the encoder. Furthermore, the encoder may include the same features as described with respect to the decoder in the feedback loop (e.g., in prediction stage 36).
[0124] Figure 7 A device 54 for decoding a predetermined block 18 of a picture using intra prediction is shown.
[0125] The device 54 is configured to read the mode index 200 from the data stream 12. The mode index 200 points to the matrix-based intra prediction mode 2051-205 n One of the MIP modes in list 204. Matrix-based intra prediction modes 2051-205 in list 204 n The number n is, for example, 12, 16 or 32. Furthermore, the embodiment described focuses on using MIP modes 2051-205 n For intra prediction, the mode index 200 may also be used to indicate that further modes are clear, like further intra prediction modes and / or inter prediction modes. The mode index 200 may be inserted into the data stream 12 by the device 14 for encoding a predetermined block 18 of a picture using intra prediction.
[0126] According to an embodiment, the matrix-based intra prediction modes 2051-205 in the matrix-based intra prediction mode list 204 are n There are 6, 8 or 16 different prediction matrices 19 associated with it.
[0127] According to an embodiment, the matrix-based intra prediction mode list 204 includes MIP modes for multiple block sizes.
[0128] According to an embodiment, there may be two or more MIP mode lists 204, wherein the two or more MIP mode lists 204 differ from each other in the block size of the predetermined block 18 associated with the MIP mode. MIP modes associated with the same or similar block sizes are included in the same list of the two or more MIP mode lists 204. For example, there may be different lists for mutually exclusive block size sets, one with 6 different matrices associated with 12 patterns of block sizes within a first block size set, one with 8 different matrices associated with 16 patterns of smaller block sizes within a second block size set, and one with 16 different matrices associated with 32 patterns of even smaller block sizes within a third block size set. This is merely an example, and it is clear that a different number of lists 204 is also possible, and each list may include MIP modes associated with a block size set different from the block size set described above.
[0129] The device 54 is configured to predict samples 108 of the predetermined block 18 by calculating a matrix-vector product 206 between an input vector 102 derived from a reference sample 17 in a neighborhood of the predetermined block 18 and a prediction matrix 19 associated with a matrix-based intra-frame prediction mode 205 pointed to by a mode index 200, and by associating components 210 of an output vector 208 obtained by the matrix-vector product 206 to sample positions 104 of the predetermined block 18.
[0130] For each matrix-based intra prediction mode 2051-205 n , all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 are represented by a fixed point representation 190 having a predetermined bit depth 192. The fixed point representation 190 may have a fractional part 194 (having n bits), an integer part 196 (having m bits), and an optional flag bit 198. Figure 7 As shown in FIG, the predetermined bit depth 192 is used for all matrix-based intra prediction modes 2051-205 in the MIP mode list 204. n are equal, compare the following constraint 1. In the case where there are two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, for example, the predetermined bit depth 192 is the same as that of all matrix-based intra prediction modes 2051-205 in the corresponding MIP mode list. n In other words, for example, the predetermined bit depth 192 is at least the same for MIP modes associated with the same set of block sizes.
[0131] According to an embodiment, the device 54 is configured to generate, for each matrix-based intra prediction mode 2051-205 n , a fixed point representation at a predetermined bit depth is stored for each entry in the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205. This is shown, for example, in the following table examples, see Listings 1 to 4.
[0132] For each matrix-based intra prediction mode 2051-205 n The device 54 is configured to generate a matrix-based intra prediction mode 2051-205 for each component 210 of the output vector 204. n A right shift 209 is performed by equal number of bits 211 to compute the matrix vector product 206 between the input vector 102 and the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205, as shown in FIG. Figure 7The number of bits 211 for the right shift is, for example, x bits. In the case where there are two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, the number of bits 211 is, for example, x bits for all matrix-based intra prediction modes 2051-205 in the corresponding MIP mode list. n are equal. In other words, the number of bits 211, for example, is at least the same for MIP modes associated with the same set of block sizes. The right shift 209, for example, is indicated by >> in equation (2) described above, and for the number of bits 211, the comparison d i , where the device 54 applies the number of bits 211 for the matrix-based intra prediction modes 2051-205 n For equality constraints, compare constraint 2 below.
[0133] For the matrix A in equation (2) i (i.e., prediction matrix 19) and parameter c i and d i The following constraints are desirable.
[0134] 1. The range of the matrix entries for the MIP is fixed, for example, the entries of the prediction matrix 19 are represented by a fixed point representation 190 with a predetermined bit depth 192. Therefore, there are predefined non-negative integers μ 1,low 、μ 1,up and μ 2,low 、μ 2,up , so that for each MIP mode i2051-205 n And for the matrix A i Each matrix item a of the prediction matrix 19 k,l ,have
[0135]
[0136] as well as
[0137]
[0138] Two specific examples of such constraints are given below.
[0139] The first example is that for a fixed positive integer μ, for all matrices A used in MIP prediction i All matrix entries a k,l ,have
[0140] 0≤a k,l ≤2 μ -1
[0141] as well as
[0142] -2μ ≤a k,l -c i ≤2 μ -1.
[0143] According to a first example, the entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 have the same sign, for example positive, and the device 54 is configured to calculate the matrix-vector product 206 for each matrix-based intra prediction mode 2051-205 n , offsetting all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 by 1-205 for the matrix-based intra prediction mode 205. n Equal offset value c i , where for each matrix-based intra prediction mode 2051-205 n , all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 are representable by an 8-bit representation of the flag. Accordingly, according to this alternative, only a 7-bit magnitude can be stored for each matrix entry. This is due to the fact that the flag bit does not have to be stored.
[0144] The second example is for each MIP mode i2051-205 n , with c i = 0, and there is a fixed positive integer v such that for each MIP mode i2051-205 n and the matrix A i Each matrix entry a k,l ,have
[0145] -2 v ≤a k,l ≤2 v -1.
[0146] According to a second example, the device 54 is configured to, for each matrix-based intra prediction mode 2051-205 n The entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 are stored in 8-bit flag values. In this second example, the device 54 does not offset the entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 by the offset value c before calculating the matrix-vector product 206. i .
[0147] 2. Shift value d i , i.e. the number of bits 211 for the right shift 209, independent of the MIP mode i 2051-205 n Therefore, there exists a positive integer d such that for each MIP mode i2051-205 n ,have
[0148] d i =d.
[0149] Optionally, the following constraints may also be desirable:
[0150] 3. Value c i , i.e., the offset to the entries of the prediction matrix 19, independent of the MIP modes i2051-205 n Therefore, there exists a positive integer c such that for each MIP mode i2051-205 n ,have
[0151] c i =c.
[0152] Thus, the device 54 may be configured to calculate the matrix vector product 206 for each matrix-based intra prediction mode 2051-205 n , for example by offsetting all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 by 1-205 for all matrix-based intra prediction modes 2051-205 of the MIP mode list 204 by addition or by subtraction. n For equal offset values, compare c.
[0153] In case of two or more MIP mode lists 204, for each of the two or more MIP mode lists 204, the offset value c is, for example, for all matrix-based intra prediction modes 2051-205 in the corresponding MIP mode list. n In other words, the offset value c is at least the same for MIP modes associated with the same set of block sizes, for example.
[0154] The reasons for imposing these constraints are as follows. Constraint 1 enables efficient implementation of the matrix-vector multiplication 206 (A i -c i )·r red , because if all matrices (A i -c i ) have a common fixed bit depth (ie, a predetermined bit depth of 192), then a specific multiplier appropriate for that bit depth may be used and applied across all MIP modes 2051-205 n are shared for computing the matrix-vector product 206. In addition, if all matrices A i All entries of can be stored with fixed precision, that is, the entries of the prediction matrix 19 are represented in fixed point representation 190, which enables efficient memory management (in the handling matrix A i Here, the important example is that all matrices A iAll entries of can be stored with 8-bit precision (ie in one byte). In other words, the entries of the prediction matrix 19 can be represented in a fixed point representation 190 (with a predetermined bit depth 192 of 8 bits).
[0155] Constraint 2 enables efficient implementation of the shift in expression (2) because if the shift value, i.e. the number of bits 211 for the right shift, does not depend on the MIP mode i, a table lookup is saved and a single fixed shift operation can be implemented for the MIP, which is beneficial for a compact SIMD implementation of equation (2) and reduces the case-dependent implementation of equation (2) in hardware implementation, i.e. the case-dependent implementation of the matrix-vector product 206 approximating the matrix-vector product with the prediction matrix in floating point precision. A particularly important example here is that the value 6 is used as the fixed shift, i.e. as the number of bits 211. The reason is that for 10-bit content, the clipping to a 10-bit range is applied to pred during the MIP prediction process. red Therefore, before the down shift by 6, i.e., performing the right shift 209 in equation (2) (with the number of bits 211 being 6), the term (A) can be stored in 16 bits (i.e., 2 bytes). i -c i )·r red +(1<<(d i -1)), where (A i -c i )·r red +(1<<(d i −1)) represents the intermediate result 108 ′ obtained by the matrix-vector product 206 for each component 210 of the output vector 208 .
[0156] According to an embodiment, the device 54 is configured to perform the matrix-vector product 206 by applying a right shift 209 to the intermediate result 108′ (eg (A)) obtained by the matrix-vector product 206 for each component 210 of the output vector 208. i -c i )·r red +(1<<(d i -1))) and the matrix-vector product 206 between the input vector 102 and the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode 205 is calculated in fixed-point arithmetic. Optionally, the intermediate result 108' is represented with a bit precision that is at least twice as high as the bit precision with which the entries of the prediction matrix 19 associated with the matrix-based intra prediction mode are stored, for example, the intermediate result 108' may be stored with 16 bits and the entries of the prediction matrix 19 may be stored as 7-bit magnitudes or as signed 8-bit representations.
[0157] According to an embodiment, the device is configured to decode the picture at a 10-bit resolution, for each matrix-based intra prediction mode 2051-205 n The magnitudes of the entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode are stored with 7-bit precision, and 6 bits are used as the number of bits 211 .
[0158] Similar to constraint 2), constraint 3) enables a more efficient implementation of equation (2) (again saving a table lookup).
[0159] The problem that this application intends to solve is how constraint 2 or constraint 3 should be satisfied together with constraint 1 so that it is not obvious that equation (2) can be used as an approximation of equation (1).
[0160] For example, suppose that by constraint 1, each MIP matrix A i All matrix entries must be stored with 8-bit precision, and by constraint 2, the fixed shift d i = 6 must be used for all MIP modes i in equation (2). Also, for simplicity assume that c i = 0. This would mean that if equation (2) should approximate equation (1), then for each floating-point matrix as a result of the training algorithm of the MIP There must be a non-negative integer c i , making Each item Must meet
[0161] where ε is appropriately small so that
[0162] Equation (2) can be performed to moderately approximate Equation (1). Here, rounding is applied to each entry of the matrix. In addition, for real numbers x, we define
[0163] clip(x, 2 7 ):=min(max(x,-2 7 ), 2 7 -1)
[0164] And by clip(A,2 7 ) indicates that A is a matrix, which is obtained by clip(-, 2 7 ) is applied to each entry of A. Note that if we assume that there is a matrix A′ with integer entries in the 8-bit range i (For which equation (2) for a variable input vector r with fixed shift 6 red Approximate equation (1)), then the matrix A defined in assignment (4) i Must be a reasonably good approximation
[0165] On the other hand, a priori, there is no reason why a training algorithm (whose output is a MIP matrix) ) should produce the corresponding matrix A defined by (4) i Approximate matrix The problem is May contain absolute values greater than , and thus the clipping in equation (4) is achieved by discarding the matrix The most significant matrix entries of introduce substantial differences between (1) and (2). Therefore, applying (4) to the posterior matrix of the training May result in the codec being specified and requires the use of the matrix A i The MIP prediction mode that performs the matrix-vector product in equation (2) deviates greatly from the “real” behavior (i.e., using the matrix Thus, the whole concept of a data-driven approach to MIP-enabled intra prediction would be violated.
[0166] In fact, it can be observed that applying Equation (4) to the training matrix underlying the MIP pattern used in the current VVC draft [1] This significantly changes the behavior of some MIP modes when compared to the base training mode, because the matrix Some of them contain terms much larger than 2.
[0167] Finally, note that solving only constraint 1 without constraint 2 is trivial, as long as each matrix The entry is at -2 7 and 2 7 -1, this is the matrix for the MIP mode that supports the current VVC draft [1] Here, it is assumed that the current c i = 0, simply define the shift value d i , making
[0168]
[0169] for Each matrix entry of holds, and (5) holds for any d′ i >d i Not true.
[0170] Further, to illustrate, the device 54 may include Figure 6.1 and 6.2The described features and / or functionalities. This means, for example, that the matrix-based intra prediction modes 2051-205 n List 204 includes one or more matrix-based intra prediction mode pairs 212. Note that list 204 may not consist exclusively of such MIP mode pairs 212 (as they are depicted in FIG. Figure 6.1 Instead of being present in the list 204 in
[0045] , there may also be other MIP modes that are applied specifically using the transposed option or the non-transposed option. For each matrix-based intra prediction mode 2051-205 n 2k+1, the prediction matrix 19 associated with the first matrix-based intra prediction mode of the corresponding matrix-based intra prediction mode pair 212 is equal to the prediction matrix 19 associated with the second matrix-based intra prediction mode of the corresponding matrix-based intra prediction mode pair 212, e.g., the same matrix 19 is used for modes 2k and 2k+1. The device is configured such that if the matrix-based intra prediction mode 205 pointed to by the mode index 200 is the first matrix-based intra prediction mode of the corresponding matrix-based intra prediction mode pair 212, the association of the reference samples 17 in the neighborhood of the predetermined block with the components 214 of the input vector 112 and the association of the sample positions 104 of the predetermined block 18 with the components 210 of the output vector 208 are transposed relative to the associations if the matrix-based intra prediction mode 205 pointed to by the mode index 200 is the second matrix-based intra prediction mode of the corresponding matrix-based intra prediction mode pair 212. That is, if in the former case a certain component of the input vector 102 is associated with the position (x, y) (where (0, 0) represents the top left corner sample of the predetermined block 18), then in the latter case it is associated with (y, x). The same applies to the components of the output vector 208. For more details, see Figure 6.1 and 6.2 Description.
[0171] Further, to illustrate, the device 54 may include Figures 5.1 to 5.4 The features and / or functionality described.
[0172] According to an embodiment, the device 54 is configured to predict samples of the predetermined block 18 offset from sample positions associated with components 210 of the output vector 208 by upsampling and / or interpolation on the basis of the output vector 208 or on the basis of reference samples 17 in the neighborhood of the output vector 208 and the predetermined block 18, such as e.g. Figures 5.1 to 5.4 shown in one of the .
[0173] According to an embodiment, the device 54 is configured to derive the input vector 102 from the reference samples 17 in the neighborhood of the predetermined block 18 by downsampling and / or pooling, such as e.g. Figures 5.1 to 5.4 shown in one of the .
[0174] According to an embodiment, the reference samples 17 in the neighborhood of the predetermined block 18 include a first reference sample 17 c above the predetermined block 18 and a second reference sample 17 a to the left of the predetermined block 18. The device 54 is configured to derive the input vector 102 from the reference samples 17 in the neighborhood of the predetermined block 18 by deriving a first intermediate component from the first reference sample 17 c by downsampling and / or pooling, deriving a second intermediate component from the second reference sample 17 a by downsampling and / or pooling, concatenating the first intermediate component and the second intermediate component to derive a preliminary input vector, and forming the input vector from the preliminary input vector.
[0175] According to an embodiment, the device 54 is configured to decode a picture with a B-bit resolution. The device 54 may be configured to decode the picture by subtracting 2 from the first component of the preliminary input vector B-1 208. The device 54 further comprises a first component of the preliminary input vector 208 and a first component of the preliminary input vector 208. The first component of the preliminary input vector 208 is then subtracted from the first component of the preliminary input vector 208 to obtain the first component of the input vector 202, and the first component of the preliminary input vector 208 is then subtracted from the first component of the preliminary input vector 208 to obtain the first component of the input vector 202. Alternatively, the device 54 may be configured to form the input vector 102 from the preliminary input vector 208 by subtracting the first component of the preliminary input vector 208 from the first component of the preliminary input vector 208 so that the input vector 102 is formed from the further component. Furthermore, the device 54 is configured to correct the output vector 208 by component-by-component addition of the first component of the preliminary input vector 208.
[0176] According to an embodiment, the matrix-based intra prediction modes 2051-205 in the matrix-based intra prediction mode list 204 are n The entries of the prediction matrix correspond to the entries in Table 2 below, see for example Listing 2. However, it is noted that another shift value (i.e., the number of another bits 211) could be chosen for listing the values in the table, and possibly the values in the table could be represented in another scale.
[0177] The following embodiments will focus on data-driven training of matrix-based intra prediction modes with predefined fixed coefficient ranges and predefined fixed shifts and their application in codecs.
[0178] The solution presented in this invention to the problem of obtaining a fixed bit depth, a fixed shift and a fixed offset is to already in the training of the MIP prediction mode, i.e. in the matrix In the derivation of , constraints 1, 2 and 3 are included. Therefore, the ranges of all matrix entries have been restricted during training, where a gradient descent algorithm is applied to continuously guide the matrix towards a (local) optimum for a predefined loss function on a large set of training data.
[0179] The simplest way to do this would be to multiply each matrix by 2 d , d as in constraint 2, then add the offset c from constraint 3 (if desired), then clip the result to the desired range of constraint 1, then subtract the offset c, and finally divide the result by 2 d However, this is not feasible because the clipping function has a gradient of zero outside the clipping range, and therefore, in such methods, every weight that falls outside the clipping range at some point in stochastic gradient descent will never be updated from that point on.
[0180] FIG8 shows a method for training matrix-based intra prediction models 2051-205 n Prediction matrix 19 of list 204 i In an embodiment of the device 310, one of the matrix-based intra prediction modes 205 i The input vector 102 derived from the reference samples 17 in the neighborhood of the predetermined block 18 and the matrix-based intra prediction mode 205 selected for the predetermined block 18 should be selected for the predetermined block 18 by calculating i Correlated prediction matrix 19 i , and predicting the samples 108 of the predetermined block 18 by associating components 210 of the output vector 208 obtained by the matrix-vector product 206 to the sample positions 104 of the predetermined block 18 .
[0181] The device 310 is configured to train 320 the matrix-based intra prediction modes 2051-205 using a gradient descent method 322 n Prediction matrix 19 of list 204 i . Prediction Matrix 19 i The training 320 is performed, for example, by using a training set of predetermined blocks 18 of known original samples and their corresponding neighborhoods 17. The matrix-based intra prediction modes 2051-205 represented in floating point representation are optimized using a cost function 324. n Prediction matrix 19 of list 204 i The representative value of the item, comparison To train 320 prediction matrix 19 i The cost function 324 depends on the prediction distortion measure 326, which is related to the prediction matrix 19. i The terms of f(x) are set to intermediate values onto which the representative values are mapped using a differentiable function 328, comparing f(x). The cost function 324 depends on the prediction distortion measure 326, for example such that the cost 325 increases with decreasing prediction quality, e.g., from This means that f(x) is applied to The prediction distortion measure 326 may define the prediction signal obtainable using a prediction matrix including intermediate values. and the deviation between the original signal associated with the predetermined block of the training set. As can be seen in Figure 8, When the signal is equal to the original signal, the predicted distortion measure 326 is zero. The entries of can be the representative values mentioned above, and the matrix The entries of can be the intermediate values mentioned above. It is noted that the gradient descent method 322 with the cost function 324 and the dependence of the cost 325 on the prediction distortion measure 326 is only schematically shown to illustrate the underlying principle of the basic device 310 .
[0182] The domain of the differentiable function 328 (e.g. Figure 9 x-axis 302 in ), and co-domain (e.g. Figure 9 The y-axis 304 in FIG. 3 is defined by a floating point representation, the image 300 of the differentiable function 328 has a predetermined dynamic range, and the differentiable function 328 is for the matrix-based intra prediction modes 2051-205 n are equal. The predetermined dynamic range is defined, for example, by max(image 300) / min(image300). According to an embodiment, max(image 300) is α+δ and min(image 300) is -α+δ, see equation (7) below. Note that, Figure 9 Only a graph of an exemplary differentiable function f(x) 328 is shown.
[0183] Furthermore, the device 310 is configured to quantize 330 the intermediate values onto the fixed point representation 190, for example after training 320, so that for each matrix-based intra prediction mode 2051-205 n , and the corresponding matrix-based intra prediction mode 205 i Correlated prediction matrix 19 i Having all terms represented by a fixed point representation 190 with a predetermined bit depth 192, such that the predetermined bit depth 192 is sufficient for matrix-based intra prediction modes 2051-205 n are equal, and such that for each matrix-based intra prediction mode 2051-205 n , and the corresponding matrix-based intra prediction mode 205 i The matrix vector product 206 between the associated prediction matrix 19 and the input vector 102 is performed by applying a matrix-based intra prediction mode 2051-205 to each component of the output vector 208. nThe number of bits 211 equal to the number of bits 211 that can be calculated by performing the right shift 209, for example, means that the non-shifted portion b of the fixed point representation x+1 to b y+x It is sufficient to represent α, see equation (7) below. Note that FIG8 shows the fixed point representation 190 as an (x+y+1)-bit signed magnitude representation. However, it is also possible that the fixed point representation 190 is an (x+y)-bit magnitude representation, for example, in the case where all intermediate values have the same sign (e.g., all intermediate values may be positive).
[0184] Therefore, as a solution, in the present invention, the clipping operation is approximated by a smooth function, such as the differentiable function 328. More precisely, from constraints 1, 2 and optionally 3, the range of unscaled matrix entries, i.e., the range of representative values of the predicted matrix entries, is calculated, and compared. If applied during training
[0185]
[0186] in represents the current matrix during training, then the result is within the bounds of constraint 1. Then, during training, each is converted to Clip to this range:
[0187]
[0188] Wherein α, β, γ and δ are real numbers that depend on the clipping range (i.e., the predetermined dynamic range). In addition, λ is a non-negative integer that can be selected experimentally. Figure 9 An example of a clipping function f(x), ie, an example of a differentiable function 328, is shown.
[0189] According to an embodiment, differentiable function 328 has a slope of 1 at the origin, is strictly monotonically increasing, and has horizontal asymptotes at the upper and lower bounds of image 300 .
[0190] According to an embodiment, the differentiable function 328 is parameterizable by a shift parameter (e.g., δ in terms of the shift of the image 300 within the common domain 304). Furthermore, the device 310 may be configured to subject the shift parameter to an optimization using a gradient descent method 322. Furthermore, the device 310 may be configured to derive an offset value c from the shift parameter to be used for offsetting, for example by addition or by subtraction, all entries of the prediction matrix 19 associated with the corresponding matrix-based intra prediction mode for each matrix-based intra prediction mode before calculating the matrix-vector product 206. The derived offset value c is for each matrix-based intra prediction mode 2051-205 n are equal.
[0191] In summary, the invention of this application is an implementation of a MIP having a portion given by equation (2) satisfying constraints 1 and 2, or an implementation of a MIP having a portion given by equation (2) satisfying constraints 1, 2 and 3, and in both cases, in a floating point matrix that is then quantized into an integer matrix In the training algorithm of , constraints 1 and 2 are employed as described in this section, and if desired, constraint 3 is also employed.
[0192] The following embodiments describe examples of storage representations of prediction matrices.
[0193] The following table shows the floating point matrices obtained from training for MIP mode for minpSizeId=2, [1] using the techniques presented in this application In detail, the parameters of the clipping function (ie the differentiable function) are chosen in such a way that the training generates matrix coefficients that can be represented using 7-bit unsigned integer numbers with a fixed shift of 6 and a fixed offset of 32.
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220] Listing 1: Floating-point matrix coefficients obtained from training
[0221] The matrix coefficients shown in Listing 1 meet the requirements presented in the previous section. To illustrate this, Listing 2 shows the matrix coefficients multiplied by 2. 6 The matrix coefficients after . The range of these coefficients deviates from the final range by only a fixed offset of 32. Therefore, the values below are the stored matrix entries in fixed point representation according to the example. The matrix is a 7×64 matrix (7 component input vector and 64 component output vector). There are 6 matrices for 6 modes. According to the embodiment, the entries can deviate from the values shown below. For example, multiplication by 2 is chosen for illustration purposes only 6 , and accordingly, when another factor is chosen, the entries of the matrix (as they are shown below) may look different.
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238]
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247] Listing 2: Scaled floating-point matrix coefficients
[0248] Now, adding a fixed offset of 32 to these matrix coefficients results in a set of matrices where all coefficients are equal to or greater than -0.5 and are therefore rounded to non-negative values. This set is shown in Listing 3.
[0249]
[0250]
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258]
[0259]
[0260]
[0261]
[0262]
[0263]
[0264]
[0265]
[0266]
[0267]
[0268]
[0269]
[0270]
[0271]
[0272]
[0273]
[0274]
[0275] Listing 3: Scaled floating-point matrix coefficients after adding a constant offset
[0276] Finally, the above matrix coefficients are rounded to integer precision, that is, the intermediate values are quantized 330 to the fixed point representation 190. Since the minimum coefficient of the above set is -0.5, the resulting integer coefficients are non-negative and therefore can be represented by unsigned integer numbers. The maximum coefficient of the above matrix set is 127.5. These coefficients would be rounded to 128 usually, but are rounded to 127 here, having introduced identical absolute rounding errors. Therefore, the resulting integer coefficients shown in List 4 are from the unsigned 7-bit scope. In addition, since the minimum rounding coefficient is 0 and the maximum rounding coefficient is 127, the resulting integer coefficients fully utilize this scope.
[0277]
[0278]
[0279]
[0280]
[0281]
[0282]
[0283]
[0284]
[0285]
[0286]
[0287]
[0288]
[0289]
[0290]
[0291] Listing 4: Unsigned 7-bit integer coefficients
[0292] Although some aspects have been described in the context of equipment, it is clear that these aspects also represent the description of the corresponding method, wherein the block or device corresponds to the feature of the method step or method step. Similarly, the aspects described in the context of the method step also represent the description of the corresponding block or item or feature of the corresponding device. Some or all of the method steps can be performed by (or using) hardware devices, such as microprocessors, programmable computers or electronic circuits. In some embodiments, one or more of the most important method steps can be performed by such devices.
[0293] The data stream of the present invention may be stored on a digital storage medium or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0294] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or software. The implementation may be performed using a digital storage medium having stored thereon electronically readable control signals, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, which cooperates with (or is capable of cooperating with) a programmable computer system to cause the execution of the corresponding method. Thus, the digital storage medium may be computer-readable.
[0295] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0296] Generally, the embodiments of the present invention can be implemented as a computer program product with a program code, when the computer program product runs on a computer, the program code is operative for performing one of the methods. The program code may, for example, be stored on a machine-readable carrier.
[0297] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0298] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0299] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, digital storage medium or recorded medium is generally tangible and / or non-transitory.
[0300] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.The data stream or the sequence of signals may, for example, be configured to be transmitted via a data communication connection, for example via the Internet.
[0301] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0302] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.
[0303] A further embodiment according to the present invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) to a receiver a computer program for performing one of the methods described herein. For example, the receiver may be a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.
[0304] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can collaborate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0305] The devices described herein may be implemented using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0306] The devices described herein or any component of a device described herein may be implemented at least partially in hardware and / or in software.
[0307] The methods described herein may be performed using a hardware device or using a computer or using a combination of a hardware device and a computer.
[0308] Any component of a method described herein or an apparatus described herein may be performed at least in part by hardware and / or by software.
[0309] The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended that the present invention be limited only by the scope of the appended patent claims and not by the specific details presented through the description and explanation of the embodiments herein.
[0310] References
[0311] [1] B. Bross et al., Versatile Video Coding (Draft 7), JVET-P2001, Geneva, October 2019.
Claims
1. An apparatus comprising at least one processor for decoding a block of a picture using intra prediction, the at least one processor being configured to: decoding a mode index from a data stream, the mode index indicating one of a plurality of matrix-based intra prediction modes based on a size of the block; deriving an input vector based on downsampled reference samples adjacent to the block; determining a matrix based on the size of the block and the pattern index; calculating a respective output for each component of a matrix-vector product between the input vector and the determination matrix, the respective output being calculated by performing a right shift by a number of bits that is independent of the matrix-based intra prediction mode indicated by the mode index; and For each component of the matrix-vector product, the corresponding output is used to predict a corresponding sample of the block. The apparatus of claim 1 , wherein the number of bits is 6.
3. The apparatus of claim 1, wherein each matrix-based intra prediction mode of the plurality of matrix-based intra prediction modes applies a corresponding matrix, all entries of the corresponding matrix being represented in a corresponding 7-bit fixed point representation.
4. The apparatus according to claim 1, wherein Based on the size of the block, there are 6, 8 or 16 matrix-based intra prediction modes among the plurality of matrix-based intra prediction modes. 5 . The apparatus of claim 1 , wherein the plurality of matrix-based intra prediction modes have 6, 8, or 16 different matrices associated therewith based on a size of the block. The apparatus according to claim 1 , wherein the reference samples adjacent to the block include a sample located at a top of the block and a sample located at a left side of the block.
7. The apparatus of claim 6 , wherein samples of the adjacent blocks located at the top of the block are downsampled to form a first vector, samples of the adjacent blocks located at the left side of the block are downsampled to form a second vector, and the input vector is derived based on a concatenation of the first vector and the second vector.
8. The apparatus of claim 1, wherein the at least one processor is configured to decode the picture at a 10-bit resolution.
9. A device comprising at least one processor for encoding a block of a picture using intra prediction, the device being configured to: encoding a mode index into a data stream, the mode index indicating one of a plurality of matrix-based intra prediction modes based on a size of the block; deriving an input vector based on downsampled reference samples adjacent to the block; determining a matrix based on the size of the block and the pattern index; calculating a respective output for each component of a matrix-vector product between the input vector and the determination matrix, the respective output being calculated by performing a right shift by a number of bits that is independent of the matrix-based intra prediction mode indicated by the mode index; and For each component of the matrix-vector product, the corresponding output is used to predict a corresponding sample of the block.
10. The apparatus of claim 9, wherein the number of bits is 6.
11. The apparatus of claim 9, wherein each matrix-based intra prediction mode of the plurality of matrix-based intra prediction modes applies a corresponding matrix, all entries of the corresponding matrix being represented in a corresponding 7-bit fixed point representation.
12. The apparatus according to claim 9, wherein Based on the size of the block, there are 6, 8 or 16 matrix-based intra prediction modes among the plurality of matrix-based intra prediction modes.
13. The apparatus of claim 9, wherein the plurality of matrix-based intra prediction modes have 6, 8, or 16 different matrices associated therewith based on a size of the block. 14 . The apparatus of claim 9 , wherein the reference samples adjacent to the block include a sample located at a top of the block and a sample located at a left side of the block.
15. The apparatus of claim 14, wherein samples of the adjacent blocks located at the top of the block are downsampled to form a first vector, samples of the adjacent blocks located at the left side of the block are downsampled to form a second vector, and the input vector is derived based on a concatenation of the first vector and the second vector.
16. The apparatus of claim 9, wherein the at least one processor is configured to decode pictures at 10-bit resolution.
17. A method for decoding a picture block using intra prediction, the method comprising: decoding a mode index from a data stream, the mode index indicating one of a plurality of matrix-based intra prediction modes based on a size of the block; deriving an input vector based on downsampled reference samples adjacent to the block; determining a matrix based on the size of the block and the pattern index; calculating a respective output for each component of a matrix-vector product between the input vector and the determination matrix, the respective output being calculated by performing a right shift by a number of bits that is independent of the matrix-based intra prediction mode indicated by the mode index; and For each component of the matrix-vector product, the corresponding output is used to predict a corresponding sample of the block. The method of claim 17 , wherein the number of bits is 6.
19. The method of claim 17, wherein each matrix-based intra prediction mode of the plurality of matrix-based intra prediction modes applies a corresponding matrix, all entries of the matrix being represented in a corresponding 7-bit fixed point representation.
20. The method of claim 17, wherein Based on the size of the block, there are 6, 8 or 16 matrix-based intra prediction modes among the plurality of matrix-based intra prediction modes.